Institute for Quantum Computing (IQC) main building, Quantum Nano Centre (QNC) at the University of Waterloo
Wednesday, September 23, 2026 - Friday, September 25, 2026 (all day)

ML4QT Symposium | Machine Learning to Advance Quantum Technologies

Fostering exchange and collaboration among experts from academia, research institutions, and industry. 

The Machine Learning to Advance Quantum Technologies (ML4QT) Symposium serves as a leading forum for researchers and developers interested in the convergence of machine learning and quantum innovation. Whether you are actively working in this space or seeking to get involved, ML4QT provides an ideal environment to discover new opportunities, foster collaboration, and help drive the next generation of quantum technologies.

This second annual gathering will be held at the Institute for Quantum Computing, within the Quantum-Nano Centre, University of Waterloo from the 23 to 25 of September 2026

Important Dates

  • July 1, 2026 |Call for Contribution Submission Deadline 
  • July 1, 2026 | General Registration Opens 

  • August 1, 2026 | Contribution Decision Notifications 
  • September 1, 2026 | Registration Closes

  • September 23 - 25 | ML4QT Symposium 

Have questions about the event? Contact us at iqc.events@uwaterloo.ca

Invited Speakers

Barry Sanders, University of Calgary

Barry Sanders headshot

Artificial Intelligence for Representing and Characterizing Quantum Systems

I review how integrating AI into quantum-system characterisation is typically based on machine learning, deep learning and language models for the core tasks of quantum-property prediction and quantum-system reconstruction with applications to certification, benchmarking, enhancing quantum information processing and identifying critical quantum phenomena. https://arxiv.org/abs/2509.04923

Connor van Rossum, University of Queensland

Headshot of Connor van Rossum

Can noise help quantum algorithms? Insights from sampling and variational methods

Quantum algorithms typically claim to be resource efficient relative to classical algorithms but it is well known that this performance differential may vanish under realistic operating conditions or noise. While the field has developed strategies to retain quantum advantage, the interaction between quantum and classical resources is often ignored but critically determines the efficacy of quantum algorithms. Our first result, in the context of noisy quantum variational learning, shows that performance depends on how noise interacts with classical optimisation. I show that commonly used subroutines in error‑mitigation strategies such as Pauli twirling can degrade variational optimisation by suppressing gradients and reducing expressivity. In contrast, biased or non‑unital noise can introduce exploitable structure that improves optimisation outcomes. Through analytical and numerical studies, this work demonstrates that preserving noise asymmetries can lead to better performance than symmetrising noise, challenging standard assumptions about mitigation in the variational setting. In recent times, quantum variational models have evolved towards sample-based quantum algorithms for chemistry. For the specific example of Sample‑based Quantum Diagonalisation (SQD), in our second result, I show that noise enables a higher unique sampling rate that allow classical stochastic algorithms to outperform quantum algorithms for naively designed chemistry demonstrations. I discuss that standard error mitigation is not useful for these algorithms, before showing how machine learning can enable better ansatze design and measurement techniques to not only increase the unique sampling rate but ensure noise-robustness and resource efficiency of these demonstrations.  

Manas Mukherjee, National University of Singapore

Manas Mukherjee Headshot

Seeking advantage in Quantum-Classical Machine Learning with trapped ion system

Quantum technologies, specifically Quantum Computing (QC), represent a paradigm shift in computational power. However, the roadmap to utility is hindered by the inherent fragility of quantum states—a hurdle that makes scaling the industry’s greatest challenge.

To address this, we have developed a quantum-classical hybrid framework designed to push the boundaries of Quantum Machine Learning (QML). By leveraging the trapped-ion platform—noted for its superior coherence times and high gate fidelities—we have demonstrated that quantum classifiers can achieve over 98% fidelity in supervised learning tasks using real-world datasets [1, 2].

While achieving parity with classical systems is a significant milestone, our current research pivots toward the ultimate goal: identifying the specific parameters and architectures where QML provides a definitive advantage over classical counterparts. Here, we will discuss detail of our findings across varied application contexts, highlighting the conditions under which quantum enhancement becomes a reality.

[1] 10.1103/physreva.106.012411

[2] 10.1016/j.isci.2025.113058

Yuval Baum, Q-CTRL

Yuval Baum Headshot

AI enhanced infrastructure software: integrating machine learning across the quantum computing stack

Quantum computers promise to revolutionize computing and have made incredible hardware strides, but they remain notoriously error-prone. While low-level control tweaks and clever circuit design can compensate for these hardware errors, current techniques require the availability of highly detailed noise models, deep algorithmic knowledge, and labor-intensive manual tuning. Because of this, they rarely generalize beyond a few narrow use cases.

In this talk, I will show how machine learning (ML) can be integrated across the entire quantum computing stack to solve this bottleneck. We will explore how ML methods outperform traditional techniques at every level: from device and gate calibration at the bottom, to compilation and error suppression in the middle, up to algorithmic optimization and quantum error correction at the top.

Kunal Sharma, IBM

Kunal Sharma Headshot

Learning Ground State Observables from Quantum Experiments

Quantum machine learning can be viewed not only as a search for speedups in classical machine tasks, but also as a way to learn from quantum data generated by quantum processors. In this talk, I will discuss this perspective through recent work on learning ground-state observables from quantum experiments. We use quantum data from approximate ground states of two-dimensional Heisenberg XXZ model, constructed using samples from IBM Heron quantum processors and classical high-performance computing, to train neural networks that predict observables across Hamiltonian parameter space. The results show accurate generalization to unseen parameters, suggesting a path toward using quantum computers as data generators for machine learning in many-body physics. 

Jem Guhit, Quantinuum

Jem Guhit Headshot

Learning Quantum State Preparation with Generative AI

Preparing quantum states efficiently remains a central challenge in quantum chemistry. This talk presents recent work on using generative AI to learn molecular ground-state preparation circuits directly from molecular Hamiltonians, replacing expensive iterative circuit construction with learned generation. I will discuss the insights gained from this approach, its current capabilities, and limitations, and conclude with an outlook on generative quantum eigensolvers and future directions for machine learning in quantum algorithm design.

Lirandë Pira, National University of Singapore

Lirande Pira Headshot

Computational Structure of Learning in Quantum Systems

How can quantum systems learn from data? Quantum learning problems arise naturally when quantum systems are used as models, as data sources, or as physical platforms.

This talk examines learning in quantum settings through two complementary perspectives: using quantum systems as learning models, and learning quantum systems themselves from data. On the modelling and designing side, a recurrent theme is the role of structure in determining learnability. Structured operator representations, together with tools from quantum linear algebra, lead to quantum models whose complexity and behavior can be analyzed more systematically.

From a complementary perspective, one can view tasks such as reconstructing states, channels, or Hamiltonians as problems of learning quantum systems. This connects naturally to ideas from system identification and operator approximation.

I will also discuss recent directions on coherent training protocols, where model parameters are treated as quantum degrees of freedom and optimized through quantum evolution. More broadly, the goal is to better understand what makes quantum systems learnable, and how these principles shape the design of quantum learning and computational models.

Pavithran Iyer, Xanadu

Pavithran Iyer Headshot

Data-Driven Belief Propagation Decoding of quantum LDPC Codes under Realistic Noise

Quantum error correction enables reliable logical computation on noisy hardware, but its guarantees typically rest on simplified noise models and incur large physical overheads. Real devices, however, exhibit correlated errors that fall outside these models. We enhance the performance of quantum low-density parity-check (qLDPC) codes, a family of codes which contain candidates with the lowest known overheads for fault tolerance, under spatially correlated noise. We leverage a correspondence between belief propagation (BP) decoding of a qLDPC code and a neural network, in which the message-passing iterations of BP unroll into the layers of the neural network. This correspondence opens a wide scope for data-driven generalizations of BP. We train the neural network weights on characterization data obtained via Cycle Error Reconstruction, yielding a decoder adapted to the error model of the underlying hardware. Our neural network decoder achieves up to a tenfold reduction in logical error rate over standard BP across diverse noise regimes.

Wolfgang Mauerer, OTH Regensburg

Wolfgang Mauerer Headshot

Machine Learning, Quanta, and the Extraction of Physical Compute Power

Quantum computing’s industrial and scientific impact depends on more than established computational advantages; it requires progress across algorithms, architecture, engineering, and foundational concepts. This talk advocates a vertically integrated, interdisciplinary systems perspective. We discuss how machine learning can help extract greater value from quantum systems by reducing unnecessary classical computation in hybrid algorithms and improving use of available quantum resources. We argue that this broader framing can reveal deeper structural origins of quantum advantage and inspire new views of learning.

Kohei Nakaji, NVIDIA

Kohei Nakaji Headshot

Toward ML for Quantum Algorithms at Practical Scale

How to apply modern machine learning techniques to quantum algorithms at practically relevant scales remains an open question. We will explore this question through the Generative Quantum Eigensolver (GQE) and its variants, which use generative machine learning to discover quantum circuits for quantum chemistry. We will then discuss MoLe, an approach for constructing molecular wave functions that can transfer across different molecules, highlighting a path toward scalable and transferable machine learning for practical quantum algorithms.

Estelle Maeve Inack

Data-enhanced self-learning projection quantum Monte Carlo with Rydberg-atom measurements

Projection quantum Monte Carlo (PQMC) provides unbiased ground-state estimates for stoquastic quantum many-body Hamiltonians, but its practical efficiency depends strongly on the quality of the guiding wavefunction used for importance sampling. We introduce a data-enhanced self-learning PQMC framework in which projective measurements from a programmable Rydberg-atom quantum simulator are used to pretrain neural guiding wavefunctions before the standard self-learning refinement. We apply the method to the square-lattice Rydberg Hamiltonian across the disordered-to-checkerboard transition, using both restricted Boltzmann machine (RBM) and autoregressive recurrent neural-network (RNN) ansätze. Pretraining on experimental measurement data substantially improves the initial guiding distribution, accelerates convergence of the mixed energy estimator, reduces finite-walker bias, and lowers the number of walker and training iterations required to reach accurate ground-state energies. The improvement is especially pronounced for RBM-guided wavefunctions, where data pretraining mitigates convergence to poor local minima, while RNN-guided wavefunctions benefit from faster convergence and improved sample efficiency. We further show that the advantage persists across system sizes, with the residual energy exhibiting a weaker walker-number scaling than in conventional self-learning PQMC. These results demonstrate that imperfect but physically informative quantum-simulator data can be used as a resource to enhance classical projector Quantum Monte Carlo simulations, suggesting a practical hybrid route for accurately studying quantum matter with near-term quantum devices.

Justyna Zwolak, NIST

Closing the Loop on Quantum Dot Control: Real-Time Feedback for Scalable Spin-Qubit Arrays

Semiconductor spin qubits have progressed from few-qubit demonstrations toward increasingly complex quantum-dot arrays fabricated in reproducible academic and industrial processes. As these systems scale, a central challenge is no longer simply finding a good operating point, but maintaining that operating point in the presence of electrostatic drift, charge rearrangements, device variability, and many interdependent controls. I will describe recent work toward feedback-enabled control of semiconductor quantum-dot devices, where machine-learning-based measurement interpretation, image-processing methods, physics-informed heuristics, and automated control routines are combined to detect changes in device state and apply corrective updates in real time.

I will focus on recent results from a 10-dot planar germanium device, including more than 48 hours of autonomous monitoring, detection of electrostatic shifts as small as 0.1 mV, real-time feedback compensation, and spatially resolved noise-correlation measurements. I will also discuss a recent live demonstration of this approach and the broader workflow requirements for deploying feedback across larger arrays. These results complement our modular control stack for quantum-dot devices, where standardized intermediate data products, explicit interfaces between tuning modules, and workflow-level performance metrics enable scalable, autonomous workflows and reliable operation.

Roman Krems, University of British Columbia

On the quantum advantage, molecular descriptors and model selection metrics for quantum machine learning

I will first demonstrate that quantum fidelity kernels can be proven to have a quantum advantage, in principle. How to take advantage of this quantum advantage in practice remains unclear. To address this challenge, I will describe several algorithms for building high-performing quantum circuits for specific machine learning tasks. I will discuss the isomorphism between quantum circuits and a subspace of polyatomic molecules and argue that molecules can be used as descriptors of quantum models. Finally, I will present an algorithm for improving quantum models using the Bayesian information citron as well as the dimension of the dynamical Lie algebra as model selection metrics.

Aida Ahmadzadegan-Shapiro, D-Wave

Beyond Optimization: Quantum Annealing for Generative Machine Learning

Quantum annealers are commonly viewed as optimization engines, but their ability to sample from complex energy-based distributions provides another natural interface with machine learning. In this talk, I will describe a hybrid architecture in which a quantum processing unit trains and samples an energy-based prior embedded within the discrete latent space of a classical generative model. 

Using molecular generation as a case study, I will discuss how continuous neural representations are mapped to binary latent variables, how a graph-restricted Boltzmann-machine prior is trained using quantum annealing, and how samples from the learned distribution are returned to a classical decoder. Compared with the classical sampling baselines studied, the QPU-trained models generated molecules with higher validity and drug-likeness. 

Rather than replacing the classical ML model, the quantum processor addresses a specific sampling task within it. I will conclude with lessons from this work on where quantum sampling may fit within existing ML workflows, how these hybrid approaches should be benchmarked, and what makes a problem a promising candidate for further exploration.

Christine Muschik, University of Waterloo

Topic TBC

Nathan Wiebe, University of Toronto

Topic TBC

Kevin Qu, University of Waterloo

Topic TBC

Virginia Frey, University of Waterloo

Topic TBC

Pierre-Emanuel Emeriau, Quandela

Topic TBC

Contributed Talks

Carlos Benavides-Riveros, IQM Quantum Computers

Neural Quantum Propagators: Operator Learning for Open Quantum Dynamics

Simulating the dynamics of open quantum systems is a central challenge for quantum technologies, yet conventional numerical methods become prohibitively expensive for large or strongly driven systems. Here, we present an operator-based machine learning framework that addresses this challenge by learning time evolution operators directly, rather than approximating individual wavefunctions or expectation values. We introduce neural quantum propagators (NQP), a universal neural network architecture for driven-dissipative quantum dynamics that handles arbitrary initial states, adapts to various external driving fields, and generalizes to long-time dynamics beyond its training window. Complementarily, within the non-Markovian quantum state diffusion formalism, we develop an operator construction algorithm that reconstructs stochastic time evolution operators from ensembles of quantum trajectories, enhancing both interpretability and transferability across related systems. We benchmark both methods on the spin-boson model across diverse spectral densities and on the three-state transition Gamma model, demonstrating strong accuracy and practical utility — including computation of absorption spectra and reconstruction of reduced density matrices at extended timescales. By shifting the learning target from states to operators, our framework unlocks a more powerful and flexible paradigm for deploying machine learning in quantum simulation.

Einar Gabbassov, University of Waterloo

Exact Stochastic Schrödinger Equations for Quantum Reverse Diffusion

A few years ago, the field of machine learning (ML) and, more broadly, our culture were shaken by advances in diffusion-based generative models such as Midjourney, DALL-E, Sora, Veo, etc. The breakthrough in the generative capabilities can be attributed mainly to the 40-year-old mathematical theory of classical reverse diffusion by B. D. Anderson. He showed that a Markov diffusion process can have a reverse diffusion process, e.g., noisy data dynamically and stochastically evolve into noise-free data.

Following advances in classical ML, the field of quantum generative modelling began to gain traction and expand. The core idea of quantum generative modelling is to define a quantum analog of a naturally occurring noisy forward process and then use hybrid techniques, including variational quantum circuits, to imitate the quantum reverse process. While these quantum ML techniques are powerful for learning a surrogate of the reverse, they do not reveal which physical principles fundamentally define a natural quantum reverse process, which, by definition, must incorporate the same noise and decoherence effects as the forward process. In other words, most current works do not realize the reverse process and instead imitate it using variationally learned and fully coherent dynamics. Although these numerical schemes imitate quantum reverse dynamics, analytical equations describing a physical quantum reverse diffusion process, on the same footing as the  forward stochastic Schrödinger equation (SSE), have so far been absent.

In our work, we provide a rigorous theoretical foundation for quantum reverse diffusion in continuously monitored quantum systems. Specifically, for forward quantum stochastic processes driven by monitored Pauli noise, we derive a family of exact stochastic Schrödinger equations that describe the corresponding reverse quantum processes. These reverse stochastic Schrödinger equations are generalizations of the forward SSE and, as such, preserve the noise and decoherence structure of the forward dynamics. Furthermore, the reverse processes are mathematically guaranteed to reverse the noise effects almost surely.  The stronger almost sure reversal of the forward dynamics can be relaxed by configuring the reverse process to steer the state onto a manifold of states; in this case, the dynamics implement a reversal in the distribution. Therefore, the presented reverse SSEs are powerful, as they enable a wide spectrum of applications, from almost sure state recovery to quantum generative modelling.

Sreeraj Rajindran Nair, University of Technology Sydney

Local tensor-train surrogates for quantum learning models

A key bottleneck in quantum machine learning is the computational cost of repeated quantum circuit evaluations during the inference phase. To address this, we present a framework for constructing fast, cheap, provably accurate classical tensor-train surrogates of fully trained quantum machine learning models within local patches of their input data space. The approach combines Taylor polynomial approximation with a tensor-train (TT) representation and embeds it in a statistical learning paradigm via empirical risk minimization. In our analysis, the Taylor-TT construction serves as a deterministic error certificate proving that the TT hypothesis class contains a good approximation; empirical risk minimization then provably recovers a surrogate with controlled generalization error and explicit bounds. This translates into three independently controllable error sources: (i) Taylor truncation error controlled by the patch radius $r$ and polynomial degree $p$, (ii) TT approximation error controlled by the bond dimension $\chi$, and (iii) statistical estimation error. While the parameter count scales polynomially in the number of data dimensions $N$, i.e., $\deff = N(p+1)\chi^2$ rather than the naive $(p+1)^N$, the worst-case constants inherit an exponential factor through the tensor-product feature norm during Taylor polynomial embedding onto TT. This cleanly separates representation complexity from feature-induced constants. Our risk bounds and sample complexity depend explicitly on the local patch radius $r$. the arXiv preprint is available at https://arxiv.org/abs/2604.25631.
 

Andrew Zhao, Sandia National Laboratories

Learning fermionic linear optics with Heisenberg scaling and physical operations

Fermionic linear optics (FLO), equivalent to fermionic Gaussian unitaries or matchgates on a line, have wide applications ranging from device benchmarking, classical shadows, quantum neural networks, and more. We revisit the problem of learning FLO circuits: given black-box query access to an unknown N-qubit FLO, produce an approximate description using as few resources as possible. Previous proposals featured 1) N^5 query complexity; 2) standard quantum limit scaling in precision; 3) unphysical operations violating fermionic superselection rules; and 4) N auxiliary quantum space. In this work, we establish efficient and experimentally friendly protocols that address all of these deficiencies: 1) we improve to N^4 scaling in general, which is further reduced to N^3 for number-conserving (passive) FLOs; 2) we achieve the optimal Heisenberg limit; 3) all operations obey superselection; and 4) we only require 1 ancilla qubit, which can be removed in most physically-relevant scenarios. This brings the task of learning FLOs closer to practical implementation, and marks the first such algorithm that attains Heisenberg scaling.

This work is joint with Aria Christensen. A preprint is available at: https://arxiv.org/abs/2602.05058

Prateek P. Kulkarni, PES University

How Fine Can You Slice It? Quantum Speedups for Certifying Noisy Linear Classifiers

Linear separability — the question of whether a labeled dataset is consistent with a halfspace classifier — is a foundational primitive in machine learning. In noisy or corrupted settings, the relevant question is not binary consistency but distance: how far is a dataset from being linearly separable? This is the tolerant testing problem, and its query complexity governs how efficiently one can certify the quality of a classifier from limited data access.

We initiate the study of quantum query complexity for tolerant geometric testing, focusing on linear separability as the central example. Our main result is a quantum algorithm achieving a quadratic speedup over the classical tight bound: we test (epsilon_1, epsilon_2)-linear separability of a labeled point set in R^d using O-tilde(sqrt(d) / sqrt(epsilon_2 - epsilon_1)) quantum queries, versus the classical Theta(d / (epsilon_2 - epsilon_1)). We also give a matching quantum lower bound and a quantum algorithm for distance estimation — approximating how far a dataset is from separable — with a polynomial improvement over classical bounds in the query complexity.

The technical engine is a density lemma for violating bases: if a dataset is epsilon-far from separable, a uniformly random (d+1)-subset is a certificate of inseparability with probability at least (epsilon/2)^(d+1). Feeding this density into quantum amplitude estimation yields the speedup. We show this proof strategy is an instance of a general paradigm — tolerant testing via density amplification — that also recovers known quantum speedups for stabilizer state testing, suggesting a unified framework for quantum-accelerated certification tasks in quantum machine learning.

We discuss implications for quantum PAC learning, quantum sample complexity, and the broader question of when quantum access to training data yields provable advantages for learning-theoretic tasks.

Jun Dai, Mila/University of Montreal

Generative Learning of Optimal Quantum Measurements

Efficiently measuring quantum states is a central challenge in quantum computing. In many applications, we need to estimate a large number of Pauli observables, many of which do not commute, while using as few measurements as possible. A standard approach is to group observables into commuting sets and measure each group together. However, general commuting groups can require relatively deep basis-change circuits, which are difficult to implement on near-term devices and may still be costly in early fault-tolerant settings. A more hardware-friendly option is to use only single-qubit rotations and qubit-wise commuting groups. This keeps the measurement circuits shallow, but usually produces smaller groups and therefore increases the number of samples needed. In this talk, I will present a new measurement scheme based on generative learning. Instead of explicitly constructing measurement groupings through NP-hard combinatorial optimization, or relying on complex derandomization procedures for classical shadows, our method directly learns an ensemble of measurement circuits tailored to a target set of observables and practical resource constraints, such as the total measurement budget and allowed circuit depth. Our numerical results show systematic improvements over state-of-the-art measurement strategies, as well as encouraging generalization beyond the training regime. Overall, this learning-based approach offers a flexible framework for practical measurement design in both near-term and early fault-tolerant quantum computing, where quantum resources remain limited and imperfect.

Youngseok Lee, NORMA

Trainability and Mode Seperation of Mixed IQP circuits

Instantaneous quantum polynomial-time (IQP) circuits are a promising route to generative modeling in the noisy intermediate-scale quantum (NISQ) era, supporting a train-on-classical, deploy-on-quantum paradigm: the low-body Pauli-Z expectation values that assemble a Maximum Mean Discrepancy (MMD) loss are classically estimable, while sampling from the trained circuit remains classically hard in the relevant regime -- so training is fully classical, yet generation retains a genuine quantum-advantage target.

This promise is limited by a tension between expressivity and trainability. An ancilla-free IQP circuit is not universal -- even on two qubits there are distributions no choice of angles reproduces -- yet restoring universality by appending an ancilla register reinstates the barren plateaus that obstruct training. We resolve this tension by working in the decomposed picture of an ancilla-augmented IQP circuit: a weighted mixture of ancilla-free IQP circuits (branches) that share one interaction graph but carry independent angles, which we call a mixed IQP.

Using the mixed IQP, we identify three data-informed initializations that keep the ancilla-augmented circuit trainable, and prove that its training succeeds only when the branches are seeded to carry distinct distributions. We show that the mixture inherits the barren-plateau avoidance of a single circuit for a polynomial number of branches, with only a polynomially small suppression of the loss curvature. Building on this, we introduce a data-partitioning cluster initialization that seeds each branch from a distinct data mode, prove it is trainable, and show it supplies the inter-branch diversity a mixture needs to surpass the ancilla-free performance -- whereas branches that collapse to near-identical distributions provably cannot, and a coincident start receives no first-order gradient to differentiate them. Branch weights are promoted to a free control set through the ancilla state, and the trained model is deployed by randomized per-shot selection of a single circuit, with no ancilla overhead at sampling time.

We demonstrate experimentally that adding ancilla branches to the mixed IQP raises its performance well beyond that of an ancilla-free IQP circuit, across four benchmarks spanning system sizes from sixteen to seven hundred eighty-four qubits and three interaction-graph families. This gain appears only when the data separate into distinct modes: on well-separated targets, routing each branch to its own mode lifts performance, while on the largest image target, whose modes overlap in Hamming space, closely spaced modes need only a small inter-branch separation to be told apart -- one the model installs on its own from any initialization, so every scheme fits the data equally well, and the mixture's gain tracks the mode structure of the data rather than the system size. Through the mixed IQP, we thus establish -- both theoretically and experimentally -- the successful training of ancilla-augmented IQP circuits.

João F. Bravo, Fraunhofer

Ravines in quantum cost landscapes: opportunities for improved VQA predictions

The geometric and topological structure of quantum cost landscapes (QCLs) governs the optimization and thus the predictive power of variational quantum algorithms (VQAs). We systematically analyze ravines — low-cost paths connecting local minima — using an adapted version of the nudged elastic band (NEB) algorithm, a method originating from theoretical chemistry. By training quantum neural networks (QNNs) to classify the concentratable entanglement of quantum states, we apply the NEB algorithm and numerically identify ravine structures in QCLs of hardware-efficient ansatzes. Beyond visualizing these ravines, we construct an ensemble prediction framework by averaging predictions from QNNs parameterized along the low-cost NEB path. We introduce a resource-light pre-training metric which quantifies local prediction variability and serves as a strong performance indicator for VQAs, even beyond the scope of this study. When base classifiers are drawn from circuit and weight initializations exhibiting high local-prediction variability, the quantum-based NEB ensembles outperform both classical and naive quantum alternatives. Moreover, a complexity analysis shows that leveraging the ravine-like structure of QCLs with the QNN NEB approach substantially reduces computational costs compared to naive QNN ensembling. A depth and qubit scaling analysis indicates that ravines persist across both scalings, and that, despite the expected growth in resource requirements with the qubit scaling, the NEB approach also accelerates convergence over the naive alternative. (arXiv:2607.01329)

Christian Tutschuku, Fraunhofer IAO

Grokking and epoch-wise double descent in quantum neural networks

Grokking, the delayed transition from memorization to generalization, is a fundamental phenomenon in gradient-based learning, yet its dynamics within variational quantum machine learning (QML) remain largely unexamined. In this work, we report the empirical observation of both the grokking transition and epoch-wise double descent in a two-qubit quantum neural network (QNN) under a complete parameterization of the SU(4) manifold. We demonstrate that overparameterization via increased circuit depth improves the probability of successful generalization. Notably, these architectures frequently exhibit an epoch-wise double descent in test error, degrading at a critical epoch before recovering into a generalizing state. Crucially, we identify a generalization decay in late-stage training, where the test error increases significantly despite a stagnant training loss. Bridging this behavior with algorithmic stability theory, our analysis reveals that this decay correlates with an unconstrained increase of the weight-norm, drifting away from sparse, phase-aligned harmonic solutions toward overfitted solutions in the Hilbert space. We analyze the underlying temporal dynamics of this transition, demonstrating how the onset of generalization is linked to optimization hyperparameters such as learning rate and weight decay. Finally, to mitigate late-stage decay, we introduce a weak explicit weight-norm regularization into the loss function. We demonstrate that this structural anchor stabilizes the post-grokking phase and permanently preserves generalization gains, providing a robust framework for training overparameterized quantum circuits.

ML4QT Organizing Committee

Achim Kempf

Achim Kempf headshot

Christian Tutschku

Christian Tutschku headshot

Christine Muschik

Christine Muschik headshot

Lirandë Pira

Lirande Pira Headshot

Virginia Frey

Virginia Frey headshot

Marco Roth

Headshot of Dr. Marco Roth

Kevin Qu