REVIEW 3 major objections 6 minor 92 references
Learning a factored voltage basis online lets multi-agent RL tune quantum-dot arrays of unseen size in nearly constant steps.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 03:06 UTC pith:CS224CXY
load-bearing objection Solid methods paper: online action factorization plus role-shared MARL gives real sim scaling for QD tuning; hardware gap is the only load-bearing soft spot. the 3 major comments →
Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
QADAPT shows that an online, Kalman-refined estimate of the local gate-to-dot capacitance matrix can reparameterize the joint voltage action space into approximately independent virtual controls. With that factorization, role-shared decentralized PPO policies trained only on four-dot systems zero-shot transfer to larger and smaller arrays while keeping the number of convergence steps nearly constant.
What carries the argument
Adaptive action-space factorization: a lightweight CNN plus Kalman filter builds a time-dependent cross-capacitance matrix Φ_t that maps physical plunger voltages into virtual voltages whose local quadratic objectives become nearly separable.
Load-bearing premise
The online capacitance estimate must stay accurate enough that each virtual voltage update mainly affects only its own dot; if real-device drift, hysteresis or closed-regime diagrams break that estimate, the claimed decoupling and size-independent scaling fail.
What would settle it
Train once on a four-dot simulator, then measure the number of charge-stability-diagram acquisitions needed to reach a fixed tuning score on real two-, six- and eight-dot devices; if steps rise sharply with size or the policy fails to transfer, the central scaling claim is false.
If this is right
- Calibration of larger quantum-dot processors can reuse the same shared policies without combinatorial retraining.
- Measurement budgets for reaching target charge and tunnel-coupling regimes scale roughly linearly rather than combinatorially with array length.
- The same factorization-plus-role-sharing pattern applies, with only input/output changes, to simultaneous-drive tuning of superconducting qubits.
- Centralized world-model or critic architectures become unnecessary once the physical action basis itself is made approximately separable.
Where Pith is reading between the lines
- If the online factorization remains stable under mild hysteresis, the same modular agents could absorb slow parameter drift without full re-optimization.
- The approach suggests a general design rule for other multi-parameter quantum control problems: first learn a physically grounded change of coordinates, then share policies across identical roles.
- Hardware validation on devices that leave the open-reservoir regime would be the decisive next experiment for practical adoption.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces QADAPT, a cooperative multi-agent RL framework for tuning electrostatically defined quantum-dot arrays. It learns an online factored action basis by estimating local gate-to-dot capacitances from charge stability diagrams with a CNN plus Kalman filter, then trains role-shared (plunger/barrier) actor-critic policies in the virtualized coordinates under a Dec-POMDP formulation reduced to a contextual bandit (γ=0). In QArray simulations the method zero-shot transfers from a 4-dot training array to 2–8-dot systems, keeps roughly constant steps-to-target while CSD acquisitions scale near-linearly in N, and outperforms black-box optimizers and several MARL/model-based baselines under a fixed measurement budget. Appendix A supplies local-linearization mismatch and preconditioning bounds; Appendix G sketches an analogous factorization for superconducting-qubit all-XY tune-up.
Significance. If the simulation results transfer, the work addresses a genuine bottleneck: combinatorial growth of joint voltage spaces and cross-talk-induced non-stationarity that currently limit autonomous tuning beyond few-dot devices. The combination of physically motivated online virtualization, role-based parameter sharing, and CTDE is a clear, modular design that is more scalable than centralized world models or joint-action critics. Strengths that should be credited include: (i) explicit mismatch/contraction analysis under local linearization (Appendix A, Thm. A.4, Prop. A.10); (ii) systematic ablations isolating virtualization, Kalman filtering, and multi-agent structure (Table 1); (iii) size sweeps and multiple external baselines (Figs. 2–3); (iv) an honest Limitations section; and (v) released code. These make the paper a useful contribution to quantum-device automation and to structure-aware MARL even if hardware validation remains future work.
major comments (3)
- [§3.2, Appendix A, Fig. 2, §5.3] The central scaling claim rests on Φ_t remaining a stable preconditioner so that C_t Φ_t^{-1}≈I and the joint quadratic objective becomes approximately separable (Appendix A, Thm. A.4 / Prop. A.10; §3.2). Fig. 2a–c and the IPPO / no-Kalman ablations (Table 1) support this only under QArray open-regime dynamics with relatively stationary couplings (Appendix B). §5.3 correctly flags hysteresis, drift, time-varying offsets, and isolated-regime CSD features as risks, but the manuscript never stress-tests them. Without at least simulated non-stationarity (e.g., process-noise sweeps, mid-episode capacitance jumps, or hysteresis), the zero-shot O(N) measurement claim and the hardware-facing abstract conclusion are under-secured. Please add such stress tests or materially scope the claim to stationary open-regime simulators.
- [§3.1, §3.4, Appendix B Table 3] Rewards are dense, immediate, and shaped from distance to a known target voltage configuration v* (and lever-arm normalization), with explicit zero/half-reward radii in Appendix B Table 3; the γ=0 bandit reduction (§3.4) relies on this. That is acceptable for controlled sim benchmarks, but it means agents are not discovering target charge/tunnel regimes from CSD-derived scores alone. For the claim of a scalable route to processor calibration, the manuscript should state clearly what replaces the oracle distance reward on hardware (classifier scores, physical cost functions, etc.) and, ideally, show that performance is not brittle when the reward is replaced by a noisier CSD-based proxy. As written, transferability of the learned policies beyond oracle-shaped sim is unclear.
- [Table 1, Fig. 3, §4] Several baselines that fail hard may not be decisive as currently reported. MADDPG and FACMAC show near-zero convergence after 500k environment steps with noted gradient pathologies (Table 1); without a hyperparameter-sensitivity or architecture-matched continuous-control setup, it is hard to attribute failure solely to the absence of action factorization versus optimization difficulty. DreamerV3 is omitted at 8 dots for compute reasons (Fig. 3). Please either strengthen these baselines (or document search budgets) or frame the comparison more carefully as “under a shared measurement budget and default continuous-control settings,” so the superiority claim remains proportionate.
minor comments (6)
- [Fig. 1] Fig. 1 caption and panel labels are dense; a short legend distinguishing physical vs virtual voltages (v_P vs u_P / vPx) would help readers who skip §3.2.
- [§3.2, Appendix A] Notation for virtual voltages switches among u_p, vPx, and Δu/Δv across §3.2, Appendix A, and the figure; unify symbols.
- [Table 1] Table 1 reports three convergence radii (2%/5%/10%) in a single cell; a multi-column layout would make effect sizes easier to read.
- [Appendix G, Figs. 6–7] The superconducting extension (Appendix G) is interesting but lightly evaluated relative to the main claims; either mark it clearly as a proof-of-concept or add a short quantitative comparison to a non-factored baseline there.
- [§1, §5.3] Minor prose: “plunger gates predominantly control… and barrier gates modulate” (§1) is clear, but “open’ regime” / ‘isolated’ regime (§5.3) should use consistent quotation style; check arXiv line breaks that split math (e.g., v^t_i).
- [§3.2, §5.1] State the CSD resolution and whether f_virt is frozen or fine-tuned when transferring across array sizes; this affects reproducibility of the zero-shot protocol.
Circularity Check
No significant circularity: empirical MARL evaluation with standard preconditioning theory; claims are not forced by definition or self-citation.
full rationale
QADAPT’s load-bearing claims (zero-shot size transfer, near-constant steps, outperformance under a fixed CSD budget) are empirical results from training shared PPO policies on a 4-dot QArray environment and evaluating against external baselines (DreamerV3, Bayesian optimization, L-BFGS, Nelder–Mead, random search, IPPO, MAPPO, MADDPG, FACMAC) and ablations (Table 1, Figs. 2–3). Appendix A derives that if Φ_t approximates the local response Jacobian so C_t Φ_t^{-1} ≈ I, the one-step quadratic objective becomes approximately separable (Thm. A.4, Prop. A.10); that is a standard linearization/preconditioning argument, not a tautology that forces the reported convergence rates. Rewards are shaped from known simulated targets—normal for controlled RL benchmarks—and do not make the scaling claim true by construction. Self-citations (e.g. QArray) supply the simulator, not a uniqueness theorem that forbids alternatives. No fitted parameter is renamed a prediction; no ansatz is smuggled in as external fact. Limitations §5.3 openly condition success on stable online Φ under real non-idealities. Score 0.
Axiom & Free-Parameter Ledger
free parameters (5)
- PPO clip ε, entropy coeff, LR, batch/minibatch sizes
- Kalman process/measurement noise (implicit)
- Reward shaping distances (zero/half reward radii, barrier reward radius)
- Convergence radii 2%/5%/10% of tuning range
- f_virt CNN architecture and CSD resolution/scan size
axioms (6)
- domain assumption Local device response is well-approximated by a Lipschitz Jacobian so one-step quadratic virtualization analysis applies (Assumption A.1, Lemma A.2).
- domain assumption Cross-talk is capturable by a (time-varying) capacitive network estimated from local CSDs up to second-nearest neighbors.
- ad hoc to paper Immediate dense rewards make γ=0 (contextual bandit) sufficient; long-horizon credit assignment is unnecessary.
- domain assumption Devices operate in an open regime with reservoir charge exchange so steady-state CSD features remain informative.
- domain assumption Plunger vs barrier role symmetry justifies two shared policy parameter sets independent of array size.
- standard math Standard Dec-POMDP / CTDE multi-agent RL formalism and PPO clipped surrogate.
invented entities (2)
-
QADAPT framework (action-factored role-shared MARL for QD tuning)
no independent evidence
-
Online virtual action basis Φ_t from CNN+Kalman capacitance estimates
no independent evidence
read the original abstract
Cooperative multi-agent reinforcement learning is well suited to problems with large parameter spaces and exploitable local structure, such as the tuning of electrostatically-defined quantum-dot arrays. However, if parameter cross-talk is strong, a non-stationary environment from the perspective of any individual agent can destabilize learning - the same effect that plagues manual tuning of such systems. We propose using a factored representation of the action space, learned online, to decouple agents and minimize their interference. Our framework, QADAPT, uses this factorization to efficiently learn shared policies based on local measurements and rewards. With this modular strategy, we achieve zero-shot generalization to unseen quantum device sizes and maintain an approximately constant number of convergence steps to reach target regimes. This work provides a scalable route toward the rapid calibration of large-scale quantum processors.
Figures
Reference graph
Works this paper leans on
-
[1]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto.Reinforcement Learning: An Introduction. MIT Press, Cambridge, MA, 2nd edition, 2018
2018
-
[2]
Human-level control through deep reinforcement learning
V olodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. Human-level control through deep reinforcement learning. nature, 518(7540):529–533, 2015
2015
-
[3]
Continuous control with deep reinforcement learning, September 15 2020
Timothy Paul Lillicrap, Jonathan James Hunt, Alexander Pritzel, Nicolas Manfred Otto Heess, Tom Erez, Yuval Tassa, David Silver, and Daniel Pieter Wierstra. Continuous control with deep reinforcement learning, September 15 2020. US Patent 10,776,692
2020
-
[4]
Experimental deep reinforcement learning for error-robust gate-set design on a superconducting quantum computer
Yuval Baum, Mirko Amico, Sean Howell, Michael Hush, Maggie Liuzzi, Pranav Mundada, Thomas Merkh, Andre RR Carvalho, and Michael J Biercuk. Experimental deep reinforcement learning for error-robust gate-set design on a superconducting quantum computer. PRX quantum, 2(4):040324, 2021
2021
-
[5]
Deep reinforcement learning quantum control on ibmq platforms and qiskit pulse
Rudy Semola, Lorenzo Moro, Davide Bacciu, and Enrico Prati. Deep reinforcement learning quantum control on ibmq platforms and qiskit pulse. In 2022 ieee international conference on quantum computing and engineering (qce), pages 759–762. IEEE, 2022
2022
-
[6]
Realizing a deep reinforcement learning agent for real-time quantum feedback
Kevin Reuer, Jonas Landgraf, Thomas Fösel, James O’Sullivan, Liberto Beltrán, Abdulkadir Akin, Gra- ham J Norris, Ants Remm, Michael Kerschbaum, Jean-Claude Besse, et al. Realizing a deep reinforcement learning agent for real-time quantum feedback. Nature Communications, 14(1):7138, 2023. 10
2023
-
[7]
Quantum feedback control with a transformer neural network architecture
Pranav Vaidhyanathan, Florian Marquardt, Mark T Mitchison, and Natalia Ares. Quantum feedback control with a transformer neural network architecture. Physical Review Research, 8(1):L012043, 2026
2026
-
[8]
Deep reinforcement learning for efficient measurement of quantum devices
V Nguyen, SB Orbell, Dominic T Lennon, Hyungil Moon, Florian Vigneau, Leon C Camenzind, Liuqi Yu, Dominik M Zumbühl, G Andrew D Briggs, Michael A Osborne, et al. Deep reinforcement learning for efficient measurement of quantum devices. npj Quantum Information, 7(1):100, 2021
2021
-
[9]
Artificial intelligence for quantum computing
Yuri Alexeev, Marwa H Farag, Taylor L Patti, Mark E Wolf, Natalia Ares, Alán Aspuru-Guzik, Simon C Benjamin, Zhenyu Cai, Shuxiang Cao, Christopher Chamberland, et al. Artificial intelligence for quantum computing. Nature Communications, 16(1):10829, 2025
2025
-
[10]
Qcaleval: Benchmarking vision- language models for quantum calibration plot understanding
Shuxiang Cao, Zijian Zhang, Abhishek Agarwal, Grace Bratrud, Niyaz R Beysengulov, Daniel C Cole, Alejandro Gómez Frieiro, Elena O Glen, Hao Hsu, Gang Huang, et al. Qcaleval: Benchmarking vision- language models for quantum calibration plot understanding. arXiv preprint arXiv:2604.25884, 2026
Pith/arXiv arXiv 2026
-
[11]
Machine learning as an enabler of qubit scalability
Natalia Ares. Machine learning as an enabler of qubit scalability. Nature Reviews Materials, 6(10):870– 871, 2021
2021
-
[12]
Data needs and challenges for quantum dot devices automation
Justyna P Zwolak, Jacob M Taylor, Reed W Andrews, Jared Benson, Garnett W Bryant, Donovan Buterakos, Anasua Chatterjee, Sankar Das Sarma, Mark A Eriksson, Eliška Greplová, et al. Data needs and challenges for quantum dot devices automation. npj Quantum Information, 10(1):105, 2024
2024
-
[13]
Environment model construction toward auto-tuning of quantum dot devices based on model-based reinforcement learning
Chihiro Kondo, Raisei Mizokuchi, Jun Yoneda, and Tetsuo Kodera. Environment model construction toward auto-tuning of quantum dot devices based on model-based reinforcement learning. APL Machine Learning, 3(1), 2025
2025
-
[14]
Metasym: A symplectic meta-learning framework for physical intelligence
Pranav Vaidhyanathan, Aristotelis Papatheodorou, Mark T Mitchison, Natalia Ares, and Ioannis Havoutis. Metasym: A symplectic meta-learning framework for physical intelligence. arXiv preprint arXiv:2502.16667, 2025
arXiv 2025
-
[15]
Meta-learning characteristics and dynamics of quantum systems
Lucas Schorling, Pranav Vaidhyanathan, Jonas Schuff, Miguel J Carballido, Dominik Zumbühl, Gerard Milburn, Florian Marquardt, Jakob Foerster, Michael A Osborne, and Natalia Ares. Meta-learning characteristics and dynamics of quantum systems. arXiv preprint arXiv:2503.10492, 2025
Pith/arXiv arXiv 2025
-
[16]
Reinforcement learning for quantum technology
Marin Bukov and Florian Marquardt. Reinforcement learning for quantum technology. arXiv preprint arXiv:2601.18953, 2026
arXiv 2026
-
[17]
Semiconductor qubits in practice
Anasua Chatterjee, Paul Stevenson, Silvano De Franceschi, Andrea Morello, Nathalie P de Leon, and Ferdinand Kuemmeth. Semiconductor qubits in practice. Nature Reviews Physics, 3(3):157–177, 2021
2021
-
[18]
Spin-qubit control with a milli-kelvin cmos chip
Samuel K Bartee, Will Gilbert, Kun Zuo, Kushal Das, Tuomo Tanttu, Chih Hwan Yang, Nard Du- moulin Stuyck, Sebastian J Pauka, Rocky Y Su, Wee Han Lim, et al. Spin-qubit control with a milli-kelvin cmos chip. Nature, pages 1–6, 2025
2025
-
[19]
Industry-compatible silicon spin-qubit unit cells exceeding 99% fidelity
Paul Steinacker, Nard Dumoulin Stuyck, Wee Han Lim, Tuomo Tanttu, MengKe Feng, Santiago Serrano, Andreas Nickl, Marco Candido, Jesus D Cifuentes, Ensar Vahapoglu, et al. Industry-compatible silicon spin-qubit unit cells exceeding 99% fidelity. Nature, pages 1–7, 2025
2025
-
[20]
Simultaneous operation of an 18-qubit modular array in germanium
Jurgen J Dijkema, Xin Zhang, Achilleas Bardakas, Daniel Bouman, Alice Cuzzocrea, David van Driel, Davide Girardi, Lucas EA Stehouwer, Giordano Scappucci, Anne-Marije J Zwerver, et al. Simultaneous operation of an 18-qubit modular array in germanium. arXiv preprint arXiv:2604.01063, 2026
arXiv 2026
-
[21]
Shared control of a 16 semiconductor quantum dot crossbar array
Francesco Borsoi, Nico W Hendrickx, Valentin John, Marcel Meyer, Sayr Motz, Floor Van Riggelen, Amir Sammak, Sander L De Snoo, Giordano Scappucci, and Menno Veldhorst. Shared control of a 16 semiconductor quantum dot crossbar array. Nature Nanotechnology, 19(1):21–27, 2024
2024
-
[22]
Fully autonomous tuning of a spin qubit
Jonas Schuff, Miguel J Carballido, Madeleine Kotzagiannidis, Juan Carlos Calvo, Marco Caselli, Jacob Rawling, David L Craig, Barnaby van Straaten, Brandon Severin, Federico Fedele, et al. Fully autonomous tuning of a spin qubit. Nature Electronics, pages 1–10, 2026
2026
-
[23]
Machine learning enables completely automatic tuning of a quantum device faster than human experts
Hyungil Moon, Dominic T Lennon, James Kirkpatrick, Nina M van Esbroeck, Leon C Camenzind, Liuqi Yu, Florian Vigneau, Dominik M Zumbühl, G Andrew D Briggs, Michael A Osborne, et al. Machine learning enables completely automatic tuning of a quantum device faster than human experts. Nature communications, 11(1):4161, 2020
2020
-
[24]
Autotuning of double-dot devices in situ with machine learning
Justyna P Zwolak, Thomas McJunkin, Sandesh S Kalantre, JP Dodson, ER MacQuarrie, DE Savage, MG Lagally, SN Coppersmith, Mark A Eriksson, and Jacob M Taylor. Autotuning of double-dot devices in situ with machine learning. Physical review applied, 13(3):034075, 2020. 11
2020
-
[25]
Marc-Antoine Roux, Joffrey Rivard, Victor Yon, Alexis Morel, Dominic Leclerc, Claude Rohrbacher, El Bachir Ndiaye, Felice Francesco Tafuri, Brendan Bono, Stefan Kubicek, et al. Rapid autotuning of a sige quantum dot into the single-electron regime with machine learning and rf-reflectometry fpga-based measurements. arXiv preprint arXiv:2509.19537, 2025
Pith/arXiv arXiv 2025
-
[26]
Automated all-rf tuning for spin qubit readout and control
Cornelius Carlsson, Jaime Saez-Mollejo, Federico Fedele, Stefano Calcaterra, Daniel Chrastina, Giovanni Isella, Georgios Katsaros, and Natalia Ares. Automated all-rf tuning for spin qubit readout and control. arXiv preprint arXiv:2506.10834, 2025
Pith/arXiv arXiv 2025
-
[27]
Cooperative multi-agent control using deep reinforcement learning
Jayesh K Gupta, Maxim Egorov, and Mykel Kochenderfer. Cooperative multi-agent control using deep reinforcement learning. In International conference on autonomous agents and multiagent systems, pages 66–83. Springer, 2017
2017
-
[28]
Computer-automated tuning of semiconductor double quantum dots into the single-electron regime
Timothy A Baart, Pieter T Eendebak, Christian Reichl, Werner Wegscheider, and Lieven MK Vandersypen. Computer-automated tuning of semiconductor double quantum dots into the single-electron regime. Applied Physics Letters, 108(21), 2016
2016
-
[29]
Experimental online quantum dots charge autotuning using neural networks
Victor Yon, Bastien Galaup, Claude Rohrbacher, Joffrey Rivard, Alexis Morel, Dominic Leclerc, Clement Godfrin, Ruoyu Li, Stefan Kubicek, Kristiaan De Greve, et al. Experimental online quantum dots charge autotuning using neural networks. Nano Letters, 25(10):3717–3725, 2025
2025
-
[30]
Machine learning techniques for state recognition and auto-tuning in quantum dots
Sandesh S Kalantre, Justyna P Zwolak, Stephen Ragole, Xingyao Wu, Neil M Zimmerman, MD Stewart Jr, and Jacob M Taylor. Machine learning techniques for state recognition and auto-tuning in quantum dots. npj Quantum Information, 5(1):6, 2019
2019
-
[31]
Tuning arrays with rays: Physics-informed tuning of quantum dot charge states
Joshua Ziegler, Florian Luthi, Mick Ramsey, Felix Borjans, Guoji Zheng, and Justyna P Zwolak. Tuning arrays with rays: Physics-informed tuning of quantum dot charge states. Physical Review Applied, 20(3):034067, 2023
2023
-
[32]
All-rf-based coarse-tuning algorithm for quantum devices using machine learning
Barnaby Van Straaten, Federico Fedele, Florian Vigneau, Joseph Hickie, Daniel Jirovec, Andrea Ballabio, Daniel Chrastina, Giovanni Isella, Georgios Katsaros, and Natalia Ares. All-rf-based coarse-tuning algorithm for quantum devices using machine learning. Physical Review Applied, 24(5):054030, 2025
2025
-
[33]
Automated tuning of double quantum dots into specific charge states using neural networks
Renato Durrer, Benedikt Kratochwil, Jonne V Koski, Andreas J Landig, Christian Reichl, Werner Wegschei- der, Thomas Ihn, and Eliska Greplova. Automated tuning of double quantum dots into specific charge states using neural networks. Physical Review Applied, 13(5):054019, 2020
2020
-
[34]
Minia- turizing neural networks for charge state autotuning in quantum dots
Stefanie Czischek, Victor Yon, Marc-Antoine Genest, Marc-Antoine Roux, Sophie Rochette, Julien Cami- rand Lemyre, Mathieu Moras, Michel Pioro-Ladrière, Dominique Drouin, Yann Beilliard, et al. Minia- turizing neural networks for charge state autotuning in quantum dots. Machine Learning: Science and Technology, 3(1):015001, 2021
2021
-
[35]
Quantum device fine-tuning using unsupervised embedding learning
Nina M van Esbroeck, Dominic T Lennon, Hyungil Moon, Vu Nguyen, Florian Vigneau, Leon C Ca- menzind, Liuqi Yu, Dominik M Zumbühl, G Andrew D Briggs, Dino Sejdinovic, et al. Quantum device fine-tuning using unsupervised embedding learning. New Journal of Physics, 22(9):095003, 2020
2020
-
[36]
Cross-architecture tuning of silicon and sige-based quantum devices using machine learning
Brandon Severin, Dominic T Lennon, Leon C Camenzind, Florian Vigneau, Federico Fedele, Daniel Jirovec, Andrea Ballabio, Daniel Chrastina, Giovanni Isella, Mathieu de Kruijf, et al. Cross-architecture tuning of silicon and sige-based quantum devices using machine learning. Scientific Reports, 14(1):17281, 2024
2024
-
[37]
Loading a quantum-dot based “qubyte” register
Christian V olk, Anne-Marije J Zwerver, Uditendu Mukhopadhyay, Pieter T Eendebak, Cornelis Jacobus van Diepen, Juan Pablo Dehollain, Toivo Hensgens, Takafumi Fujita, Christian Reichl, Werner Wegscheider, et al. Loading a quantum-dot based “qubyte” register. npj Quantum Information, 5(1):29, 2019
2019
-
[38]
Cnn-assisted automatic cross-capacitance matrix update for virtual-gate control of quantum dot arrays
Tatsuo Tsuzuki, Hideaki Yuta, Yui Muto, Arne Ludwig, Andreas Dirk Wieck, Akira Oiwa, and Takafumi Fujita. Cnn-assisted automatic cross-capacitance matrix update for virtual-gate control of quantum dot arrays. Japanese Journal of Applied Physics, 65(1):01SP03, 2026
2026
-
[39]
Modular autonomous virtualization system for two-dimensional semiconductor quantum dot arrays
Anantha S Rao, Donovan Buterakos, Barnaby van Straaten, Valentin John, Cécile X Yu, Stefan D Oost- erhout, Lucas Stehouwer, Giordano Scappucci, Menno Veldhorst, Francesco Borsoi, et al. Modular autonomous virtualization system for two-dimensional semiconductor quantum dot arrays. Physical Review X, 15(2):021034, 2025
2025
-
[40]
Yui Muto, Michael R Zielewski, Motoya Shinozaki, Kosuke Noro, and Tomohiro Otsuka. Automatic detection of single-electron regime of quantum dots and definition of virtual gates using u-net and clustering. arXiv preprint arXiv:2501.05878, 2025. 12
arXiv 2025
-
[41]
Automated charge transition detection in quantum dot charge stability diagrams
Fabian Hader, Fabian Fuchs, Sarah Fleitmann, Karin Havemann, Benedikt Scherer, Jan V ogelbruch, Lotte Geck, and Stefan Van Waasen. Automated charge transition detection in quantum dot charge stability diagrams. IEEE Transactions on Quantum Engineering, 2025
2025
-
[42]
End-to-end analysis of charge stability diagrams with transformers
Rahul Marchand, Lucas Schorling, Cornelius Carlsson, Jonas Schuff, Barnaby van Straaten, Taylor L Patti, Federico Fedele, Joshua Ziegler, Parth Girdhar, Pranav Vaidhyanathan, et al. End-to-end analysis of charge stability diagrams with transformers. arXiv preprint arXiv:2508.15710, 2025
Pith/arXiv arXiv 2025
-
[43]
Auto- matic virtual voltage extraction of a 2x2 array of quantum dots with machine learning
Giovanni A Oakes, Jingyu Duan, John JL Morton, Alpha Lee, Charles G Smith, and M Zalba. Auto- matic virtual voltage extraction of a 2x2 array of quantum dots with machine learning. arXiv preprint arXiv:2012.03685, 2020
Pith/arXiv arXiv 2012
-
[44]
A review of cooperative multi-agent deep reinforcement learning
Afshin Oroojlooy and Davood Hajinezhad. A review of cooperative multi-agent deep reinforcement learning. Applied Intelligence, 53(11):13677–13722, 2023
2023
-
[45]
Multi-agent reinforcement learning: Independent vs
Ming Tan et al. Multi-agent reinforcement learning: Independent vs. cooperative agents. In Proceedings of the tenth international conference on machine learning, pages 330–337, 1993
1993
-
[46]
Value-decomposition networks for cooperative multi-agent learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al. Value-decomposition networks for cooperative multi-agent learning. arXiv preprint arXiv:1706.05296, 2017
Pith/arXiv arXiv 2017
-
[47]
Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder De Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. Monotonic value function factorisation for deep multi-agent reinforcement learning. Journal of Machine Learning Research, 21(178):1–51, 2020
2020
-
[48]
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi. Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning. In International conference on machine learning, pages 5887–5896. PMLR, 2019
2019
-
[49]
Deep multi-agent reinforcement learning for decentralized continuous cooperative control
Christian Schroeder de Witt, Bei Peng, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson. Deep multi-agent reinforcement learning for decentralized continuous cooperative control. arXiv preprint arXiv:2003.06709, 19, 2020
Pith/arXiv arXiv 2003
-
[50]
Coun- terfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. Coun- terfactual multi-agent policy gradients. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018
2018
-
[51]
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. Multi-agent actor-critic for mixed cooperative-competitive environments. Advances in neural information processing systems, 30, 2017
2017
-
[52]
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian Schroeder de Witt, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson. Facmac: Factored multi-agent centralised policy gradients. Advances in neural information processing systems, 34:12208–12221, 2021
2021
-
[53]
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar, Rob Fergus, et al. Learning multiagent communication with backpropagation. Advances in neural information processing systems, 29, 2016
2016
-
[54]
Graph policy gradients for large scale robot control
Arbaaz Khan, Ekaterina Tolstaya, Alejandro Ribeiro, and Vijay Kumar. Graph policy gradients for large scale robot control. In Conference on robot learning, pages 823–834. PMLR, 2020
2020
-
[55]
Scalable multi-agent reinforcement learning through intelligent information aggregation
Siddharth Nayak, Kenneth Choi, Wenqi Ding, Sydney Dolan, Karthik Gopalakrishnan, and Hamsa Balakrishnan. Scalable multi-agent reinforcement learning through intelligent information aggregation. In International conference on machine learning, pages 25817–25833. PMLR, 2023
2023
-
[56]
Multi-agent deep reinforcement learning with extremely noisy observations
Ozsel Kilinc and Giovanni Montana. Multi-agent deep reinforcement learning with extremely noisy observations. arXiv preprint arXiv:1812.00922, 2018
Pith/arXiv arXiv 2018
-
[57]
Bridging marl to sarl: An order-independent multi-agent transformer via latent consensus
Zijian Zhao, Jing Gao, and Sen Li. Bridging marl to sarl: An order-independent multi-agent transformer via latent consensus. arXiv preprint arXiv:2604.13472, 2026
Pith/arXiv arXiv 2026
-
[58]
Multi- agent reinforcement learning is a sequence modeling problem.Advances in Neural Information Processing Systems, 35:16509–16521, 2022
Muning Wen, Jakub Kuba, Runji Lin, Weinan Zhang, Ying Wen, Jun Wang, and Yaodong Yang. Multi- agent reinforcement learning is a sequence modeling problem.Advances in Neural Information Processing Systems, 35:16509–16521, 2022
2022
-
[59]
Hierarchical consensus-based multi-agent reinforcement learning for multi-robot cooperation tasks
Pu Feng, Junkang Liang, Size Wang, Xin Yu, Xin Ji, Yiting Chen, Kui Zhang, Rongye Shi, and Wenjun Wu. Hierarchical consensus-based multi-agent reinforcement learning for multi-robot cooperation tasks. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 642–649. IEEE, 2024. 13
2024
-
[60]
Scaling multi-agent reinforcement learning with selective parameter sharing
Filippos Christianos, Georgios Papoudakis, Muhammad A Rahman, and Stefano V Albrecht. Scaling multi-agent reinforcement learning with selective parameter sharing. In International Conference on Machine Learning, pages 1989–1998. PMLR, 2021
1989
-
[61]
Christian Schroeder De Witt, Tarun Gupta, Denys Makoviichuk, Viktor Makoviychuk, Philip HS Torr, Mingfei Sun, and Shimon Whiteson. Is independent learning all you need in the starcraft multi-agent challenge? arXiv preprint arXiv:2011.09533, 2020
Pith/arXiv arXiv 2011
-
[62]
The sur- prising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. The sur- prising effectiveness of ppo in cooperative multi-agent games. Advances in neural information processing systems, 35:24611–24624, 2022
2022
-
[63]
Trust region policy optimisation in multi-agent reinforcement learning
Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang. Trust region policy optimisation in multi-agent reinforcement learning. arXiv preprint arXiv:2109.11251, 2021
Pith/arXiv arXiv 2021
-
[64]
Heterogeneous- agent reinforcement learning
Yifan Zhong, Jakub Grudzien Kuba, Xidong Feng, Siyi Hu, Jiaming Ji, and Yaodong Yang. Heterogeneous- agent reinforcement learning. Journal of Machine Learning Research, 25(32):1–67, 2024
2024
-
[65]
Hypermarl: Adaptive hypernetworks for multi-agent rl
Kale-ab Abebe Tessera, Arrasy Rahman, Amos Storkey, and Stefano V Albrecht. Hypermarl: Adaptive hypernetworks for multi-agent rl. arXiv preprint arXiv:2412.04233, 2024
arXiv 2024
-
[66]
Qarray: A gpu-accelerated constant capacitance model simulator for large quantum dot arrays
Barnaby van Straaten, Joseph Hickie, Lucas Schorling, Jonas Schuff, Federico Fedele, and Natalia Ares. Qarray: A gpu-accelerated constant capacitance model simulator for large quantum dot arrays. SciPost Physics Codebases, page 035, 2024
2024
-
[67]
Codebase release 1.3 for qarray
Barnaby van Straaten, Joseph Hickie, Lucas Schorling, Jonas Schuff, Federico Fedele, and Natalia Ares. Codebase release 1.3 for qarray. SciPost Physics Codebases, page 035, 2024
2024
-
[68]
The complexity of de- centralized control of markov decision processes
Daniel S Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein. The complexity of de- centralized control of markov decision processes. Mathematics of operations research, 27(4):819–840, 2002
2002
-
[69]
Learning to utilize shaping rewards: A new approach of reward shaping
Yujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang, Yingfeng Chen, Jianye Hao, Feng Wu, and Changjie Fan. Learning to utilize shaping rewards: A new approach of reward shaping. Advances in Neural Information Processing Systems, 33:15931–15941, 2020
2020
-
[70]
Bayesian filtering: From kalman filters to particle filters, and beyond
Zhe Chen et al. Bayesian filtering: From kalman filters to particle filters, and beyond. Statistics, 182(1):1– 69, 2003
2003
-
[71]
Nonlinear bayesian estimation: From kalman filtering to a broader horizon.IEEE/CAA Journal of Automatica Sinica, 5(2):401– 417, 2018
Huazhen Fang, Ning Tian, Yebin Wang, MengChu Zhou, and Mulugeta A Haile. Nonlinear bayesian estimation: From kalman filtering to a broader horizon.IEEE/CAA Journal of Automatica Sinica, 5(2):401– 417, 2018
2018
-
[72]
Actor-critic algorithms
Vijay Konda and John Tsitsiklis. Actor-critic algorithms. Advances in neural information processing systems, 12, 1999
1999
-
[73]
Scaling laws for a multi-agent reinforcement learning model
Oren Neumann and Claudius Gros. Scaling laws for a multi-agent reinforcement learning model. arXiv preprint arXiv:2210.00849, 2022
Pith/arXiv arXiv 2022
-
[74]
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017
Pith/arXiv arXiv 2017
-
[75]
Survey on applications of multi-armed and contextual bandits
Djallel Bouneffouf, Irina Rish, and Charu Aggarwal. Survey on applications of multi-armed and contextual bandits. In 2020 IEEE congress on evolutionary computation (CEC), pages 1–8. IEEE, 2020
2020
-
[76]
Experience-driven networking: A deep reinforcement learning based approach
Zhiyuan Xu, Jian Tang, Jingsong Meng, Weiyi Zhang, Yanzhi Wang, Chi Harold Liu, and Dejun Yang. Experience-driven networking: A deep reinforcement learning based approach. In IEEE INFOCOM 2018-IEEE conference on computer communications, pages 1871–1879. IEEE, 2018
2018
-
[77]
A simplex method for function minimization
John A Nelder and Roger Mead. A simplex method for function minimization. The Computer Journal, 7(4):308–313, 1965
1965
-
[78]
Gaussian Processes for Machine Learning
Carl Edward Rasmussen and Christopher K I Williams. Gaussian Processes for Machine Learning. MIT Press, 2006
2006
-
[79]
Taking the human out of the loop: A review of bayesian optimization
Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P Adams, and Nando de Freitas. Taking the human out of the loop: A review of bayesian optimization. Proceedings of the IEEE, 104(1):148–175, 2016. 14
2016
-
[80]
Updating quasi-newton matrices with limited storage
Jorge Nocedal. Updating quasi-newton matrices with limited storage. Mathematics of Computation, 35(151):773–782, 1980
1980
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.