Pith. sign in

REVIEW 3 major objections 5 minor 34 references

A vision-language agent can close the full transmon calibration loop and specialize to one drifting chip without weight updates, by growing a short natural-language device note.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 04:13 UTC pith:CKCBUUNI

load-bearing objection Careful systems paper: co-designed physics env + VL calibration loop + truth-free device-note specialization; adaptation numbers are thin but the design is honest and referee-worthy. the 3 major comments →

arxiv 2607.03193 v1 pith:CKCBUUNI submitted 2026-07-03 quant-ph cs.AIcs.LG

Self-Specializing Vision-Language Transmon Chip Calibration in a Physics-Grounded Environment

classification quant-ph cs.AIcs.LG
keywords transmon calibrationvision-language agentssuperconducting qubitsgradient-free adaptationdevice notephysics-grounded simulationflux distortiononline specialization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Calibrating a superconducting transmon chip is a sequential decision problem under noise, drift, and a finite measurement budget: choose experiments, read ambiguous plots, judge fit quality, and revise stale beliefs as the device ages. This paper asks whether a vision-language agent can run that loop end to end without ever seeing hidden ground truth, and whether it can specialize itself to one physical device by appending a small, human-readable note rather than by fine-tuning weights. The answer is built from three co-designed pieces: a physics-grounded simulator (circuit-quantized truth, flux-line distortion, wall-time and mid-scan drift, gate leakage), a plot-reading agent that maintains a structured notebook and submits parameters under evidence gates, and a gradient-free reflector that proposes note edits accepted only when they beat the base agent on the same frozen device snapshot. Under budget pressure, six adaptation iterations raised worst-case measured CZ fidelity from 0.678 to 0.787 and cut variance; a single accepted note lifted CZ from 0.678 to 0.913 on its paired snapshot. A planted-fault study shows the note can name a hardware fault truth-free, with its main value as failure-floor insurance rather than defect-specific repair. A sympathetic reader cares because the agent, scoring, and reward are designed to transfer to real hardware via a measurement-backend swap, leaving only the accept gate as a simulation affordance.

Core claim

A tool-using vision-language agent can calibrate a whole transmon chip without hidden truth and, without any weight updates, specialize to one persistent drifting device by growing an interpretable device note from truth-free anomaly signatures. On a hard-tier chip under budget pressure, six online iterations raised worst-case measured CZ fidelity from 0.678 to 0.787 and reduced variance (reproduced at four-qubit scale); a single accepted note raised CZ from 0.678 to 0.913 on the identical frozen snapshot. Planted-fault contrast confirms the note is causal diagnosis and failure-floor insurance, not confabulation or mean-fidelity repair.

What carries the argument

The device note plus paired-snapshot accept gate: a short, human-readable prompt slot grown by a reflector from truth-free fit-quality and fidelity anomalies, admitted only if the challenger note beats the incumbent on the same frozen drifted snapshot using measured gate fidelity, submission validity, and budget—isolating strategy improvement from device drift without touching model weights.

Load-bearing premise

That the gains isolated by freezing and replaying the same drifted device snapshot in simulation will still show up when that gate is replaced on real hardware by a held-out chip slice or by noisy, sample-costly repeat-and-average runs under residual drift.

What would settle it

Deploy the same agent and reflector on a real multi-qubit transmon chip with a held-out-slice or back-to-back repeat-and-average accept gate: if worst-case measured two-qubit fidelity and variance do not improve over the empty-note baseline under a tight call budget, or if the note fails to produce distinctive truth-free diagnoses under planted or known hardware faults, the specialization claim does not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper asks whether a vision-language agent can close the full superconducting-transmon calibration loop without hidden truth and specialize to one persistent drifting chip without weight updates. It co-designs three artifacts: (i) a physics-grounded, ablatable simulator (scqubits circuit truth, LTI flux distortion with FIR predistortion, wall-time and mid-scan drift, CZ leakage, measured gate quality); (ii) a multimodal tool-using agent that reads plots, maintains a notebook, and commits under fit-quality gates; and (iii) gradient-free online specialization via a human-readable device note, admitted by a paired-snapshot accept gate on a truth-free objective (measured fidelities, submission validity, budget). Headline results: under budget pressure, six iterations raise worst-case CZ fidelity from 0.678 to 0.787 (reproduced at four-qubit scale); a single accepted note lifts CZ from 0.678 to 0.913 on a frozen snapshot; a planted-fault study shows truth-free diagnosis of a dead coupler, with principal value framed as failure-floor insurance rather than defect-specific repair. Agent, scoring, and reward are argued to transfer via a measurement-backend swap; only the accept gate is a simulation affordance.

Significance. If the specialization results hold under stronger statistics and a hardware-realistic accept gate, this is a substantial contribution to agentic quantum control: an end-to-end, auditable, weight-free specialization loop scored against measured gate quality rather than hidden parameters. The environment itself is a first-class deliverable—circuit-quantized truth, wall-time drift, mid-scan aging, and residual-RMS predistortion scoring are independently ablatable and motivated by transfer, and the scalar-deprived vision and toy-vs-full environment ablations cleanly isolate subclaims that many agent papers leave untested. The truth-free reward design (with a regression pin against coverage/truth leakage) and the interpretable device-note mutation surface are methodologically careful and deployment-relevant. Even with the present sample-size caveats, the co-designed stack is a useful controlled laboratory for experimental policy under drift and budget.

major comments (3)
  1. The load-bearing specialization claim (worst-case CZ floor 0.678 o0.787 over six iterations; single paired accept +0.235) rests on Table 3 (one hard-tier trajectory) and Fig. 4 (n=6 iterations per budget point, three paired seeds, labeled preliminary). Setup states agent runs use temperature 0.4 and no API seed, so rollouts are non-deterministic even when physics is seeded. With that stochasticity and seed count, the reported floor-lift and variance cut can be driven by a few unlucky base-agent episodes rather than a stable note effect. Because the authors correctly reject mean fidelity as the summary statistic and rest the reliability claim on floor/variance, those quantities need multi-seed error bars (or bootstrap over independent agent seeds) and a pre-registered accept threshold before the headline can be treated as established. The planted-fault arm cleanly supports diagnosis langu
  2. Section 4.4 and the Discussion correctly identify freeze-and-replay as a simulation-only affordance and propose held-out-slice or repeat-and-average as hardware forms. The weakest transfer assumption is that the causal floor-lift measured under identical-snapshot pairing will remain meaningful once residual drift and noise-correlated repeats (or spatial non-transfer of the note) enter. The manuscript does not quantify how much of the paired gain survives under a simulated held-out-slice gate or under controlled inter-repeat drift. A short ablation that re-admits notes under those degraded gates on the same chips would substantially strengthen the deployment claim; without it, the specialization half of the abstract overreaches relative to what is demonstrated.
  3. Section 5.1 reports a four-qubit existence proof (median edge CZ 0.887, note a no-op at comfortable budget) and a brief budget-pressure floor-lift at chip scale (0.700 o0.810). Controlled adaptation, planted-defect, and component ablations are confined to two-qubit edges by design. That isolation is methodologically sound for edge-local effects, but the abstract’s “reproducing at four-qubit scale” phrasing invites reading the adaptation result as chip-scale specialization. Either expand the four-qubit adaptation statistics (seeds, variance, note content) or narrow the abstract/claim language so that chip-scale reproduction is limited to whole-chip calibration capability and a single budget-pressure floor check, not the full online loop.
minor comments (5)
  1. Table 5’s plain no_vision null is well explained by text mirroring, but the main text could flag earlier (Section 3.2) that the default tool responses are information-redundant with the plots, so readers do not misread the later scalar-deprived test as a post-hoc rescue.
  2. Fig. 4 annotations (+0.00, +0.11, …) would be clearer with error bars or min–max ranges across the three paired seeds rather than only worst-case point estimates.
  3. Equation (1) is the standard dispersive ZZ; a brief note that the environment uses full two-transmon diagonalization (already stated in prose) next to the equation would avoid readers treating (1) as the simulator’s truth.
  4. Appendix decision traces and notebook dumps are valuable; cross-reference which hard-tier seed/config they come from so they can be tied to scorecard numbers if code is released.
  5. References [13] and [14] are concurrent agent/VLM calibration works; a one-sentence contrast table (skill library vs. plot-reading policy; static QA vs. closed loop) in the Introduction would help placement.

Circularity Check

0 steps flagged

No significant circularity: specialization claims rest on measured, truth-free device objectives with paired same-snapshot controls, not on fitted inputs renamed as predictions or self-definitional loops.

full rationale

This is an empirical systems paper, not a first-principles derivation. The agent never sees hidden ground truth; the online reward is restricted to quantities measured on the device (gate fidelities, valid submission fraction, budget), and a regression test is cited that poisoning coverage/truth fields leaves the reward bit-identical. The paired-snapshot accept gate freezes the device and compares incumbent vs. challenger notes on the identical snapshot precisely to isolate strategy from drift, rather than defining improvement by construction. The planted-fault contrast is an external causal check (defect-on vs. matched defect-off at the same seed), not a fit-then-predict loop. Environment observables are derived from scqubits circuit parameters and scored by measured RB/Bell/CZ on the truth device; ablations recover simpler behavior bit-for-bit. Citations are to standard circuit-QED and agent literature, not load-bearing self-citations of uniqueness theorems or smuggled ansätze. No step reduces a claimed prediction to its own fitted input or definition. Residual concerns (few seeds, non-deterministic agent rollouts, simulation-only freeze-and-replay) are statistical/deployment risks, not circularity.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 3 invented entities

The central claim rests on a stack of domain models (circuit quantization, LTI flux distortion, multi-channel wall-time drift, leakage) treated as adequate proxies for transfer, plus operational choices (budgets, accept gate, model temperatures) that shape the measured floor-lift. No new physical particle or force is invented; the main invented control objects are the device note and the paired-snapshot accept gate. Free parameters are mostly simulation and loop hyperparameters rather than fits to a claimed universal constant.

free parameters (6)
  • Per-sub-task call budgets (e.g. 16/22, 14/18–36/50)
    Budget pressure is the regime where the note's floor-lift appears; thresholds are chosen experimentally and the effect is inert below a minimum budget.
  • noiseScale and difficulty-tier circuit/noise ranges
    Hard-tier and noiseScale settings define the operating point of all headline fidelity numbers.
  • Agent/reflector temperatures (0.4 / 0.7) and model choice
    Non-deterministic policy behavior and reflector proposals depend on these hand-chosen settings; no API seed is used.
  • Accept-gate threshold and paired-seed count
    Whether a note edit is admitted depends on the truth-free objective threshold and multi-seed pairing design.
  • Flux-distortion kernel magnitudes (τ ~ 10–25 ns, poles, AWG ring)
    Calibrated to a published hardware regime rather than in-house telemetry; they set residual-RMS and gate-fidelity ceilings.
  • Composite scalar 0.5·coverage + 0.5·measured_fidelity (when used)
    Outside the scorecard, this hand-chosen mix is used for triage/accept decisions when a single scalar is needed.
axioms (6)
  • domain assumption Circuit-quantized scqubits parameters yield self-consistent calibration observables (f01, α, χ, g, ζ) adequate for policy transfer studies.
    Section 2.2; replaces independent uniform sampling of observables.
  • domain assumption Wall-time-scaled multi-channel drift, mid-scan drift, LTI flux distortion, and CZ leakage are the dominant failure modes a calibration policy must confront for hardware transfer.
    Section 2 and Table 1; motivates environment fidelity as a co-equal contribution.
  • domain assumption Measured gate fidelities (RB/Bell/CZ/iSWAP) on the truth device with installed parameters are a valid truth-free objective for online adaptation.
    Sections 2.5 and 4.3; basis of the accept gate and hardware-transfer reward claim.
  • ad hoc to paper A small natural-language device note is a sufficient mutation surface for useful per-device specialization without weight updates.
    Section 4.1; structural choice of the adaptation method.
  • ad hoc to paper Paired same-snapshot evaluation isolates note improvement from device drift in simulation, and held-out-slice/repeat-and-average are adequate hardware substitutes.
    Section 4.4; load-bearing for causal attribution and deployment story.
  • standard math Standard open-system simulation tools (QuTiP Lindblad, scqubits diagonalization) and FIR/Tikhonov predistortion are valid backends for the claimed observables.
    Sections 2.1–2.3; conventional numerical methods.
invented entities (3)
  • Device note (isolated additive prompt slot with per-qubit/per-edge strategy) no independent evidence
    purpose: Serve as the sole human-readable mutation surface for gradient-free online specialization without weight updates.
    Introduced in Section 4 as the adaptation object; independent evidence is only the paper's own accept-gate experiments, not an external benchmark.
  • Paired-snapshot accept gate no independent evidence
    purpose: Admit note edits only when they improve a truth-free objective on an identical frozen device state, isolating strategy from drift.
    Section 4.4; simulation-native mechanism later reduced to held-out-slice or repeat-and-average forms.
  • Physics-grounded ablatable transmon calibration environment (as a co-designed research artifact) no independent evidence
    purpose: Provide hardware-transfer-relevant failure modes so learned policies are not trained only against toy noise.
    Contribution 1 / Section 2; not a digital twin of a specific fridge, magnitudes from literature.

pith-pipeline@v1.1.0-grok45 · 26734 in / 4182 out tokens · 33804 ms · 2026-07-12T04:13:32.125712+00:00 · methodology

0 comments
read the original abstract

Calibrating a superconducting transmon chip is a sequential decision problem under noise, drift, and a finite budget: an expert must choose experiments, read ambiguous plots, judge fit quality, and revise stale beliefs as the chip drifts. We study whether a vision-language agent can close this loop and specialize itself to one physical device without weight updates, via three co-designed artifacts. The first is a physics-grounded simulation environment for transmon chips: calibration observables derive from circuit-quantized parameters via scqubits, with realistic flux-line distortion, wall-time-scaled and mid-scan drift, and gate leakage, concerns a toy simulator would omit; each tool call advances a modeled clock so drift accrues by wall time, not call count. The second is a vision-language agent that runs the loop end to end, calling tools, reading plots, maintaining a structured notebook, and submitting parameters without hidden truth, scored against hidden parameters and gate fidelities measured on the device. The third is gradient-free online adaptation: a reflector reads truth-free anomaly signatures from past attempts and grows a small, human-readable device note appended to the prompt, admitted by a paired-snapshot accept gate that isolates strategy improvement from drift. On a hard-tier chip under budget pressure, six iterations raised the worst-case CZ fidelity from 0.678 to 0.787 and cut its variance, reproducing at four-qubit scale; a single accepted note raised CZ fidelity from 0.678 to 0.913 on its paired snapshot. A planted-fault study confirms the note is causal, diagnosing a hardware fault truth-free, its principal value raising the failure floor and cutting variance. The agent, scoring, and reward transfer to real hardware via a measurement-backend swap; only the accept gate is a simulation affordance, reducing to a held-out-slice or repeat-and-average form.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 2 canonical work pages

  1. [1]

    Yu, Jay Gambetta, A

    Jens Koch, Terri M. Yu, Jay Gambetta, A. A. Houck, D. I. Schuster, J. Majer, Alexandre Blais, M. H. Devoret, S. M. Girvin, and R. J. Schoelkopf. Charge-insensitive qubit design derived from the Cooper pair box.Physical Review A, 76(4):042319, 2007. doi: 10.1103/PhysRevA.76.042319

  2. [2]

    Grimsmo, S

    Alexandre Blais, Arne L. Grimsmo, S. M. Girvin, and Andreas Wallraff. Circuit quantum electrodynamics. Reviews of Modern Physics, 93(2):025005, 2021. doi: 10.1103/RevModPhys.93.025005

  3. [3]

    Orlando, Simon Gustavsson, and William D

    Philip Krantz, Morten Kjaergaard, Fei Yan, Terry P. Orlando, Simon Gustavsson, and William D. Oliver. A quantum engineer’s guide to superconducting qubits.Applied Physics Reviews, 6(2):021318, 2019. doi: 10.1063/1.5089550

  4. [4]

    Motzoi, J

    F. Motzoi, J. M. Gambetta, P. Rebentrost, and F. K. Wilhelm. Simple pulses for elimination of leakage in weakly nonlinear qubits.Physical Review Letters, 103(11):110501, 2009. doi: 10.1103/PhysRevLett.103.110501

  5. [5]

    J. M. Gambetta, F. Motzoi, S. T. Merkel, and F. K. Wilhelm. Analytic control methods for high-fidelity unitary operations in a weakly nonlinear oscillator.Physical Review A, 83(1):012308, 2011. doi: 10.1103/PhysRevA.83. 012308

  6. [6]

    Easwar Magesan, J. M. Gambetta, and Joseph Emerson. Scalable and robust randomized benchmarking of quantum processes.Physical Review Letters, 106(18):180504, 2011. doi: 10.1103/PhysRevLett.106.180504

  7. [7]

    Cross, Lev S

    Andrew W. Cross, Lev S. Bishop, Sarah Sheldon, Paul D. Nation, and Jay M. Gambetta. Validating quantum computers using randomized model circuits.Physical Review A, 100(3):032328, 2019. doi: 10.1103/PhysRevA.1 00.032328

  8. [8]

    Isakov, Vadim N

    Sergio Boixo, Sergei V. Isakov, Vadim N. Smelyanskiy, Ryan Babbush, Nan Ding, Zhang Jiang, Michael J. Bremner, John M. Martinis, and Hartmut Neven. Characterizing quantum supremacy in near-term devices. Nature Physics, 14:595–600, 2018. doi: 10.1038/s41567-018-0124-x

  9. [9]

    Bardin, Rami Barends, et al

    Frank Arute, Kunal Arya, Ryan Babbush, Dave Bacon, Joseph C. Bardin, Rami Barends, et al. Quantum supremacy using a programmable superconducting processor.Nature, 574:505–510, 2019. doi: 10.1038/s41586-0 19-1666-5

  10. [10]

    Gambetta, and Blake R

    Andrew Wack, Hanhee Paik, Ali Javadi-Abhari, Petar Jurcevic, Ismael Faro, Jay M. Gambetta, and Blake R. Johnson. Quality, speed, and scale: three key attributes to measure the performance of near-term quantum computers, 2021. URLhttps://arxiv.org/abs/2110.14108. 13

  11. [11]

    Egger, Stefan Filipp, Frank K

    Nicolas Wittler, Federico Roy, Kevin Pack, Max Werninghaus, Anurag Saha Roy, Daniel J. Egger, Stefan Filipp, Frank K. Wilhelm, and Shai Machnes. Integrated tool set for control, calibration, and characterization of quantum devices applied to superconducting qubits.Physical Review Applied, 15(3):034080, 2021. doi: 10.1103/PhysRevApplied.15.034080

  12. [12]

    Swarnadeep Majumder, Leonardo Andreta de Castro, and Kenneth R. Brown. Real-time calibration with spectator qubits.npj Quantum Information, 6:19, 2020. doi: 10.1038/s41534-020-0251-y

  13. [13]

    Vibe calibration: Autonomous bring- up of a 112-qubit superconducting quantum processor by a skill-orchestrating language agent, 2026

    Huikai Xu, Jiaxiu Han, Shigang Ou, Cheng Ye, Zisong Shen, Jing Gao, Yijia Wang, Tianrui Che, Yu Song, Weiyang Liu, Lei Wang, Lin-Feng Zhang, Pan Zhang, and Hai-Feng Yu. Vibe calibration: Autonomous bring- up of a 112-qubit superconducting quantum processor by a skill-orchestrating language agent, 2026. URL https://arxiv.org/abs/2606.22376

  14. [14]

    Beysengulov, Daniel C

    Shuxiang Cao, Zijian Zhang, Abhishek Agarwal, Grace Bratrud, Niyaz R. Beysengulov, Daniel C. Cole, Alejandro Gómez Frieiro, Elena O. Glen, Hao Hsu, Gang Huang, Raymond Jow, Greshma Shaji, Tom Lubowe, Ligeng Zhu, Luis Mantilla Calderón, Nicola Pancotti, Joel Pendleton, Brandon Severin, Charles Etienne Staub, Sara Sussman, Antti Vepsäläinen, Neel Rajeshbhai...

  15. [15]

    ReAct: Synergizing reasoning and acting in language models

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. InInternational Conference on Learning Representations,

  16. [16]

    URLhttps://openreview.net/forum?id=WE_vluYUL-X

  17. [17]

    Toolformer: Language models can teach themselves to use tools

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems, volume 36, 2023. URLhttps://proceedings.neurips.cc /paper_files/paper/2023/hash/d842425e4bf79ba...

  18. [18]

    Patil, Tianjun Zhang, Xin Wang, and Joseph E

    Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. Gorilla: Large language model connected with massive APIs. InAdvances in Neural Information Processing Systems, volume 37, 2024. URLhttps: //proceedings.neurips.cc/paper_files/paper/2024/hash/e4c61f578ff07830f5c37378dd3ecb0d-Abstr act-Conference.html

  19. [19]

    ToolLLM: Facilitating large language models to master 16000+ real-world APIs

    Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. ToolLLM: Facilitating large language models to master 16000+ real-world APIs. InInternational Conference on Learning Representat...

  20. [20]

    Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D

    Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D. White, and Philippe Schwaller. Augmenting large language models with chemistry tools.Nature Machine Intelligence, 6:525–535, 2024. doi: 10.1038/s42256-024-00832-8

  21. [21]

    Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes

    Daniil A. Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. Autonomous chemical research with large language models.Nature, 624(7992):570–578, 2023. doi: 10.1038/s41586-023-06792-0

  22. [22]

    J. R. Johansson, P. D. Nation, and Franco Nori. QuTiP: An open-source Python framework for the dynamics of open quantum systems.Computer Physics Communications, 183(8):1760–1772, 2012. doi: 10.1016/j.cpc.2012.02. 021

  23. [23]

    J. R. Johansson, P. D. Nation, and Franco Nori. QuTiP 2: A Python framework for the dynamics of open quantum systems.Computer Physics Communications, 184(4):1234–1240, 2013. doi: 10.1016/j.cpc.2012.11.019

  24. [24]

    scqubits: a Python package for superconducting qubits.Quantum, 5:583,

    Peter Groszkowski and Jens Koch. scqubits: a Python package for superconducting qubits.Quantum, 5:583,

  25. [25]

    doi: 10.22331/q-2021-11-17-583

  26. [26]

    M. A. Rol, L. Ciorciaro, F. K. Malinowski, B. M. Tarasinski, R. E. Sagastizabal, C. C. Bultink, Y. Salathe, N. Haandbaek, J. Sedivy, and L. DiCarlo. Time-domain characterization and correction of on-chip distortion of control pulses in a quantum processor.Applied Physics Letters, 116(5):054001, 2020. doi: 10.1063/1.5133894. 14

  27. [27]

    Dharun Venkateswaran, Felice Francesco Tafuri, Yuanzheng Paul Tan, Bruno Aznar Martinez, Alisa Danilenko, Likai Yang, Arnaud Carignan-Dugas, Christoph Hufnagel, Rainer Dumke, Philip Krantz, and Eric T. Holland. Digital predistortion for flux control of tunable superconducting qubits, 2026. URLhttps://arxiv.org/abs/ 2604.15895

  28. [28]

    P. V. Klimov, J. Kelly, Z. Chen, M. Neeley, A. Megrant, B. Burkett, R. Barends, K. Arya, B. Chiaro, Y. Chen, et al. Fluctuations of energy-relaxation times in superconducting qubits.Physical Review Letters, 121(9):090502,

  29. [29]

    doi: 10.1103/PhysRevLett.121.090502

  30. [30]

    Ithier, E

    G. Ithier, E. Collin, P. Joyez, P. J. Meeson, D. Vion, D. Esteve, F. Chiarello, A. Shnirman, Y. Makhlin, J. Schriefl, and G. Schön. Decoherence in a superconducting quantum bit circuit.Physical Review B, 72(13):134519, 2005. doi: 10.1103/PhysRevB.72.134519

  31. [31]

    Reflexion: Language agents with verbal reinforcement learning

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. InAdvances in Neural Information Processing Systems, volume 36,

  32. [32]

    URL https://proceedings.neurips.cc/paper_files/paper/2023/hash/1b44b878bb782e6954cd888 628510e90-Abstract-Conference.html

  33. [33]

    Optimizing generative AI by backpropagating language model feedback.Nature, 639:609–616, 2025

    Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Pan Lu, Zhi Huang, Carlos Guestrin, and James Zou. Optimizing generative AI by backpropagating language model feedback.Nature, 639:609–616, 2025. doi: 10.1038/s41586-025-08661-4

  34. [34]

    I know / I need / Next

    Gemini Team Google. Gemini: A family of highly capable multimodal models, 2023. URLhttps://arxiv.org/ abs/2312.11805. A Appendix A.1 Representative Tool-Call Plots Section 3 states that the agent’s primary evidence is a rendered plot, not a scalar summary; this appendix shows what that evidence actually looks like. All panels are drawn from hard-tier two-...