Pith. sign in

REVIEW 4 major objections 6 minor 47 references

ORBIT-Q shows that research-grade quantum coding is an agent–framework co-performance problem: TensorCircuit-NG leads under agents, yet frontier agents still lag expert implementations.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 04:51 UTC pith:LXUUTEVB

load-bearing objection Solid dual-axis agent×framework benchmark with real tables and an open suite; TC’s “win” is agent-mediated co-performance under a TC-authored task set, not a clean expert framework championship. the 4 major comments →

arxiv 2607.03105 v1 pith:LXUUTEVB submitted 2026-07-03 quant-ph

ORBIT-Q: Dual-axis benchmarking of autonomous agents in scientific quantum programming

classification quant-ph
keywords ORBIT-Qautonomous coding agentsquantum software frameworksdifferentiable quantum computingdual-axis benchmarkingartifact runtimeframework-agent synergyscientific code generation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Scientific code generation cannot be scored like ordinary programming. A script that prints a plausible number can still violate physics, skip differentiability, or bypass the required quantum software stack. ORBIT-Q packages twelve research-level quantum workflows—variational algorithms, noise calibration, large shallow circuits, matrix-product targets, and related pipelines—into a dual-axis benchmark that holds either the agent or the framework fixed. A three-stage verifier (functional tests, language-model semantic audit, human recheck) rejects framework bypasses and surrogate objectives, then records both agent cost and the runtime of the generated artifact against expert TensorCircuit-NG references. Under that protocol, TensorCircuit-NG completes the most tasks with the best relative artifact speed among the tested stacks, Codex with GPT-5.5 is the strongest agent configuration on that stack, and a clear gap to expert code remains on unsolved tasks and slowdowns.

Core claim

Under a unified multi-tier verifier, dual-axis evaluation shows TensorCircuit-NG has the highest agent-driven completion and artifact efficiency among TensorCircuit-NG, PennyLane, TorchQuantum, and MindQuantum (10/12 versus 8/12, 4/12, and 4/12 with a fixed Codex GPT-5.5 agent), Codex with GPT-5.5 is the strongest tested agent on TensorCircuit-NG (10/12), and a substantial gap remains versus expert TensorCircuit-NG references in unsolved tasks and geometric-mean artifact slowdown (about 2.2× on passed TensorCircuit-NG tasks).

What carries the argument

ORBIT-Q’s dual-axis agent–framework matrix, backed by a three-tier verifier (deterministic functional check, source-level semantic audit for framework fidelity and problem match, human recheck) plus separate agent-side cost and artifact-side runtime metrics.

Load-bearing premise

That how well agents succeed and how fast their code runs under one shared harness is a fair measure of a framework’s real capability, even though expert-optimized baselines exist only for one of the compared stacks.

What would settle it

Have independent expert teams implement the same twelve tasks natively in each framework under the same verifier and hardware budget, then check whether the ranking by valid solutions and geometric-mean artifact runtime still places TensorCircuit-NG first and still shows the reported agent-to-expert slowdown gap.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Quantum software APIs will be judged not only by expert power but by whether agents can discover and compose their performant native paths from docs and local exploration.
  • Benchmarks for scientific coding must score framework-native fidelity and artifact runtime, not only unit-test pass/fail.
  • Economic comparisons of coding agents should use cost per valid scientific solution, not nominal price per million tokens.
  • The same task suite can host a second leaderboard of expert-optimized implementations ranking end-to-end framework performance without the agent layer.
  • Prompt-side performance checklists can cut agent token use and solve time without closing the expert-level artifact-efficiency gap.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Framework designers who optimize only for human experts may still lose on agent-driven research workflows if primitives are hard to discover or poorly differentiated.
  • Safety-layer false refusals during local scientific exploration can dominate end-to-end agent reliability even when the base model can write the code.
  • Expanding beyond a compact twelve-task sample will be needed before rankings can be treated as stable across the full space of quantum programming paradigms.
  • The dual-axis template (agent fixed / stack fixed, multi-tier physical-fidelity verifier) is portable to other differentiable scientific domains beyond quantum software.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. ORBIT-Q introduces a dual-axis benchmark for autonomous coding agents on research-grade quantum workflows. The suite comprises 12 framework-agnostic tasks (variational algorithms, MPS workflows, noise calibration, large-scale sampling/optimization) evaluated under a multi-tier verifier: deterministic functional checks, GPT-5.5 source-level semantic audit against framework bypass and problem mismatch, and human recheck. Along one axis the framework is fixed and agent/model configurations vary; along the other the agent is fixed and TensorCircuit-NG (TC), PennyLane, TorchQuantum, and MindQuantum are compared. Reported results: with Codex+GPT-5.5, TC passes 10/12, PennyLane 8/12, TorchQuantum and MindQuantum 4/12 each, with TC also showing the lowest geometric-mean artifact slowdown relative to expert TC references; on TC, Codex+GPT-5.5 is strongest among tested agents (10/12), with a remaining gap to expert TC code (~2.2× geometric mean on passed tasks). The paper further reports agent-side wall time, tokens, and cost, and finds that a TC performance-checklist prompt reduces agent cost without improving artifact efficiency.

Significance. If the protocol and measurements hold, this is a timely and useful contribution: standard coding benchmarks are inadequate for scientific quantum software, where physical fidelity, differentiability, and framework-native semantics matter. The dual-axis design, three-tier validity pipeline, and explicit separation of agent-side cost from artifact-side runtime are genuine methodological advances and are supported by detailed task-level logs (Supplementary Tables S2–S10), failure-mode attribution (Table S1), and an open repository. The work is also valuable as a reusable task suite for framework-performance leaderboards independent of agents. These strengths make the paper of clear interest to quantum software and AI-for-science communities, provided comparative claims about frameworks are scoped to what the agent-mediated design actually measures.

major comments (4)
  1. [Abstract; Results (Fig. 2a); Discussion; Methods] Abstract and Results (Fig. 2a; “TC exhibits the highest capability and performance efficiency…”) present a framework ranking that the Discussion correctly scopes as agent-mediated usability, not best expert implementations. Expert TC references alone set the runtime denominator for all frameworks (Methods; grey diamond in Fig. 2), so non-TC slowdowns are never compared to expert-optimized PL/TQ/MQ code. Competing Interests disclose that the authors created TC. Please either (i) add expert baselines for at least one competing framework on a subset of tasks, or (ii) systematically rephrase Abstract/Results/Conclusions so every ranking claim is explicitly “agent-discoverable capability under the fixed harness,” and report absolute artifact runtimes alongside TC-relative ratios so the denominator choice is transparent.
  2. [Table I; Supplementary Note 1; Results (Fig. 3); Supplementary Table S10] Table I tasks (MPS input/refinement, 2D sampling, 512-qubit local observables, MPS-target overlap, Haldane/qutrit) align strongly with tensor-network and native-operator paths that TC emphasizes. Supplementary Table S10 Task 01 fails on MindQuantum with an explicit external-MPS obstruction. Without a documented task-selection protocol or a sensitivity analysis (e.g., which tasks drive the 10/12 vs 4/12 gap), the framework-axis ranking risks being partly a match to TC’s design center rather than a general capability measure. Please document how tasks were chosen, which framework primitives each task is intended to stress, and discuss selection bias as a limitation with quantitative impact (pass rates with/without TN-heavy tasks).
  3. [Results (three-stage verification); Supplementary Note 6; Fig. 1b] The semantic audit that invalidates many non-TC submissions as framework bypasses is performed by GPT-5.5 (Supplementary Note 6), the same model family that produces the leading Codex solutions. This couples the strongest agent configuration to the validity gate. Please report inter-auditor agreement (second model or fully human double-review on all borderline cases), the fraction of functional-pass / semantic-fail decisions by framework, and whether re-auditing with an independent model changes pass counts in Fig. 2a. Without this, the multi-tier pipeline’s independence is not established for the central comparative claim.
  4. [Fig. 2b; Supplementary Table S1; Abstract] Agent-axis comparisons mix harnesses: GPT-5.5 uses Codex; Opus-4.8, GLM-5.2, and Sonnet-4.6 use Claude Code (Fig. 2b caption; Supplementary Fig. S1). Safety-refusal failures for Opus (Table S1: two cyber-safeguard refusals) are product-level and harness-dependent. The claim that “Codex with GPT-5.5 is the strongest tested agent configuration on TC” is therefore confounded by harness × model. Either evaluate at least one model under both harnesses, or restate the agent-axis result as configuration (harness+model) ranking and avoid model-only language in Abstract/Results.
minor comments (6)
  1. [Fig. 2; Methods (Artifact runtime…)] Fig. 2 and Fig. 3 use geometric-mean runtime ratios but do not state in the main text how many timed passed tasks enter each mean, or how timeouts/missing timings are handled. Add n and a short Methods sentence.
  2. [Supplementary Listing 1; Methods] Line-count (~200 non-empty non-comment lines) and 300 s solution runtime caps (Supplementary Listing 1) are free parameters that can favor concise framework-native APIs. Mention sensitivity or justify these thresholds in Methods.
  3. [Methods (Framework inclusion…)] Qiskit/Cirq exclusion is reasonable under the autodiff-native policy (Supplementary Note 2), but a one-sentence main-text note would help readers who expect those ecosystems in a quantum-software benchmark.
  4. [Fig. 1c] Fig. 1c matrix shows incomplete cells (only TC fully populated across agents). Clarify in the caption that off-diagonal agent×framework cells were not run, to avoid implying a full factorial design.
  5. [Fig. 4; Methods] Cost accounting excludes verifier-side audit tokens (Methods). State this also in the Fig. 4 caption so economic comparisons are not misread as full end-to-end cost.
  6. [Introduction; Methods] Typos/spacing: “quantumbenchmarksfocus”, “artifact-levelefficiencyeval-”, “agentevaluationruntime” and similar join errors appear in the Introduction and Methods; a full proofread pass is needed.

Circularity Check

0 steps flagged

Empirical dual-axis benchmark; no derivation reduces to its inputs by construction—only mild COI/self-citation asymmetry around TC baselines.

full rationale

ORBIT-Q is an empirical agent–framework benchmark, not a first-principles derivation. Load-bearing claims are measured pass counts (e.g., TC 10/12, PL 8/12, TQ/MQ 4/12 under fixed Codex GPT-5.5; Codex GPT-5.5 10/12 on TC) and evaluator-timed artifact runtimes, including geometric-mean slowdowns versus expert TC references. Those quantities are not fitted parameters renamed as predictions, nor are they defined in terms of the ranking they support: agent TC solutions still show ~2.2× expert slowdown and two failures, so the expert baseline does not force TC to win by construction. Relative runtime uses a common expert-TC denominator for all frameworks, which is equivalent to ranking absolute measured times and does not make non-TC scores tautological. Self-citations to TensorCircuit / TensorCircuit-NG and related author papers document the evaluated stack and prior domain work; they do not supply a uniqueness theorem or ansatz that forces the dual-axis results. Competing Interests and Discussion correctly disclose TC authorship and that the framework axis measures agent-mediated usability rather than expert-optimized implementations for every stack—structural asymmetry and COI, not Eq. X = Eq. Y. Task design favoring TN-native workflows and GPT-5.5 dual role as solver/auditor are fairness/methodology risks outside circularity. Score 1 only for non-load-bearing author–framework self-reference; no circular steps meet the quote-and-reduce bar.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 1 invented entities

The central ranking claims rest on a constructed evaluation protocol rather than free physical constants. Load-bearing choices are the 12-task sample, framework-native validity policy, multi-tier audit rules, expert-TC-only runtime baseline, and the assumption that agent success under documentation/tool use fairly ranks frameworks. No new physical entities are postulated; ORBIT-Q itself is the invented measurement instrument.

free parameters (4)
  • Task suite composition (12 challenges)
    Hand-curated research workflows define difficulty and coverage; different task choices would change pass rates and framework rankings.
  • Agent time budget (1800 s) and solution runtime caps (score decay after 180 s, zero at 300 s)
    Hard cutoffs in Supplementary Listing 1 and Methods directly affect which solutions count as valid/efficient.
  • Line-count limit (~200 non-empty non-comment lines)
    Static policy treats longer solutions as failed strategy, shaping agent behavior independent of physics correctness.
  • Semantic-audit model and human adjudication thresholds
    GPT-5.5 audit plus human recheck decide borderline framework-bypass cases; different auditors could flip some passes (e.g., tasks 08/12).
axioms (4)
  • ad hoc to paper Framework-native automatic differentiation and core quantum APIs are required; custom NumPy/JAX simulators or surrogate objectives are invalid even if numerical checks pass.
    Stated throughout Introduction, Methods, and Supplementary Note 6; defines validity beyond functional tests.
  • domain assumption Agent-mediated completion and artifact runtime under a shared harness are informative proxies for framework expressivity, discoverability, and performance.
    Discussion and Supplementary Note 5 explicitly treat agent results as proxy for framework quality while noting expert leaderboards would isolate pure stack performance.
  • domain assumption Qiskit/Cirq/TensorFlow Quantum are out of scope because they lack comparable native end-to-end autodiff or platform fit for this suite.
    Methods and Supplementary Note 2; shapes the framework ranking by inclusion criteria.
  • standard math Geometric mean of (solution runtime / expert TC runtime) over passed timed tasks is a portable efficiency statistic.
    Methods: Artifact runtime aggregation; standard relative-performance summary, not a physical law.
invented entities (1)
  • ORBIT-Q benchmark (task suite + dual-axis protocol + multi-tier verifier) independent evidence
    purpose: Provide a discriminative testbed for agent–framework co-performance on research quantum workflows.
    The paper’s primary contribution is this constructed instrument; independent evidence is the public repo and reported evaluation logs, not external physics confirmation.

pith-pipeline@v1.1.0-grok45 · 28325 in / 3328 out tokens · 35636 ms · 2026-07-12T04:51:33.458392+00:00 · methodology

0 comments
read the original abstract

Autonomous coding agents perform well on many conventional programming tasks, but scientific computing demands a rigorous validation paradigm that extends beyond simple functional test completion: generated code must preserve physical fidelity, differentiable workflows, framework-native semantics, and scalable representations. We introduce Open Research Benchmark for Integrated Tasks in Quantum Computing (ORBIT-Q) to address this gap. At its core, ORBIT-Q contributes a carefully curated suite of complex, research-level quantum workflows that serves as a challenging testbed for modern scientific programming. ORBIT-Q combines a rigorous multi-tier verification pipeline to support two orthogonal comparisons: different agent harness and model configurations at a fixed quantum software framework, and different quantum software frameworks at a fixed agent. In our systematic evaluations, TensorCircuit-NG (TC) exhibits the highest capability and performance efficiency among the evaluated quantum software frameworks under agent-driven programming, and Codex with GPT-5.5 is the strongest tested agent configuration on TC. However, a significant performance and design gap remains between frontier autonomous agents and human expert reference implementations. We further evaluate two efficiency dimensions: agent-side resource use and artifact-side runtime. Together, these results establish ORBIT-Q as a rigorous benchmark for autonomous scientific programming, framework-agent synergy, and quantum software performance.

Figures

Figures reproduced from arXiv: 2607.03105 by Shi-Xin Zhang, Yu-Qin Chen.

Figure 1
Figure 1. Figure 1: FIG. 1 [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4 [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

47 extracted references · 9 linked inside Pith

  1. [1]

    OpenAI, GPT-4 technical report, arXiv preprint arXiv:2303.08774 (2023)

  2. [2]

    DeepSeek-AI, DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning, Nature645, 633 (2025)

  3. [3]

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brock- man, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Win- ter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herber...

  4. [4]

    Austin, A

    J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton, Program synthesis with large language mod- els, arXiv preprint arXiv:2108.07732 (2021)

  5. [5]

    Hendrycks, S

    D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt, Measuring coding challenge compe- tence with APPS, in Proceedings of the Neural Informa- tion Processing Systems Track on Datasets and Bench- marks, Vol. 1 (2021). 9

  6. [6]

    C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. R. Narasimhan, SWE-bench: Can lan- guage models resolve real-world GitHub issues?, in Inter- national Conference on Learning Representations (2024)

  7. [7]

    J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. R. Narasimhan, and O. Press, SWE-agent: Agent- computer interfaces enable automated software engineer- ing, in Advances in Neural Information Processing Sys- tems 37 (2024) pp. 50528–50652

  8. [8]

    N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica, Live- CodeBench: Holistic and contamination free evaluation of large language models for code, in International Con- ference on Learning Representations (2025)

  9. [9]

    Schuld, V

    M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, and N. Killoran, Evaluating analytic gradients on quantum hardware, Physical Review A99, 032331 (2019)

  10. [10]

    X. Yuan, J. Sun, J. Liu, Q. Zhao, and Y. Zhou, Quantum simulation with hybrid tensor networks, Physical Review Letters127, 040501 (2021)

  11. [11]

    Zhang, Z.-Q

    S.-X. Zhang, Z.-Q. Wan, C.-K. Lee, C.-Y. Hsieh, S. Zhang, and H. Yao, Variational quantum-neural hy- brid eigensolver, Physical Review Letters128, 120502 (2022)

  12. [12]

    Orús, A practical introduction to tensor networks: Matrix product states and projected entangled pair states, Annals of Physics349, 117 (2014)

    R. Orús, A practical introduction to tensor networks: Matrix product states and projected entangled pair states, Annals of Physics349, 117 (2014)

  13. [13]

    U.Schollwöck,Thedensity-matrixrenormalizationgroup in the age of matrix product states, Annals of Physics 326, 96 (2011)

  14. [14]

    S. R. White, Density matrix formulation for quantum renormalization groups, Physical Review Letters69, 2863 (1992)

  15. [15]

    Vidal, Efficient classical simulation of slightly entan- gled quantum computations, Physical Review Letters91, 147902 (2003)

    G. Vidal, Efficient classical simulation of slightly entan- gled quantum computations, Physical Review Letters91, 147902 (2003)

  16. [16]

    I. L. Markov and Y. Shi, Simulating quantum computa- tion by contracting tensor networks, SIAM Journal on Computing38, 963 (2008)

  17. [17]

    D. S. Steiger, T. Häner, and M. Troyer, ProjectQ: An open source software framework for quantum computing, Quantum2, 49 (2018)

  18. [18]

    Javadi-Abhari, M

    A. Javadi-Abhari, M. Treinish, K. Krsulich, C. J. Wood, J. Lishman, J. Gacon, S. Martiel, P. D. Nation, L. S. Bishop, A. W. Cross, B. R. Johnson, and J. M. Gam- betta, Quantum computing with Qiskit, arXiv preprint arXiv:2405.08810 (2024)

  19. [19]

    Broughton, G

    M. Broughton, G. Verdon, T. McCourt, A. J. Martinez, J. H. Yoo, S. V. Isakov, P. Massey, R. Halavati, M. Y. Niu, A. Zlokapa, E. Peters, O. Lockwood, A. Skolik, S. Jerbi, V. Dunjko, M. Leib, M. Streif, D. Von Dollen, H. Chen, S. Cao, R. Wiersema, H.-Y. Huang, J. R. McClean, R. Babbush, S. Boixo, D. Bacon, A. K. Ho, H. Neven, and M. Mohseni, TensorFlow Quan...

  20. [20]

    J. R. Johansson, P. D. Nation, and F. Nori, QuTiP: An open-source Python framework for the dynamics of open quantum systems, Computer Physics Communica- tions183, 1760 (2012)

  21. [21]

    Gray, quimb: A Python package for quantum informa- tionandmany-bodycalculations,JournalofOpenSource Software3, 819 (2018)

    J. Gray, quimb: A Python package for quantum informa- tionandmany-bodycalculations,JournalofOpenSource Software3, 819 (2018)

  22. [22]

    Gray and S

    J. Gray and S. Kourtis, Hyper-optimized tensor network contraction, Quantum5, 410 (2021)

  23. [23]

    A. W. Cross, L. S. Bishop, S. Sheldon, P. D. Nation, and J. M. Gambetta, Validating quantum computers us- ing randomized model circuits, Physical Review A100, 032328 (2019)

  24. [24]

    Tomesh, P

    T. Tomesh, P. Gokhale, V. Omole, G. S. Ravi, K. N. Smith, J. Viszlai, X.-C. Wu, N. Hardavellas, M. R. Martonosi, and F. T. Chong, SupermarQ: A scalable quantum benchmark suite, in 2022 IEEE International Symposium on High-Performance Computer Architec- ture (HPCA) (2022) pp. 587–603

  25. [25]

    A. Li, S. Stein, S. Krishnamoorthy, and J. Ang, QASM- Bench: A low-level quantum benchmark suite for NISQ evaluation and simulation, ACM Transactions on Quan- tum Computing4, 1 (2023)

  26. [26]

    Quetschlich, L

    N. Quetschlich, L. Burgholzer, and R. Wille, MQT Bench: Benchmarking software and design automation tools for quantum computing, Quantum7, 1062 (2023)

  27. [27]

    Biamonte, P

    J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, and S. Lloyd, Quantum machine learning, Nature549, 195 (2017)

  28. [28]

    Song, D.-L

    L.Hu, S.-H.Wu, W.Cai, Y.Ma, X.Mu, Y.Xu, H.Wang, Y. Song, D.-L. Deng, C.-L. Zou, and L. Sun, Quan- tum generative adversarial learning in a superconducting quantum circuit, Science Advances5, eaav2761 (2019)

  29. [29]

    Chen and S.-X

    Y.-Q. Chen and S.-X. Zhang, Superior resilience to poi- soning and amenability to unlearning in quantum ma- chine learning, Nature Communications17, 3716 (2026)

  30. [30]

    Chen and S.-X

    Y.-Q. Chen and S.-X. Zhang, Intrinsic preservation of plasticity in continual quantum learning, arXiv preprint arXiv:2511.17228 (2025)

  31. [31]

    Zhang and Y.-Q

    S.-X. Zhang and Y.-Q. Chen, Quantum subliminal learn- ing, arXiv preprint arXiv:2605.29557 (2026)

  32. [32]

    Preskill, Quantum computing in the NISQ era and beyond, Quantum2, 79 (2018)

    J. Preskill, Quantum computing in the NISQ era and beyond, Quantum2, 79 (2018)

  33. [33]

    Bharti, A

    K. Bharti, A. Cervera-Lierta, T. H. Kyaw, T. Haug, S. Alperin-Lea, A. Anand, M. Degroote, H. Heimonen, J. S. Kottmann, T. Menke, W.-K. Mok, S. Sim, L.- C. Kwek, and A. Aspuru-Guzik, Noisy intermediate- scale quantum algorithms, Reviews of Modern Physics 94, 015004 (2022)

  34. [34]

    Cerezo, A

    M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, and P. J. Coles, Variational quantum algo- rithms, Nature Reviews Physics3, 625 (2021)

  35. [35]

    Peruzzo, J

    A. Peruzzo, J. McClean, P. Shadbolt, M.-H. Yung, X.-Q. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’Brien, A variational eigenvalue solver on a photonic quantum processor, Nature Communications5, 4213 (2014)

  36. [36]

    J. R. McClean, J. Romero, R. Babbush, and A. Aspuru- Guzik, The theory of variational hybrid quantum- classical algorithms, New Journal of Physics18, 023023 (2016)

  37. [37]

    Farhi, J

    E. Farhi, J. Goldstone, and S. Gutmann, A quan- tum approximate optimization algorithm, arXiv preprint arXiv:1411.4028 (2014)

  38. [38]

    Zhou, S.-T

    L. Zhou, S.-T. Wang, S. Choi, H. Pichler, and M. D. Lukin, Quantum approximate optimization algorithm: Performance, mechanism, and implementation on near- term devices, Physical Review X10, 021067 (2020)

  39. [39]

    L.Cheng, Y.-Q.Chen, S.-X.Zhang, andS.Zhang,Quan- tum approximate optimization via learning-based adap- tive optimization, Communications Physics7, 83 (2024). 10

  40. [40]

    Zhang, J

    S.-X. Zhang, J. Allcock, Z.-Q. Wan, S. Liu, J. Sun, H. Yu, X.-H. Yang, J. Qiu, Z. Ye, Y.-Q. Chen, C.-K. Lee, Y.-C. Zheng, S.-K. Jian, H. Yao, C.-Y. Hsieh, and S. Zhang, TensorCircuit: a quantum software framework for the NISQ era, Quantum7, 912 (2023)

  41. [41]

    Zhang, Y.-Q

    S.-X. Zhang, Y.-Q. Chen, W. Li, J. Sun, W.-G. Ma, P.-L. Zheng, Y.-X. Huang, Q.-X. Wang, H. Yu, Z. Li, X. Huang, Z.-L. Li, Z.-Q. Wan, S. Liu, J. Qiu, J. Miao, Z. Song, Y. Yan, K. Tsuoka, P. Zhang, L. Wang, H. Fan, C.-Y. Hsieh, H. Yao, and T. Xiang, TensorCircuit-NG: A universal, composable, and scalable platform for quan- tum computing and quantum simulati...

  42. [42]

    Bergholm, J

    V. Bergholm, J. Izaac, M. Schuld, C. Gogolin, S. Ahmed, V. Ajith, M. S. Alam, G. Alonso-Linaje, B. Akash- Narayanan, A. Asadi, J. M. Arrazola, U. Azad, S. Ban- ning, C. Blank, T. R. Bromley, B. A. Cordier, J. Ceroni, A. Delgado, O. Di Matteo, A. Dusko, T. Garg, D. Guala, A. Hayes, R. Hill, A. Ijaz, T. Isacsson, D. Ittah, S. Ja- hangiri, P. Jain, E. Jiang,...

  43. [43]

    MIT Han Lab, TorchQuantum: A PyTorch-based frame- work for quantum simulation and quantum machine learning (2022), software repository

  44. [44]

    X. Xu, J. Cui, Z. Cui, R. He, Q. Li, X. Li, Y. Lin, J. Liu, W. Liu, J. Lu, M. Luo, C. Lyu, S. Pan, M. Pavel, R. Shu, J. Tang, R. Xu, S. Xu, K. Yang, F. Yu, Q. Zeng, H. Zhao, Q. Zheng, J. Zhou, X. Zhou, Y. Zhu, Z. Zou, A. Bayat, X.Cao, W.Cui, Z.Li, G.Long, Z.Su, X.Wang, Z.Wang, S. Wei, R.-B. Wu, P. Zhang, and M.-H. Yung, Mind- Spore Quantum: A user-friendl...

  45. [45]

    C. R. Harris, K. J. Millman, S. J. van der Walt, R. Gom- mers, P. Virtanen, D. Cournapeau, E. Wieser, J. Tay- lor, S. Berg, N. J. Smith, R. Kern, M. Picus, S. Hoyer, M. H. van Kerkwijk, M. Brett, A. Haldane, J. F. del Río, M. Wiebe, P. Peterson, P. Gérard-Marchant, K. Shep- pard, T. Reddy, W. Weckesser, H. Abbasi, C. Gohlke, and T. E. Oliphant, Array prog...

  46. [46]

    Bradbury, R

    J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. Vander- Plas, S. Wanderman-Milne, and Q. Zhang, JAX: com- posable transformations of Python+NumPy programs (2018), software repository

  47. [47]

    Cirq Developers, Cirq (2023), software repository. Supplementary Information ORBIT-Q: Dual-axis benchmarking of autonomous agents in scientific quantum programming Shi-Xin Zhang ∗ Institute of Physics, Chinese Academy of Sciences, Beijing 100190, China Yu-Qin Chen† Graduate School of China Academy of Engineering Physics, Beijing 100193, China Supplementar...