{"id":"c2170e2a-a45d-41b1-b02c-9d9c03a8c022","arxiv_id":"2607.03105","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A dual-axis agent–framework benchmark finds TensorCircuit-NG and Codex with GPT-5.5 strongest on research quantum workflows, yet agent artifacts remain slower and less complete than expert TC code.","lead":"ORBIT-Q is a 12-task benchmark that tests AI coding agents on research-grade quantum software workflows, scoring both whether the code is scientifically valid and how fast it runs. It shows TensorCircuit-NG and Codex/GPT-5.5 lead among tested stacks, but agents still lag expert human code.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Framework ranking rests on agent-mediated scores for a TC-authored suite without expert baselines for competing frameworks.","rationale":"The paper’s dual-axis design, multi-tier verifier, open task suite, and detailed supplementary tables are real contributions; agent-axis rankings on fixed TC and the remaining expert-TC gap are comparatively well supported. The load-bearing soft spot is exactly the reader’s weakest assumption: treating agent-mediated pass rates and TC-relative runtimes as evidence that TC has the highest framework capability/efficiency. Task composition, sole expert baseline on TC, disclosed TC authorship, and GPT-5.5’s dual solver/auditor role jointly weaken any stronger reading of the framework axis. The manuscript already hedges in Discussion, so the right stance remains CONDITIONAL (agent–framework co-performance on this suite, not a definitive expert-optimized framework championship)—no verdict shift. A concrete expert-PL (etc.) reimplementation would settle whether the concern lands or is overstated.","tokens_in":24294,"tokens_out":640,"duration_ms":30293,"concrete_test":"Commission independent expert implementations of all 12 tasks in PennyLane (and, if feasible, TQ/MQ) under the same containerized verifier, framework-native policy, and timed run_solution protocol; recompute pass counts and geometric-mean runtime ratios against the existing expert TC references. If expert-PL reaches ≥10/12 valid and geometric-mean runtime within ~2× of expert TC on shared passed tasks, the agent-mediated TC lead cannot be read as framework capability or end-to-end performance superiority.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparative claim—that TC has the highest capability and artifact efficiency among TC/PennyLane/TorchQuantum/MindQuantum under agent-driven programming (Fig. 2a: 10/12 vs 8/12/4/12/4/12; smallest geometric-mean slowdown vs expert TC)—requires that fixed-harness agent success and runtime be a fair proxy for framework quality. That proxy is least secure where: (1) Competing Interests disclose the authors created TC; (2) expert references and the runtime denominator exist only for TC (Methods; Fig. 2 grey diamond), so non-TC efficiency is never compared to expert-optimized PL/TQ/MQ; (3) several tasks in Table I (MPS input/refinement, 2D sampling, 512-qubit local observables, MPS-target overlap) align with TC-native tensor-network paths, and MQ Task 01 fails with an explicit missing external-MPS obstruction (Table S10); (4) GPT-5.5 both generates the leading Codex solutions and performs the semantic audit that invalidates many non-TC submissions as framework bypasses. Discussion correctly scopes the comparison as agent-mediated usability, but abstract/results still present TC as exhibiting the highest capability and performance efficiency among the evaluated frameworks—stronger than the dual-axis protocol alone can establish.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"ORBIT-Q introduces a dual-axis benchmark for autonomous coding agents on research-grade quantum workflows. The suite comprises 12 framework-agnostic tasks (variational algorithms, MPS workflows, noise calibration, large-scale sampling/optimization) evaluated under a multi-tier verifier: deterministic functional checks, GPT-5.5 source-level semantic audit against framework bypass and problem mismatch, and human recheck. Along one axis the framework is fixed and agent/model configurations vary; along the other the agent is fixed and TensorCircuit-NG (TC), PennyLane, TorchQuantum, and MindQuantum are compared. Reported results: with Codex+GPT-5.5, TC passes 10/12, PennyLane 8/12, TorchQuantum and MindQuantum 4/12 each, with TC also showing the lowest geometric-mean artifact slowdown relative to expert TC references; on TC, Codex+GPT-5.5 is strongest among tested agents (10/12), with a remaining gap to expert TC code (~2.2× geometric mean on passed tasks). The paper further reports agent-side wall time, tokens, and cost, and finds that a TC performance-checklist prompt reduces agent cost without improving artifact efficiency.","tokens_in":24649,"tokens_out":1695,"duration_ms":30866,"significance":"If the protocol and measurements hold, this is a timely and useful contribution: standard coding benchmarks are inadequate for scientific quantum software, where physical fidelity, differentiability, and framework-native semantics matter. The dual-axis design, three-tier validity pipeline, and explicit separation of agent-side cost from artifact-side runtime are genuine methodological advances and are supported by detailed task-level logs (Supplementary Tables S2–S10), failure-mode attribution (Table S1), and an open repository. The work is also valuable as a reusable task suite for framework-performance leaderboards independent of agents. These strengths make the paper of clear interest to quantum software and AI-for-science communities, provided comparative claims about frameworks are scoped to what the agent-mediated design actually measures.","major_comments":[{"comment":"Abstract and Results (Fig. 2a; “TC exhibits the highest capability and performance efficiency…”) present a framework ranking that the Discussion correctly scopes as agent-mediated usability, not best expert implementations. Expert TC references alone set the runtime denominator for all frameworks (Methods; grey diamond in Fig. 2), so non-TC slowdowns are never compared to expert-optimized PL/TQ/MQ code. Competing Interests disclose that the authors created TC. Please either (i) add expert baselines for at least one competing framework on a subset of tasks, or (ii) systematically rephrase Abstract/Results/Conclusions so every ranking claim is explicitly “agent-discoverable capability under the fixed harness,” and report absolute artifact runtimes alongside TC-relative ratios so the denominator choice is transparent.","section":"Abstract; Results (Fig. 2a); Discussion; Methods"},{"comment":"Table I tasks (MPS input/refinement, 2D sampling, 512-qubit local observables, MPS-target overlap, Haldane/qutrit) align strongly with tensor-network and native-operator paths that TC emphasizes. Supplementary Table S10 Task 01 fails on MindQuantum with an explicit external-MPS obstruction. Without a documented task-selection protocol or a sensitivity analysis (e.g., which tasks drive the 10/12 vs 4/12 gap), the framework-axis ranking risks being partly a match to TC’s design center rather than a general capability measure. Please document how tasks were chosen, which framework primitives each task is intended to stress, and discuss selection bias as a limitation with quantitative impact (pass rates with/without TN-heavy tasks).","section":"Table I; Supplementary Note 1; Results (Fig. 3); Supplementary Table S10"},{"comment":"The semantic audit that invalidates many non-TC submissions as framework bypasses is performed by GPT-5.5 (Supplementary Note 6), the same model family that produces the leading Codex solutions. This couples the strongest agent configuration to the validity gate. Please report inter-auditor agreement (second model or fully human double-review on all borderline cases), the fraction of functional-pass / semantic-fail decisions by framework, and whether re-auditing with an independent model changes pass counts in Fig. 2a. Without this, the multi-tier pipeline’s independence is not established for the central comparative claim.","section":"Results (three-stage verification); Supplementary Note 6; Fig. 1b"},{"comment":"Agent-axis comparisons mix harnesses: GPT-5.5 uses Codex; Opus-4.8, GLM-5.2, and Sonnet-4.6 use Claude Code (Fig. 2b caption; Supplementary Fig. S1). Safety-refusal failures for Opus (Table S1: two cyber-safeguard refusals) are product-level and harness-dependent. The claim that “Codex with GPT-5.5 is the strongest tested agent configuration on TC” is therefore confounded by harness × model. Either evaluate at least one model under both harnesses, or restate the agent-axis result as configuration (harness+model) ranking and avoid model-only language in Abstract/Results.","section":"Fig. 2b; Supplementary Table S1; Abstract"}],"minor_comments":[{"comment":"Fig. 2 and Fig. 3 use geometric-mean runtime ratios but do not state in the main text how many timed passed tasks enter each mean, or how timeouts/missing timings are handled. Add n and a short Methods sentence.","section":"Fig. 2; Methods (Artifact runtime…)"},{"comment":"Line-count (~200 non-empty non-comment lines) and 300 s solution runtime caps (Supplementary Listing 1) are free parameters that can favor concise framework-native APIs. Mention sensitivity or justify these thresholds in Methods.","section":"Supplementary Listing 1; Methods"},{"comment":"Qiskit/Cirq exclusion is reasonable under the autodiff-native policy (Supplementary Note 2), but a one-sentence main-text note would help readers who expect those ecosystems in a quantum-software benchmark.","section":"Methods (Framework inclusion…)"},{"comment":"Fig. 1c matrix shows incomplete cells (only TC fully populated across agents). Clarify in the caption that off-diagonal agent×framework cells were not run, to avoid implying a full factorial design.","section":"Fig. 1c"},{"comment":"Cost accounting excludes verifier-side audit tokens (Methods). State this also in the Fig. 4 caption so economic comparisons are not misread as full end-to-end cost.","section":"Fig. 4; Methods"},{"comment":"Typos/spacing: “quantumbenchmarksfocus”, “artifact-levelefficiencyeval-”, “agentevaluationruntime” and similar join errors appear in the Introduction and Methods; a full proofread pass is needed.","section":"Introduction; Methods"}],"recommendation":"major_revision","confidential_remarks":"The Competing Interests disclosure is appropriate, but the combination of TC authorship, TC-only expert baselines, and TC-favoring task emphasis is the main editorial risk: reviewers in the quantum-software community may read the framework ranking as self-serving even if the dual-axis protocol is carefully designed. Requiring either external expert baselines or consistently agent-usability-scoped language throughout Abstract/Results would substantially reduce that risk. Scope fit for a serious quant-ph / scientific-computing venue is good if claims are tightened; the benchmark infrastructure itself is the durable contribution."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a usable methods paper for people who care about coding agents on real scientific stacks. The new piece is not another unit-test coding suite; it is a 12-task research-grade quantum workflow set plus a dual-axis protocol (fix agent, vary framework; fix framework, vary agent) and a multi-tier validity filter that catches framework bypasses and surrogate physics. They also measure artifact runtime against expert TC references and log agent tokens/cost. That combination is worth having.\n\nWhat they do well is concrete. Pass rates, runtime ratios, failure-mode tags, and cost panels are in the main figures and Supplementary Tables S2–S10. The agent axis on TC is the cleanest cut: Codex/GPT-5.5 at 10/12, Opus 9/12, Sonnet 7/12, GLM 6/12, all short of expert TC on efficiency. The checklist experiment is honest: it cuts agent cost without fixing the two hard fails or the ~2.2× geometric slowdown. Discussion correctly says the framework axis is agent-mediated usability, not best-expert code in every stack. Repo pointers are there.\n\nSoft spots, in proportion. Competing interests: the authors wrote TC and TC leads the framework axis (10/12 vs 8/4/4). Expert baselines and the runtime denominator exist only for TC, so non-TC “slowdown” is not compared to expert PL/TQ/MQ. Several tasks lean on MPS/TN/local-observable paths that play to TC’s strengths; MQ Task 01 even dies on missing external-MPS support. GPT-5.5 is both the strongest solver and the semantic auditor that kills many bypasses. None of that makes the tables fake; it means abstract/results language that TC has the “highest capability and performance efficiency among the evaluated frameworks” should be read as co-performance under this suite, not a definitive framework ranking. Twelve tasks and free parameters (time budgets, line limits, audit thresholds) keep confidence moderate.\n\nWho it is for: quantum software people, agent-eval people, and anyone building scientific coding harnesses. Not a physics result. Math is not the load-bearing part; the data and protocol are. Citations look normal for the niche.\n\nI would send it to peer review. Treat TC’s lead as scoped, push for expert baselines on other frameworks or softer abstract wording, and keep the dual-axis idea. Worth engaging if you work near agents or quantum software; not mandatory for pure algorithm theory.","headline":"Solid dual-axis agent×framework benchmark with real tables and an open suite; TC’s “win” is agent-mediated co-performance under a TC-authored task set, not a clean expert framework championship.","tokens_in":25255,"tokens_out":631,"would_cite":true,"duration_ms":9794,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"ORBIT-Q shows that research-grade quantum coding is an agent–framework co-performance problem: TensorCircuit-NG leads under agents, yet frontier agents still lag expert implementations.","keywords":["ORBIT-Q","autonomous coding agents","quantum software frameworks","differentiable quantum computing","dual-axis benchmarking","artifact runtime","framework-agent synergy","scientific code generation"],"falsifier":"Have independent expert teams implement the same twelve tasks natively in each framework under the same verifier and hardware budget, then check whether the ranking by valid solutions and geometric-mean artifact runtime still places TensorCircuit-NG first and still shows the reported agent-to-expert slowdown gap.","tokens_in":25147,"feed_emoji":"⚛️","tokens_out":984,"duration_ms":16574,"temperature":0.7,"pith_summary":"Scientific code generation cannot be scored like ordinary programming. A script that prints a plausible number can still violate physics, skip differentiability, or bypass the required quantum software stack. ORBIT-Q packages twelve research-level quantum workflows—variational algorithms, noise calibration, large shallow circuits, matrix-product targets, and related pipelines—into a dual-axis benchmark that holds either the agent or the framework fixed. A three-stage verifier (functional tests, language-model semantic audit, human recheck) rejects framework bypasses and surrogate objectives, then records both agent cost and the runtime of the generated artifact against expert TensorCircuit-NG references. Under that protocol, TensorCircuit-NG completes the most tasks with the best relative artifact speed among the tested stacks, Codex with GPT-5.5 is the strongest agent configuration on that stack, and a clear gap to expert code remains on unsolved tasks and slowdowns.","feed_headline":"Top agents still lag experts on quantum coding tasks","feed_subtitle":"ORBIT-Q ranks agent–framework pairs on research workflows, fidelity, and runtime","key_machinery":"ORBIT-Q’s dual-axis agent–framework matrix, backed by a three-tier verifier (deterministic functional check, source-level semantic audit for framework fidelity and problem match, human recheck) plus separate agent-side cost and artifact-side runtime metrics.","core_discovery":"Under a unified multi-tier verifier, dual-axis evaluation shows TensorCircuit-NG has the highest agent-driven completion and artifact efficiency among TensorCircuit-NG, PennyLane, TorchQuantum, and MindQuantum (10/12 versus 8/12, 4/12, and 4/12 with a fixed Codex GPT-5.5 agent), Codex with GPT-5.5 is the strongest tested agent on TensorCircuit-NG (10/12), and a substantial gap remains versus expert TensorCircuit-NG references in unsolved tasks and geometric-mean artifact slowdown (about 2.2× on passed TensorCircuit-NG tasks).","pith_inferences":["Framework designers who optimize only for human experts may still lose on agent-driven research workflows if primitives are hard to discover or poorly differentiated.","Safety-layer false refusals during local scientific exploration can dominate end-to-end agent reliability even when the base model can write the code.","Expanding beyond a compact twelve-task sample will be needed before rankings can be treated as stable across the full space of quantum programming paradigms.","The dual-axis template (agent fixed / stack fixed, multi-tier physical-fidelity verifier) is portable to other differentiable scientific domains beyond quantum software."],"forward_implications":["Quantum software APIs will be judged not only by expert power but by whether agents can discover and compose their performant native paths from docs and local exploration.","Benchmarks for scientific coding must score framework-native fidelity and artifact runtime, not only unit-test pass/fail.","Economic comparisons of coding agents should use cost per valid scientific solution, not nominal price per million tokens.","The same task suite can host a second leaderboard of expert-optimized implementations ranking end-to-end framework performance without the agent layer.","Prompt-side performance checklists can cut agent token use and solve time without closing the expert-level artifact-efficiency gap."],"fun_headline_variants":["Agents still lag experts on ORBIT-Q quantum coding","TensorCircuit-NG leads frameworks under agent coding","Codex GPT-5.5 tops agents on TensorCircuit-NG tasks","ORBIT-Q finds agent-expert gap on quantum workflows","Dual-axis ORBIT-Q: TC best, agents trail expert code"],"cache_read_input_tokens":0,"weakest_assumption_plain":"That how well agents succeed and how fast their code runs under one shared harness is a fair measure of a framework’s real capability, even though expert-optimized baselines exist only for one of the compared stacks.","fun_headline_variants_meta":{"raw":{"variants":["Agents still lag experts on ORBIT-Q quantum coding","TensorCircuit-NG leads frameworks under agent coding","Codex GPT-5.5 tops agents on TensorCircuit-NG tasks","ORBIT-Q finds agent-expert gap on quantum workflows","Dual-axis ORBIT-Q: TC best, agents trail expert code"]},"model":"grok-4.5","effort":"low","cost_usd":0.006638,"raw_usage":{"total_tokens":1727,"prompt_tokens":835,"num_sources_used":0,"completion_tokens":90,"cost_in_usd_ticks":66380000,"prompt_tokens_details":{"text_tokens":835,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":802,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":835,"tokens_out":90,"duration_ms":6661,"temperature":1.0,"reasoning_tokens":802,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T04:51:33.458392+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Have independent expert teams implement the same twelve tasks natively in each framework under the same verifier and hardware budget, then check whether the ranking by valid solutions and geometric-mean artifact runtime still places TensorCircuit-NG first and still shows the reported agent-to-expert slowdown gap.","supporting_citations":[],"review_version":1}