{"id":"5a194736-9514-4ce2-b45b-b4ef423977db","arxiv_id":"2607.17498","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"An LLM-driven workflow proposes and validates quantum control protocols by simulation, beating literature baselines on Rydberg MIS, XXZ chains, and random Ising models, and transferring learned counterdiabatic coefficients to larger systems via a graph neural network.","lead":"An LLM acts as an automated quantum-control co-scientist: it proposes pulse shapes and helper terms, tests them by simulation, and archives winners and failures across three model problems, reporting higher fidelities than the literature baselines it starts from. Read this for a concrete test of whether AI agents can replace weeks of hand-tuning in quantum hardware — and where that claim currently falls short.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central claim of autonomous LLM-driven discovery is unsupported by shipped evidence: candidate families are author-enumerated (SI S6–S9, S13–S15), selection is mechanical best-of-N (SI Note 1C), and no prompts, logs, or code are released. A matched non-LLM baseline would settle causality","rationale":"The reader's weakest assumption is exactly the load-bearing concern. The paper's claimed novelty is that the LLM autonomously proposes structural hypotheses and writes code, but every described candidate family is human-enumerated and the selection rule is mechanical. Without prompts, agent logs, or code, the most parsimonious explanation is a well-structured grid search with best-of-N selection. This is not an accusation of misconduct; the physics is plausible and the SI is unusually detailed. But the causal attribution to the LLM is not supported by the shipped evidence, and the data-availability statement directly contradicts the 'auditable workflow' promise. The secondary issue of under-strength baselines is real but secondary; the primary gap is the missing non-LLM control that would demonstrate the LLM's added value. Given the reproduction instructions and detailed SI, the correct verdict remains CONDITIONAL, conditioned on releasing the repository and running a matched-budget non-LLM baseline. Therefore UNCHANGED is appropriate.","tokens_in":20045,"tokens_out":2762,"duration_ms":26399,"concrete_test":"Release the repository with prompts, agent interaction logs, and code under a public commit hash; then run the same C6 62-slot and XXZ η/α candidate pools with a non-LLM driver such as random search or differential_evolution using the identical evaluation budget. If a non-LLM driver matches the reported fidelities within statistical error, the LLM is not causally necessary for the discovered protocols; if it falls short, the LLM-attribution claim gains support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To make the central claim 'the workflow autonomously discovers ... outperform literature baselines' true, the LLM must be the driver that proposes the structural families and writes validation code. The artifacts actually described undermine this: Methods states the search 'combines heuristics, low-dimensional scans and randomized candidate pools'; SI Note 1C says the selection is 'the candidate with the largest final ground-subspace fidelity'; SI Eqs. S6–S9 and S13–S15 list the knot positions, sampling ranges, and η/α grids — all author-enumerated. No prompts, agent logs, or code are shipped; the data-availability statement only offers 'upon reasonable request' and code 'upon publication,' which conflicts with the 'auditable workflow' promise. If the LLM simply formatted these pre-specified pools and picked the best simulation, the reported fidelities are best-of-N selection, not autonomous discovery. The secondary premise that the reproduced literature baselines (Refs. 9, 32, 33 — two 2026 preprints) are faithful is also unverified, and an under-strength baseline would inflate every improvement. Neither issue makes the physics implausible; it makes the causal attribution unestablished.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces QOC-Workbench, an LLM-driven workflow for quantum optimal control that is claimed to autonomously propose structural hypotheses, write and run simulation code, and accumulate reusable control motifs across Hamiltonian families. Three case studies are presented: (1) Rydberg-atom MIS pulse design, where the workflow reportedly discovers hardware-compliant auxiliary Y-quadrature pulses and piecewise schedules outperforming an ACQC baseline; (2) an XXZ spin-chain benchmark, where target catalysts, retuned commutator-CD corrections, and nonlinear schedule deformations improve on a reproduced OI+AH+CD baseline; and (3) random TFIM instances, where a GNN is trained on small systems and applied zero-shot to larger ones to generate weighted-CD coefficients. The headline claims are that the workflow 'autonomously discovers' superior protocols and that the discovered strategies were not hard-coded or explicitly prompted.","tokens_in":20304,"tokens_out":5281,"duration_ms":50984,"significance":"If the central causal claim were established, the paper would represent a significant advance: an auditable LLM-driven discovery loop that proposes new control functional forms, transfers motifs across physical platforms, and amortizes per-instance optimization into neural generators. The paper has tangible strengths: the reported fidelities come from standard exact ODE simulation; the SI is unusually transparent about candidate pools, per-family slot counts, and full evaluation traces; and the authors explicitly acknowledge limitations (e.g., the ACQC Y-coefficient is only a noninteracting-sector approximation, the XXZ CD term is 'not a complete CD construction', and the GNN's apparent edge over its teacher is explicitly disclaimed in SI Note 3C). These features make the numerical results credible as simulations. However, the manuscript's central attribution of the results to autonomous LLM-driven proposal generation is not currently supported by the shipped evidence; the search as described is a structured, human-enumerated best-of-N parameter scan. Because the paper's novelty and stated contribution depend on that attribution, the current evidence is insufficient for the claims","major_comments":[{"comment":"The central claim that the LLM 'autonomously parses physics literature, proposes structural hypotheses, and writes code' (Abstract) is not supported by the paper's own description of the search. The C6 candidate pool is a fixed 62-slot construction: seven scaled-ACQC slots, seven smooth five-knot Y slots, 45 random piecewise slots, and three targeted piecewise slots, with all knot positions and sampling ranges enumerated in Eqs. (S6)–(S9). Selection is explicitly 'the candidate with the largest final ground-subspace fidelity at that particular total time' (SI Note 1C). The XXZ search likewise uses author-specified grid scans over η and α and a fixed family of schedule deformations (Eqs. S13–S15). This is a structured best-of-N search, not an open-ended proposal loop. No prompts, agent logs, or code are shipped (Data and code availability), so an independent reader cannot distinguish 'LLM","section":"Methods, 'Search and parameter selection'; SI Note 1C; Eqs. (S6)–(S9)"},{"comment":"All headline 'outperform literature baseline' claims are relative to reproduced baselines from Refs. [9], [32], and [33], two of which are 2026 preprints. The manuscript does not provide the baseline parameters, reproduction code, or a comparison between the reproduced values and the values in the original papers. If the reproduced baselines are under-strength, every reported improvement is overestimated. This is a secondary load-bearing premise. The paper should ship the baseline artifacts and, ideally, a table with reproduced vs. published baseline fidelities/energies so that readers can verify the comparison. The current 'upon reasonable request' and 'upon publication' statements do not meet the 'auditable workflow' standard promised in the Introduction.","section":"Data and code availability; baseline reproduction (Refs. 9, 32, 33)"},{"comment":"The Discussion asserts: 'None of the discovered strategies was hard-coded or explicitly prompted.' The only prompt-related artifact in the paper is the generic kick-start template in SI Note 4, which is a user-facing template for future tasks, not the actual prompts used in Cases 1–3. Without the task-specific prompts, the LLM's system-level instructions, and the chronological agent logs, the paper cannot substantiate that the proposals emerged from the Workbench's accumulated experience rather than from the authors' enumeration of candidate families. This is a specific, fixable omission, but it must be addressed: either provide the full artifact log as supplementary material, or soften the 'autonomous' claims to what the current evidence supports.","section":"Discussion; SI Note 4"}],"minor_comments":[{"comment":"The SI honestly states that the GNN's final-fidelity advantage over the teacher (0.680 vs 0.654) 'should not be interpreted as evidence that the GNN has learned a superior control path.' The main text should carry this caveat in the same place where the fidelity comparison is reported: the GNN is 'comparable to' the teacher, not better in a meaningful sense; the transfer claim rests on coefficient R², not on fidelity superiority.","section":"SI Note 3C; main text Case 3"},{"comment":"The abbreviations 'C6' and 'C10' are used in the main text and figure captions but only defined in SI Note 1A. Define them at first use in the main text for readers who do not consult the SI.","section":"Fig. 2 caption; main text"},{"comment":"The energy shift E_K(λ) is said to be 'chosen above the instantaneous spectrum by grid search.' Specify the grid or give the chosen shift values for the reported instances; otherwise the 'weighted-CD teacher' is not fully reproducible.","section":"Eq. (S17)"},{"comment":"The best C10 beta-bump parameters (A=0.15, p=1.6, q=2.4) appear only in the Fig. S5 caption. State them in SI Note 1D near the family definition, since they are the main C10 result.","section":"SI Note 1D / Fig. S5"},{"comment":"TensorCircuit-NG/JAX is identified, but no version numbers or integrator tolerances are reported. For a paper whose results are all numerical, these details matter for reproducibility.","section":"Methods, 'Time-evolution simulation'"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the journal's scope and the numerical simulations appear carefully done. The decisive issue is the mismatch between the 'autonomous LLM-driven discovery' framing and the actual search evidence. I would require, as a condition of acceptance, (1) full release of task-specific prompts and agent logs, or at minimum a detailed chronological record of proposals with code snapshots, and (2) a matched non-LLM baseline (e.g., the same candidate families and best-of-N selection without any LLM) to establish that the LLM adds value beyond automated parameter scanning. If the authors cannot supply those, the manuscript should be reframed as a structured computational search with LLM assistance, which would substantially reduce its novelty claim. The baseline reproduction issue (Refs. 9, 32, 33) should also be addressed with concrete verification tables."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the physics and the supplementary information are genuinely good. The discovered protocols — the early-biased beta-bump Y-pulse, the η=3.0 catalyst with retuned α, the mixed-power schedules — are concrete and new, and the zero-shot GNN transfer on weighted-CD coefficients (N=6→N=8, R²=0.961, fidelity 0.680 vs teacher 0.654) is a real numerical result. The paper is unusually honest about its own approximations: it calls the CD ansatz “not a complete CD construction,” and in SI Note 3C explicitly warns not to interpret the GNN’s edge over its teacher as evidence of a superior control path. That kind of candor is rare and should be credited.\n\nSecond, the central claim — that an LLM autonomously drove these discoveries — is not supported by what is actually shipped. The Methods say the search “combines heuristics, low-dimensional scans and randomized candidate pools.” The SI enumerates every candidate family: knot positions, sampling ranges, η/α grids, all specified by the authors (Eqs. S6–S9, S13–S15). Selection is mechanical best-of-N — “the candidate with the largest final ground-subspace fidelity.” No prompts, agent logs, or code are released; the data-availability statement says “available from corresponding authors upon reasonable request” and the repository comes only “upon publication,” which directly contradicts the “auditable workflow” promise in the abstract. The stress-test note is right: with the evidence in hand, an independent reader cannot distinguish “LLM-driven discovery” from “structured human-designed search with best-of-N selection.” That is a load-bearing flaw for the paper’s narrative, even though the protocols themselves are not in doubt.\n\nThe secondary concern about baseline reproduction is real but milder. The baselines come from two 2026 preprints and one 2026 paper; the authors reproduce them rather than take numbers on faith, but none of us can independently verify those reproductions. An under-strength baseline would inflate all gains. I’d flag it as a question to the authors, not a fatal objection.\n\nProportionately: the paper is solid as a quantum control paper, shaky as an AI-autonomy paper. The honest fix is to deposit the repository with prompts, logs, and code, and to run the same candidate pools with a non-LLM driver (random search or a standard optimizer) so readers can see the delta. Without that, the “automated co-scientist” framing should be softened.\n\nWho it’s for: quantum optimal control researchers, especially anyone working on Rydberg arrays, spin chains, or counterdiabatic driving. They will find the protocol families and the GNN transfer genuinely useful. It deserves a serious referee — send it to review — but with the expectation of major revision on the attribution claims. I’d cite the GNN transfer result and the protocol families, not the LLM framing.","headline":"Solid, honestly reported quantum control results and a real GNN transfer result are wrapped in an over-claimed LLM-autonomy story; the physics deserves review, the attribution does not.","tokens_in":775,"tokens_out":1245,"would_cite":true,"duration_ms":23424,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"QOC-Workbench replaces per-instance expert protocol design with an auditable LLM-driven search loop that proposes structural control changes, validates them by simulation, and accumulates reusable motifs, outperforming literature baselines","keywords":["quantum optimal control","large language models","counterdiabatic driving","Rydberg atom arrays","XXZ spin chains","transverse-field Ising model","graph neural networks","autonomous scientific discovery"],"falsifier":"Replace the LLM proposal engine with a script that randomly samples from the same candidate families (same knot positions, sampling ranges, η/α grids, and schedule deformation formulas) and selects the best candidate by final fidelity; if the reported best fidelities are reproduced, the LLM's structural hypotheses are not the causal source of the gains. Conversely, if the random baseline yields substantially worse results, the LLM's proposals are load-bearing.","tokens_in":19748,"feed_emoji":"⚛️","tokens_out":2312,"duration_ms":24733,"temperature":0.7,"pith_summary":"This paper introduces QOC-Workbench, a workflow in which a large language model acts as an autonomous quantum control co-scientist. Instead of tuning parameters within a fixed control formula, the LLM reads physics literature, proposes structural modifications to the time-dependent Hamiltonian, writes code to test them via exact simulation, and stores successes and failures in a persistent memory. The authors demonstrate the loop across three settings: Rydberg-atom arrays for maximum-independent-set problems, XXZ spin chains with catalyst and schedule co-design, and random transverse-field Ising models where per-instance optimization is replaced by a graph-neural-network generator that transfers to larger systems. In the first two cases the discovered protocols exceed the reproduced literature baselines in final fidelity and energy metrics, and in the third the generator matches the teacher fidelity on unseen larger lattices without retraining. If true, this changes the unit of progress in quantum optimal control from hand-crafted per-instance pulses to accumulating, transferable design knowledge.","feed_headline":"LLM co-scientist discovers quantum control protocols that beat baselines","feed_subtitle":"Three case studies show the loop improving Rydberg, spin-chain, and Ising-model evolution — and transferring learned paths to larger systems","key_machinery":"The load-bearing mechanism is the auditable generate–simulate–record loop: a knowledge base of literature priors, an LLM proposal engine that formulates physically motivated structural hypotheses, an inner solver or neural generator that turns hypotheses into explicit time-dependent Hamiltonians, exact simulation for evaluation, and timestamped artifact memory that archives both successes and failures. Cross-paradigm transfer is carried by reusable control motifs — e.g., the analytic counterdiabatic waveform from single-atom physics is relaxed into multi-knot and beta-bump envelopes for interacting Rydberg arrays, and a target catalyst with commutator-CD is combined with schedule deformation","core_discovery":"The central claim is that an LLM-driven, artifact-accumulating design loop can autonomously discover quantum control protocols that outperform existing literature baselines across different physical paradigms. The authors show that the workflow, starting from reproduced baselines, proposes and validates structural innovations: empirical relaxations of analytic counterdiabatic waveforms, endpoint-vanishing target catalysts, nonlinear schedule deformations, and amortized neural generators. In Case 1, the loop improves average Rydberg fidelity from 0.882 to 0.931 and finds an early-biased beta-bump pulse that beats the ACQC baseline on a ten-site instance. In Case 2, combining a target catalyst","pith_inferences":["The causal attribution of the reported gains to the LLM's autonomous proposals is not fully established by the manuscript: the method description shows that every candidate family (knot positions, sampling ranges, η/α grids) is enumerated by the authors, and the winner is selected mechanically as the best final fidelity. An ablation that replaces the LLM proposal engine with a random or template-b","The GNN's zero-shot fidelity slightly exceeding the exact weighted-CD teacher on the N=8 test set hints that the teacher's variational objective (weighted action with K=3) and the actual final-fidelity objective are not identical; a smoother graph-conditioned approximation can therefore accidentally improve finite-time evolution. This suggests a testable extension: train the generator directly on ","The workflow's auditability promise depends on releasing the actual agent logs, prompts, and code snapshots; the current availability statement offers data and parameter files only upon request. If the full artifact trail is made public, independent readers could verify that the reported candidate evaluations were produced by LLM proposals rather than by the human-enumerated grids.","The cross-paradigm claim would be strengthened by a quantitative transfer metric: e.g., measuring how often a motif discovered in Case 1 is directly reused (with minimal adaptation) in Case 2 or a new platform, rather than only showing that similar motifs appear in both cases."],"forward_implications":["If the workflow is correct, new quantum control protocols for a given platform could be produced in hours rather than weeks of expert trial-and-error, because the LLM loop automates the structural search.","Control motifs learned in one Hamiltonian family (e.g., endpoint-vanishing catalysts, schedule deformations) could be retrieved and re-expressed in another platform's native control language, enabling genuine cross-paradigm reuse.","Per-instance optimization, the bottleneck of variational counterdiabatic driving, could be replaced by amortized graph-conditioned generators that generalize to larger systems without retraining.","Protocol design becomes an auditable, reproducing enterprise where each reported result is linked to executable code, parameter files, and archived failures, allowing the community to accumulate design knowledge instead of isolated pulses."],"fun_headline_variants":["LLM co-scientist beats quantum control baselines","AI designs quantum control protocols that outperform baselines","LLM writes and validates code for better quantum control","Cross-paradigm LLM loop improves Rydberg, spin-chain, and Ising control"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim that the LLM autonomously discovered the winning protocols rests on the assumption that the performance gains are caused by the LLM's proposals rather than by the human-enumerated candidate families and the mechanical best-of-N selection among them.","fun_headline_variants_meta":{"raw":{"variants":["LLM co-scientist beats quantum control baselines","AI designs quantum control protocols that outperform baselines","LLM writes and validates code for better quantum control","Cross-paradigm LLM loop improves Rydberg, spin-chain, and Ising control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000335,"raw_usage":{"total_tokens":1726,"prompt_tokens":807,"completion_tokens":919,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":847}},"tokens_in":551,"tokens_out":919,"duration_ms":8818,"temperature":1.0,"reasoning_tokens":847,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T17:49:03.944212+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace the LLM proposal engine with a script that randomly samples from the same candidate families (same knot positions, sampling ranges, η/α grids, and schedule deformation formulas) and selects the best candidate by final fidelity; if the reported best fidelities are reproduced, the LLM's structural hypotheses are not the causal source of the gains. Conversely, if the random baseline yields substantially worse results, the LLM's proposals are load-bearing.","supporting_citations":[],"review_version":1}