{"id":"ed30742b-8225-4d81-9abf-815c7d1a7ef5","arxiv_id":"2602.03466","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An LLM with memory, score feedback, and restart-from-best finds high-entanglement quantum circuits, reaching Meyer-Wallach 1.0 on 25 qubits within 45 queries.","lead":"This paper tests a memory-augmented test-time optimization loop in which a large language model edits quantum circuits and receives score feedback, aiming to maximize Meyer–Wallach entanglement. The method reached perfect entanglement scores on 25-qubit circuits within 45 evaluations, far beating a random hill-climbing baseline, though with small sample sizes and a weak comparison.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No same-model no-feedback control on 25-qubit task; causal claim for feedback/restart is untested.","rationale":"The missing no-feedback control is the most load-bearing concern because it directly tests the proposed mechanism (score-difference feedback and restart) that the abstract credits for the headline result. Without it, the Q=1.0 runs could be explained by the model's prior knowledge of high-entanglement stabilizer-state families rather than by test-time learning from evaluator feedback. This is exactly the reader's identified weakest assumption, and it remains unresolved. The additional mismatch in baseline initialization further weakens the comparative claim, but the causal attribution is the more fundamental issue. A focused ablation—removing only feedback and restart while keeping the same model, budget, and starting circuits—would settle whether the mechanism works. Since the reader already issued a CONDITIONAL verdict based on this concern, no adjustment is needed.","tokens_in":12068,"tokens_out":7154,"duration_ms":64206,"concrete_test":"Run GPT5.2 on 25-qubit circuits under the exact protocol of Table II (three queries of 15 steps, 45 oracle calls) but remove the score-difference feedback sentence from the prompt and disable restart-from-best, starting from the same initial circuits as rows A–H of Table II. Repeat for at least 5 seeds per row. If any no-feedback run reaches Q>0.9, the causal attribution to feedback/restart is not supported; if all plateau below ~0.7, the claim is strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'feedback and restart mechanisms enable multiple runs to reach Q(ψ)=1.0 within 45 oracle calls' is not supported by a controlled comparison. The paper reports no GPT5.2 no-feedback run on the 25-qubit task under the same 45-evaluation budget. The only no-feedback results are GPT5.1 on 20 qubits (Table I), and the prose statements about plateauing at Q≈0.48 or ~0.7 on 25 qubits do not specify the model, prompt, or budget. Consequently, the Q=1.0 successes in Table II could be due to GPT5.2's stronger prior knowledge of high-entanglement stabilizer/graph-state structures (the authors themselves note the synthesized states are products of Bell pairs and GHZs) rather than to the feedback/restart mechanism. In addition, the random hill-climbing baseline is not initialization-matched: LLM runs in Table II start from Q=0.05–0.47, while the baseline is only run from Q≈0.02 and Q≈0.08; its 'stall below 0.29' may reflect poor starting circuits rather than inferior sample efficiency. Thus the causal and comparative claims both rest on untested counterfactuals.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a memory-augmented test-time optimization framework for LLM-driven quantum circuit synthesis, using the Meyer–Wallach (MW) global entanglement measure as the black-box objective. The method combines episodic memory of high-scoring candidates, score-difference feedback, and restart-from-best sampling. On 20-qubit circuits with GPT5.1 and no feedback, the framework reaches Q=0.99. On 25-qubit circuits with GPT5.2, feedback and restarts are reported to enable multiple runs to reach Q=1.0 within 45 oracle calls, while a budget-matched random hill-climbing baseline is said to stall below Q≈0.29. The paper also analyzes the structure of synthesized states, finding that high-MW solutions are often stabilizer or graph-state-like constructions (products of Bell pairs and GHZ states).","tokens_in":12396,"tokens_out":9336,"duration_ms":85857,"significance":"If the central claims were fully supported, the work would demonstrate that LLMs with simple episodic memory and scalar feedback can efficiently optimize a black-box quantum objective at 25 qubits, substantially larger than typical prior ML-based circuit synthesis. The paper's transparency is a strength: it discloses the full 11-run collection in Table II, computes the MW objective externally (no circularity), and provides prompt templates in the appendix. However, the comparative and causal claims are not yet established because the key ablations and controls are missing.","major_comments":[{"comment":"The claim that 'feedback and restart mechanisms enable multiple runs to reach Q(ψ)=1.0' is not supported by a controlled comparison. No GPT5.2 no-feedback run is reported at 25 qubits under the same 45-evaluation budget. The only no-feedback data are GPT5.1 on 20 qubits (Table I), and the prose plateaus (≈0.48 in §III B; ≈0.7 in the Introduction) are inconsistent and do not specify model, prompt, or budget. Without this control, the Q=1.0 outcomes could be due to GPT5.2's prior knowledge of high-entanglement stabilizer/graph states (the authors themselves note the solutions are products of Bell pairs and GHZs) rather than to the feedback/restart mechanism. Please add the missing ablation or reframe the causal claim.","section":"§III B, Table II; Abstract"},{"comment":"The random hill-climbing baseline is not initialization-matched to the LLM runs. The text says the baseline starts from 'the same random initialization used in Table II' at Q≈0.02, and from Q=0.08, whereas Table II reports initial MW values from 0.05 to 0.47 (rows A and C use pre-sampled circuits). The baseline ceiling is stated as ≈0.16 and ≈0.29 depending on the start, while the abstract uses only ≈0.29. Because the LLM runs begin from substantially better initial circuits, the comparison does not isolate sample efficiency. Run the baseline from each initial circuit in Table II and report the full distribution.","section":"§III B, baseline paragraph; Table II"},{"comment":"The paper claims LLMs can 'sample optimal circuits beyond repeated sampling approaches' but does not implement a repeated-sampling (best-of-n) baseline with the same model. Given that the high-MW solutions are simple stabilizer states (products of Bell pairs and GHZ), GPT5.2 might produce them from memory in a single sample without memory or feedback. A same-model best-of-45 control (or at least a random-search baseline over the same gate set) is needed to attribute the performance to the proposed framework rather than to model priors. Otherwise, the causal claims should be substantially softened.","section":"§III A; §I"}],"minor_comments":[{"comment":"The RY angle set is inconsistent: Section II C states angles {3,7,25}, while the prompt in Appendix V A and the Fig. 1 caption use {3.0,10.0,25.0}. This ambiguity prevents exact reproduction; please reconcile.","section":"§II C vs Appendix V A; Fig. 1"},{"comment":"The paper reports no error bars or confidence intervals. The full 11-run collection is disclosed, which is good, but the success rate is not quantified. Please report e.g. bootstrap CIs for the fraction of runs reaching relevant thresholds (Q≥0.9, Q=1.0).","section":"§III B, Table II"},{"comment":"The abstract's 'multiple runs reach Q(ψ)=1.0' is supported by only 2 of 11 runs in Table II (rows A and C1); rows C2, E, F, H reach 0.91–0.99. Please clarify the threshold and report exact counts.","section":"Abstract vs Table II"},{"comment":"The naive-querying plateau on 25 qubits is reported as '∼0.7' in the Introduction and '∼0.48' in Section III B, with no model/budget specification. Please make these numbers consistent and define what 'naive querying' includes.","section":"Introduction vs §III B"},{"comment":"The sentence 'it is worth highlighting that this metric does not produce a W state' is unclear — Q is a function, not a state generator. Please rephrase (e.g., to describe the value of Q for W states).","section":"§II B"},{"comment":"The manuscript contains several typos and incomplete sentences (e.g., 'This yielded to the theory of naive sampling approaches', missing articles). An editorial pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The central evidence relies on proprietary GPT5.1/5.2 API calls, which limits reproducibility; the authors should be asked to release interaction logs or reproduce the key ablation with an open-weight model. The comparison to a single random hill-climbing baseline is weak; stronger baselines (random search, best-of-n sampling, evolutionary search) would be expected. The benchmark itself is arguably too simple because the objective is saturated by product states of Bell pairs and GHZ, which the model may know from pretraining; the authors should either strengthen the benchmark or temper the causal claims. That said, the paper is honest about limitations and discloses the full run collection, which is to its credit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: the paper shows GPT5.2 with feedback and restart-from-best can find circuits with Meyer–Wallach Q=1.0 on 25 qubits within 45 evaluations, and Table II reports all 11 runs, not cherry-picks. But the claim that feedback/restart is what does the work is not actually tested.\n\nWhat's new: applying a dynamic-cheatsheet-style memory plus RL-style score-difference feedback to quantum circuit synthesis, with a gate set restricted to discrete RY angles. The authors also analyze the final states and show that high-Q solutions are mostly products of Bell pairs and GHZ states—disconnected stabilizer states. That is a useful observation about the MW objective, and it tempers the practical significance. They also test non-Clifford RY angles.\n\nSoft spots, in proportion: the main one is the missing control. There is no GPT5.2 no-feedback run on the 25-qubit task under the same 45-evaluation budget. The prose says naive querying plateaus at ~0.48 or ~0.7 (the text gives both numbers), but that is not tabulated with the same protocol. Without this control, the successes in Table II could come from GPT5.2's prior knowledge of stabilizer/graph-state structure—the solved states look like textbook Bell/GHZ factories—rather than from the feedback loop. The baseline is also not initialization-matched: rows start at Q=0.05–0.47, while the baseline starts at 0.02 and 0.08, so the 0.29 stall may partly be a starting-point effect. Sample sizes are tiny (11 runs, 4 runs), no error bars. Code and exact prompts are not shipped, and the model is proprietary, limiting reproducibility. The authors are honest about limitations and disclose the full run collection, which earns them credit. A minor inconsistency: the abstract implies a clean 45-call budget, and the protocol does match that, but the prose describes \"3 different experiments\" while Table II has 11 rows—likely a wording slip.\n\nVerdict: the empirical phenomenon is plausible but under-supported. The central causal claim needs a same-model no-feedback ablation and an initialization-matched baseline. Still, this is a legitimate demonstration, and the quantum circuit synthesis benchmark with MW could be useful to the community. I would send it to peer review with the expectation that revision must add those controls. For my own work I wouldn't cite the strong claim yet; maybe cite as a data point in a survey.","headline":"Useful demo of LLM test-time optimization for circuit synthesis, but the headline causal claim about feedback/restart lacks a controlled test.","tokens_in":12845,"tokens_out":2634,"would_cite":false,"duration_ms":27740,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68","68Q12"],"pacs":["03.67.-a","03.67.Mn"],"model":"deepseek-v4-flash","headline":"A memory-and-feedback loop turns an LLM into a dependable quantum-circuit optimizer, reaching perfect entanglement on 25-qubit circuits within 45 oracle calls where a random hill-climber stalls below 0.29.","keywords":["quantum circuit synthesis","test-time learning","large language models","Meyer–Wallach entanglement","black-box optimization","feedback prompting","restart-from-best","memory-augmented search"],"falsifier":"Run the same 25-qubit setup with the same budget and the same initial circuits, but remove only the score-difference feedback from the prompt while keeping memory and restarts; tabulate success rates. If Q = 1.0 is reached just as often, the paper's central claim about feedback's role is falsified. Also, if a random hill-climber with restart-from-best over the same 45-evaluation budget reaches Q = 1.0, the LLM's specific contribution would be called into question.","tokens_in":11982,"feed_emoji":"⚛️","tokens_out":4852,"duration_ms":47026,"temperature":0.7,"pith_summary":"The paper claims that large language models can act as black-box optimizers for quantum circuit synthesis if they are embedded in a closed test-time loop: keep the best circuits in an episodic memory, feed back the signed change in the Meyer–Wallach entanglement score, and restart from the best circuit when progress stalls. On 20-qubit circuits a feedback-free version reaches Q = 0.99 in one favorable 35-gate case, but with only about a 10% chance of exceeding 0.8. On 25-qubit circuits, adding score-difference feedback and restart-from-best makes several runs reach Q = 1.0 within 45 oracle calls, while a budget-matched random hill-climbing baseline never exceeds roughly 0.29. A sympathetic reader cares because this suggests expensive scientific design problems can be tackled without training data or fine-tuning, using only inference-time memory and evaluator feedback.","feed_headline":"Feedback-guided LLM reaches perfect 25-qubit entanglement in 45 calls","feed_subtitle":"Memory and score feedback let a language model design circuits that a hill-climbing baseline never approaches.","key_machinery":"The load-bearing mechanism is an episodic memory plus feedback loop. At each iteration the LLM receives the current gate list, the signed change in Meyer–Wallach value (e.g. 'you improved by 0.12'), and—on restart—the best circuit seen so far. The Meyer–Wallach measure, a permutation-invariant number in [0,1] built from single-qubit purities, is the objective being maximized. The loop enforces a fixed gate count drawn from {CNOT, H, RY} with a small fixed set of RY angles, so the LLM can only rearrange connectivity, not inflate the gate count. This turns circuit design into a sequence of creative local edits constrained by a textual reward signal and protected against trivial solutions.","core_discovery":"The central claim is that test-time memory and evaluator feedback, not model scale alone, carry the LLM from random circuits to globally entangled states. On 20 qubits, a feedback-free prompt with episodic memory reaches Q = 0.99 in one 35-gate run, but with a low success rate: roughly 10% of runs exceed Q > 0.8. On 25 qubits, the full recipe—score-difference feedback plus restart-from-best—lets multiple independent runs hit Q = 1.0 within three queries of 15 candidate samples each. The winning circuits are written out explicitly and decompose into tensor products of Bell pairs and GHZ states, i.e. stabilizer (Clifford) states, sometimes with RY rotations. The same oracle budget given to a r","pith_inferences":["The paper never runs a same-model, same-budget control without feedback on the 25-qubit task; the 'naive querying' plateau is described in prose, not tabulated, so the causal share of feedback versus restart versus model scaling is not yet isolated—a direct ablation would settle it.","Because the MW objective saturates on disconnected Bell-pair products, this benchmark may be easier than it looks; objectives that reward graph-state connectivity or genuine multipartite entanglement could show a different LLM-versus-baseline gap.","Restart-from-best is a simple evolutionary strategy; keeping a diverse set of memory traces (more than one prior best) might overcome the 'stuck topology' failures the authors report for rows D2/D3.","A budget-matched deterministic local search using the same gate set and MW evaluator could clarify how much of the success comes from the LLM's prior knowledge of quantum circuits rather than from the feedback loop itself."],"forward_implications":["If the procedure holds up, LLM-based synthesis can produce maximally entangled 25-qubit circuits with far fewer oracle calls than repeated random search.","High-scoring circuits discovered by the method are structurally simple—tensor products of Bell pairs and GHZ states—suggesting the Meyer–Wallach objective is largely saturated by disconnected stabilizer states.","Score-difference feedback and restart-from-best are presented as the levers that break the Q ≈ 0.5–0.7 plateau observed under naive querying; the paper attributes the gain to these mechanisms.","The method requires no fine-tuning and no training data, only a black-box simulator for evaluation, so it transfers in principle to other expensive black-box design objectives."],"fun_headline_variants":["Memory and feedback drive LLM to perfect 25-qubit entanglement in 45 calls","LLM with episodic memory hits Q=1.0 on 25 qubits in 45 calls","Restart-from-best plus feedback beats hill-climbing on 25-qubit design","45 oracle calls turn LLM into perfect 25-qubit circuit designer","Test-time memory and feedback let LLM design perfect 25-qubit states"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's headline contrast assumes the improvement comes from feedback and restart, but it does not report a same-model, same-budget control without feedback on the 25-qubit task; the 'naive querying' plateau is described in prose rather than tabulated, and some successful runs start from pre-sampled circuits, so the causal role of feedback and restart is not fully isolated.","fun_headline_variants_meta":{"raw":{"variants":["Memory and feedback drive LLM to perfect 25-qubit entanglement in 45 calls","LLM with episodic memory hits Q=1.0 on 25 qubits in 45 calls","Restart-from-best plus feedback beats hill-climbing on 25-qubit design","45 oracle calls turn LLM into perfect 25-qubit circuit designer","Test-time memory and feedback let LLM design perfect 25-qubit states"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000839,"raw_usage":{"total_tokens":3491,"prompt_tokens":740,"completion_tokens":2751,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":484,"completion_tokens_details":{"reasoning_tokens":2643}},"tokens_in":484,"tokens_out":2751,"duration_ms":19810,"temperature":1.0,"reasoning_tokens":2643,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T04:58:07.580472+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 25-qubit setup with the same budget and the same initial circuits, but remove only the score-difference feedback from the prompt while keeping memory and restarts; tabulate success rates. If Q = 1.0 is reached just as often, the paper's central claim about feedback's role is falsified. Also, if a random hill-climber with restart-from-best over the same 45-evaluation budget reaches Q = 1.0, the LLM's specific contribution would be called into question.","supporting_citations":[],"review_version":1}