{"id":"07e18a57-9bc3-4170-8e32-de92967cabf5","arxiv_id":"2411.19914","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A variational protocol that learns measurement-and-feedback strategies prepares AKLT states with less pre-measurement entanglement than the analytic Smith et al. protocol and discovers a protocol for a specific AKLT edge-mode state.","lead":"The paper trains variational quantum circuits that include measurements and classical feedback, and shows they can learn feedback strategies for the AKLT state that a purely unitary circuit would need more depth to match. A generalist might read it because it suggests a path toward automatically discovering measurement-based state preparation protocols, including one for a specific AKLT state with no known deterministic low-depth protocol.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strongest result—preparing the AKLT |↑↑⟩ edge-mode state—rests on one 'lucky' seed; the paper's own teacher–student data in App. D says shallow circuits on entangled intermediates are rife with local minima, so no trainability claim backs this.","rationale":"The reader's verdict is CONDITIONAL and the weakest_assumption is the expressibility and trainability of the fixed ansätze. The paper is a numerical study with code and data, so the central claim is not falsified by the single-seed issue, but the strongest advertised result—the deterministic preparation of the |↑↑⟩ AKLT state—needs a multi-seed or basin-analysis demonstration. The verification step is computationally feasible because the paper's own MPS-based Julia code is published (Zenodo, GitHub). The depth and mutual-information comparisons are somewhat overstated but are secondary. I do not see an internal inconsistency or a fatal flaw; the concern is robustness of the headline numerical evidence. Hence CONDITIONAL, unchanged from the reader's verdict, with the recommended action being additional multi-seed and code-check evidence rather than a change of verdict.","tokens_in":22285,"tokens_out":2192,"duration_ms":19366,"concrete_test":"Re-run the Sec. 7 optimization for the AKLT |↑↑⟩ edge-mode state with, say, 100 random seeds under the published update-frequency and ancilla-regularization schedule; report the fraction of runs reaching the claimed infidelity threshold and whether the successful protocol's feedback angles coincide with one basin. If the success rate is below ~5% and no parameter-independent signature of the basin exists, then the single-lucky-run demonstration cannot support a deterministic-protocol claim. Second, check the published mVQE.jl code for whether the violet 'lucky' run's reported infidelity is computed on the same cost function as the other seeds.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline claim that a learned measurement-feedback protocol deterministically prepares the spin-up/spin-up AKLT edge-mode state (Sec. 7, Fig. 6(a)) is supported only by a single 'lucky' run (violet line), while the blue, orange, and green seeds get trapped in local minima. This is not merely a statistics concern: the paper itself (Sec. 3, App. D) shows that the sparse two-qubit feedback ansatz of depth 5 applied to entangled intermediate states has a loss landscape riddled with local minima, and that shallow circuits on entangled initial states are hard to optimize. The reader's weakest_assumption—that the fixed ansätze are expressive and trainable enough—is exactly the unproven premise. There is no analytic or reproducibility argument that the lucky run is a stable fixed point of the landscape rather than a rare escape. A related concern is the comparison against Smith et al.: the 'less mutual information before measurement' claim (Sec. 5, Fig. 4(g)) is not tied to a provable circuit-depth reduction; the paper only notes that the learned U1 needs depth 7 versus depth 8 to mimic Smith, while the minimum theoretical depth for the A-S-S-S-S-A pattern is 6. The depth claim is therefore weaker than the abstract's 'reducing circuit depth.' RNN scaling in Sec. 6 shows the bidirectional RNN does not learn optimal feedback for large sizes, with per-site infidelity growing unfavorably, so 'scalability' is demonstrated only in the weak sense of a fixed pre-measurement unitary plus an imperfect RNN. These are addressable by better numerical evidence rather than by a flaw in the central construction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a variational framework that combines projective measurements, ancilla resets, and classical conditional feedback with parameterized circuits to learn measurement-based state-preparation protocols. The target benchmark is the spin-1 AKLT chain (16 physical qubits, encoded as two qubits per spin), with the analytic constant-depth fusion protocol of Smith et al. as reference. The main technical claims are: (i) naive optimization of the measurement-feedback objective is plagued by measurement-induced local minima characterized by collapse of the measurement-outcome distribution P(M) to a delta function (or low entropy); (ii) two heuristics, unequal parameter update frequencies and an ancilla-distribution regularizer, mitigate these minima; (iii) the learned U1 protocol requires less pre-measurement mutual information than the Smith protocol, with circuit depth 7 versus 8 for the pre-measurement unitary; (iv) a translationally invariant ansatz with an RNN feedback module extrapolates to larger system sizes, with per-site infidelity growing unfavorably for large Ns; and (v) a single 'lucky' optimization run prepares a specific AKLT edge-mode state (|up up>) at high fidelity, a task with no known deterministic low-depth protocol. The paper is openly self-critical about its limitations: optimization remains non-convex, the RNN underperforms for large systems, and the edge-mode success rests on one seed.","tokens_in":22705,"tokens_out":4705,"duration_ms":37096,"significance":"If the central claims hold, the paper makes a useful contribution by demonstrating that variational search over measurement-and-feedback protocols can rediscover, and modestly improve on, analytic constant-depth preparation strategies for a nontrivial SPT state, and by identifying a concrete optimization failure mode specific to measurement-based variational circuits. The protocol equations (2)-(3) are clean and the fidelity objective (6)-(7) is externally defined, so the main numerical claims are not circular. The paper ships code and data on Zenodo and GitHub, which is a genuine strength. However, the headline 'learned protocol for AKLT |up up>' is supported by one seed only, and the scaling section explicitly reports deterioration with system size, so the significance is moderate and mostly proof-of-principle rather than a demonstrated robust new state-preparation method.","major_comments":[{"comment":"The central claimed result, a learned deterministic protocol preparing the AKLT |up up> edge-mode state, rests on a single 'lucky' seed (violet curve), while the three other seeds shown remain trapped in local minima. The paper states this explicitly but does not provide any reproducibility evidence, such as a second successful seed, a robustness check of the final fidelity to small parameter perturbations, or a test that the protocol is a stable fixed point rather than a rare escape. Given that Section 4 and Appendix D document a loss landscape riddled with local minima for exactly this ansatz class, the claim 'the results demonstrated the possibility of learning such a protocol' needs at least one more independent success and ideally a stability analysis before it can be treated as robust.","section":"Sec. 7, Fig. 6(a)"},{"comment":"The scalability claim is materially weaker than the abstract suggests. The RNN feedback does not learn the optimal correction for large sizes: per-site infidelity grows with Ns, and the paper notes that the infidelity per site for the further-optimized 'optimal correction' is nearly size-independent, proving that the RNN is the bottleneck. The text candidly admits that the RNN was not optimized to convergence and that the unidirectional RNN underperformed without a clear explanation. This is acceptable as a proof-of-concept, but the abstract's phrase 'demonstrating scalability' should be qualified to something like 'demonstrating extrapolation at moderate sizes with performance that degrades with system size'.","section":"Sec. 6, Fig. 5(c)"},{"comment":"The claim of 'reducing circuit depth' is supported only by depth 7 versus 8 for a specific mimicked pre-measurement unitary, not by an end-to-end circuit comparison. The learned U1 needs depth 7 versus depth 8 to replicate the Smith protocol's pre-measurement unitary; the minimum theoretical depth for the ASSSSA pattern is 6, so the learned protocol is not shown to be shallower than optimal, only shallower than a depth-8 replication. The conclusion wording in Sec. 8 ('potential for shallower circuits') is appropriately hedged, but the abstract's 'measurement-based shortcuts to reduce circuit depth' may overstate what is demonstrated. Please either state explicitly that the depth reduction is from 8 to 7 for the pre-measurement unitary only, or provide an end-to-end depth comparison.","section":"Sec. 5, Fig. 4(g) and depth discussion"},{"comment":"The two mitigation strategies are presented as based on a conjecture in Sec. 4.1 ('We conjecture that the extreme sharpening...') and on a regularizer with hyperparameter c in Sec. 4.2. The numerical evidence in Figs. 2 and 3 is consistent with the conjecture, but the paper does not provide a controlled test isolating the mechanism, such as showing that the same update-frequency schedule with a different optimizer or a different ansatz fails, or that the regularizer's benefit is not simply due to enlarging the search space. As these methods are explicitly ad hoc, a small ablation study or additional seed statistics would materially strengthen the claim that the identified learning-rate-imbalance mechanism is correct.","section":"Secs. 4.1 and 4.2"}],"minor_comments":[{"comment":"The abstract and Sec. 6 claim 'scalability', but Sec. 6 itself reports unfavorable per-site error growth; please harmonize the wording, for example by saying 'demonstrates extrapolation to larger sizes with expected trade-offs'.","section":"Abstract and Sec. 6"},{"comment":"Figure 4(g) uses a logarithmic y-axis for mutual information; the claim of a 'maximum mutual information length of 6 before measurement' is based on the decay plot, but the exact extraction of a 'length' from the plot is not described. Please define how the length is estimated.","section":"Sec. 5, Fig. 4(g)"},{"comment":"The regularizer depends on the window c and ratio r; the text says the window width c was chosen so that if lR = 0 then max P / min P < r, but Eq. (13) uses c in the threshold while Eq. (12) divides by Na. Please clarify whether c is dimensionless or normalized by Na, as the current notation is ambiguous.","section":"Sec. 4.2, Eqs. (12)-(14)"},{"comment":"In Eq. (26), the CiRX(theta) matrix has an extra comma after the last row entry; this is a typographical error. Also, the claim that the ansatz is capable of representing any two-qubit gate at a depth of five would benefit from a citation or a brief proof sketch, since it is load-bearing for the feedback ansatz choice.","section":"App. D, Eq. (26)"},{"comment":"There is a typo in the caption of Fig. 6(a): 'nad' should be 'and'. In addition, the code repository link in Ref. [55] contains a typo in the repository name: 'varaitional' should be 'variational'.","section":"Sec. 7, Fig. 6 caption and code block"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and well-scoped, but the central 'discovery' result in Sec. 7 is a single-seed lucky run, and the scaling section in Sec. 6 undermines the abstract's 'scalability' claim. The authors should be asked to: (i) provide at least one additional successful seed for the AKLT |up up> preparation, ideally with a stability check; (ii) reword the abstract and conclusion to match the actual evidence; and (iii) strengthen the local-minima mechanism claim with a brief ablation or additional seed statistics. These are fixable within the scope of a revision. The code and data availability is a point in favor. The paper has already been accepted in Quantum, so the revision would presumably be a post-acceptance adjustment or an erratum; if that is the case, the editor may wish to weigh whether the remaining single-seed evidence is acceptable for the journal's standards."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line up front: this is a real contribution, and the authors are unusually candid about its limits. The new thing is that they learn all feedback unitaries simultaneously for a variational circuit with projective measurements, rather than the greedy per-round methods in [26,27], and they back it up with code, data, and MPS simulations. The two local-minima mitigations (feedback update frequency and ancilla regularization) are clearly explained and demonstrated; the GHZ appendix with 50 runs per lambda gives some statistical backbone. The mutual-information comparison to Smith et al. is a nice diagnostic, and the finding that the learned pre-measurement unitary needs less mutual information than the analytic protocol is a real, if modest, result.\n\nWhere it gets soft: the abstract overclaims. 'Reducing circuit depth' is really 7 versus 8 layers for mimicking Smith's protocol, with a theoretical minimum of 6; that is not a depth reduction to boast about. 'Demonstrating scalability' is also generous: the bidirectional RNN in Sec. 6 does not learn optimal feedback for large sizes, and per-site infidelity grows with system size. The paper says this in the body, but the abstract does not.\n\nThe bigger caveat is the specific AKLT edge-mode state in Sec. 7. The headline protocol for |↑↑⟩ is a single 'lucky' run obtained with gradually decreasing noise. Four seeds, one success. The paper is honest about calling it lucky and says it 'warrants further analysis,' but a lone seed is not evidence of a stable or reproducible protocol. Worse, the paper's own teacher-student analysis in App. D shows that the sparse feedback ansatz on entangled intermediate states is riddled with local minima, which is exactly why the other seeds fail. So the claim 'no known deterministic protocol exists, and we found one' is really 'one optimization run found something that looks like one; we do not know if it is a stable fixed point or a rare escape.' That is a legitimate proof-of-principle, not a demonstrated protocol.\n\nMinor: the local-minima explanation is explicitly conjectural, and the left-correctability check is a heuristic stability test rather than a hard characterization. The main AKLT runs are single-seed without error bars. None of these are fatal, but they are precisely the places a referee should push.\n\nWho this is for: anyone working on measurement-based state preparation or variational circuits with feedback. It deserves serious refereeing, and with the abstract toned down it would be a comfortable accept. I would cite it if I were doing adaptive state prep. End recommendation: send to review, but insist the authors either add seeds/statistics or soften the abstract; the body is already close to honest.","headline":"Honest numerical paper with a genuinely new learning framework; just don't let the abstract's 'scalability' and 'depth reduction' claims outrun what the body shows.","tokens_in":23198,"tokens_out":4727,"would_cite":true,"duration_ms":39952,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68"],"pacs":["03.67.Lx","03.67.Mn"],"model":"deepseek-v4-flash","headline":"A variational circuit with measurement and learned feedback prepares the 16-qubit AKLT state with high fidelity, using less pre-measurement mutual information than analytic fusion.","keywords":["variational quantum circuits","measurement-based state preparation","quantum feedback","AKLT state","local minima","recurrent neural network","mutual information","mid-circuit measurement"],"falsifier":"Re-run the spin-up edge-mode optimization from many random seeds under the paper's update-frequency and ancilla-regularization settings, and count how many runs reach the reported low infidelity; if only the reported lucky seed succeeds, the claimed learned protocol is a single-run optimization event rather than a reproducible protocol.","tokens_in":22100,"feed_emoji":"⚛️","tokens_out":20403,"duration_ms":168630,"temperature":0.7,"pith_summary":"The paper tries to establish that a variational circuit can learn to treat projective measurement and classical feedback as resources for state preparation, rather than as corrections bolted onto a unitary ansatz. On the spin-1 Affleck-Kennedy-Lieb-Tasaki (AKLT) benchmark at 16 qubits, the learned protocol reaches high fidelity with a fixed shallow circuit and needs less mutual information before the measurement than the analytic constant-depth fusion protocol, which points toward shallower implementable circuits. To get there, the authors identify a family of measurement-induced local minima, characterized by collapse of the ancilla measurement distribution to low entropy, and suppress them by updating feedback parameters faster than the pre-measurement unitary and by adding an ancilla regularization term. A translationally invariant ansatz with recurrent-neural-network feedback extends the approach to systems of 8 to 32 qubits, with average per-site infidelity around $3\\times10^{-3}$ though not optimal at large sizes. The paper also reports a learned deterministic preparation of the AKLT state with both edge modes spin-up, a task for which no known short deterministic circuit exists, as evidence that learning can discover new protocols.","feed_headline":"Learned feedback prepares 16-qubit AKLT states more cheaply","feed_subtitle":"Less pre-measurement mutual information means the learned AKLT protocol may use shallower circuits than analytic fusion.","key_machinery":"The object that carries the argument is the feedback function $\\theta_2 = f(M;W)$: a map from the ancilla measurement outcome $M$ to the angles of the post-measurement unitary $U_2$. For small ancilla spaces it is a tabular lookup $W_M$; for larger systems it becomes a recurrent neural network. Around this function sits the protocol $\\rho_1 = U_1(\\theta_1)\\rho_0 U_1^\\dagger$, $\\rho_M = |0\\rangle_A\\langle M| \\rho_1 |M\\rangle_A \\langle 0|$, and $\\rho_2 = \\sum_M U_2(f(M;W)) \\rho_M U_2^\\dagger$, which lets a single round of measurement and correction implement a completely positive trace-preserving (CPTP) map. The paper's diagnostic for the newly identified traps is the Shannon entropy $H$ of the measurement distribution $P(M)$ together with the system-ancilla entanglement entropy $S$; at the traps $H$ collapses to low integer values, meaning the optimizer has learned to bypass the measurement. The two mitigation mechanisms both act on this diagnosis: updating $W$ more often than $\\theta_1$ keeps the feedback competitive with the pre-measurement unitary, and the ancilla regularization term pushes $P(M)$ toward uniform, preventing $H$ from collapsing.","core_discovery":"The central claim is that a parameterized feedback protocol, in which ancillas are projectively measured and the measurement outcome $M$ selects the angles $\\theta_2 = f(M;W)$ of a subsequent unitary through a learned function $f$, can prepare the four-fold AKLT manifold with high fidelity without prior knowledge of the analytic protocol. The optimization is non-greedy: every unitary and every feedback response is learned at the same time over all possible measurement outcomes, so the protocol can realize any completely positive trace-preserving map together with the correct conditional corrections. The learned protocol is compared with the analytic fusion protocol through two-qubit mutual information; it produces a similar block-like entanglement structure but requires less mutual information before measurement, and a shallower pre-measurement circuit of depth 7 suffices where mimicking the analytic unitary requires depth 8. The same machinery, with a translationally invariant ansatz and recurrent neural network feedback, prepares AKLT states for systems up to 32 qubits and extrapolates to sizes it was not trained on, albeit with non-optimal large-size corrections. For the specific AKLT state with both edge modes spin-up, a target with no known deterministic low-depth protocol, one optimization run reaches low infidelity, and the resulting protocol is left/right correctable even though no such constraint was imposed.","pith_inferences":["A testable extension the paper leaves implicit: if learned feedback protocols consistently need less mutual information before the measurement, then the amount of pre-measurement entanglement a target state requires could be treated as a resource, with learned protocols providing upper bounds for states like AKLT.","The spin-up edge-mode result comes from a single favorable optimization run, so the natural next step is to distill that run into a fixed gate sequence and verify it from independent seeds, turning an existence result into a reproducible protocol.","Because the recurrent neural network plateaus on large systems, swapping it for a transformer or state-space model, an option the paper names, would test whether the scalability bottleneck is the feedback architecture or the pre-measurement ansatz."],"forward_implications":["Measurement and feedback enter the variational optimization as trainable components, so the learned protocol represents a non-unitary CPTP map rather than a unitary circuit; this is what lets it match a constant-depth fusion protocol.","For the 16-qubit AKLT manifold, the learned protocol needs less two-qubit mutual information before the measurement than the analytic fusion protocol, and the pre-measurement circuit is shallower, with depth 7 versus depth 8 for replicating the analytic unitary.","The measurement-induced local-minima traps, diagnosed by collapse of the Shannon entropy of the ancilla outcomes, can be suppressed by updating feedback parameters more often and by adding the ancilla regularization term; the same strategies improve GHZ preparation in the appendix.","With a translationally invariant ansatz and RNN feedback, the protocol trains on sizes 8 to 32 qubits and extrapolates to untrained sizes, reaching average per-site infidelity around $3\\times10^{-3}$ across the trained sizes, though corrections are not optimal for large systems.","A deterministic learned protocol is reported for the AKLT state with both edge modes spin-up, a target with no known deterministic low-depth protocol, indicating that learning can discover state-preparation strategies beyond known analytic constructions."],"supporting_citations":[{"why":"Supplies the analytic constant-depth fusion protocol for AKLT that the learned protocol is benchmarked against and must match.","marker":"[12]"},{"why":"Defines the spin-1 AKLT state whose four-fold edge-mode manifold is the target state throughout.","marker":"[41]"},{"why":"Documents that shallow variational circuits are swamped with local minima, the optimization obstacle the paper classifies and mitigates.","marker":"[19]"},{"why":"Shows how cost functions create barren plateaus in shallow parameterized circuits, a related landscape trap the feedback approach must avoid.","marker":"[20]"},{"why":"Defines variational quantum algorithms, the framework being extended by measurement and feedback.","marker":"[6]"}],"fun_headline_variants":["Learned feedback finds shallower circuits for AKLT states","Feedback learning beats analytic fusion for AKLT preparation","Scalable learned feedback prepares AKLT states to 32 qubits","Measurement feedback discovers new AKLT state preparation protocols"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the fixed circuit shapes chosen for the pre-measurement unitary and the feedback step are expressive and optimizable enough for the AKLT targets; the paper gives no general guarantee, and its own teacher-student simulations show that shallow circuits applied to entangled intermediate states are riddled with local minima.","fun_headline_variants_meta":{"raw":{"variants":["Learned feedback finds shallower circuits for AKLT states","Feedback learning beats analytic fusion for AKLT preparation","Scalable learned feedback prepares AKLT states to 32 qubits","Measurement feedback discovers new AKLT state preparation protocols"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000515,"raw_usage":{"total_tokens":2515,"prompt_tokens":976,"completion_tokens":1539,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":1474}},"tokens_in":592,"tokens_out":1539,"duration_ms":12067,"temperature":1.0,"reasoning_tokens":1474,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:40:59.991507+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the spin-up edge-mode optimization from many random seeds under the paper's update-frequency and ancilla-regularization settings, and count how many runs reach the reported low infidelity; if only the reported lucky seed succeeds, the claimed learned protocol is a single-run optimization event rather than a reproducible protocol.","supporting_citations":[{"cited_title":"Smith, Eleanor Crane, Nathan Wiebe, and S.M","cited_arxiv_id":null,"evidence_quote":"Supplies the analytic constant-depth fusion protocol for AKLT that the learned protocol is benchmarked against and must match."},{"cited_title":"Rigorous results on valence-bond ground states in antiferromag- nets","cited_arxiv_id":null,"evidence_quote":"Defines the spin-1 AKLT state whose four-fold edge-mode manifold is the target state throughout."},{"cited_title":"Cost function dependent barren plateaus in shallow parametrized quantum cir- cuits","cited_arxiv_id":null,"evidence_quote":"Shows how cost functions create barren plateaus in shallow parameterized circuits, a related landscape trap the feedback approach must avoid."},{"cited_title":"Variational quantum algo- rithms","cited_arxiv_id":null,"evidence_quote":"Defines variational quantum algorithms, the framework being extended by measurement and feedback."}],"review_version":1}