{"id":"6b9b08a4-69dc-4067-8d64-5675a19a96ff","arxiv_id":"2509.08890","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"An unsupervised attention network trained on measurement outcomes detects measurement-induced entanglement between distant qubits in 34-qubit and 36-qubit cluster states, and its learning difficulty peaks near the expected measurement-induced phase transition.","lead":"Using a superconducting quantum processor, the authors trained an unsupervised neural network to predict the state of two untouched qubits from the random outcomes of measurements on many other qubits, then used the network's predictions to reveal entanglement created by those measurements. The approach could let experimenters observe measurement-induced quantum effects without postselecting rare outcomes or knowing the state preparation circuit in advance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed learnability transition at large θ may be a finite-capacity/training-budget artifact; the gate-based model shows m→ρ_m structure exists while the NN fails within 20 epochs, so the phase-transition narrative needs a scaling test.","rationale":"In good faith, the paper's core rigorous contribution is the cross-correlation bound of Eq. (2): a positive measured N_SC implies positive true measurement-averaged negativity regardless of the quality of the computational model. This makes the 1D detection and the 2D intermediate-θ detection solid. The exposed soft spot is not the bound but the interpretation of the NN's large-θ failure as a learnability transition related to a measurement-induced phase transition. The paper itself contrasts NN and gate-based models (Fig. 3B, Fig. 4A) and states that at large θ the NN cannot approximate m↦ρ_m even though structure is present; this is precisely where a fixed 20-epoch budget or a particular architecture can mimic a phase boundary. The SI provides no scaling of model width, depth, or training time, so the intrinsic nature of the failure is unestablished. A controlled retraining on clean simulated data with larger capacity and longer training would settle whether the failure is intrinsic or practical. If it is removable, the abstract's 'transition in the ability of a classical agent' would need to be rephrased as a property of the chosen learner, and the claim of observing the MIPT without advance knowledge would not be supported from the large-θ side; the intermediate-θ MIE detection would remain valid. This is the same concern as the reader's weakest assumption, so the existing CONDITIONAL verdict should stand unchanged. The 'without postselection' wording should also be clarified in light of the SI's error-detection postselection, but that is secondary to the learnability concern.","tokens_in":23961,"tokens_out":10432,"duration_ms":124144,"concrete_test":"Using the released training code, retrain the attention model on noiseless simulated 6×6 cluster states at θ/π=0.5 with 100 epochs and 4× hidden dimension, keeping the rest of the protocol identical. Evaluate held-out S_QC and N_QC. If S_QC falls materially below 2 bits or N_QC becomes positive, the large-θ failure in Figs. 3B/4A is a finite-capacity/training-budget artifact; if S_QC remains ≈2 even with this increased capacity and clean data, the intrinsic-learner interpretation is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing premise is that the NN's failure at large θ is an intrinsic learnability transition rather than a consequence of the chosen architecture and the fixed 20-epoch training budget. The paper's own Fig. 3B shows that at θ/π≈0.5 the gate-based model yields S_QC significantly below the NN's ≈2 bits, so the relation m↦ρ_m is not information-theoretically absent: a classical model with gate knowledge can capture it. The claim that the observed change is 'closely related to a measurement-induced phase transition' therefore rests on the assumption that no larger or longer-trained network would learn the large-θ mapping. No scaling in model size or training time is reported; the training curves in Fig. 3A stop at t=20. Consequently the peak in D_KL reduction (Fig. 3C) and the vanishing NN negativity bound at large θ (Fig. 4A) could be practical model-saturation effects rather than a physical transition. The rigorous cross-correlation bound in Eq. (2) protects the intermediate-θ MIE detection, since any model gives a valid lower bound, but the broader 'observe MIPT without advance knowledge' claim is weakened if the large-θ failure is removable by more resources. Separately, the abstract's 'without postselection' is stronger than the SI's error-detection postselection, though this is secondary to the main concern.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports experiments on Google's Sycamore and Willow processors in which post-measurement states of two probe qubits are characterized from the outcomes of many other measurements. Cluster states are prepared in 1D (L up to 34) and 2D (6x6) arrays; all non-probe qubits are measured, and the probe qubits are measured in random Pauli bases to generate classical shadows. A generative transformer NN is trained unsupervised, mapping outcome sets m to estimated probe states rho^C_m; cross-correlations with independent test shadows, via the inequalities S^SC >= S_m and N^SC <= N_m (Eqs. 1-2), yield bounds on average entropy and negativity. Main claims: (i) in 1D, the NN and tensor-network models give positive negativity bounds comparable to a gate-based model, up to L=34; (ii) in 2D, a peak in the detected negativity and in the amount of information learned appears at intermediate measurement-basis angle theta, while at large theta the NN produces near-maximally-mixed estimates and fails to detect MIE although the gate-based model still does; this failure is interpreted as a learnability transition related to the measurement-induced phase transition, observable without advance knowledge of the quantum state and without postselection.","tokens_in":24302,"tokens_out":13567,"duration_ms":149944,"significance":"The methodological core is sound and valuable. The cross-correlation bounds (Eqs. (1) and (2)) are derived self-contained in the SI and are rigorous for any model rho^C_m; the train/test split makes the ML evaluation non-circular; the code and data are released (Refs. [47,48]); and the 1D experiment gives clean evidence of MIE up to L=34 from unsupervised models. The 2D results provide a concrete finite-size demonstration that a data-driven model can certify positive negativity in a window of intermediate measurement angles, and the row-flip sensitivity analysis (SI Fig. 12A) is a falsifiable diagnostic of nonlocal dependence on outcomes. The principal weakness is the interpretive claim that the NN's failure at large theta is a physical learnability transition: the paper's own Fig. 3B shows that the gate-based model captures structure that the 20-epoch NN misses, so without a scaling study the observed transition may reflect model capacity and training budget rather than an information-theoretic MIPT signature. I therefore view the MIE detection claims as supported, and the transition narrative as requiring additional evidence or substantially weakened wording.","major_comments":[{"comment":"The learnability-transition claim rests on the NN's failure at large theta, but the evidence is one architecture trained for a fixed t=20 epochs. Fig. 3B shows the gate-based model yields S^QC well below 2 bits at theta/pi ~ 0.5, so the m -> rho_m structure is present in the data, and the text concedes the NN cannot approximate it. The peak in D_KL reduction (Fig. 3C) and the vanishing negativity bound (Fig. 4A) could then reflect model capacity or the 20-epoch budget rather than an intrinsic transition. Please report (a) training curves beyond 20 epochs (at least for theta/pi ~ 0.3-0.5) or saturation evidence; (b) a scaling test in model size/training-set size showing the ~2-bit plateau persists; (c) the Fig. 3C peak location relative to the gate-based finite-size crossing (SI Fig. 10A). Without these, the abstract's transition claim should be weakened to 'this NN under the fixed traini","section":"Two dimensions; Figs. 3A-C, 4A"},{"comment":"The abstract's 'without postselection' is internally inconsistent with the SI, which discards runs flagged by error-detection qubits and states 'This corresponds to post-selecting on error-free repeats of experiment.' This is not postselection on the exponentially many outcome strings m, so the method's advantage over Ref. [16] survives, but the wording should be 'without postselecting on measurement outcomes' and the discarded fraction should be stated. Also, the NN receives positional encodings and a 2D causal masking schedule; 'without advance knowledge of the quantum state' should be scoped to exclude knowledge of the measurement layout.","section":"Abstract vs SI Sec. I A"}],"minor_comments":[{"comment":"Both rho_i1 and rho_i2 are defined with averages over R1; the second should average over R2.","section":"SI Sec. V B, Eq. (22)"},{"comment":"The axes and captions use 'mu' (mu/deg) while the text uses theta for the measurement basis; unify the notation, preferably with theta/pi as in the text.","section":"Figs. 3-4"},{"comment":"The caption states the plotted quantity is S^QC_m(t) - S^QC_m(20); clarify that this equals D_KL(t) - D_KL(20) and describe how non-monotonic training would appear.","section":"Fig. 3C"},{"comment":"For the classification estimates in Eq. (23), report the class sizes R1, R2 and whether depolarization was needed; the text says it was not, but the convergence to the post-measurement ensemble average should be stated more explicitly.","section":"SI Sec. V B"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern is legitimate. The learnability transition is the weakest point; if scaling tests (longer training, larger models) cannot be performed, the authors should explicitly repackage the claim as a practical learning failure in a fixed architecture rather than a MIPT-related transition. The MIE detection claims themselves are rigorous and well executed; the gap between the abstract's claims and the SI's error-detection postselection should also be corrected. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, the core detection result is solid: the cross-correlation bounds in Eqs. (1) and (2) are rigorous, the 1D experiment shows positive negativity bounds from unsupervised models up to L=34, and the 2D intermediate-theta peak in the negativity bound is a genuine lower bound on measurement-induced entanglement. They also ship code and data. Second, the broader claim about observing a measurement-induced phase transition through a learnability transition is weaker than the abstract suggests. At theta/pi near 0.5, the gate-based model sees entanglement structure that the NN fails to learn within 20 epochs. The paper calls this a learnability transition, but the data are equally consistent with the NN being too small or undertrained for that harder region. There is no scaling test in model size or training time, so the transition in D_KL reduction and the vanishing NN negativity bound could be a model-capacity artifact. That does not undermine the intermediate-theta detection, because Eq. (2) gives a valid bound for any model, but it does undermine the phase-transition narrative. Also, the abstract's 'without postselection' is overstated: the SI explicitly says they discard runs flagged by error-detection CNOTs. That is postselection, just not exponentially costly postselection. What is new: unsupervised training of an attention-based network to map measurement outcomes to post-measurement probe states, then using those models to form cross-correlation bounds. The 1D result that the NN matches a gate-based model without being told the gates is genuinely nice. The 2D peak in KL reduction at intermediate theta is suggestive, and the nonlocal dependence on row-5 but not row-6 outcomes is a useful sanity check. Bottom line: the experimental method and the 1D/2D negativity bounds deserve a serious referee and probably publication after revision. The learnability-transition interpretation needs a scaling test or a more careful statement that it is a finite-resource learning effect, not established as the MIPT. I would send it to review, ask for the scaling data and rewording, and cite it for the cross-correlation-based detection method.","headline":"The negativity detection is real and worth taking seriously; the learnability-transition story at large theta is oversold and needs a scaling test.","tokens_in":727,"tokens_out":2939,"would_cite":true,"duration_ms":46696,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.65.Ud","03.67.-a","03.67.Lx"],"model":"deepseek-v4-flash","headline":"A generative neural network trained only on measurement outcomes reveals entanglement induced by measuring many qubits, with no model of the state preparation and no postselection on outcome strings.","keywords":["measurement-induced entanglement","unsupervised learning","neural networks","classical shadows","entanglement negativity","measurement-induced phase transition","cluster states","learnability transition"],"falsifier":"Run the same 2D experiment at large measurement angle with a much larger, better-trained network (or an exact classical simulation of the outcome-to-state map). If a scalable model then reproduces the gate-based negativity bound, the claimed learnability transition is an artifact of finite model capacity; if even an optimal learner fails while the gate-based model succeeds, the transition is intrinsic.","tokens_in":23873,"feed_emoji":"🧠","tokens_out":7894,"duration_ms":74246,"temperature":0.7,"pith_summary":"Many measurements create long-range entanglement between the qubits that were not measured, but the post-measurement state is different for each random outcome string, so the entanglement is normally invisible without postselecting on those outcomes. This paper shows that the entanglement can be certified from raw data alone: an unsupervised attention-based neural network maps observed outcome strings to estimates of the two-qubit post-measurement state, and cross-correlating those estimates with held-out classical shadows gives positive lower bounds on measurement-averaged entanglement negativity. The evidence includes one-dimensional cluster chains up to 34 qubits and 6×6 two-dimensional arrays near an intermediate measurement angle. In the 2D data, sweeping the measurement basis produces a sharp change in how much the network can learn, which the paper identifies as a learnability transition closely related to a measurement-induced phase transition. The reason to care: if measurement-induced effects can be decoded from unlabelled data, they become observable in systems whose preparation is not accurately known, including prospective quantum error correction and control settings.","feed_headline":"Unsupervised AI detects entanglement hidden in quantum measurements","feed_subtitle":"Without knowing the state or picking outcome strings, it certifies entanglement on up to 34 qubits.","key_machinery":"The load-bearing objects are the quantum-classical cross-correlation inequalities. Given any computational model ρ^C_m for the two-probe post-measurement state, the quantum Kullback-Leibler divergence ensures S^QC_m = −Tr[ρ_m log ρ^C_m] ≥ S_m, and the projector onto negative eigenvalues of the partial transpose ensures N^QC_m = −Tr[(ρ_m)^TA Π((ρ^C_m)^TA)] ≤ N_m. These bounds become experimentally accessible because weighted averages of classical shadows over repeats converge to the same quantities: the shadow noise has zero mean. The second ingredient is the unsupervised generative network: an attention-based transformer that takes the outcome string m as input and outputs a valid 4×4 densit","core_discovery":"Central claim: measurement-induced entanglement can be revealed from measurement data alone—no model of the prepared state, no postselection on outcome strings. For any model m→ρ^C_m, cross-correlations give S^QC ≥ S_m and N^QC ≤ N_m, and classical-shadow averages realize these bounds because shadow noise averages to zero. An unsupervised attention-based network learns such a map from outcomes. Cross-correlating its predictions with held-out shadows yields positive negativity bounds in 1D cluster chains up to L=34 (matching a gate-based model) and in 2D 6×6 arrays near an intermediate angle. The network's sharp failure at larger angles is read as a learnability transition tied to the measure","pith_inferences":["A sharper test the paper leaves implicit: retrain at large angles with larger capacity or longer schedules; if the entropy bound drops toward the gate-based value, the observed transition is partly a model-capacity artifact rather than an intrinsic property of the data.","The learned model's insensitivity to outcomes in the farthest row, combined with sensitivity to the adjacent row, suggests the trained network itself can be used as a diagnostic of the 'light cone' of measurement-induced correlations—an experimental tool not developed in the paper.","If the same unsupervised protocol is applied to monitored random circuits, the peak in KL-divergence reduction during training may serve as a generic finite-size order parameter for measurement-induced criticality, independent of any chosen entanglement measure."],"forward_implications":["If a trained network can certify measurement-induced entanglement without a state-preparation model, then measurement-induced phase transitions can in principle be located in any experimental platform with repeated state preparation and single-shot readout.","Because the same cross-correlation bounds apply to coherent information and general observables, the scheme transfers directly to quantum error correction, where syndrome-to-logical-state maps could be learned rather than assumed.","The data show that the detectability window of the network (intermediate angles) is narrower than the entanglement window seen by the gate-based model, so a 'learnability transition' does not coincide exactly with the physical transition at finite size; this distinction matters for interpreting measurement-induced phase transition experiments.","In 1D, learned models match gate-based models for negativity bounds even though they ignore the preparation circuit, suggesting the method remains useful as systems grow and gate-level descriptions become unreliable."],"fun_headline_variants":["Unsupervised AI exposes entanglement hidden in measurements","Neural nets reveal quantum entanglement from measurement data","Measurement-induced entanglement spotted by unsupervised learning","AI certifies entanglement without postselecting outcomes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claim would collapse if the network's failure at large measurement angles were due to the fixed 20-epoch training budget or the chosen architecture rather than to an intrinsic difficulty of learning the measurement-to-state map.","fun_headline_variants_meta":{"raw":{"variants":["Unsupervised AI exposes entanglement hidden in measurements","Neural nets reveal quantum entanglement from measurement data","Measurement-induced entanglement spotted by unsupervised learning","AI certifies entanglement without postselecting outcomes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000242,"raw_usage":{"total_tokens":1345,"prompt_tokens":713,"completion_tokens":632,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":575}},"tokens_in":457,"tokens_out":632,"duration_ms":7172,"temperature":1.0,"reasoning_tokens":575,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T20:02:50.094341+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 2D experiment at large measurement angle with a much larger, better-trained network (or an exact classical simulation of the outcome-to-state map). If a scalable model then reproduces the gate-based negativity bound, the claimed learnability transition is an artifact of finite model capacity; if even an optimal learner fails while the gate-based model succeeds, the transition is intrinsic.","supporting_citations":[],"review_version":1}