{"id":"bf9fdd8e-dd79-46a3-9d0d-c5116c58874d","arxiv_id":"2608.11330","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Quantum autoencoders and data-reuploading classifiers show smaller output-score shifts and better retention of discrimination than standard classical baselines under feature-level detector smearing in two collider benchmarks.","lead":"Parameterised quantum circuits trained on collider data keep their outputs more stable than standard neural networks when detector inputs are artificially smeared, while matching accuracy on clean data. The study suggests quantum models could aid trigger systems, but the comparison does not yet isolate the quantum origin of the effect.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The robustness advantage may reflect unmatched model capacity, training-set size, and regularization rather than a quantum-specific inductive bias; matched classical controls are needed before the central claim is supported.","rationale":"The paper is a carefully structured empirical study, and the raw observation that the tested quantum models show smaller score shifts under smearing is supported by the figures. The reader's CONDITIONAL verdict is appropriate because the central attribution to a quantum-specific inductive bias rests on a comparison in which classical baselines differ substantially in training-sample size, parameter count, and regularization. Appendix A partially addresses the training-sample-size confound for the unsupervised setting, but it does not control parameter count, and no such control is provided for the supervised setting. The supervised setting also has an internal inconsistency: pT observables are smeared while the derived kinematic variables computed from them are held fixed, so the test does not model a fully consistent detector shift. These are addressable with additional experiments rather than fatal flaws, so the correct verdict remains CONDITIONAL. I do not see a reason to move the verdict to REJECT or ACCEPT: the paper's explicit limitations and the partial sample-size scan in Appendix A show good faith, but the stronger claim about quantum robustness as an inductive bias is not yet established.","tokens_in":22443,"tokens_out":4995,"duration_ms":48100,"concrete_test":"Retrain the supervised comparison with a classical MLP regularized to match the effective smoothness and capacity of the four-layer quantum classifier: use the same 1000-event balanced training set, roughly the same parameter budget (about 211 trainable parameters), and weight decay or early stopping chosen so that the MLP's score-shift profile on an un-smeared validation set matches the quantum model's. Then re-run the smearing scan in Figures 9 and 11. If the regularized MLP reproduces the quantum classifier's mean-squared score deviation and AUC-retention curves, the quantum-specific robustness attribution is not supported; if the quantum classifier remains substantially more stable, the inductive-bias claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the abstract and Conclusions, is that parameterised quantum models exhibit a useful robustness inductive bias under detector-induced smearing. The experiments support the weaker statement that the specific quantum models tested are less sensitive than the specific classical baselines, but they do not yet establish that the quantum circuit structure is the cause. In the unsupervised study, each QAE has 16 trainable circuit parameters and is trained on 500 embedded background events, while the AE has 312 parameters and is trained on 10^6 events (Section 2.2); the VAE and AE also differ in objective and implicit regularization. Appendix A retrains the classical AE on 500 events and still sees larger score shifts, but parameter count and regularization remain unmatched, so the result remains consistent with a general 'lower-capacity or smoother model is more stable' explanation. In the supervised study, the four-layer quantum classifier has 7Lnq + 2nq + 1 = 211 trainable parameters for L=4, nq=7, trained on 1000 events, whereas the MLP has 81 parameters trained on 10^6 events and the linear model has 8 parameters (Section 3.2); the text's description of the MLP as 'approximately parameter-matched' is therefore inaccurate. The linear baseline, which has the smallest score shift under smearing, shows that low sensitivity can simply reflect weak input dependence. Without a classical model matched in capacity, training-set size, and regularization, the robustness gap could be a general property of constrained or smooth function classes rather than a property specific to parameterised quantum circuits. The supervised smearing test also holds the derived variables E_T^miss, M_TR, and M_TDelta fixed while smearing their constituent pT inputs (Section 3.3); a real detector shift would recompute these derived variables, so this leg of the comparison is evaluated under an internally inconsistent input distribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports two empirical studies comparing parameterized quantum models with classical baselines for collider-event selection under detector-induced smearing. In the unsupervised setting, quantum autoencoders (QAEs) trained on 500 embedded background events are compared with classical and variational autoencoders trained on 10^6 events; in the supervised setting, one- to four-layer data-reuploading classifiers trained on 1000 labelled events are compared with linear and multilayer-perceptron classifiers trained on 10^6 events. All models are evaluated on clean test data and on features smeared multiplicatively, with model parameters and preprocessing transformations held fixed. The paper finds competitive clean-data AUC for the quantum models and generally smaller mean-squared score shifts under smearing, and interprets this as evidence for a quantum-specific robustness inductive bias.","tokens_in":22672,"tokens_out":4355,"duration_ms":42545,"significance":"Robustness under distribution shift is practically important for collider triggers, and the paper provides a clearly described benchmark protocol with public datasets, exact circuit simulation, and a useful Appendix A sample-size scan. If the central attribution were supported, the result would be significant for quantum machine learning in high-energy physics. However, the experiments as presented support only the weaker statement that the specific quantum models tested are less sensitive than the specific classical baselines; the causal attribution to the parameterized-quantum-circuit hypothesis class is confounded by unequal training-set sizes, parameter counts, and regularization. The paper is well structured and the measurements are transparently reported, but the central claim needs additional matched classical controls before it can be accepted.","major_comments":[{"comment":"The unsupervised robustness comparison does not control for model capacity or training-data size. Each QAE has 16 trainable circuit parameters and is trained on 500 embedded background events, while the AE has 312 parameters and is trained on 10^6 events (Section 2.2). Appendix A retrains the AE on 500 events, but parameter count and regularization remain unmatched, so the smaller score shifts in Figure 13 remain fully consistent with the explanation that a lower-capacity or smoother model is more stable. A classical AE matched in parameter count and trained on 500 events, or a classical model with equivalent implicit regularization, is needed before the robustness can be attributed to the quantum model family rather than to the training configuration.","section":"Section 2.2, Appendix A"},{"comment":"The statement that the four-layer data-reuploading classifier is 'approximately parameter-matched' to the MLP is inaccurate. The L=4, nq=7 circuit contains 7Lnq = 196 circuit parameters plus 15 readout parameters, for 211 total, while the MLP has 81 parameters, and the quantum model is trained on 1000 labelled events versus 10^6 for the MLP. The linear classifier, with eight parameters, exhibits the smallest score shift but has near-random AUC, demonstrating that low sensitivity can simply reflect weak input dependence. The supervised comparison therefore requires classical baselines matched in parameter count, training-sample size, and regularization before the observed robustness can be attributed to a quantum inductive bias.","section":"Section 3.2"},{"comment":"The smearing protocol in the supervised study applies the multiplicative perturbation to the pT observables but does not recompute the derived kinematic inputs EmissT, MTR, and MTDelta, even though these are kinematically derived from the smeared momenta. The shifted input vectors therefore mix smeared low-level features with reference-derived features, which is not a consistent detector-variability scenario and may artificially reduce the measured score shift. The authors should either recompute the derived variables from the smeared momenta or explicitly characterize the test as a partial-feature perturbation and limit the robustness claim accordingly.","section":"Section 3.3"}],"minor_comments":[{"comment":"The product upper limit 'j=1+1' appears to be a typo for 'j=i+1'; please correct this and check the surrounding index notation.","section":"Equation (5)"},{"comment":"The word 'ansätz' should be 'Ansatz' in the text and figure captions.","section":"Throughout"},{"comment":"The text says 'each unmasked input feature' is smeared, but the description immediately before Equation (16) refers specifically to pT features; please clarify precisely which features are perturbed in the unsupervised study.","section":"Section 2.3"},{"comment":"The abbreviations MTR and MTDelta are used without definition; a one-sentence definition of these SUSY kinematic variables would improve accessibility.","section":"Section 3.2"},{"comment":"The shaded bands represent variation across smearing realizations only; the paper does not report variation across independent training seeds or initializations, so it is unclear whether the qualitative ordering of the models is stable. Please state this limitation or add a multiple-seed analysis.","section":"Figures 4, 5, 9, and 11"},{"comment":"No code or data-availability statement is provided; given the exact-simulation pipeline and public datasets, releasing code would substantially aid reproducibility.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and addresses a timely application of quantum machine learning. The main issue is that the central robustness claim is framed as a property of the quantum hypothesis class, while the experiments compare models that differ simultaneously in model family, capacity, training-set size, and regularization. I believe this is fixable with matched classical controls and a more cautious interpretation, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: solid proof-of-principle empirical study of a real question, but the headline claim overreaches. The quantum models tested are less sensitive to detector smearing than the classical baselines tested; that is not the same as showing parameterised quantum circuits provide a robustness inductive bias, and the paper's own numbers leave the stronger claim unsupported.\n\nWhat is actually new: this is the first comparison I know of between QAEs and data-reuploading classifiers under stochastic feature-level detector smearing for collider event selection, on public ADC2021 and SUSY data. The setup is clear: controlled smearing, frozen preprocessing and parameters, AUC plus score-deviation metrics, ten smearing realisations. The observation that compact QAEs stay competitive in clean-data AUC while showing smaller score shifts is worth having. Appendix A partially controls for training-set size in the unsupervised case and shows the classical AE still shifts more at N=500; that is a useful check.\n\nThe problems are all about attribution. The classical baselines are not matched in capacity, training-set size, or regularization. QAE: 16 parameters on 500 events; AE: 312 parameters on 1e6. Supervised: four-layer quantum has 211 parameters on 1000 events; MLP has 81 parameters on 1e6; linear has 8. The text calls the MLP 'approximately parameter-matched', which is inaccurate. With those confounds, the smaller score shifts are exactly what you would expect from lower-capacity or more constrained models; the linear baseline, which has the smallest shift of all, shows how low sensitivity can just mean weak input dependence. Appendix A does not fix this because it only varies N for the unsupervised models and leaves parameter count and regularization unmatched.\n\nA smaller but real issue: in the supervised smearing test, the derived variables E_T^miss, M_TR, and M_TDelta are held fixed while their constituent pT inputs are smeared. A real detector shift would recompute these derived variables, so that leg of the comparison runs on internally inconsistent inputs.\n\nNone of this kills the paper. The measurements look internally consistent, and the authors are upfront about simplifications. But the central conclusion needs matched classical controls, including equal parameter counts or explicit regularization sweeps, supervised sample-size controls, and recomputed derived variables, before it supports 'quantum robustness.'\n\nI would send it to referees, expecting major revision. It is a legitimate empirical contribution for the QML/HEP community, and I would cite the measured robustness comparison with the caveat.","headline":"The empirical robustness result is real for the specific models tested, but the stronger claim of a quantum-specific inductive bias is not yet supported by the unmatched classical baselines.","tokens_in":23320,"tokens_out":2948,"would_cite":true,"duration_ms":28339,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Parameterised quantum circuits provide a robustness inductive bias for collider-event selection, shifting their scores less than classical neural networks when detector-induced smearing is applied.","keywords":["quantum machine learning","collider event selection","distribution shift robustness","quantum autoencoder","data reuploading","anomaly detection","trigger systems","detector smearing"],"falsifier":"Train the classical autoencoder and multilayer perceptron on the same number of events as the quantum models and with matched parameter counts, apply the same smearing with frozen preprocessing, and compare the mean-squared score shifts; if the classical models then shift as little as the quantum models, the claimed quantum-specific robustness inductive bias is refuted.","tokens_in":22157,"feed_emoji":"⚛️","tokens_out":11884,"duration_ms":110717,"temperature":0.7,"pith_summary":"This paper asks whether the constrained structure of parameterised quantum circuits makes them more stable than classical neural networks when the input distribution shifts after deployment, as happens when collider detectors age or drift in calibration. Using two collider-event benchmarks, it trains quantum autoencoders for unsupervised anomaly detection and data-reuploading quantum classifiers for supervised signal-versus-background selection, then applies controlled feature-level smearing to the inputs with all parameters and preprocessing frozen. Under that shift, the quantum models generally produce smaller changes in their output scores and retain their signal-background discrimination better than the more expressive classical autoencoders and multilayer perceptron, while remaining competitive on clean inputs. The paper concludes that collider-event selection, and trigger systems in particular, is a candidate setting for robust quantum machine learning.","feed_headline":"Under detector drift, quantum models stay steadier than neural nets","feed_subtitle":"Small quantum circuits trained on fewer events shift less than neural networks when detector data drifts.","key_machinery":"The objects that carry the argument are two parameterised quantum circuit families. The first is the quantum autoencoder, whose per-event anomaly score is $1 - f_{\\mathrm{vac}}(x;\\phi)$, the infidelity of the trash register with the vacuum state after a learned unitary compresses background states into a one-qubit latent register; because the score comes from a quantum compression map measured by a swap test, it has a different functional form from classical reconstruction error. The second is the data-reuploading classifier, which alternates trainable angle-encoding layers (affine maps $w x + b$ feeding $R_Y$ and $R_Z$ rotations), a fixed cyclic entangling layer, and trainable single-qubit rotations, then reads out local and nearest-neighbour qubit correlations through a classical sigmoid. Repeated reuploading layers enlarge the accessible frequency spectrum of the decision function without adding qubits, giving a controlled dial for expressivity. The mechanism proposed is that the restricted, structured hypothesis class of these circuits limits how sensitively the learned output can track small input perturbations, so detector drift produces smaller score displacements while discrimination is retained.","core_discovery":"The central claim is that compact parameterised quantum models provide a useful robustness inductive bias for collider-event selection. In the unsupervised study, quantum autoencoders trained on 500 background events with 16 circuit parameters, against a classical autoencoder with 312 parameters trained on $10^6$ events, achieve competitive anomaly detection, especially at the low false-positive rates relevant to triggers, and under smearing their anomaly-score distributions move far less than the classical baselines' while their ROC AUC stays stable. In the supervised study, a four-layer data-reuploading classifier trained on 1000 labelled events reaches a clean-input AUC of 75.5% versus 74.3% for the multilayer perceptron trained on $10^6$ events, yet its output scores shift less under smearing and it retains the largest AUC in the strongly shifted regime, whereas the MLP degrades steadily. The paper reads these results as evidence that the hypothesis class of parameterised quantum circuits, through its encoding, entangling geometry, and fidelity-based compression, resists detector-induced shifts better than highly flexible classical models.","pith_inferences":["If the stability comes from low effective model complexity rather than quantum mechanics itself, then classical models with strong regularisation or early stopping might reproduce the robustness; this is testable and would not diminish the operational relevance.","The same smearing protocol should be repeated with a realistic detector-response model, finite measurement shots, and device noise, since the simulations here are noiseless and the robustness ordering could change under those conditions.","Because the fidelity-based quantum anomaly score and the reuploading classifier are structurally different, observing robustness in both suggests the effect is tied to quantum hypothesis classes broadly; testing other encodings would sharpen that conclusion."],"forward_implications":["Compact quantum autoencoders with only local entangling connectivity match the discrimination and robustness of all-to-all connected circuits, pointing toward hardware-efficient trigger implementations.","A four-layer data-reuploading classifier reaches clean-input AUC comparable to a multilayer perceptron while degrading less under strong smearing, placing quantum classifiers in a favourable robustness-performance trade-off.","Collider trigger models, which cannot be retrained frequently as detector conditions drift, could benefit from models whose event-level scores shift less under calibration changes.","Robustness under distribution shift should be reported alongside accuracy when comparing quantum and classical models for high-energy physics."],"supporting_citations":[{"why":"Supplies the quantum-autoencoder compression construction and the fidelity-based anomaly score.","marker":"[17]"},{"why":"Supplies the data-reuploading architecture used for the supervised quantum classifier.","marker":"[18]"},{"why":"Establishes that repeated data encoding enlarges the expressive power of the quantum model.","marker":"[19]"},{"why":"Applies quantum autoencoders to high-energy-physics anomaly detection, the setting extended here.","marker":"[11]"},{"why":"Provides the classical autoencoder baseline used in the unsupervised trigger-oriented comparison.","marker":"[45]"},{"why":"Supplies the unsupervised anomaly-detection benchmark and its signal samples.","marker":"[47]"},{"why":"Supplies the supervised signal-versus-background collider dataset.","marker":"[48]"},{"why":"Provides the feature-selection methodology for the supervised SUSY benchmark.","marker":"[10]"},{"why":"Provides the variational-autoencoder baseline whose robustness is compared with the quantum autoencoders.","marker":"[53]"}],"fun_headline_variants":["Quantum circuits beat neural nets at detector-drift endurance","Small quantum models shrug off collider detector drift","Quantum autoencoders resist detector shifts better than classical","Robustness edge: quantum models under detector smearing","Quantum classifiers stay accurate when detectors drift"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on the assumption that the robustness gap comes from the quantum circuit structure and not from the fact that the classical baselines were trained on roughly a thousand times more data and with many more parameters; the paper's appendix checks training-set size only for the unsupervised case, and does not match parameter counts or the supervised training sets.","fun_headline_variants_meta":{"raw":{"variants":["Quantum circuits beat neural nets at detector-drift endurance","Small quantum models shrug off collider detector drift","Quantum autoencoders resist detector shifts better than classical","Robustness edge: quantum models under detector smearing","Quantum classifiers stay accurate when detectors drift"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000712,"raw_usage":{"total_tokens":3239,"prompt_tokens":1013,"completion_tokens":2226,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":2153}},"tokens_in":629,"tokens_out":2226,"duration_ms":14218,"temperature":1.0,"reasoning_tokens":2153,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:46.421463+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the classical autoencoder and multilayer perceptron on the same number of events as the quantum models and with matched parameter counts, apply the same smearing with frozen preprocessing, and compare the mean-squared score shifts; if the classical models then shift as little as the quantum models, the claimed quantum-specific robustness inductive bias is refuted.","supporting_citations":[{"cited_title":"Event classification with quantum machine learning in high-energy physics,","cited_arxiv_id":null,"evidence_quote":"Provides the feature-selection methodology for the supervised SUSY benchmark."}],"review_version":1}