{"id":"bc033e63-f1ef-4ed5-a284-ea1313e72d53","arxiv_id":"2608.06846","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An interface-matched factorial finds no consistent accuracy or stability benefit from replacing a classical map with a parameterized quantum circuit in a hybrid classifier.","lead":"This paper tests whether a small quantum circuit inside a hybrid classifier improves accuracy over a classical replacement with the same inputs and outputs. Across controlled paired experiments on a breast-cancer dataset, it finds no consistent quantum advantage, and it shows why hybrid model performance should not be credited to the quantum layer without such controls.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the paper's scoped negative claim (“no consistent PQC contribution across the tested BCW settings”) is supported by the reported paired factorial, and the acknowledged limitations do not undermine it.","rationale":"The paper's strongest claim is a conditional: if the factorial shows no consistent switch-A effect, then hybrid performance should not be attributed to the quantum layer. The antecedent is supported by the data: one of four contrasts is positive before correction, fails BH, and changes sign at n_q=8. The experiment is small and saturated, but the paper does not overclaim equivalence or general absence of effect; it says 'not replicated across the tested settings.' The parameter-count mismatch between surrogate and PQC, which the reader identified as the weakest assumption, actually works in favor of the paper's attribution claim: a larger classical map matching the PQC shows that the quantum layer is not necessary for the observed performance. Thus the concern about masking a true effect would only affect generalization to other tasks, not the validity of the scoped conclusion. The one non-zero contrast is handled appropriately with BH correction and an explicit 'result to replicate' caveat. I find no internal inconsistency or unsupported leap in the central argument. The administrative incomplete competing-interests declaration is noted by the reader and does not affect the scientific verdict.","tokens_in":668,"tokens_out":1013,"duration_ms":231997,"concrete_test":"Independently recompute the four paired quantum-minus-classical contrasts and their 95% CIs from the deposited per-seed CSVs (Table 2); verify that the n_q=4 attention contrast remains the only interval excluding zero and that the Benjamini-Hochberg corrected p-value is 0.100. If any recomputed CI crosses zero differently, the 'no consistent effect' reading changes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"I find no load-bearing concern that would change the verdict. The central claim is deliberately scoped to the tested BCW settings and is a statement about replication, not equivalence. The classical surrogate tanh(Wξb) has more parameters than the PQC (Table 6), so it is a stronger rather than weaker control; the absence of a consistent positive quantum-minus-classical contrast therefore cannot be explained by a too-weak classical baseline. The single positive contrast at n_q=4 attention (p=0.025, BH p=0.100) is correctly identified as non-replicated, and the paper explicitly declines to infer equivalence (Sections 5.1 and 6.6). Low power and near-ceiling accuracy are acknowledged limitations that weaken any general claim but do not falsify the scoped 'not replicated' conclusion. I therefore have no significant objection.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper tests whether a parameterized quantum circuit (PQC) improves a hybrid quantum-classical model's performance on classical datasets, using an interface-matched classical map as the control while holding all other components fixed. The architecture, Quantum-Embedded Attention (QEA), consists of a classical backbone, a learnable projector, a shallow PQC with Pauli readout, and a classical attention decoder. The central experiment is a 2x2 factorial on Breast Cancer Wisconsin at n_q in {4,8}, where the PQC is independently swapped for a classical map and the attention decoder for a linear head, across five paired seeds per cell. Three of four paired quantum-minus-classical confidence intervals include zero, and the one positive contrast at n_q=4 with the attention decoder reverses sign at n_q=8 and does not survive Benjamini-Hochberg correction. The paper concludes that the PQC contribution is not replicated across the tested settings, does not claim equivalence, and reports an exploratory five-dataset grid with a large deficit on CIFAR-10. The manuscript is unusually careful with run accounting, leakage caveats, and the distinction between current Pauli-readout and legacy probability-readout protocols.","tokens_in":12563,"tokens_out":7094,"duration_ms":74889,"significance":"If the scoped negative claim is taken as the contribution, this is a valuable methodological result: it demonstrates how controlled component attribution should be performed in hybrid quantum-classical machine learning and provides a template for reporting negative results with paired seeds, confidence intervals, multiplicity correction, and complete run accounting. The paper explicitly avoids overclaiming: it does not infer equivalence from overlapping intervals, does not attribute exploratory-grid performance to the quantum layer, and discloses the parameter-count mismatch between the PQC and its classical surrogate. The main limitations—a single tabular dataset for the factorial, only five paired seeds, near-ceiling accuracy, a higher-capacity classical control, and exact statevector simulation—are all acknowledged and do not undermine the narrow conclusion that no consistent switch-A effect was observed. The availability of code, configuration files, and run-level CSV files further strengthens reproducibility.","major_comments":[],"minor_comments":[{"comment":"The competing interests section currently contains the placeholder text 'This declaration must be completed and approved by all authors before resubmission'; it needs to be replaced with an actual statement before the manuscript can be accepted.","section":"Competing interests"},{"comment":"The dataset name is typeset inconsistently: 'CIF AR-10' appears in the abstract, Table 1, and Section 4.1, while Figure 3 uses 'CIFAR-10'; please unify the spelling.","section":"Throughout"},{"comment":"The parameter audit shows that the classical surrogate has considerably more trainable parameters than the PQC (50 vs 8 at n_q=4 and 324 vs 16 at n_q=8); the paper's term 'interface-matched' is honest and the discussion correctly avoids claiming equivalence, but readers should be reminded in the abstract or conclusion that the null result is relative to this higher-capacity classical control.","section":"Section 4.2 / Table 6"},{"comment":"In Eq. (6), the phrase 'with depth L' is ambiguous because L appears both as the number of ansatz layers and as the outer product index; please state explicitly that L is the number of repeated blocks in the hardware-efficient ansatz.","section":"Section 3.3, Eq. (6)"}],"recommendation":"minor_revision","confidential_remarks":"I agree with the reader that the central scoped claim is sound and the negative result is carefully supported. The only reason I recommend minor revision rather than accept is the placeholder competing interests statement and a few local presentation issues; these are trivially fixable and do not affect the scientific content."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a carefully scoped negative result, and the methodology is the real story. The interface-matched factorial on BCW — paired seeds, a classical surrogate with the same input/output dimensions, and full run accounting — is a genuinely useful template for QML evaluation. The central claim, that the PQC shows no consistent contribution across the tested settings, holds up against the reported statistics. The one positive contrast at n_q=4 attention is small, flips sign at n_q=8, and does not survive correction; the paper correctly treats it as a replication candidate, not a finding.\n\nWhat it does well: the authors are unusually honest. They disclose the parameter mismatch between the PQC and the classical surrogate, the shared scaler leakage, the clip-level BirdCLEF split, the collapsed runs, and the exploratory status of the five-dataset grid. They do not claim equivalence, only non-replication, which is the right epistemic stance for a low-power experiment. The distinction between the interface-matched factorial and the descriptive residual variant (QEA-R) is clear and prevents over-attribution.\n\nSoft spots, in proportion: the controlled experiment is limited to one near-ceiling dataset with five seeds, so power to detect small effects is low. That is an acknowledged limitation, but it means the negative result is narrow. The classical surrogate tanh(Wξ+b) has many more parameters than the PQC, so it is a stronger rather than weaker control, but it is not a parameter-matched control; the authors say this, and it does not undermine the null. The exploratory grid is not interface-matched, so its failures on CIFAR-10 and SUSY are descriptive, not causal. The incomplete competing interests statement is an administrative issue, not a scientific one. I see no load-bearing flaw.\n\nThe citation pattern is sound: they engage with the relevant literature and position their work against their own prior single-qubit study, which they are explicitly testing rather than relying on. This is a model of how to report a negative result with the rigour that QML needs.\n\nWho this is for: anyone evaluating hybrid quantum-classical models, or designing controlled comparisons of quantum components. It deserves a serious referee, and I would cite it as an example of honest negative-result reporting. Send it to peer review, with encouragement rather than skepticism.","headline":"A carefully scoped negative result whose methodology, not the quantum model, is the real contribution; it deserves peer review.","tokens_in":13085,"tokens_out":1717,"would_cite":true,"duration_ms":19799,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A controlled 2×2 factorial finds no consistent quantum-minus-classical effect from embedding a PQC in a hybrid classifier on BCW.","keywords":["quantum machine learning","hybrid quantum-classical","parameterized quantum circuits","component attribution","interface-matched controls","negative results","quantum-embedded attention","BCW factorial"],"falsifier":"A replication run of the same switch-A factorial on a non-saturated task such as CIFAR-10, with a prespecified smallest effect of interest, a residual-free core QEA, and at least ten paired seeds, would settle the claim: if the PQC shows a positive paired contrast whose confidence interval excludes zero at both $n_q = 4$ and $n_q = 8$ and both decoders after correction, the 'not replicated' conclusion fails.","tokens_in":12219,"feed_emoji":"⚛️","tokens_out":12051,"duration_ms":113842,"temperature":0.7,"pith_summary":"This paper sets out to test whether a parameterized quantum circuit (PQC) improves a hybrid quantum-classical classifier once everything else is held fixed, and reports no consistent PQC contribution across the settings tested. It inserts a small PQC between a classical feature projector and a classical attention decoder, then independently swaps the circuit for a classical map with the same input/output dimensions (switch A) and the attention decoder for a linear head (switch B). On Breast Cancer Wisconsin, three of four paired quantum-minus-classical confidence intervals include zero, and the one positive contrast (+1.63 points with attention at four qubits) reverses sign at eight qubits and does not survive multiplicity correction. The paper's positive contribution is methodological: it shows why hybrid-model performance should not be credited to the quantum layer without an interface-matched control, and it reports all planned runs including collapses and incomplete cells.","feed_headline":"Quantum layer shows no consistent edge over classical map","feed_subtitle":"Interface-matched swaps on breast-cancer data leave three of four paired effects including zero.","key_machinery":"The load-bearing object is the interface-matched $2 \\times 2$ factorial built on the Quantum-Embedded Attention (QEA) pipeline, in which a classical backbone and projector feed an angle vector $\\xi$ to a data-dependent PQC $Q_\\omega$ whose one- and two-body Pauli expectations form the readout vector $m$, consumed by a classical attention decoder. Switch A replaces the circuit with $\\tanh(W\\xi+b)$, whose input and output dimensions match the PQC's, and switch B replaces the attention decoder with a linear head; the projector, readout interface, seeds, and training budget are held fixed. The statistical object that carries the argument is the within-seed, paired quantum-minus-classical contrast at $n_q = 4$ and $n_q = 8$, not the marginal interval overlap.","core_discovery":"The paper's central claim is a null result under controlled component attribution: on the interface-matched BCW factorial with $n_q = 4$ and $n_q = 8$, independently replacing the PQC with the classical surrogate $\\tanh(W\\xi+b)$ and the attention decoder with a linear head produces no consistent paired quantum-minus-classical effect. Three of the four paired 95% confidence intervals include zero; the $n_q = 4$ attention contrast is $+1.63$ percentage points with interval $[0.34, 2.92]$ but reverses sign at $n_q = 8$, has an unadjusted $p = 0.025$, and does not survive Benjamini-Hochberg correction across the four contrasts. The authors conclude that the PQC contribution is not replicated across the tested settings, explicitly decline to claim equivalence, and present the five-dataset grid as descriptive because those columns are not interface-matched.","pith_inferences":["If the same null pattern appears on a non-saturated task, the likely explanation for many hybrid-QML gains is the classical projector or decoder rather than the circuit; future claims should be framed around hardware-specific advantages such as native entangling operations or sampling cost.","A sharper control would match parameter counts or capacity between PQC and surrogate; since the surrogate has 50 parameters vs 8 at $n_q = 4$ (Table 6), the current test is conservative only if extra classical parameters actually help on BCW.","Applying the same switch-A factorial to CIFAR-10 with a residual-free core QEA and prespecified effect sizes could determine whether the bottleneck is the angle projector rather than the circuit itself.","Reporting collapsed and incomplete runs, as this paper does, may become a useful norm; otherwise publicly reported hybrid-QML averages are vulnerable to selection on completed seeds."],"forward_implications":["A claim that a hybrid model's performance comes from its quantum layer now requires the same interface-matched control; a well-performing hybrid pipeline is evidence about the full pipeline, not the circuit.","The single positive contrast at $n_q = 4$ with attention ($+1.63$ points) is a replication target, not a stable effect: it is small relative to the ~96% ceiling, reverses sign at $n_q = 8$, and is one of four contrasts with five seeds.","Comparable point accuracies on AG News, BCW, and BirdCLEF in the cross-modality grid cannot be credited to the PQC, because the QEA-R variant includes a classical angle-residual bypass and the columns differ in readout width.","The CIFAR-10 failure (40.28% for QEA-R vs 84.11% for the plain classical model) is a pipeline-level observation, not an isolated circuit effect, but it argues against any general performance benefit from this quantum embedding.","The data do not establish equivalence: five paired seeds on one saturated dataset are too imprecise to conclude the PQC is interchangeable with a classical map."],"supporting_citations":[{"why":"The predecessor single-qubit, single-dataset study whose approximately three-point BirdCLEF improvement motivates the need for an interface-matched factorial.","marker":"Chen et al. [2024]"},{"why":"Provides the data re-uploading construction and the comparison of a single-qubit classifier with a one-hidden-layer network, used to argue that $n_q = 1$ accuracy is not evidence of a quantum contribution.","marker":"Pérez-Salinas et al. [2020]"},{"why":"Analyzes encoding-dependent variational models as truncated Fourier series, grounding the paper's claim that encoding repetitions change the accessible frequencies of the quantum map.","marker":"Schuld et al. [2021]"},{"why":"Identifies barren-plateau trainability risks, justifying the shallow ansatz depth used in the QEA architecture.","marker":"McClean et al. [2018]"},{"why":"Large-scale benchmark of quantum kernel models finding no universal advantage and emphasizing classical controls; motivates the component-attribution design even though QEA is not a kernel method.","marker":"Schnabel and Roth [2025]"},{"why":"Survey of quantum transformers that distinguishes PQC-based hybrids from fault-tolerant proposals, used to locate QEA and to warn that error-corrected execution would not automatically create an algorithmic advantage.","marker":"Zhang and Zhao [2025]"}],"fun_headline_variants":["Quantum-embedded attention shows no consistent gain","Controlled swap finds no edge for quantum layer","Quantum layer fails controlled test on classical data","No consistent PQC benefit in matched hybrid test","Quantum-embedded transformer: no advantage in controlled runs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison rests on the unverified assumption that the classical surrogate $\\tanh(W\\xi+b)$ is a fair stand-in for the PQC at the same input/output interface even though it has many more parameters (50 vs 8 at $n_q = 4$), and that the near-saturated BCW task with five paired seeds is sensitive enough to reveal a real quantum-classical difference if one exists.","fun_headline_variants_meta":{"raw":{"variants":["Quantum-embedded attention shows no consistent gain","Controlled swap finds no edge for quantum layer","Quantum layer fails controlled test on classical data","No consistent PQC benefit in matched hybrid test","Quantum-embedded transformer: no advantage in controlled runs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000249,"raw_usage":{"total_tokens":1612,"prompt_tokens":1066,"completion_tokens":546,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":682,"completion_tokens_details":{"reasoning_tokens":486}},"tokens_in":682,"tokens_out":546,"duration_ms":5981,"temperature":1.0,"reasoning_tokens":486,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:41:20.772573+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication run of the same switch-A factorial on a non-saturated task such as CIFAR-10, with a prespecified smallest effect of interest, a residual-free core QEA, and at least ten paired seeds, would settle the claim: if the PQC shows a positive paired contrast whose confidence interval excludes zero at both $n_q = 4$ and $n_q = 8$ and both decoders after correction, the 'not replicated' conclusion fails.","supporting_citations":[],"review_version":1}