{"id":"c7776009-2a32-4c13-8154-2aaeca0df6e8","arxiv_id":"2608.00139","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A transverse-field Ising quantum reservoir is more robust than a classical echo-state network on a simple oscillator benchmark, but less accurate on simulated EEG.","lead":"This paper tests a quantum reservoir computer—a quantum circuit used as a fixed dynamical memory—for forecasting simple oscillator and simulated EEG time series. It reports that the quantum reservoir beats a classical echo-state network on average on the simple benchmark and runs on IBM hardware, but is less accurate on the EEG-like task.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quantum 'overall' advantage on benchmark vanishes under best-vs-best comparison; claim rests on pooling classical hyperparameter configurations.","rationale":"The reader's weakest assumption correctly identifies the core problem: the quantum advantage claim is an artifact of comparing against a pooled classical baseline that includes many poorly performing configurations, rather than against the best-tuned classical model. The paper explicitly uses Pareto optimization to select configurations for hardware and EEG experiments, so the appropriate comparison for the benchmark should also be between selected configurations. Under that comparison, Tables I and II show the classical ESN is substantially better. The lack of a validation split compounds this issue, as the reported metrics for the selected parameters are computed on the same data used for selection. The paper does include an honest negative result on EEG and a hardware demonstration, so a conditional acceptance with required revisions is appropriate, not rejection. My read does not change the reader's verdict, hence UNCHANGED.","tokens_in":9738,"tokens_out":5474,"duration_ms":61399,"concrete_test":"Recompute the superimposed-oscillator comparison using only the Pareto rank-1 configuration (or the single best-NMSE configuration) from each model, evaluated on a hold-out validation set that was not used in the Pareto optimization. If the classical ESN's NMSE remains at ~0.01 versus the quantum reservoir's ~0.05, then the 'overall' claim in the Abstract and §III.A must be revised or qualified. Additionally, re-run the Pareto selection with a three-way train/validation/test split of the 25 repetitions to check whether the reported best parameters generalize out-of-sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central benchmark claim—that the quantum reservoir outperforms the classical ESN 'overall' (§III.A, Abstract)—depends on an aggregation choice that the paper itself does not justify. The comparison pools the classical reservoir's 40 hyperparameter configurations into a single median baseline, then reports that each of the 40 quantum configurations beats that pooled median. However, Tables I and II show that the best-tuned classical configuration achieves NMSE 0.0116 (rank 3) while the best-tuned quantum configuration achieves NMSE 0.0556 (rank 5). If instead one compares the configurations that the paper's own Pareto optimization would select, the classical ESN is clearly superior. The 'overall' framing therefore conflates 'typical random configuration' with 'performance after tuning,' which is not the relevant comparison for a method whose hyperparameters are explicitly optimized in Section K. Additionally, the Pareto selection in Section K uses the same repetitions on which the final metrics are reported, with no train/validation/test split described; this risks overfitting the reported numbers for the chosen configurations. The EEG negative result and hardware demonstration are informative, but the headline claim of quantum advantage on the benchmark is not robust to a best-vs-best comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces a quantum reservoir computing (QRC) approach based on a transverse-field Ising Hamiltonian with nearest-neighbor couplings, multi-basis Pauli measurements, quadratic ridge regression readout, and temporal multiplexing. The method is evaluated on two tasks: a superimposed-oscillator benchmark (using 4/5 qubits versus 250/375-node echo-state networks) and a simulated multi-frequency EEG signal with a parallel reservoir architecture. On the benchmark, the authors report that the quantum simulator has lower median NMSE, DTW, and 1-PLV than the classical ESN for all 40 quantum configurations, and they demonstrate execution of the same task on IBM Heron R2 hardware. On the EEG task, the classical reservoir outperforms the quantum simulator. The abstract concludes that 'the quantum reservoir outperforms a classical counterpart overall' on the standard benchmark.","tokens_in":10018,"tokens_out":3439,"duration_ms":39650,"significance":"If the headline benchmark claim were robust, this would be a useful contribution to the small-data neural time-series literature and would strengthen the case for QRC on near-term hardware. The paper has notable strengths: it uses matched preprocessing, training, and readout pipelines for both implementations; it reports a hardware demonstration with noise mitigation; and it honestly reports the negative EEG result. However, the 'outperforms overall' claim depends on an unconventional aggregation of the classical hyperparameter grid and on excluding a large fraction of classical repetitions, and the Pareto-based parameter selection may overfit the reported metrics. As it stands, the paper is best assessed as a feasibility study and baseline for QRC on neural forecasting, with the cross-model comparison requiring substantial re-analysis before the central claim can be accepted.","major_comments":[{"comment":"The claim that the quantum reservoir 'outperforms' the classical counterpart overall is not supported by a best-vs-best or tuning-selected comparison. The manuscript pools the classical reservoir's 40 hyperparameter configurations into a single median, excludes 44.7% of classical repetitions with NMSE≥1, and then shows that each of 40 quantum configurations beats that pooled median. However, Tables I and II show that the best classical configuration achieves NMSE 0.0116, DTW 0.9305, and 1-PLV 0.0266, while the best quantum configuration achieves NMSE 0.0556, DTW 1.3884, and 1-PLV 0.0446. The classical best is clearly better. The 'overall' framing therefore conflates 'typical untuned classical configuration' with 'performance after the Pareto optimization described in Section II.K'. Please report the comparison that the paper's own tuning procedure would select, or justify why the pooled","section":"Section III.A, Tables I and II"},{"comment":"The Pareto-frontier selection in Section II.K uses the same median DTW, NMSE, and 1-PLV values computed over the same repetitions that are later reported as the final results in Section III.B. No train/validation/test split or nested cross-validation is described. This makes the 'best parameter' results fitted rather than predictive. The reported numbers for the selected configurations are therefore optimistically biased, and the same issue affects both the quantum and classical results. A separate validation set (or a hold-out repetition set) must be used for parameter selection before reporting final metrics.","section":"Section II.K / Section III.B"},{"comment":"The robustness comparison is not informative as presented. The manuscript states that no quantum repetition failed to converge, whereas 44.7% of classical repetitions had NMSE≥1 and were excluded from the aggregate comparison. The subsequent 'quantum was more robust' statement is then based only on the variance among converging runs. This is not an apples-to-apples comparison of overall performance: a method with a 44.7% failure rate is very different from one with a 0% failure rate even if the medians of the successful runs are similar. Please report results on all repetitions, or combine failure rate and error into a single metric, so that 'overall' reflects both accuracy and reliability.","section":"Section III.A and Section J"}],"minor_comments":[{"comment":"The expression for c_k is typeset awkwardly; the division by 0.5 appears to be outside the sum and is unclear. Please rewrite to make the normalization explicit.","section":"Eq. (3)"},{"comment":"The evolution time dt is described as being in 'arbitrary units'. Please specify whether this is a dimensionless parameter or a physical time scale, since the comparison with ESN leak rate depends on this interpretation.","section":"Section II.D"},{"comment":"Typographical issues: 'MANOV A' should be 'MANOVA', 'ibm quebec' should be 'IBM Quebec', and 'P IN Q2' should be 'PINQ2'.","section":"Section III.B / Acknowledgments"},{"comment":"The abstract says the quantum reservoir 'outperforms a classical counterpart overall' on the benchmark, but the EEG section reports the opposite. The qualifier 'overall' is doing substantial work; please define it explicitly in the abstract or soften the wording to match the nuanced result.","section":"Abstract / Introduction"}],"recommendation":"major_revision","confidential_remarks":"The headline quantum-advantage claim is likely to be scrutinized and, as written, is not robust to a best-vs-best comparison. The paper would be stronger if the central narrative were repositioned around feasibility, reliability, and the EEG baseline, with the benchmark comparison reported transparently for all configurations and with a proper validation split. The current aggregation choices could be seen as overclaiming, but the underlying data and hardware demonstration are valuable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline claim that the quantum reservoir outperforms the classical one on the benchmark is not as robust as it looks. Tables I and II show the best classical configuration (NMSE 0.0116) is clearly better than the best quantum one (0.0556), so the 'overall' advantage comes from pooling the classical reservoir's 40 hyperparameter configurations into a single median, after excluding 44.7% of classical repetitions as non-converging. That is a legitimate robustness statement — quantum is more consistent across untuned configurations — but it is not the same as saying QRC is more accurate. The abstract and Section III.A overstate the case.\n\nWhat is genuinely new and useful: the specific combination of a transverse-field Ising reservoir, multi-basis readout, quadratic ridge regression, and a parallel reservoir architecture applied to simulated EEG, plus a demonstration on IBM hardware. The EEG result is an honest negative: the classical parallel ESN wins, and the quantum implementation produces stable but less accurate forecasts. That kind of null result is valuable for the QRC community.\n\nThe main soft spot beyond the aggregation issue is the hyperparameter selection. Section II.K picks Pareto-optimal configurations using the same DTW/NMSE/1-PLV metrics and the same repetitions that are then reported as results in Section III. There is no train/validation/test split described, so the numbers in the tables are fitted, not predicted. This is not fatal for the comparative robustness claim, but it means the quantitative results are optimistic for both implementations.\n\nA minor point: the hardware experiment uses a different protocol (sliding window, no intermediate readouts, 5,000 shots) than the simulator, so the NMSE of 0.0715 on hardware is not directly comparable to the simulator numbers. That is fine for a feasibility demo, but it should be labeled as such.\n\nWho is this for: researchers working on quantum reservoir computing, especially those interested in neural time-series or small-data regimes. It deserves a serious referee, but the benchmark claim needs to be reframed and the validation protocol clarified. I would recommend peer review with major revision.","headline":"The benchmark 'overall advantage' claim does not survive best-vs-best comparison; the paper is otherwise a careful, honest empirical study with a useful negative EEG result and a real hardware demonstration.","tokens_in":10515,"tokens_out":2410,"would_cite":true,"duration_ms":28963,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68","68T07","62M10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A transverse-field Ising quantum reservoir outperforms a classical echo-state network overall on a standard forecasting benchmark, but not on simulated EEG.","keywords":["quantum reservoir computing","neural time-series forecasting","transverse-field Ising model","echo state network","electroencephalography","near-term quantum hardware","small-data forecasting","polynomial ridge regression"],"falsifier":"Recompute the benchmark comparison without discarding non-converging classical runs, or compare the single best-tuned classical configuration (NMSE 0.0116) against the single best-tuned quantum configuration (NMSE 0.0556) from the paper's own tables; if the classical then wins on the majority of configurations, the claimed quantum advantage fails. Alternatively, run the same EEG task with the full signal fed to one reservoir instead of three independent sub-reservoirs and check whether quantum accuracy reaches classical levels.","tokens_in":9604,"feed_emoji":"🧠","tokens_out":5646,"duration_ms":54629,"temperature":0.7,"pith_summary":"This paper asks whether quantum reservoir computing can forecast neural time series from short recordings. On a standard superimposed-oscillator benchmark, a 4- or 5-qubit transverse-field Ising reservoir with multi-basis readout and quadratic ridge regression achieved lower median error than a classical echo-state network across all 40 tested configurations, with no non-converging runs. The same reservoir also ran on superconducting quantum hardware with modest error. On a more realistic simulated EEG signal, however, the classical reservoir was clearly more accurate, and the authors attribute this to limited per-sub-reservoir capacity and loss of cross-frequency structure. The paper concludes that near-term quantum reservoirs can converge on EEG-like data but do not yet surpass classical methods on complex neural signals.","feed_headline":"Quantum reservoir beats classical net on oscillator forecast","feed_subtitle":"On simulated EEG, though, the classical echo-state network still wins; the quantum edge is reliability, not accuracy.","key_machinery":"The central object is a transverse-field Ising Hamiltonian that acts as the reservoir: qubits coupled in a ring with nearest-neighbor X-X interactions of strength J_{ij} = (i+j)^k / c_k, plus a uniform Z-field. Inputs are encoded in one qubit's amplitude; the system evolves for a time dt per input under one Trotter step; and the feature vector contains the expectation values <X>, <Y>, <Z> on every qubit, augmented by all pairwise products. The coupling exponent k and evolution time dt are the tuned parameters that shape information mixing and memory. Temporal multiplexing divides dt into virtual nodes to expand the feature dimension without adding qubits, and a rewind protocol re-encodes the","core_discovery":"The paper's central claim is that a quantum reservoir built from a transverse-field Ising model, with X, Y, Z readout and polynomial ridge regression, can outperform a classical echo-state network on a small-data forecasting benchmark, and that the advantage comes largely from robustness and convergence rather than from superior accuracy in every regime. On all 40 quantum configurations, the median NMSE, DTW, and 1-PLV were lower than the pooled classical baseline, with median improvements of 35.8%, 22.3%, and 36.0%. The quantum reservoir converged in all repetitions, whereas 44.7% of classical repetitions had to be discarded as non-converging. On simulated EEG with a parallel three-reservoi","pith_inferences":["The claimed 'overall' quantum advantage rests on pooling all classical hyperparameter configurations into one median after discarding 44.7% of classical repetitions that failed to converge; a best-vs-best comparison from the paper's own tables favours the classical ESN (NMSE 0.0116 vs 0.0556).","A fairer comparison would keep all classical repetitions in the distribution; the classical non-convergence may reflect hyperparameter sensitivity rather than a capacity limit.","Because the EEG test used 12 qubits in three independent 4-qubit reservoirs with only linear readout, a single larger reservoir with richer readout might behave differently on EEG-like data.","The quantum reservoir's consistent convergence could matter in clinical settings that prioritize stability over peak accuracy, but that would need validation on recorded, not simulated, EEG."],"forward_implications":["If quantum reservoirs are at least as accurate and more reliable than classical ESNs on simple oscillatory signals, short-recording forecasting tasks with many repeated trials may benefit, especially when reliability matters.","Prediction accuracy depends strongly on reservoir parameters k and dt, so parameter selection is a key determinant of practical success.","A parallel architecture of independent sub-reservoirs reduced accuracy on multi-frequency EEG because cross-frequency structure could only be recovered at the linear readout.","The same forecasting pipeline runs on current superconducting hardware, with errors only modestly worse than in simulation.","On current evidence, clinical EEG forecasting would still favour classical reservoirs; quantum advantage on complex neural signals remains undemonstrated."],"fun_headline_variants":["Quantum reservoir tops classical on small-data forecast","Quantum reservoir converges where classical fails","Quantum reservoir wins small data, stable on EEG"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claimed advantage depends on how the classical comparison is aggregated: the paper pools 40 classical hyperparameter configurations into a single median after discarding 44.7% of classical repetitions that failed to converge, rather than comparing the best classical configuration against the best quantum one.","fun_headline_variants_meta":{"raw":{"variants":["Quantum reservoir tops classical on small-data forecast","Quantum reservoir converges where classical fails","Quantum reservoir wins small data, stable on EEG"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000832,"raw_usage":{"total_tokens":3479,"prompt_tokens":761,"completion_tokens":2718,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":2675}},"tokens_in":505,"tokens_out":2718,"duration_ms":20819,"temperature":1.0,"reasoning_tokens":2675,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T01:10:56.549208+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the benchmark comparison without discarding non-converging classical runs, or compare the single best-tuned classical configuration (NMSE 0.0116) against the single best-tuned quantum configuration (NMSE 0.0556) from the paper's own tables; if the classical then wins on the majority of configurations, the claimed quantum advantage fails. Alternatively, run the same EEG task with the full signal fed to one reservoir instead of three independent sub-reservoirs and check whether quantum accuracy reaches classical levels.","supporting_citations":[],"review_version":1}