{"id":"0640d529-1348-4298-8de0-c6d7c2dc4a43","arxiv_id":"2607.09445","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Scalable ORS expressivity (top-K output probabilities vs Haar) plus effective feature rank jointly diagnose quantum-reservoir quality independent of Hilbert dimension and under hardware noise.","lead":"A new pair of diagnostics scores fixed quantum reservoirs for machine learning without exponential cost or circuit training. The order-statistics expressivity score plus feature-matrix rank predict which random quantum systems will actually learn, including under noise and on IBM hardware.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The central empirical claim holds under the paper's own evidence: multi-basis ORS ranks reservoir families consistently with established complexity measures, while R_eff tracks when that expressivity reaches the linear readout. The global-depolarizing model is an acknowledged approximation confined to the hardware section and does not underwrite the synthetic or real-data performance results. Because the concern does not threaten the strongest claim, no verdict adjustment is warranted.","tokens_in":15411,"tokens_out":385,"duration_ms":4053,"concrete_test":"Recompute multi-basis G_ORS and R_eff for the G3 and D2,XZY families on the Fourier and NARMA tasks at n=6 with exact (noiseless) probabilities only; confirm that the expressivity hierarchy and the ORS–R_eff–MSE alignment remain unchanged when the hardware-aware correction is entirely omitted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest-assumption note on the global-depolarizing hardware correction is real but not load-bearing for the central claim. The paper's strongest claim is the joint predictive power of multi-basis ORS (task-independent expressivity hierarchy) and R_eff (task-dependent coverage) across QELM/QRC benchmarks; that claim is supported by noiseless synthetic results (Figs. 2–3), real-data appendices, and small-system cross-checks against KL, level-spacing, and Krylov diagnostics (Fig. 1). The depolarizing correction (Eqs. 6–11) is used only for the secondary hardware-compatibility demonstration (Table II); the authors already flag that real noise is not exactly global depolarizing. No hidden circularity, derivation gap, or unsupported leap in the main argument is present.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes a two-axis diagnostic for quantum reservoirs (QRC and QELM). The first axis is a task-independent order-statistics (ORS) expressivity score that compares only the top-K output probabilities of a reservoir ensemble to an analytical Haar order-statistics baseline (Eqs. 3–5), with a multi-basis extension (Eq. 15) and a closed-form global-depolarizing correction (Eqs. 6–11). The second axis is the task-dependent effective rank R_eff of the feature matrix (Eq. 16). Small-system validation against KL fidelity divergence, level-spacing ratio, and Krylov complexity (Fig. 1) shows a consistent transition to Haar-like behavior. Noise-corrected ORS remains discriminative under simulated depolarizing noise (Table I) and on IBM Aachen hardware (Table II). Across synthetic Fourier/NARMA and real LiH/EMSIG benchmarks (Figs. 2–5), multi-basis ORS ranks reservoir families by intrinsic expressivity while R_eff indicates when that expressivity becomes usable for the linear readout.","tokens_in":15616,"tokens_out":938,"duration_ms":8436,"significance":"If the results hold, the work supplies a practical, scalable alternative to full-distribution or full-unitary diagnostics that become unusable as Hilbert-space dimension grows. The analytical Haar baseline, cost independence of D, multi-basis extension that exposes commuting IQP structure, and closed-form depolarizing correction are concrete technical contributions that make the diagnostic usable on near-term hardware. Explicit cross-checks against KL, spectral chaos, and Krylov complexity, plus both synthetic and real QELM/QRC tasks, strengthen the claim that the two axes jointly explain performance. The framework is architecture-agnostic and therefore useful for comparing gate-based and analog reservoirs.","major_comments":[{"comment":"Sec. IIC and Table II: the hardware demonstration relies on an effective fidelity f_Bq estimated from calibration data under a global-depolarizing model (Eqs. 12–13). The paper correctly notes that real device noise is not exactly global depolarizing. Because Table II is presented as evidence that ORS remains informative on hardware, a short quantitative check of residual sensitivity (e.g., comparison of corrected vs uncorrected gaps, or a simple coherent-error simulation) would make the hardware claim more robust; without it the hardware result remains supportive but secondary.","section":null},{"comment":"Sec. III C and Figs. 2–3: for the non-commuting D2 extensions the multi-basis ORS approaches the Haar reference while R_eff and MSE remain suboptimal. The text attributes this to residual correlations not captured by top-K ranks and to memory effects in QRC. A brief ablation (larger K, more bases, or a simple memory-capacity diagnostic) would clarify whether the observed decoupling is fundamental or an artifact of the chosen K and B; the central claim that the two axes are complementary is otherwise well supported.","section":null}],"minor_comments":[{"comment":"Appendix A: the large-D asymptotic form of Pk(x) is used throughout; a short statement of the n range where the approximation remains accurate (or a finite-D correction) would help readers applying ORS at small n.","section":null},{"comment":"Fig. 1: the multiple right-hand axes for different K make visual comparison slightly crowded; a single normalized gap or an inset could improve readability.","section":null},{"comment":"Sec. IIF: the mean-degree parameter d of the Erdős–Rényi interaction graph is introduced without an explicit formula for the expected number of ZZ terms; a one-line clarification would aid reproducibility.","section":null},{"comment":"Notation: GORS, G(f)_ORS and Gmb(B) are used interchangeably in places; a consistent symbol table or early definition would reduce minor ambiguity.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a solid methods contribution for near-term quantum ML. The global-depolarizing hardware correction is the weakest link but is already flagged by the authors and is not load-bearing for the main expressivity–coverage claim. Fit for a specialized quant-ph or quantum-ML venue is good; no novelty or citation concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a practical paper that removes a real bottleneck. Existing QR diagnostics (full KL to Haar fidelity, level statistics, Krylov) blow up with system size; they adapt Micklitz’s order-statistics idea into a reservoir-ensemble score that only needs the top-K bitstring probabilities, stays independent of Hilbert-space dimension, and comes with a closed-form depolarizing correction. Pairing that with the ordinary effective rank of the feature matrix is the useful move: ORS ranks families by intrinsic expressivity, R_eff tells you when that expressivity actually reaches the linear readout.\n\nWhat they do well is concrete. At n=6 they show ORS tracks the same transition as KL, consecutive-spacing ratio, and Krylov (Fig. 1). The multi-basis extension cleanly exposes the basis dependence of commuting IQP-style reservoirs that single-basis scores miss. Noise-corrected gaps stay discriminative under simulated depolarizing noise up to n=25 and on IBM Aachen hardware (Table II). Across synthetic Fourier/NARMA and real LiH/EMSIG tasks the hierarchy is consistent: G3 and the non-commuting D2 extensions look better on both axes and predict better; restricted Clifford-like and pure D2 families do not. No circularity—ORS is task-independent, R_eff is just the participation ratio of the observed features, performance is held-out MSE/R^{2}.\n\nSoft spots are minor and mostly flagged by the authors. The hardware correction assumes an effective global depolarizing channel; real device noise is messier, so the IBM numbers are supportive rather than definitive. Free parameters (K, number of bases B, observable pool size) are chosen reasonably but not exhaustively swept. No code or data release, so re-implementation cost is non-zero. None of that undercuts the central empirical claim.\n\nThis is for people who actually pick or design quantum reservoirs on near-term hardware. Math and citations look solid; the adaptation from Micklitz is properly credited. I would send it to referees without hesitation and would cite the multi-basis ORS + R_eff pair when I next need a scalable QR diagnostic.","headline":"Solid, usable two-axis diagnostic for quantum reservoirs: multi-basis ORS for expressivity plus R_eff for coverage, with hardware-compatible noise correction that actually works on IBM data.","tokens_in":16273,"tokens_out":550,"would_cite":true,"duration_ms":6184,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Two scalable scores tell which fixed quantum reservoirs will actually learn: one measures how Haar-like their outputs are, the other how many usable feature directions reach the classical readout.","keywords":["quantum reservoir computing","quantum extreme learning machines","expressivity","order statistics","effective rank","Haar measure","depolarizing noise","quantum machine learning"],"falsifier":"Run the same G3 versus G1 or commuting versus non-commuting IQP families on a device whose noise is known to be highly structured (or under a simulated non-depolarizing channel) and check whether the noise-corrected multi-basis ORS gap still correctly ranks the families while the effective-rank / test-error relationship collapses.","tokens_in":16283,"feed_emoji":"⚛️","tokens_out":705,"duration_ms":11473,"temperature":0.7,"pith_summary":"Quantum reservoirs process data with fixed random quantum dynamics and a cheap classical linear readout, so performance hinges on which random family is chosen. Existing quality checks either explode with system size or only work for special models. This paper gives a two-axis diagnosis that stays practical as the number of qubits grows and works on real hardware. The first axis is an order-statistics (ORS) expressivity score: keep only the few largest output probabilities of each reservoir instance and compare them to the known analytic distribution for Haar-random states; a closed-form correction removes the trivial effect of depolarizing noise. The second axis is the effective rank of the feature matrix actually seen by the readout, which counts how many independent, input-dependent directions are available for learning. On both synthetic and real extreme-learning and reservoir-computing tasks, multi-basis ORS ranks reservoir families by intrinsic expressivity while the effective rank shows when that expressivity becomes usable predictive power. The same noise-corrected ORS gap continues to separate expressive from restricted families under simulated noise and on IBM hardware.","feed_headline":"Two scores pick which quantum reservoirs will learn","feed_subtitle":"Order statistics measure Haar-like expressivity; effective rank shows when it reaches the readout","key_machinery":"The order-statistics (ORS) expressivity gap: the harmonic-weighted log-likelihood of the K largest output probabilities of a reservoir ensemble, measured against the closed-form Haar order-statistics density (with optional multi-basis average and depolarizing correction). It is paired with the participation-ratio effective rank of the column-centred feature matrix seen by the linear readout.","core_discovery":"A reservoir family is useful when two complementary diagnostics align: its multi-basis order-statistics gap to Haar is near zero (intrinsic expressivity) and the effective rank of the measured feature matrix is large (task-dependent coverage). ORS never needs the full output distribution, is independent of Hilbert-space dimension for fixed top-K ranks, and admits an exact depolarizing correction that remains informative on real devices.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Two scalable scores diagnose quantum reservoirs by expressivity and coverage","ORS expressivity and Reff coverage select viable quantum reservoirs","Haar-gap order stats plus feature rank flag usable quantum reservoirs","Task-free ORS score plus effective rank diagnose reservoirs at scale","Quantum reservoir families ranked by order-stats expressivity and rank"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The hardware correction treats device noise as a single global depolarizing channel whose fidelity can be estimated from gate and readout calibration data; if real noise is strongly coherent, correlated or non-Markovian, the corrected gap can mis-rank families.","fun_headline_variants_meta":{"raw":{"variants":["Two scalable scores diagnose quantum reservoirs by expressivity and coverage","ORS expressivity and Reff coverage select viable quantum reservoirs","Haar-gap order stats plus feature rank flag usable quantum reservoirs","Task-free ORS score plus effective rank diagnose reservoirs at scale","Quantum reservoir families ranked by order-stats expressivity and rank"]},"model":"grok-4.5","effort":"low","cost_usd":0.005774,"raw_usage":{"total_tokens":1537,"prompt_tokens":766,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":57740000,"prompt_tokens_details":{"text_tokens":766,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":703,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":766,"tokens_out":68,"duration_ms":6248,"temperature":1.0,"reasoning_tokens":703,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T02:56:47.512469+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same G3 versus G1 or commuting versus non-commuting IQP families on a device whose noise is known to be highly structured (or under a simulated non-depolarizing channel) and check whether the noise-corrected multi-basis ORS gap still correctly ranks the families while the effective-rank / test-error relationship collapses.","supporting_citations":[],"review_version":1}