{"id":"25f1a88a-9eec-4a22-a730-a2dc29d2a659","arxiv_id":"2607.17705","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Ten-class MNIST is classified on a real 127-qubit IBM Eagle processor using 12 qubits, with four-way circuit packing giving a ~3.8-4x inference speedup at no mean accuracy cost.","lead":"Researchers ran a 12-qubit quantum circuit on IBM's 127-qubit Eagle processor to classify ten handwritten digits, reaching about 74-79% accuracy. They show packing four circuits per run gives near-4x inference speedup at no accuracy cost, and that fine-tuning on hardware adds nothing, so the practical recipe is simulator training plus hardware inference.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"On-hardware fine-tuning null is underpowered: two COBYLA iterations cannot support the inference-only workflow claim.","rationale":"The reader's weakest assumption identifies the hardware fine-tuning null as the load-bearing point; I agree. The feasibility claim that a 12-qubit classifier runs on hardware and the QMP noiseless equivalence are well supported by the factorization argument and the shared-weight design. However, the workflow recommendation — train on simulator, use hardware only for inference — is exactly what the underpowered null is asked to justify, and the paper's own noise threshold makes the +3.11 pt contrast indistinguishable from a real effect. The test-set leakage in the five-ansatz pilot (selection on the same 75-sample test set used in Table 3) is also a genuine concern, but it primarily affects the absolute accuracy numbers rather than the workflow recommendation; the fine-tuning null is the single most load-bearing condition for the paper's central practical claim. A conditional verdict remains appropriate: the paper is an honest and useful feasibility study, but the headline workflow recommendation should be re-measured before being called unambiguous.","tokens_in":17356,"tokens_out":5734,"duration_ms":67433,"concrete_test":"Repeat the Hybrid Phase 2 comparison with all 20 COBYLA epochs executed on ibm_yonsei (instead of only epochs 19–20), using, e.g., 10 seeds and a larger test set (Ntest=300) while keeping all other cells fixed. If the mean Cell 2→4 contrast becomes ≥3 pts with a confidence interval excluding zero, the inference-only workflow is not supported; if the contrast remains within sampling noise with a tight interval, the paper's null is credible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central workflow recommendation — train on a simulator and reserve quantum hardware for inference only (Sec. 4.4) — depends entirely on the null result that on-hardware fine-tuning yields no accuracy gain. That null is underpowered by the paper's own criteria. The test set has Ntest=75, so one sample is 1.33 pts and the paper treats differences below roughly two samples (~2.7 pts) as within sampling noise (Sec. 4.1). The key contrast Cell 2→4 is +3.11 pts (~2 samples), which is therefore statistically indistinguishable from a real fine-tuning benefit of the same size. The hardware phase is restricted to B=20 COBYLA evaluations, i.e. one iteration per epoch for only epochs 19–20 (Sec. 2.3.2), so the optimizer has almost no chance to move the loss; the flat Phase 2 curves in Fig. 3 are equally consistent with an optimizer that was barely run. Calling the conclusion 'unambiguous' (Sec. 4.4) overstates what a three-seed, two-iteration null can support. If a true benefit of ~3 pts exists, the recommended inference-only workflow would be suboptimal. This is the load-bearing condition for the practical workflow claim, and the paper's own noise budget shows the evidence cannot distinguish the null from a modest real effect.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a hybrid classical-quantum classifier for ten-class MNIST executed end-to-end on a 127-qubit IBM Eagle processor. The main technical contributions are a two-phase training protocol (Adam on classical encoder/readout layers with frozen quantum parameters, followed by COBYLA optimization of quantum parameters), the first application of Quantum Multi-Programming (QMP) to a trained quantum classifier at K=4 packing, and a five-cell controlled comparison designed to isolate hardware noise, QMP packing effects, and on-hardware fine-tuning. The headline empirical claims are: a 12-qubit VQC reaches about 74–79% accuracy on hardware with a modest, reproducible noise penalty; QMP packing leaves mean accuracy statistically unchanged while amplifying cross-seed variance and reducing job submissions roughly fourfold; and two epochs of on-hardware COBYLA fine-tuning produce no measurable accuracy gain, motivating an inference-only deployment strategy. The paper also benchmarks against a matched-capacity classical MLP and finds no per-parameter quantum advantage, framing the contribution as a feasibility-and-workflow demonstration.","tokens_in":17603,"tokens_out":7578,"duration_ms":86266,"significance":"If the results hold, this is a useful feasibility-and-workflow result: it demonstrates a non-trivial multi-class quantum image classifier on current superconducting hardware, provides a clean shared-weight five-cell design that isolates hardware noise and packing effects, and gives a simple noiseless proof that QMP is equivalent to serial execution when no inter-circuit gates are applied. The matched-capacity classical baseline and the honest framing as a feasibility study are strengths. The central limitations are that the model-selection pilot and the final evaluation share the same 75-sample test set, and that the null result for on-hardware fine-tuning is probed by only two COBYLA hardware iterations per seed. These issues currently prevent the absolute accuracy numbers and the strong workflow recommendation from being taken at face value.","major_comments":[{"comment":"The model-selection pilot in §3.2 reports test accuracy on the Ntest=75 subset and selects VQC ZigZag as the highest-accuracy model. The five-cell evaluation in Table 3 then reports test accuracy on the same Ntest=75 subset. Because the test set has already been used to choose the architecture, the absolute accuracies in Table 3 and the comparison to the classical MLP in §4.5 are not unbiased estimates of generalization. This does not destroy the within-group cell contrasts, which compare bit-identical weights, but it is load-bearing for the feasibility claim that a 12-qubit classifier achieves 74–79% on ten-class MNIST. Please hold out a separate test set for final evaluation, or report the pilot selection metric on the validation set and reserve the test set for the final five-cell results.","section":"§3.2, Table 1; §3.1, Table 3"},{"comment":"The claim that on-hardware fine-tuning yields no accuracy gain, described as 'unambiguous' in §4.4, is supported by only two COBYLA iterations per seed on the QPU (B=20, one iteration per epoch, epochs 19–20). The key contrast Cell 2→4 is +3.11 pts, which the paper itself treats as approximately two test samples and within noise. With three seeds and no formal power analysis, this cannot rule out a modest real benefit of the same size. The flat Phase 2 curves in Fig. 3 are equally consistent with an optimizer that was barely run. Please either run more hardware iterations with a proper uncertainty quantification, or reframe the conclusion as 'no detectable benefit within the tested budget' and temper the inference-only workflow recommendation accordingly.","section":"§4.4, §2.3.2"},{"comment":"The paper's 'within sampling noise' criterion is based on the 75-sample test-set granularity (~1.33% per sample), but the contrasts in Table 4 are differences of three-seed means. The standard error of a three-seed mean difference is typically about 1–2 pts for these cells; for Cell 2→4, the difference of 3.11 pts is roughly 3 standard errors of the difference if the reported cross-seed standard deviations are used. Calling this 'within noise' conflates per-sample test-set resolution with seed-to-seed variability. A formal paired or unpaired test, or a bootstrap over seeds, should be reported before concluding that the hardware-training effect is zero.","section":"Table 4, §4.1"}],"minor_comments":[{"comment":"The caption contains a duplicated sentence about post-transpilation metrics on the IBM basis; please remove the repetition.","section":"Table 1 caption"},{"comment":"The text states that COBYLA 'moves the loss without changing a single classification decision,' but Fig. 3 shows validation accuracy, not loss. Please clarify whether the loss curves are shown or state that the claim is inferred from accuracy trajectories.","section":"Fig. 3"},{"comment":"The noiseless QMP verification is reported for K=1 and K=2 only. The factorization argument in Eq. (3) is general, but stating why K=4 was not simulated (statevector size) would help avoid the impression that the verification is incomplete.","section":"§2.4.1"},{"comment":"The pilot reports VQC ZigZag test accuracy as 78.7±0.0% across three seeds. Perfect zero spread is surprising given that seeds change the data subset and optimizer initialization; please report the individual seed values or explain the source of determinism.","section":"§3.2, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid engineering/feasibility study, and the QMP equivalence and shared-weight five-cell design are strong points. The two major issues are fixable in scope: separate model selection from final test evaluation, and either add more hardware fine-tuning iterations or substantially soften the 'unambiguous' no-benefit claim. There is also a notable pattern of self-citation to the corresponding author's QMP-related papers, but this does not affect the technical content and is not grounds for concern by itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what to know: this paper actually runs ten-class MNIST on a 127-qubit Eagle, uses QMP to pack four identical 12-qubit circuits into one job, and does it with a controlled five-cell design that holds bit-identical weights across most contrasts. That is genuinely new in the hardware QML literature, which has mostly stopped at binary or low-accuracy multi-class. The central feasibility claim — a 12-qubit classifier running at 74–79% on real hardware — is supported. The QMP equivalence is derived from state factorization and verified to 1e-6 on a noiseless simulator; the matched-capacity classical baseline is an honest calibration; the negative results are framed as feasibility, not advantage. Credit where due: the experimental design, the seeded reproducibility, and the frank discussion of the 75-sample test set are all above average for this area.\n\nThe soft spots, in proportion:\n\n1. The ansatz pilot selected VQC ZigZag using the same 75-sample test set that later appears in the headline table. That introduces a small optimistic bias in the architecture choice. It doesn't break the feasibility claim, but the reported accuracies are, strictly, post-selection. The authors should say this more plainly or re-measure on a fresh split.\n\n2. The stress-test note is right about the fine-tuning null. The \"train on simulator, deploy for inference only\" workflow is the paper's most practically interesting recommendation, but it is built on two COBYLA hardware iterations per seed (epochs 19–20, B=20). The key contrast, Cell 2→4, is +3.11 pts, which is about 2 samples on a 75-sample test set — inside the paper's own noise threshold. Calling the conclusion \"unambiguous\" overstates what a three-seed, two-iteration null can support. A real fine-tuning benefit of similar size would not be detected. The recommendation should be softened or re-measured with more hardware iterations.\n\n3. Minor: the wall-time speedup (3.76x vs 4x ceiling) comes from single runs without repeated-job error bars, and device load was not held fixed. The effect is large, so this is a detail, but worth a caveat.\n\nOverall the paper is a serious, readable feasibility study. The main advertised result — that multi-class classification can run end-to-end on current hardware with QMP packing — holds. The workflow recommendation is appropriately conditional after the stress test, but that's a revision, not a rejection.\n\nWho benefits: anyone working on QML hardware deployment, NISQ workflows, or QMP. I'd bring it to a reading group and I'd cite it as a state-of-the-art feasibility demonstration. It deserves a serious referee and, with revisions, publication.","headline":"A solid, honestly-scoped feasibility study: the ten-class MNIST deployment on 127-qubit hardware holds up, but the 'no benefit to hardware fine-tuning' claim rests on a two-iteration null and should be softened.","tokens_in":18189,"tokens_out":1777,"would_cite":true,"duration_ms":21812,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 12-qubit quantum classifier can run ten-class MNIST end-to-end on a current 127-qubit superconducting processor, with parallel circuit packing delivering a near-fourfold inference speedup at no mean accuracy cost.","keywords":["quantum machine learning","image classification","quantum multi-programming","NISQ","variational quantum classifier","MNIST","two-phase training","IBM Eagle"],"falsifier":"Run the hybrid Phase 2 with all 20 COBYLA epochs on the quantum processor across the same three seeds and compare K=1 test accuracy against the simulator-only checkpoint: if the mean improves by more than approximately two test samples (about 2.7 percentage points), the paper's conclusion that on-hardware fine-tuning yields no measurable benefit is refuted.","tokens_in":17169,"feed_emoji":"⚛️","tokens_out":5465,"duration_ms":50663,"temperature":0.7,"pith_summary":"This paper claims that ten-class MNIST image classification can be run end-to-end on a current 127-qubit superconducting quantum processor, and that a practical workflow for noisy intermediate-scale quantum devices is to train the quantum model on a classical simulator and reserve the quantum hardware for inference only. The authors show this by combining a two-phase training protocol that avoids the costly parameter-shift gradients of on-hardware training, a 12-qubit circuit selected from a five-ansatz comparison for its accuracy–compilation-cost trade-off, and Quantum Multi-Programming, which packs four copies of the trained circuit onto one device so that four samples are inferred per job submission. They report that QMP keeps mean test accuracy statistically unchanged while cutting job submissions fourfold and delivering a ~3.8–4x end-to-end speedup, and that two epochs of on-hardware fine-tuning do not measurably improve accuracy. The paper is careful to note that the quantum module shows no per-parameter accuracy advantage over a matched classical network at this scale, so it frames the result as a feasibility-and-workflow demonstration rather than a claim of quantum advantage.","feed_headline":"Ten-class MNIST runs end-to-end on a 127-qubit IBM processor","feed_subtitle":"Four-way circuit packing gives ~4x inference speedup with no mean accuracy loss — and simulator training suffices.","key_machinery":"The load-bearing mechanism is the combination of a two-phase training protocol and Quantum Multi-Programming (QMP). Phase 1 trains only the classical encoder and readout with Adam on a noiseless simulator while the quantum parameters stay frozen at random initialization; Phase 2 optimizes only the 15 quantum parameters with the gradient-free COBYLA algorithm, either on a simulator or on hardware, avoiding the 2N circuit executions per step that parameter-shift gradients would require on a real device. QMP packs K identical logical circuits onto disjoint, routing-disconnected regions of the 127-qubit heavy-hex lattice; because no entangling gate crosses circuit boundaries, the joint state fac","core_discovery":"On its own terms, the paper's central discovery is that a nontrivial 12-qubit parameterized quantum circuit can serve as the feature map of a ten-class image classifier executed on real IBM Eagle hardware, with a modest and stable accuracy cost: the same trained weights that score 79.11% on a noiseless simulator score 73.78% on the device, a hardware-noise penalty of 5.33 percentage points that reproduces across seeds. The second discovery is that Quantum Multi-Programming — packing K=4 copies of the 12-qubit circuit onto 48 physically separated qubits — is mathematically equivalent to serial execution in the noiseless limit and on hardware leaves the mean test accuracy unchanged within samp","pith_inferences":["If the variance amplification seen at K=4 is dominated by per-day calibration drift across the four circuits, then averaging over multiple circuit assignments (or over more seeds) should recover the K=1 mean with a smaller standard error, making QMP effectively free at higher packing factors — a test the paper did not run.","The result that two hardware COBYLA epochs did not move accuracy suggests an even stronger conjecture: for these shallow circuits, the Phase-2 loss landscape is so flat that the classical encoder/readout alone determines the decision boundary. A direct test would be to train only the classical layers (no Phase 2 at all) and compare accuracy; the paper's flat Phase 2 curves already hint this would ","The workflow of training on a noiseless simulator and deploying for inference on hardware transfers naturally to other NISQ-era learning tasks — kernel-based quantum models, quantum generative models, or hybrid models where the quantum circuit is a fixed feature map — as long as the circuit is small enough to simulate faithfully during training.","A testable extension: push QMP to K=8 or higher on the same device. The paper's factorization argument predicts the mean should stay within sampling noise; if instead a mean degradation appears, it would mark the onset of noise-driven cross-circuit interference that the current K=4 study cannot detect."],"forward_implications":["If the workflow is correct, multi-class quantum image classification on current hardware is feasible without on-hardware training, removing the parameter-shift cost that has confined most hardware QML to binary tasks.","QMP gives a structural throughput gain for inference: with per-job overhead dominating wall time on cloud-accessed processors, packing K circuits reduces job submissions K-fold and delivers a near-K-fold speedup at no mean accuracy cost.","The null result for on-hardware fine-tuning implies that for small quantum parameter counts, the trained quantum circuit can be treated as a fixed nonlinear feature map, so simulator training plus hardware inference is the appropriate deployment pattern on NISQ devices.","The five-ansatz comparison with post-transpilation depth and CNOT counts provides a template for choosing classifier circuits for hardware deployment based on accuracy–cost trade-offs rather than accuracy alone.","Because the paper finds no per-parameter accuracy advantage over a matched classical MLP at this scale, the contribution is a controlled experimental template for isolating hardware noise, training location, and packing effects — a benchmark others can reuse."],"fun_headline_variants":["10-class MNIST on real 127-qubit IBM Eagle","Quantum Multi-Programming: 4x parallel inference, no accuracy loss","Simulator-trained quantum classifier runs on IBM hardware","12-qubit feature map handles 10-class image recognition","End-to-end ten-class classification on current IBM quantum processors"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the absence of a measurable accuracy gain from two epochs of on-hardware COBYLA fine-tuning (tested on a 75-sample test set, three seeds, where the paper's own noise threshold is about two samples) justifies the recommendation to train on a simulator and use hardware only for inference; if more hardware epochs moved accuracy, that recommendation would collapse.","fun_headline_variants_meta":{"raw":{"variants":["10-class MNIST on real 127-qubit IBM Eagle","Quantum Multi-Programming: 4x parallel inference, no accuracy loss","Simulator-trained quantum classifier runs on IBM hardware","12-qubit feature map handles 10-class image recognition","End-to-end ten-class classification on current IBM quantum processors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000501,"raw_usage":{"total_tokens":2292,"prompt_tokens":755,"completion_tokens":1537,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":1468}},"tokens_in":499,"tokens_out":1537,"duration_ms":15148,"temperature":1.0,"reasoning_tokens":1468,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T17:12:59.470081+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the hybrid Phase 2 with all 20 COBYLA epochs on the quantum processor across the same three seeds and compare K=1 test accuracy against the simulator-only checkpoint: if the mean improves by more than approximately two test samples (about 2.7 percentage points), the paper's conclusion that on-hardware fine-tuning yields no measurable benefit is refuted.","supporting_citations":[],"review_version":1}