{"id":"60671f33-40d5-4855-8d1b-26591908b019","arxiv_id":"2507.14116","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A supervised quantum Boltzmann machine with parallel annealing reaches small-CNN-level accuracy on PneumoniaMNIST and BreastMNIST in fewer epochs, while cutting QPU time by 69.65% versus sequential annealing.","lead":"This paper trains quantum Boltzmann machines on two medical image datasets using a parallel annealing scheme that embeds ten model copies into one quantum chip run. The authors report accuracy close to small convolutional neural networks in far fewer training epochs, with a roughly 70% reduction in quantum processing time.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SA-based hyperparameter selection is the load-bearing risk for the QBM(QA) comparisons; the three-seed QA runs do not establish that settings chosen by simulated annealing transfer to D-Wave hardware.","rationale":"The reader's weakest assumption identifies exactly the most load-bearing risk: the QBM(QA) comparison uses hyperparameters selected by SA and only three seeds. The paper's own Sec. IV limitation statement confirms that SA 'does not necessarily return the exact same parameters' as QA, and that ideal optimization should be executed using QA. The central classification claim is therefore conditional on a transfer assumption that has not been tested. The speedup claim, by contrast, is measured directly on QPU time in Sec. III-C and Fig. 5, and the PQA embedding strategy is described concretely enough that the 69.65% figure is plausible and less vulnerable to this concern. I also considered whether the CNN baseline's reported 'vanishing gradients' could make the comparison unfair, and whether the 'fewer epochs' claim is misleading because a QBM epoch is far more expensive than a CNN epoch. These are real secondary issues, but the SA-to-QA transfer is more foundational: if the QA model was trained with settings optimized for the wrong sampler, no amount of rerunning with the same settings would fix the comparison. The conditional verdict is therefore appropriate, and the proposed concrete test would either validate the transfer assumption or force a re-evaluation of the headline claim.","tokens_in":17617,"tokens_out":13578,"duration_ms":640658,"concrete_test":"Run a bounded QA-based hyperparameter validation on both datasets: sample 10 configurations from the Table I search space and train each with PQA on D-Wave using 10 seeds, then compare (a) the best validation composite score and (b) test ACC/AUC at the selected epoch against the SA-selected configuration's three-seed QA results. If the QA-selected configuration's test metrics fall outside the SA-selected configuration's seed-induced spread, or if the ranking of configurations changes substantially between SA and QA, the transfer assumption in Sec. III-B.1 fails and the QBM(QA) comparison in Fig. 4 is not a fair measure of the proposed approach.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central comparison—QA-trained QBMs using PQA reaching CNN-comparable accuracy with markedly fewer epochs—rests on exactly one QA configuration per dataset. In Sec. III-B.1, all QBM hyperparameters (hidden units, learning rate, epochs, batch size, sample count) were selected using Simulated Annealing because QPU time was limited, and only those configurations were then retrained on D-Wave hardware with three random seeds (Sec. III-C, Fig. 4). The authors concede in Sec. IV that SA 'does not necessarily return the exact same parameters' as QA. Since D-Wave sampling is an approximate, hardware-specific Boltzmann sampler with an instance-dependent effective temperature (Sec. II-B, [7], [28]), the SA-optimal setting need not be near-optimal, or even stable, on the QPU. If the QA-optimal hyperparameters differ materially, the QBM(QA) curves in Fig. 4—the only direct evidence for the 'comparable to CNNs' and 'markedly smaller epochs' claims—could misstate what the PQA-trained model can actually achieve. The three-seed averaging compounds the problem: the authors explicitly say the QA standard deviation is not representative enough for conclusions, yet Fig. 4 compares these QA curves against 10-seed CNN and SA curves. The 69.65% QPU-time speedup claim is separate and much better supported, so this concern does not undermine the whole paper; it undermines the fairness of the headline classification comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents an improved parallel quantum annealing (PQA) scheme for supervised training of quantum Boltzmann machines (QBMs) on the D-Wave Advantage system. The authors partition the Pegasus graph into ten isolated subgraphs, embed one QBM instance per subgraph, and draw ten samples per annealing cycle. They evaluate QBMs on PneumoniaMNIST and BreastMNIST, using simulated-annealing-selected hyperparameters for all QBM runs and a small subset (three seeds) of QA retraining; they compare against similarly sized CNNs and report per-epoch test accuracy and AUC curves. They also measure QPU time for PQA versus sequential QA and report a 69.65% reduction. The paper concludes that QBMs reach CNN-comparable accuracy with markedly fewer epochs and that PQA yields a large QPU-time reduction.","tokens_in":17840,"tokens_out":7137,"duration_ms":83147,"significance":"If the results hold, the paper provides one of the first demonstrations of supervised QBM training on real annealer hardware for medical image benchmarks, with a concrete technique (controlled subgraph placement with buffer zones) that mitigates crosstalk in parallel embeddings. The QPU-time measurement is direct and the 'fewer epochs' claim is supported by per-epoch curves on both datasets. The significance is moderate: the classification accuracies are not competitive with large CNNs, the QA portion of the comparison rests on three seeds with SA-selected hyperparameters, and the authors themselves note that no clear general conclusion can yet be drawn about QBM versus CNN performance. The contribution is nevertheless a useful step toward practical PQA-based QBM training.","major_comments":[{"comment":"The QA-based QBM results, which are the only direct evidence for the headline 'comparable to CNNs with markedly fewer epochs' claim, are produced with hyperparameters selected by SA rather than by QA. All QBM hyperparameters (hidden units, learning rate, batch size, sample count, epochs) were optimized with SA because QPU time was limited, and only the single best configuration per dataset was retrained on hardware with three seeds. The authors concede in Sec. IV that SA 'does not necessarily return the exact same parameters' as QA. Since D-Wave sampling is hardware-specific and has an instance-dependent effective temperature (Sec. II-B), the SA-optimal settings may be far from QA-optimal. The three-seed averaging is also explicitly acknowledged as not representative. This does not invalidate the PQA speed-up measurement, but it does mean the Fig. 4 comparisons between QBM(QA) and CNN/QBM(SA) are not yet a reliable basis for the abstract's claims. Please provide a sensitivity analysis on the QPU (e.g., vary learning rate and hidden-unit count around the SA optimum on a validation subset) or re-scope the claims to 'QBM(SA)' and report QA results as preliminary.","section":"Sec. III-B.1, III-C, Fig. 4"},{"comment":"The abstract states that QBMs 'achieve reasonable results, comparable to those of similarly-sized CNNs, with markedly smaller numbers of epochs,' but Sec. III-C explicitly says 'we do not see any clear conclusions that can be drawn from these two experiments about the general (medical) image classification performance of QBMs in comparison to CNNs just yet.' The per-epoch comparison in Fig. 4 is based on one selected hyperparameter configuration per model class, not on the distribution of configurations shown in Fig. 3. Please either align the abstract and conclusion with the more cautious statement in Sec. III-C, or provide a statistical comparison across multiple configurations and seeds that supports the stronger claim.","section":"Abstract vs. Sec. III-C"},{"comment":"The input encoding is underspecified. The text assigns 'one input unit to each of the 784 pixel values' but does not state how the 28x28 grayscale pixel values (presumably in [0,255] or normalized [0,1]) are mapped to the v_d values used in Eq. (8). If real-valued inputs are used directly as conditional biases, this should be stated explicitly together with the normalization; if the pixels are binarized, the threshold should be given. This detail is needed to reproduce the parameter counts (1568 input weights plus 2 for one hidden unit) and the experiments. It also affects the claim that 'input units do not necessitate specific hardware resources,' since arbitrary real-valued biases are still a modeling choice that must be documented.","section":"Sec. II-D, Eq. (8)"},{"comment":"The 69.65% QPU-time speedup is a headline quantitative claim, but the description reports only a single measurement campaign ('we tracked the QPU time ... in seconds for 3 mini-batches') without stating the number of repeated runs, the variance across configurations, or whether the sequential baseline includes the same number of samples under identical embedding conditions. Please report per-configuration raw times and at least a standard deviation or interquartile range, and clarify whether the times are QPU access times only or include programming/readout overhead. As written, the precision of '69.65%' is not assessable.","section":"Sec. III-C, Fig. 5"}],"minor_comments":[{"comment":"The phrase 'Despite employing this alternative as as a workaround' contains a duplicated 'as'.","section":"Sec. III-B.1"},{"comment":"The caption spells 'PneunomiaMNIST' instead of 'PneumoniaMNIST', and Sec. III-C contains 'both of theses questions' instead of 'these questions'.","section":"Fig. 3 caption, Sec. III-C"},{"comment":"The caption reads 'best identified hyperparameters settings'; it should be 'best identified hyperparameter settings'. Given the authors' own caveat, the QBM(QA) standard deviation should be visually distinguished or annotated as based on three seeds.","section":"Fig. 4 caption"},{"comment":"Reference [29] contains 'Accesed' instead of 'Accessed'; please also check the formatting of the URLs and DOIs in Refs. [1], [20], [21] for consistency.","section":"References"},{"comment":"For BreastMNIST, the authors should clarify the label convention ('normal'/'benign' as positive and 'malignant' as negative) against the original dataset's class definitions, since this affects the direction of the AUC interpretation.","section":"Sec. III-A"}],"recommendation":"major_revision","confidential_remarks":"The strongest result in the paper is the PQA QPU-time reduction; the classification comparison is less robust. I would be comfortable with acceptance only after the abstract and conclusions are aligned with the evidence and the SA-to-QA transfer risk is either mitigated or explicitly scoped. The paper is within scope for a quantum-machine-learning venue; novelty over Noe et al. is incremental but acceptable given the supervised setting and MedMNIST benchmark."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is a genuine step forward for practical QBM training. It applies parallel quantum annealing to supervised discriminative QBMs on real medical image benchmarks, and the buffer-separated subgraph embedding is a sensible, engineering-level fix for the crosstalk issue Pelofske et al. flagged. The 69.65% QPU-time speedup is measured, clearly explained, and holds across the configurations shown; that is the part I trust most. The paper also deserves credit for what it does not claim: no quantum advantage, and the discussion explicitly walks through the limitations.\n\nThe soft spots are real but mostly acknowledged. The classification comparison, especially the \"comparable to CNNs with fewer epochs\" headline, rests on a single QA hyperparameter configuration per dataset and three seeds. Since those hyperparameters were selected with SA, and the authors concede SA 'does not necessarily return the exact same parameters' on the QPU, the QA curves in Fig. 4 could be biased. That is a legitimate concern, but it is an admitted limitation, not a hidden one. The paper's own framing is appropriately cautious.\n\nTwo things bother me more. The pixel-to-binary encoding for MedMNIST images is never specified, which makes the QBM experiments hard to reproduce. And I saw no code or data availability statement; for an empirical paper like this, that should be a requirement.\n\nThe CNN comparison is also a bit apples-to-oranges, as the authors recognize: CNNs exploit spatial structure, QBMs get flattened pixels. The per-epoch observation (QBMs reach decent accuracy in a few epochs) is interesting even if final accuracies are similar, but with three seeds it remains suggestive.\n\nWho is this for: quantum annealing practitioners, QBM researchers, and people building hybrid classical-quantum pipelines. I would send it to peer review. A good referee can push for the encoding details, more seeds, and a clearly stated SA-to-QA transfer caveat in the abstract. The PQA contribution is solid and the reporting is honest.","headline":"A useful, honestly hedged engineering contribution: PQA with buffer-separated subgraphs trains supervised QBMs on medical images, and the ~70% QPU-time speedup is the strongest result; the CNN-comparable accuracy claim is suggestive but rests on thin QA evidence.","tokens_in":18497,"tokens_out":2382,"would_cite":true,"duration_ms":28230,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that supervised Quantum Boltzmann Machines trained with parallel quantum annealing can classify MedMNIST medical images as accurately as similarly sized CNNs while needing far fewer epochs, and that the parallel scheme…","keywords":["Quantum Boltzmann Machines","Parallel Quantum Annealing","Medical image classification","MedMNIST","Supervised Boltzmann machine training","Quantum annealing","Boltzmann sampling","Near-term quantum machine learning"],"falsifier":"Repeat the QBM(QA) training with hyperparameters tuned directly on the quantum annealer rather than via simulated annealing, using many random seeds; if accuracy drops relative to the three-seed SA-tuned runs or the early-epoch advantage disappears, the central practical claim fails. A complementary check is to compare sample quality and classification accuracy from parallel embeddings with and without buffer zones, which would reveal whether the 69.65% speed-up comes at a hidden quality cost.","tokens_in":17349,"feed_emoji":"🩻","tokens_out":9663,"duration_ms":100355,"temperature":0.7,"pith_summary":"Training a Quantum Boltzmann Machine normally costs a lot of quantum-annealer time, because every gradient step needs many Boltzmann samples and each sample has traditionally come from one problem instance at a time. This paper tries to make supervised QBM training practical by embedding several independent copies of the same network in one annealing run, separated on the chip so they do not interfere, and by keeping the input images out of the qubit encoding entirely: clamped inputs act only as biases. On PneumoniaMNIST and BreastMNIST, the authors report that the resulting QBMs reach test accuracy comparable to similarly sized CNNs while needing far fewer epochs, and that the parallel embedding reduces quantum-hardware time by 69.65% compared with sequential annealing. They present this as a step toward real-world, near-term quantum machine learning, while noting that the hardware results use only three random seeds and that hyperparameters were chosen with simulated annealing as a stand-in for the quantum process.","feed_headline":"Parallel annealing cuts quantum Boltzmann training time by 70%","feed_subtitle":"Supervised QBMs match similarly sized CNNs on MedMNIST while needing far fewer epochs.","key_machinery":"The central object is the parallel embedding of a QUBO, the binary energy formulation of the QBM, onto the annealer's hardware graph. Because the input pixels are clamped, the QBM's energy for the free units reduces to an effective bias term, so only hidden units and the label unit need physical qubits; the paper encodes each model as a QUBO with at most 21 logical qubits. The hardware graph is partitioned into ten subgraphs with buffer zones of removed nodes between them, the same QUBO is embedded into each subgraph using an automatic embedding routine, and a single annealing cycle draws ten samples at once. This spatial separation is the part of the argument that is supposed to preserve sample quality while delivering the measured 69.65% reduction in processing time.","core_discovery":"The paper's central claim is that an annealing-based Quantum Boltzmann Machine can be trained for supervised binary image classification on current hardware with a parallel embedding scheme that makes the training time competitive with classical baselines. Only the hidden units and one label unit are embedded as qubits; the 784 input pixels are clamped to the network and enter only through effective biases, so the embedded model stays small regardless of image size. Ten copies of the quadratic unconstrained binary optimization (QUBO) model are placed in ten separated subgraphs of the annealer's graph, and one annealing cycle returns ten Boltzmann samples. On PneumoniaMNIST the QA-trained QBM reaches 84.03% test accuracy and an AUC of 0.7996; on BreastMNIST it reaches 76.28% and 0.5946. The authors do not claim a decisive accuracy win over CNNs; their claim is that this near-classical accuracy is reached within the first five to eight epochs and that the parallel annealing gives a 69.65% reduction in quantum-hardware time, which together make supervised QBM training a plausible near-term option.","pith_inferences":["A useful next experiment is to measure end-to-end wall-clock time including graph partitioning and embedding, not just annealer time, to see whether the 69.65% advantage survives in practice.","A classical fully connected Boltzmann Machine or discriminative Restricted Boltzmann Machine trained with the same update rule would isolate the quantum contribution better than the CNN baseline; the paper itself notes that such a comparison is needed.","If simulated-annealing hyperparameters do not transfer to the annealer, the three-seed QA results may underestimate the model: QA-tuned hyperparameters could close the gap to the larger ResNet baselines.","For multi-class medical datasets, label units would scale linearly with the number of classes, which should keep the parallel embedding scheme usable for realistic clinical label sets."],"forward_implications":["QPU-time budgets for QBM experiments can be cut by roughly two-thirds, allowing more seeds, more hyperparameter trials, or larger datasets for the same cost.","Because input pixels enter only as biases, moving to larger images does not change the number of embedded qubits, so the approach scales to bigger inputs without needing bigger hardware.","The supervised formulation extends parallel-annealing QBM training from the unsupervised setting of Noe et al. to classification, the setting needed for medical diagnostics.","On scarce medical datasets like BreastMNIST, all tested models struggle to generalize, so the practical benefit here is faster training rather than higher accuracy."],"supporting_citations":[{"why":"Introduced parallel quantum annealing for unsupervised QBM training; this paper extends that idea to supervised learning.","marker":"[1]"},{"why":"Supplies the MedMNIST data sets (PneumoniaMNIST and BreastMNIST) used for all experiments and the ResNet reference results.","marker":"[2]"},{"why":"Establishes the discriminative QBM training rule, the clamped-input energy reduction, and the mapping from QBM to an Ising Hamiltonian.","marker":"[7]"},{"why":"Provides the precedent of using simulated annealing as a proxy for quantum annealing in QBM training and prior medical-image QBM work.","marker":"[20]"},{"why":"Introduces parallel quantum annealing and documents the sample-quality drop that motivates this paper's buffer-zone separation.","marker":"[22]"},{"why":"Shows that raw QA samples without effective-temperature correction can be used for Boltzmann machine training, which the paper relies on.","marker":"[27]"},{"why":"Documents the prohibitive QPU-time cost of annealing-based QBM training, which the parallel approach is designed to reduce.","marker":"[21]"}],"fun_headline_variants":["Parallel annealing cuts QBM training time by 70% on medical images","Quantum Boltzmann machine matches CNN accuracy on MedMNIST in fewer epochs","70% quantum hardware time saved by parallel annealing for supervised QBMs","Parallel annealed QBMs rival CNNs on medical imaging with fewer training epochs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that hyperparameters chosen with simulated annealing transfer to the quantum annealer, since the paper concedes the two sampling processes do not necessarily return the same parameters and the hardware results use only three seeds.","fun_headline_variants_meta":{"raw":{"variants":["Parallel annealing cuts QBM training time by 70% on medical images","Quantum Boltzmann machine matches CNN accuracy on MedMNIST in fewer epochs","70% quantum hardware time saved by parallel annealing for supervised QBMs","Parallel annealed QBMs rival CNNs on medical imaging with fewer training epochs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000864,"raw_usage":{"total_tokens":3762,"prompt_tokens":979,"completion_tokens":2783,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":2705}},"tokens_in":595,"tokens_out":2783,"duration_ms":22127,"temperature":1.0,"reasoning_tokens":2705,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:01:00.238104+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the QBM(QA) training with hyperparameters tuned directly on the quantum annealer rather than via simulated annealing, using many random seeds; if accuracy drops relative to the three-seed SA-tuned runs or the early-epoch advantage disappears, the central practical claim fails. A complementary check is to compare sample quality and classification accuracy from parallel embeddings with and without buffer zones, which would reveal whether the 69.65% speed-up comes at a hidden quality cost.","supporting_citations":[{"cited_title":"Quantum parallel training of a boltzmann machine on an adiabatic quantum computer,","cited_arxiv_id":null,"evidence_quote":"Introduced parallel quantum annealing for unsupervised QBM training; this paper extends that idea to supervised learning."},{"cited_title":"MedMNIST v2-a large-scale lightweight benchmark for 2d and 3d biomedical image classification,","cited_arxiv_id":null,"evidence_quote":"Supplies the MedMNIST data sets (PneumoniaMNIST and BreastMNIST) used for all experiments and the ResNet reference results."},{"cited_title":"Quantum boltzmann machine,","cited_arxiv_id":null,"evidence_quote":"Establishes the discriminative QBM training rule, the clamped-input energy reduction, and the mapping from QBM to an Ising Hamiltonian."},{"cited_title":"Towards transfer learning for large-scale image classification using annealing-based quantum boltzmann machines,","cited_arxiv_id":null,"evidence_quote":"Provides the precedent of using simulated annealing as a proxy for quantum annealing in QBM training and prior medical-image QBM work."},{"cited_title":"Parallel quantum annealing,","cited_arxiv_id":null,"evidence_quote":"Introduces parallel quantum annealing and documents the sample-quality drop that motivates this paper's buffer-zone separation."},{"cited_title":"Benchmarking Quantum Hardware for Training of Fully Visible Boltzmann Machines","cited_arxiv_id":"1611.04528","evidence_quote":"Shows that raw QA samples without effective-temperature correction can be used for Boltzmann machine training, which the paper relies on."},{"cited_title":"Exploring unsupervised anomaly detection with quantum boltzmann machines in fraud detection,","cited_arxiv_id":null,"evidence_quote":"Documents the prohibitive QPU-time cost of annealing-based QBM training, which the parallel approach is designed to reduce."}],"review_version":1}