{"id":"13b73802-e312-49cd-9122-0085cbb5d942","arxiv_id":"2608.11373","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Optimized quantum classifiers do not beat AutoML-tuned classical neural networks on standard oncological benchmarks, and resource estimates indicate the datasets are too small to make quantum advantage plausible.","lead":"The authors compared optimized quantum and classical machine learning models on cancer classification datasets and found no evidence that quantum models perform better. The result suggests moving to more complex, higher-dimensional biological data before expecting practical quantum advantage.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No-advantage claim extrapolates from quantum models capped at 10 qubits while several spatial datasets have Q_dataset(1.0) above 50; the predicted-advantage regime is untested.","rationale":"The reader's weakest assumption identifies the same load-bearing gap: quantum models are stopped at 10 qubits even though Q_dataset(1.0) for several spatial datasets is 26.8-66.2. I find this to be the single most important limitation because it sits directly between the empirical comparison and the abstract's general conclusion. The paper is otherwise careful: training accuracy is driven to 100%, loss landscapes are addressed via sub-net initialization, classical baselines are AutoML-optimized, and the no-advantage claim is phrased as absence of evidence rather than proof of absence. The spatial limitation is explicitly acknowledged in Section III-C, which is a credit to the authors, but acknowledgment does not remove the fact that the recommendation to prioritize higher-dimensional datasets is partly an extrapolation. I would not change the reader's CONDITIONAL verdict: the concern is real and testable, but it does not invalidate the results for the qubit range that was run. If the high-qubit test is performed and still shows no advantage, the paper's central claim would be substantially stronger; until then, conditional on release of framework/code and high-qubit results is the right stance.","tokens_in":16986,"tokens_out":5399,"duration_ms":50968,"concrete_test":"Run the Red Cedar quantum classifier on ResNet18-preprocessed PathMNIST (Q_dataset(1.0)=56.5, about four label qubits) at 20, 30, 40, and where feasible 50 feature qubits, using the same stratified five-fold splits and TPOT2 MLP baselines already reported. If quantum held-out accuracy stays within one standard error of the classical baseline at every size, the no-advantage claim extends into the predicted high-qubit regime; if quantum accuracy exceeds the classical baseline by more than two standard errors at any size, the central claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that no quantum advantage is found depends on quantum models trained only at 4, 6, 8, and 10 qubits (Section II-C4). For the spatial datasets the paper's own resource metric places theoretical saturation much higher: flattened BreastMNIST Q_dataset(1.0)=66.2±11.0 and ResNet18 PathMNIST Q_dataset(1.0)=56.5±1.4 (Section III-C). The paper concedes: 'since we only ran the quantum model up to ten qubits and the spatial datasets have much larger Q_dataset(1.0)s, it is difficult to make broader observations.' The classical AutoML baselines, by contrast, are evaluated across the full bit range up to Q_dataset(1.0) (Section II-D), so the comparison is asymmetric exactly where a quantum advantage could appear. Moreover, Q_dataset(1.0) is acknowledged to be optimistic because unseen test samples absent from the training set are counted as correct (Section II-C2), so the theoretical estimates cannot substitute for the missing high-qubit runs. The abstract's 'no evidence' is formally safe, but the accompanying recommendation to move beyond current benchmarks relies on a null result in a regime that was never actually tested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a benchmarking methodology for quantum versus classical machine learning on oncological data, using the Red Cedar bit-bit encoding framework for quantum models and TPOT2-optimized classical neural networks as baselines. Experiments cover WDBC, three TCGA/MLOmics genomics datasets, and three MedMNIST spatial datasets, each with modality-specific preprocessing. The quantum models are trained using sub-net initialization and exact coordinate updates, with all completed runs reaching 100% training accuracy; test accuracies are compared against classical AutoML models and a theoretical test-accuracy estimate derived from collision statistics. The paper's central claim is that no evidence of quantum advantage is found on these benchmarks, and it recommends that the field move toward higher-dimensional, more biologically realistic datasets.","tokens_in":17270,"tokens_out":5882,"duration_ms":58644,"significance":"If the result holds, this is a useful contribution to QML benchmarking: it attempts to control for optimization failures by driving quantum models to 100% training accuracy, uses stratified five-fold splits with standard errors, and compares against AutoML-tuned classical models rather than fixed hand-tuned baselines. The resource-estimation idea is falsifiable and is supported by a public web-based estimator. However, the significance is limited by three issues: the quantum models are run only up to ten qubits while classical models are run across the full bit range; the theoretical test-accuracy metric is acknowledged to be optimistic; and the 50-qubit 'potential for quantum advantage' threshold is inherited from the authors' prior framework without independent calibration. These gap prevent the broad 'no evidence' conclusion from being fully supported for the spatial datasets.","major_comments":[{"comment":"The comparison is asymmetric in exactly the regime where quantum advantage could appear. Quantum models are trained only at 4, 6, 8, and 10 qubits, while classical AutoML models are trained for all bit counts up to Q_dataset(1.0). For the spatial datasets, Q_dataset(1.0) is reported as 66.2±11.0 for flattened BreastMNIST and 56.5±1.4 for ResNet18-preprocessed PathMNIST, and the paper itself concedes in Section III-C that 'since we only ran the quantum model up to ten qubits and the spatial datasets have much larger Q_dataset(1.0)s, it is difficult to make broader observations.' The abstract's 'no evidence of quantum advantage' and the discussion's recommendation to move beyond current benchmarks are therefore broader than the experiments support. The claim should be explicitly restricted to the tested qubit range, or the authors should provide additional quantum runs at higher qubit counts and a matched-qubit classical comparison at the same sizes.","section":"§II-C4, §II-D, §III-C"},{"comment":"The theoretical test-accuracy metric cannot substitute for the missing high-qubit experiments because it is acknowledged to be optimistic. Section II-C2 states that test samples not present in the training encoded set are counted as correctly classified in the theoretical accuracy, making the estimate 'realistically too optimistic.' Since Q_dataset(1.0) is computed under the same convention, the theoretical curves for spatial datasets do not establish what a quantum model would actually achieve at 50 or more qubits. The paper should either correct this metric to account for unseen test samples or present empirical results in the high-Q_dataset regime before drawing conclusions about the absence of advantage.","section":"§II-C2, §III-C"},{"comment":"The 50-qubit threshold for 'potential for quantum advantage' is inherited from the authors' prior resource-estimation framework and from Ref. [60], but Ref. [60] concerns simulation hardness of random quantum circuits, not the trainability or generalization of quantum classifiers. As used here, Q_dataset(1.0)>50 is a heuristic, and the paper should state this limitation explicitly and ideally calibrate the threshold against learning performance. Otherwise, the classification of WDBC and omics datasets as 'below the threshold' is not a strong independent reason to expect no quantum advantage.","section":"§II-C2, §IV"},{"comment":"The manuscript reports that some quantum models did not finish running within the five-day time limit, but it gives no count of completed runs per dataset, preprocessing method, or fold. If the completed subset is not representative, the 'no evidence' conclusion could be biased by censoring. The authors should report completion rates and, where feasible, analyze whether timed-out runs differ systematically from completed runs.","section":"§III-C, §II-C4"}],"minor_comments":[{"comment":"The exact computational definition of theoretical test accuracy is not given; since it is central to the resource-estimation metric, include the precise collision/overlap formula or pseudocode.","section":"§II-C2"},{"comment":"The classical bit range is [1, Q_dataset(1.0)−⌈log2(classes)⌉], while Q_dataset(1.0) is said to include label bits; please clarify and ensure the x-axis labels in Figures 2–4 ('number of bits allocated to features') are consistent with the text.","section":"§II-D, Figure captions"},{"comment":"The removal of two colliding samples from PathMNIST and BreastMNIST should specify whether this was done before or after the stratified split, and how the conflicting labels were resolved.","section":"§II-A3"},{"comment":"There is a typo: 'is possible that the AutoML algorithm' should read 'it is possible that the AutoML algorithm.'","section":"§IV"},{"comment":"The quantum training software is described as available only upon request; given that the experiments are central to the paper, a more detailed algorithmic description or pseudocode would improve reproducibility.","section":"§V"}],"recommendation":"major_revision","confidential_remarks":"The paper is authored by employees of Cascade Quantum, the developer of the Red Cedar framework that is both the subject and the primary tool of the benchmarking study. The framework is not openly available, which may limit independent verification. The editor may wish to ensure that reviewers can access the software or that the authors provide sufficient algorithmic detail. The central empirical finding is plausible, but the over-broad interpretation of the null result in the spatial-data regime is the main obstacle to acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful, honest negative benchmark. It applies the Red Cedar bit-bit encoding and resource estimation to WDBC, three TCGA omics datasets, and three MedMNIST spatial datasets, and measures quantum versus AutoML-optimized classical models across preprocessing strategies. The experimental care is real: stratified five-fold splits, standard errors, Mann-Whitney tests, and quantum models that all reach 100% training accuracy. The paper also admits its main limitations directly, which is more than most QML benchmarks do.\n\nWhat is actually new: the Q_dataset(1.0) resource estimates for these datasets, and the systematic comparison of quantum and classical models on bit-bit encoded data. The overall no-advantage finding is consistent with prior negative QML benchmarks, but this paper adds a reproducible methodology and a useful template.\n\nThe soft spots are real but not fatal. The quantum models are capped at ten qubits. Several spatial datasets have Q_dataset(1.0) above 50, and the paper itself concedes it is difficult to make broader observations there. So the abstract's 'no evidence' is safe, but the forward-looking recommendation—that the field should move to higher-dimensional data—leans on a null result in a regime that was never actually tested. The theoretical test accuracy metric is also optimistic by design: unseen test samples not in the training set count as correctly classified. And the Q_dataset threshold comes from the authors' own prior work, so the 'potential for quantum advantage' standard is internally defined. That is not disqualifying, but it means the resource estimates should be treated as a heuristic, not an external ground truth.\n\nWho this is for: anyone working on QML benchmarking, especially in biomedical or biological applications. It will be most useful as a methodological reference and a caution against toy benchmarks. The lack of open-source code for the quantum framework (available only upon request) is a limitation, but the resource estimation tool is public.\n\nVerdict: send it to peer review. The concerns are addressable—higher-qubit runs, corrected theoretical metric, released code—and the paper deserves referee time despite the current limits.","headline":"A careful negative QML benchmark on oncology data, but the strongest conclusion is bounded by a 10-qubit cap and an optimistic resource metric from the authors' own framework.","tokens_in":17774,"tokens_out":1788,"would_cite":true,"duration_ms":16239,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A benchmarking study that drives quantum classifiers to 100% training accuracy finds no evidence that they outperform classical networks on oncological datasets spanning tabular, genomic, and imaging data.","keywords":["quantum machine learning","quantum advantage","oncological data","cancer classification","benchmarking","bit-bit encoding","resource estimation","variational quantum classifiers"],"falsifier":"Run the same bit-bit encoded quantum classifier on one of the imaging datasets whose $Q_{\\text{dataset}}(1.0)$ exceeds 50 qubits at 20, 30, and 40 qubits and compare its test accuracy with the automated classical baseline over the same five folds; a statistically significant quantum advantage at any of those sizes would overturn the paper's no-advantage conclusion.","tokens_in":16788,"feed_emoji":"⚛️","tokens_out":10479,"duration_ms":118529,"temperature":0.7,"pith_summary":"This paper tries to settle whether quantum machine learning models are actually better than classical models on the oncological datasets that dominate the quantum-machine-learning literature. It constructs a comparison in which every quantum classifier is trained to 100% training accuracy, so optimization failure cannot explain any gap, and pits them against classical neural networks tuned by an automated search on the same encoded data. Across the Wisconsin breast-cancer dataset, three cancer-subtype genomics datasets, and three medical imaging datasets, quantum and classical test accuracies are statistically indistinguishable at the qubit counts tested, with no clear advantage on the imaging data either. The paper's practical conclusion is that claims of quantum advantage on these well-worn small benchmarks are unsupported, and that the field should move toward higher-dimensional, more biologically realistic datasets.","feed_headline":"No quantum advantage found in oncological benchmark tests","feed_subtitle":"Trained to 100% accuracy, quantum classifiers match—not beat—classical networks on breast, genomics, and imaging data.","key_machinery":"The central machinery is bit-bit encoding paired with a resource-estimation quantity derived from it. Bit-bit encoding discretizes each classical feature into a user-specified number of bits, allocates those bits in proportion to how much mutual information each feature shares with the class, and loads the resulting bitstrings into data qubits in the computational basis; because the encoding has a universal-approximation property, the quantum model can be trained to 100% training accuracy. Training uses an incremental warm start that grows the model from four to ten qubits and coordinate updates that require no classical optimizer, which avoids landing in flat, hard-to-optimize regions of the loss landscape. From the encodings, the framework computes $Q_{\\text{dataset}}(1.0)$, the smallest number of qubits at which no two samples with different class labels collide and perfect train and test accuracy is theoretically possible; values above 50 are treated as indicating a dataset large enough that a quantum model might plausibly beat classical simulation. The classical baseline is an automated neural-network search run on the same discretized inputs, so both paradigms face the same information budget.","core_discovery":"The paper's central claim is that no evidence of practical quantum advantage appears on oncological classification problems when both model classes are given their best shot. On the Wisconsin breast-cancer dataset, quantum and classical models reach statistically indistinguishable test accuracies at every qubit count tested; on the three genomics datasets the same holds for the great majority of qubit counts under both raw and feature-selected inputs; and on the imaging datasets the quantum runs are sparser but show no advantage. Every quantum model studied reaches 100% training accuracy, which the paper argues rules out training failures, flat optimization landscapes, and limited expressivity as explanations for the results. The resource-estimation metric $Q_{\\text{dataset}}(1.0)$ — the number of encoded qubits at which perfect train and test accuracy becomes theoretically possible — is below 50 for the tabular and genomics datasets, and only some imaging datasets approach or exceed the 50-qubit threshold the framework associates with potential quantum advantage.","pith_inferences":["The theoretical test-accuracy metric treats test samples that never appear in the training set as correctly classified, so $Q_{\\text{dataset}}(1.0)$ is an optimistic bound; datasets that barely clear 50 qubits may actually sit in the no-advantage regime once this optimism is corrected.","A sharper test than rerunning the same benchmarks would be to generate synthetic oncological-style datasets with controlled higher-order feature interactions and enough samples to separate memorization from generalization, then ask whether advantage appears exactly where $Q_{\\text{dataset}}(1.0)$ passes the threshold.","The fact that removing 99.8% of omics features left $Q_{\\text{dataset}}(1.0)$ essentially unchanged suggests the resource estimate tracks intrinsic problem complexity rather than preprocessing details; if this holds across modalities, $Q_{\\text{dataset}}(1.0)$ could serve as a cheap pre-screening tool before any quantum training is attempted."],"forward_implications":["On the tabular and genomics benchmarks, the resource estimate alone predicts no quantum advantage, so future studies reporting advantage on these datasets need to show why the estimate does not apply.","Preprocessing improves classical generalization in several cases, but it does not create a statistically significant quantum-classical gap.","Because every quantum model reaches 100% training accuracy, the absence of advantage should be attributed to the information content of the benchmarks rather than to optimization failures in the quantum training procedure.","The imaging datasets with $Q_{\\text{dataset}}(1.0)$ near or above 50 qubits are the only candidates left where quantum advantage could still appear, but the ten-qubit training ceiling leaves that possibility open."],"supporting_citations":[{"why":"Supplies the bit-bit encoding and incremental warm-start training method that lets every quantum model reach 100% training accuracy.","marker":"[26]"},{"why":"Defines the $Q_{\\text{dataset}}(1.0)$ resource-estimation metric and the 50-qubit threshold used to judge a dataset's quantum-advantage candidacy.","marker":"[27]"},{"why":"Provides the automated machine-learning classical neural-network baselines on the same discretized inputs.","marker":"[28]"},{"why":"Supplies the cleaned cancer-subtype genomics datasets used for the omics experiments.","marker":"[44]"},{"why":"Supplies the standardized medical imaging benchmark datasets used for the spatial experiments.","marker":"[47]"},{"why":"Establishes the need for rigorous comparison protocols in benchmarking quantum against classical models.","marker":"[16]"},{"why":"Supplies the Wisconsin breast-cancer tabular dataset on which quantum and classical accuracies are statistically indistinguishable.","marker":"[29]"},{"why":"Identifies omics and spatial oncological data as promising candidates for quantum advantage, motivating the modalities tested.","marker":"[13]"}],"fun_headline_variants":["Quantum ML fails to beat classical on cancer data","No quantum edge on oncological datasets","Quantum models match, not beat, classical on cancer data","Fair test finds no quantum advantage in oncology ML"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The no-advantage conclusion assumes that training the quantum model only up to ten qubits is enough to expose quantum advantage; several imaging datasets would need roughly 27 to 66 qubits by the paper's own resource estimate, so a larger quantum model could in principle change the result.","fun_headline_variants_meta":{"raw":{"variants":["Quantum ML fails to beat classical on cancer data","No quantum edge on oncological datasets","Quantum models match, not beat, classical on cancer data","Fair test finds no quantum advantage in oncology ML"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000575,"raw_usage":{"total_tokens":2691,"prompt_tokens":901,"completion_tokens":1790,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":1731}},"tokens_in":517,"tokens_out":1790,"duration_ms":27559,"temperature":1.0,"reasoning_tokens":1731,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:20.751265+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same bit-bit encoded quantum classifier on one of the imaging datasets whose $Q_{\\text{dataset}}(1.0)$ exceeds 50 qubits at 20, 30, and 40 qubits and compare its test accuracy with the automated classical baseline over the same five folds; a statistically significant quantum advantage at any of those sizes would overturn the paper's no-advantage conclusion.","supporting_citations":[{"cited_title":"Quantum machine learning in medical image analysis: A survey,","cited_arxiv_id":null,"evidence_quote":"Supplies the standardized medical imaging benchmark datasets used for the spatial experiments."},{"cited_title":"How many qubits does a machine learning problem require?","cited_arxiv_id":null,"evidence_quote":"Defines the $Q_{\\text{dataset}}(1.0)$ resource-estimation metric and the 50-qubit threshold used to judge a dataset's quantum-advantage candidacy."},{"cited_title":"Ribeiro, A","cited_arxiv_id":null,"evidence_quote":"Provides the automated machine-learning classical neural-network baselines on the same discretized inputs."},{"cited_title":"Mlomics: Cancer multi-omics database for machine learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the cleaned cancer-subtype genomics datasets used for the omics experiments."},{"cited_title":"Quantum computing for oncology,","cited_arxiv_id":null,"evidence_quote":"Identifies omics and spatial oncological data as promising candidates for quantum advantage, motivating the modalities tested."}],"review_version":1}