{"id":"fa456fc4-75c9-4058-8f17-6b61eb11bddf","arxiv_id":"2505.04234","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A trainable quantum feature map plus a support-vector-based quantum SVM improves simulated IRIS classification accuracy and distinguishability over least-squares QSVM.","lead":"The authors propose a trainable quantum feature mapping and a support-vector-based quantum classifier that they say improves classification accuracy and reduces measurement overhead on the IRIS dataset. The work is a candidate incremental improvement for noisy intermediate-scale quantum machine learning, but its theoretical support is partly heuristic.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Appendix D's lower bound a ≥ 1/sqrt(5M_j) rests on a zero-mean CLT model for nonnegative kernel overlaps, so the multi-class readout guarantee is unsupported.","rationale":"The reader's weakest assumption names exactly this CLT/mean-1/3 step. I agree. The concern is load-bearing because the multi-class readout is one of the paper's advertised advances and Section V B explicitly relies on the Appendix D iteration bound. The flaw is internal: a zero-mean distribution is impossible for a nonnegative overlap average, and the '∫_0^a g(x)dx = 1/4' step confuses a random variable with a density. This is not a dispute with consensus; it is a mathematical inconsistency in a proof. The SV-QSVM distinguishability argument, while heuristic, is supported by the IRIS comparison and the amplitude-normalization intuition, so I would not reject the paper. The correct response is to require the authors to replace Appendix D with a valid lower bound or with an explicit assumption on the minimum same-class overlap; the numerical part of Section VI D can then serve as evidence for the specific IRIS kernel. This leaves the conditional verdict unchanged.","tokens_in":16639,"tokens_out":12770,"duration_ms":136144,"concrete_test":"Re-derive the amplitude lower bound using the correct CLT model g ~ N(μ, σ²/M_j) with μ=E[k]>0, and compute the minimum target amplitude a_j = (1/M_j)Σ_i k(x_i,x) over all 150 IRIS points for the kernels used in Section VI D. If any a_j is below 1/√(5M_j), or if the quantile equation does not yield the Appendix D expression, then the R interval in Section V B does not establish reliable multi-class readout.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The weakest link is the proof of the iteration bound for the quantum iterative multi-classifier in Appendix D, which Section V B depends on. The derivation assumes that g(x)=Σ_i α_i k(x_i,x)/√M_j follows N(0, 1/(3M_j)) and then sets ∫_0^a g(x)dx = 1/4 to obtain a ≈ 1/√(5M_j). This is not a valid probability statement, and the zero-mean normal model is inconsistent with k(x_i,x) ∈ [0,1]: the amplitude stored for a class is an average of nonnegative kernel overlaps, whose CLT limit has positive mean μ, not zero. If μ is positive, a tends to a constant as M_j grows rather than to 0, so the O(√M_j) iteration count is not a worst-case bound; if μ is very small, a can fall below 1/√(5M_j) and the prescribed number of Grover iterations is insufficient. Thus the claimed high-probability readout for the multi-class framework does not follow from the stated analysis. The IRIS simulations show large post-iteration amplitudes, but that does not supply the missing general guarantee.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a trainable quantum feature mapping (TQFM) built from data re-uploading circuits, and uses it in three ways: as an explicit classifier, as an ensemble model, and as a source of quantum kernels for support vector machines. The main algorithmic contribution is the SV-QSVM, which constructs classification trial states from support vectors only, rather than from all training samples, in order to improve the distinguishability of the decision value. The paper also introduces a quantum iterative multi-classifier based on amplitude amplification to read out one-versus-one or one-versus-rest results, with a proposed iteration bound in Appendix D. Numerical simulations on the IRIS dataset are reported for each component, showing improved clustering and classification accuracy compared to a standard quantum feature map and to the LS-QSVM baseline.","tokens_in":16935,"tokens_out":5826,"duration_ms":59956,"significance":"If the theoretical claims were rigorously established, the SV-QSVM construction would be a useful and intuitive contribution to quantum kernel methods: restricting the trial state to support vectors is a natural way to amplify the decision value and could reduce measurement overhead. The numerical demonstrations on IRIS support the qualitative story, and the paper is honest about the simplicity of the experiments. The work also points to an interesting direction in using amplitude amplification for multi-class readout. However, several of the paper's central theoretical justifications—Theorem 1, Lemma 2, and the Appendix D iteration bound—contain gaps that prevent the results from being accepted as rigorous. The empirical results are promising but do not by themselves supply the missing proofs.","major_comments":[{"comment":"The proof of Theorem 1 assumes that the kernel values k(x_i,x) are independent and identically distributed and then invokes the Central Limit Theorem to conclude f(x) ~ N(0, 1/M). This is not justified: for a fixed trained feature map, the k(x_i,x) are not independent, and more importantly they are nonnegative (or at least not zero-mean in general), so the CLT limit has a positive mean rather than zero mean. Consequently the claimed scaling f(x) = O(1/sqrt(M)) is not established. This theorem is load-bearing because it motivates the entire SV-QSVM construction as a fix for poor distinguishability. The statement should be rephrased as a conditional statement under explicit assumptions on the kernel-value distribution, or replaced by a rigorous bound.","section":"Section III C, Theorem 1"},{"comment":"The lower bound a ≈ 1/sqrt(5 M_j) on the target amplitude is derived by modeling g(x) = sum_i alpha_i k(x_i,x)/sqrt(M_j) as a normal N(0, 1/(3 M_j)) and then writing ∫_0^a g(x) dx = 1/4. This is not a valid probability calculation: g(x) is a random variable, not a probability density, and the zero-mean normal model is inconsistent with k(x_i,x) ∈ [0,1] because the mean of a sum of nonnegative overlaps is nonnegative. If the true mean is positive, a tends to a constant rather than to 0 as M_j grows, so the iteration bound R ∈ [1, (sqrt(5 M_j) π − 1)/2] is not a worst-case bound. If the mean is very small, a can be smaller than 1/sqrt(5 M_j) and the prescribed number of Grover iterations is insufficient. Thus the claim that the quantum iterative multi-classifier reads out results with high probability is unsupported by the stated analysis.","section":"Appendix D and Section V B"},{"comment":"The proof of Lemma 2 is incomplete. The bound ∥δ∥_2 ≤ O(M/sqrt(R)) introduces a quantity R that is never defined, and the step from the perturbation bound in Lemma 1 to this expression is not shown. The argument that λ_min(K + (1/γ)I) ≥ 1/γ is a 'constant lower bound' is true for the noiseless matrix, but Lemma 1 applies to the noisy matrix K', and the effect of noise on the eigenvalue should be stated explicitly. The final shot count O(M^4/ε^2) therefore does not follow from the given reasoning. This lemma is important for the claimed error analysis of SV-QSVM, so the proof needs to be made rigorous or the claim weakened.","section":"Section IV D, Lemma 2"},{"comment":"The derivation of the misclassification rate is internally inconsistent. The text first defines p(x) as the probability density of h(x) and states ∫_{-1}^{1} p(x) dx = 1, but then writes the maximum misclassification rate as ∫_0^ε p(x) dx, and later uses integrals such as 1 − ∫_{-ε}^{ε} p(x) dx and ∫_{-ε}^{-ε} p(x) dx. Since x is the trial datum and h(x) is a function of x, the integral limits in x-space do not correspond to decision errors in h-space. The probabilistic model needs to be clarified, for example by defining the distribution of h(x) directly and expressing the error probability as an integral over the decision threshold.","section":"Section IV D, misclassification-rate analysis"},{"comment":"The sentence 'using the trained kernel matrix for machine learning algorithms, such as SVM, will definitely outperform the untrained one in terms of performance' is an overclaim. The simulations demonstrate improvement on the IRIS dataset for one specific circuit layout, but they do not establish a general guarantee. This statement should be rephrased as an empirical observation about the tested cases.","section":"Section VI A, closing paragraph"}],"minor_comments":[{"comment":"The phrase 'partially evenly weighted trail state' appears to be a typo for 'trial state'; the same term is used correctly elsewhere in the paper.","section":"Section IV D"},{"comment":"The symbol R is used in the proof without definition; if it denotes the number of measurement shots, this should be stated explicitly and consistently with the rest of the paper.","section":"Section IV D, Lemma 2"},{"comment":"In the row for Class 2 with r=1, the average probability after iteration is printed as '0747', which appears to be a typo for '0.747'.","section":"Table II"},{"comment":"The preparation of the oracles U_L and U_x and the state U_x|j-1⟩|0...0⟩ assumes the ability to coherently superpose training samples and the new datum; the resource cost of these oracles is not discussed, and this assumption should be stated explicitly when claiming a reduced readout burden.","section":"Section V B and Appendix D"}],"recommendation":"major_revision","confidential_remarks":"The paper contains several promising ideas and the numerical experiments are a useful proof of concept, but the theoretical sections need substantial repair. The issues in Theorem 1, Lemma 2, and Appendix D are not mere presentation problems: they concern the core statements about distinguishability, error bounds, and the multi-class readout guarantee. I would be willing to reconsider a revised version that either supplies rigorous proofs under clearly stated assumptions or honestly demotes these statements to heuristics supported only by the simulations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The honest summary: this is an incremental but potentially useful quantum machine learning paper. The genuinely new piece is the support-vector-selected partial trial state combined with a Grover-amplified multi-class readout. The TQFM itself is largely a repackaging of data re-uploading and tailored quantum kernels, and the authors mostly acknowledge the lineage. The IRIS simulations are transparent and do support the qualitative claim that SV-QSVM gives better class separation than LS-QSVM. That is the paper's real contribution: an intuitive, numerically plausible way to improve the distinguishability of quantum classifiers by using only support vectors.\n\nThe theory, however, is the weak spot. Theorem 1 assumes the kernel values are i.i.d. and then invokes the CLT to get a zero-mean normal for f(x). That is not justified, and the same mistake infects Appendix D, which is load-bearing for the multi-class readout claim. The lower bound a ≥ 1/sqrt(5M_j) is derived by modeling the class amplitude as zero-mean normal, but the amplitude is an average of nonnegative kernel overlaps. The CLT limit for such an average has positive mean, not zero. If the mean is positive, the amplitude does not shrink like 1/sqrt(M_j) and the Grover iteration count is not a worst-case guarantee. If the mean is tiny, the bound fails. Either way, the claimed guarantee does not follow. This is not a minor gap; it undermines the main multi-classifier claim. The misclassification-rate analysis in Section IV D also has inconsistent formulas, with the roles of ε and ϵ getting tangled. Lemma 2's shot count relies on a constant lower bound on λ_min that is asserted, not derived. And on the numerics side, the reported accuracy on all 150 IRIS points includes the training samples, which inflates the numbers; no code or data are shipped.\n\nNone of these flaws kill the core intuition—support vectors should improve distinguishability, and the simulations support that. But they do kill the paper's stated theoretical guarantees. The simulations are not enough to validate the general claims, especially the multi-class readout guarantee.\n\nBottom line: worth a serious referee, but with a clear request to rewrite the theory. Either derive a real bound on the class amplitude using the nonnegativity of kernels, or present the multi-class readout as a heuristic backed by numerics rather than a theorem. As is, I would not cite it in my own work, but I would want to see the revision.","headline":"A useful but overclaimed QML paper: the support-vector trial state idea is sound and the IRIS numerics are honest, but the theoretical guarantees—especially the multi-class readout bound—are not supported by the analysis.","tokens_in":17432,"tokens_out":2243,"would_cite":false,"duration_ms":23599,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that training the quantum feature map and classifying from support-vector-only trial states amplifies the decision value and lifts quantum classifiers out of the poor-distinguishability regime of least-squares quantum…","keywords":["quantum machine learning","quantum support vector machine","trainable quantum feature mapping","quantum kernel methods","amplitude amplification","multi-class classification","distinguishability","variational quantum circuits"],"falsifier":"Compute the class-wise overlap $a=\\sum_i \\alpha_i k(x_i,x)/\\sqrt{M_j}$ for held-out samples of a trained mapping; if any correctly classified sample has $a<1/\\sqrt{5M_j}$ yet the multi-class circuit still reads out correctly after the prescribed $R$ iterations, the Appendix D bound is not the operative mechanism. If every such sample has $a$ below that bound and readout fails, the paper's multi-class readout claim is falsified.","tokens_in":16449,"feed_emoji":"⚛️","tokens_out":13856,"duration_ms":109411,"temperature":0.7,"pith_summary":"Quantum kernel classifiers lose separation as the training set grows, because the decision value $f(x)=\\sum_{i=1}^{M}\\alpha_i y_i k(x_i,x)/\\sqrt{M}$ shrinks like $O(1/\\sqrt{M})$, so the sign is easily buried in measurement noise. This paper proposes two coupled fixes: a trainable quantum feature mapping $U(x,\\theta)$ whose parameters are optimized so that same-class samples cluster, and a support-vector-based quantum SVM (SV-QSVM) that builds the trial state only from support vectors, replacing the $1/\\sqrt{M}$ normalization with $1/\\sqrt{m_s}$. The paper claims this amplifies the decision value and improves both accuracy and distinguishability, and it verifies the gain numerically against the least-squares quantum SVM. It then extends the same amplification idea to a quantum iterative multi-classifier, using amplitude-amplification iteration to read out one-versus-one and one-versus-rest results with high probability. A reader should care because the bottleneck addressed—readout of a tiny inner-product sign—is the step that dominates practical quantum kernel classification.","feed_headline":"Amplify quantum classifier decisions with support-vector trial states","feed_subtitle":"Trainable maps plus support-vector states keep quantum kernels readable as datasets grow","key_machinery":"The load-bearing objects are the trainable quantum feature mapping $U(x,\\theta)=U_l(x,\\theta_l)U_{\\mathrm{ent}}\\cdots U_2(x,\\theta_2)U_{\\mathrm{ent}}U_1(x,\\theta_1)$, a data re-uploading circuit optimized by minimizing the loss $E(\\theta)=1-\\frac{1}{L}\\sum_{j=1}^{L}\\frac{1}{M_j}\\sum_{i=1}^{M_j}|\\langle\\psi(x_i^j,\\theta)|y_j\\rangle|^2$, and the partially evenly weighted trial state $|\\upsilon\\rangle=\\sum_i c_i |i-1\\rangle|\\psi(x,\\theta^*)\\rangle/\\sqrt{m_s}$ over support vectors, built from the even superposition by amplitude amplification. The training loss pulls same-class feature states toward one label vector per class, which makes the kernel matrix well-conditioned for downstream SVM training. The partial trial state is the mechanism that turns the decision-value shrinkage into an amplification: replacing the $1/\\sqrt{M}$ normalization of the least-squares quantum SVM with $1/\\sqrt{m_s}$ increases $|f(x)|$ and moves values away from zero, where the sign readout is fragile. For multi-class readout the same amplitude-amplification operator iterates $R\\in[1,(\\sqrt{5M_j}\\pi-1)/2]$ times to concentrate amplitude on the class with the largest overlap, converting a probability that starts near the uniform $1/L$ into one above $0.5$.","core_discovery":"The central claim is that the poor distinguishability of quantum kernel classifiers on large datasets is not an unavoidable feature of quantum kernels but can be engineered away. The least-squares quantum SVM treats every training sample as a support vector, so its trial state is an evenly weighted superposition over all $M$ samples and its decision value $f(x)=\\sum_i\\alpha_i y_i k(x_i,x)/\\sqrt{M}$ shrinks as $O(1/\\sqrt{M})$ by the central limit theorem. SV-QSVM instead learns the support-vector set and classifies with a partially evenly weighted trial state $|\\upsilon\\rangle=\\sum_i c_i |i-1\\rangle|\\psi(x,\\theta^*)\\rangle/\\sqrt{m_s}$, where $c_i$ marks support vectors; this raises the trial amplitude from $1/\\sqrt{M}$ to $1/\\sqrt{m_s}$ and therefore amplifies $f(x)$. Equipped with a trainable quantum feature mapping whose parameters $\\theta^*$ are optimized to cluster same-class feature states, the paper claims the resulting classifier is strictly more distinguishable than the least-squares quantum SVM and numerically achieves higher accuracy. The paper also claims that iterative amplitude amplification on the stored class overlaps makes multi-class readout reliable with few measurements, provided each class overlap is at least about $1/\\sqrt{5M_j}$.","pith_inferences":["The $O(1/\\sqrt{M})$ collapse of the evenly weighted decision value is a generic property of any quantum classifier that superposes all training samples; if the paper's diagnosis is right, the support-vector subsetting trick could be applied to other kernel-based quantum algorithms, not just SVMs.","The Appendix D bound assumes squared kernel overlaps have mean $1/3$ and behave like independent random variables; this is testable directly by histogramming $k(x_i,x)^2$ on a trained mapping, and the multi-class advantage would be on firmer ground if the same readout worked when that bound fails.","One could replace the amplitude-amplification iteration with quantum amplitude estimation to estimate each class overlap to a chosen precision, trading the threshold assumption for standard phase-estimation overhead.","The numerical comparison is on a small, well-clustered benchmark; the claimed accuracy gap should be probed on larger and noisier kernels before treating the distinguishability gain as universal."],"forward_implications":["A trained TQFM kernel matrix should outperform the untrained kernel matrix for any kernel-based classifier, because the training objective directly separates classes in feature space.","SV-QSVM's decision value scales as $1/\\sqrt{m_s}$ instead of $1/\\sqrt{M}$, so the measurement shots needed to fix the sign of $f(x)$ drop whenever support vectors are a small fraction of the dataset.","Total kernel-estimation shots of $O(M^4/\\epsilon^2)$ suffice to keep the noisy classifier within $\\epsilon$ of the noiseless one, giving a concrete resource bound for the whole SV-QSVM pipeline.","Amplitude-amplification iteration over stored class overlaps lets a multi-class label be read out with probability above $0.5$, reducing the readout burden relative to measuring every pairwise overlap separately.","The same trained feature mapping can serve as an explicit classifier, an ensemble preparation, and a kernel-matrix source, so the optimization cost is amortized across several classification strategies."],"supporting_citations":[{"why":"Supplies the quantum feature-space encoding and the inversion-test method for estimating kernel overlaps that the TQFM and SV-QSVM build on.","marker":"[16]"},{"why":"Defines the least-squares quantum SVM baseline whose evenly weighted trial state and decision value SV-QSVM is designed to beat.","marker":"[10]"},{"why":"Provides the data re-uploading strategy that the trainable quantum feature mapping layout is based on.","marker":"[22]"},{"why":"Provides amplitude amplification, used both to build the partially weighted support-vector trial state and to read out multi-class results.","marker":"[2]"},{"why":"Supplies the variational SVM formulation and fault-tolerant classification equation that SV-QSVM adapts.","marker":"[42]"},{"why":"Gives the stability bound for definite quadratic programs used in the error analysis of SV-QSVM.","marker":"[43]"},{"why":"Provides the reduced-overhead overlap estimation protocol used to train the feature mapping without expensive swap tests.","marker":"[24]"},{"why":"Serves as the unified quantum classification framework whose reported accuracy the multi-class simulation is compared with.","marker":"[31]"}],"fun_headline_variants":["Support-vector trial states amplify quantum kernel decisions","Trainable kernels and support vectors raise quantum classifier accuracy","Partially weighted trial states improve quantum SVM distinguishability","Quantum SVM with support-vector states improves classification","Amplified trial states sharpen trainable quantum classifiers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The multi-class readout advantage rests on the assumption that after training, the summed kernel overlap for the true class is at least about $1/\\sqrt{5M_j}$, which the paper derives by treating squared kernel overlaps as independent random variables with mean $1/3$; if trained kernels yield smaller overlaps, the prescribed number of amplitude-amplification iterations will not push the readout probability above $0.5$.","fun_headline_variants_meta":{"raw":{"variants":["Support-vector trial states amplify quantum kernel decisions","Trainable kernels and support vectors raise quantum classifier accuracy","Partially weighted trial states improve quantum SVM distinguishability","Quantum SVM with support-vector states improves classification","Amplified trial states sharpen trainable quantum classifiers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000779,"raw_usage":{"total_tokens":3442,"prompt_tokens":942,"completion_tokens":2500,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":2426}},"tokens_in":558,"tokens_out":2500,"duration_ms":19157,"temperature":1.0,"reasoning_tokens":2426,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:33:41.669809+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the class-wise overlap $a=\\sum_i \\alpha_i k(x_i,x)/\\sqrt{M_j}$ for held-out samples of a trained mapping; if any correctly classified sample has $a<1/\\sqrt{5M_j}$ yet the multi-class circuit still reads out correctly after the prescribed $R$ iterations, the Appendix D bound is not the operative mechanism. If every such sample has $a$ below that bound and readout fails, the paper's multi-class readout claim is falsified.","supporting_citations":[{"cited_title":"Havl´ıˇ cek, A","cited_arxiv_id":null,"evidence_quote":"Supplies the quantum feature-space encoding and the inversion-test method for estimating kernel overlaps that the TQFM and SV-QSVM build on."},{"cited_title":"P ´erez-Salinas, A","cited_arxiv_id":null,"evidence_quote":"Provides the data re-uploading strategy that the trainable quantum feature mapping layout is based on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides amplitude amplification, used both to build the partially weighted support-vector trial state and to read out multi-class results."},{"cited_title":"Xu, X.-Y","cited_arxiv_id":null,"evidence_quote":"Supplies the variational SVM formulation and fault-tolerant classification equation that SV-QSVM adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the stability bound for definite quadratic programs used in the error analysis of SV-QSVM."},{"cited_title":"Cincio, Y","cited_arxiv_id":null,"evidence_quote":"Provides the reduced-overhead overlap estimation protocol used to train the feature mapping without expensive swap tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Serves as the unified quantum classification framework whose reported accuracy the multi-class simulation is compared with."}],"review_version":1}