{"id":"3cb132ba-5c40-4b59-b635-372ef003b8db","arxiv_id":"2607.19318","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"On a unified VQE benchmark, QNBAD noise-induced attacks amplify energy error up to 8.84x, QTrojan up to 7.52x, and QDoor at most 1.37x.","lead":"This paper introduces VQE-AdvBench, a benchmark that runs seven adversarial attacks on the Variational Quantum Eigensolver over simulated IBM backends, finding that attacks corrupting zero-noise extrapolation give the largest relative energy errors. It is useful as a first standardized red-teaming setup for VQE-as-a-service, but the severity ranking rests on inconsistent baselines and is published without code or error bars.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Severity ordering is an artifact of comparing QNBAD's ZNE-clean baseline (e.g., 0.0166 on CAI-H2) with the higher single-noisy baseline (0.0770) used for other attacks; under a common baseline QNBAD falls below QTrojan on all five backends.","rationale":"The reader's weakest assumption identifies the inconsistent clean baselines as a potential cause of the ranking; our stress-test confirms and sharpens this into a decisive defect. Using the paper's own numbers, the QNBAD peak (8.84x on CAI) is inflated by a ZNE-extrapolated clean denominator that is 4.6x lower than the single-noisy baseline used for QTrojan on the same backend. Recomputing with a common baseline reduces QNBAD's CAI amplification to ~1.4x, below QTrojan's 4.48x, and QTrojan already dominates QNBAD on four of five backends under the paper's mixed baselines. Therefore the central severity ordering is not merely unverified but contradicted by the available data. This moves the verdict from CONDITIONAL to REJECT: the main claimed finding should not be accepted without a common-baseline recomputation, and the abstract's comparison of maxima across different backends further undermines the conclusion. No ad hominem is intended; the issue is purely methodological.","tokens_in":8443,"tokens_out":9979,"duration_ms":88839,"concrete_test":"Recompute all amplification factors with a single shared baseline: divide Table II QNBAD absolute errors by Table I single-noisy clean values for the same backend/molecule (e.g., CAI QF H2: 0.1038/0.0770=1.35x), and if possible re-run QTrojan/FGSM/PGD through the same ZNE pipeline used for QNBAD. If QNBAD's GEO-Mean falls below QTrojan's on every backend and the abstract's ordering reverses, the central claim is refuted.","verdict_should_be":"REJECT","load_bearing_attack":"Section V explicitly assigns different clean baselines: QTrojan/FGSM/PGD are amplified relative to a single-noisy clean estimate, while QNBAD is amplified relative to a ZNE-extrapolated clean estimate. These denominators are not commensurable—on CAI-H2, the ZNE clean is 0.0166 vs 0.0770 in Table I, a 4.6x gap. Re-expressing QNBAD's CAI absolute errors against Table I's single-noisy baseline drops QF's GEO-Mean from 8.84x to about 1.40x, QM to about 1.32x, and QS to about 0.72x—all below QTrojan's 4.48x on the same backend. The abstract's 'up to 8.84x' vs '7.52x' also compares maxima on different backends (CAI vs MON). Within each backend, QTrojan's GEO-Mean exceeds every QNBAD variant on four of five backends even using the paper's own mixed baselines; on the fifth (CAI), the reversal depends entirely on the inconsistent ZNE baseline. Thus the headline claim that noise-induced attacks are the most damaging is a baseline/comparison artifact, not a robust result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces VQE-AdvBench, a unified red-teaming benchmark for the Variational Quantum Eigensolver, and uses it to compare seven attack scenarios: the QTrojan circuit-level backdoor, the QDoor parameter backdoor, parameter-space FGSM/PGD, and three QNBAD noise-induced variants. Evaluations are performed on H2 and H3+ across five IBM Quantum fake backends under a fixed molecule–ansatz–backend–metric protocol. The central claim is a severity ordering: noise-induced attacks on the ZNE pipeline are the most damaging (up to 8.84× error amplification), followed by QTrojan (7.52×), while QDoor is the least effective (up to 1.37×). The paper also sketches defenses, and explicitly leaves their empirical evaluation to future work.","tokens_in":8790,"tokens_out":7404,"duration_ms":78714,"significance":"A common red-teaming protocol for variational quantum algorithms is timely and practically useful; the paper correctly identifies that prior attack evaluations are heterogeneous and incomparable. The manuscript is transparent about its scope (small molecules, fixed ansatz, fake backends, unevaluated defenses), and the protocol is a valuable starting point. However, the headline severity ordering is not supported as printed because the two result tables use different clean baselines. Re-baselining the QNBAD numbers against the single-noisy clean baseline used for the other attacks reverses the key comparison on the backend where the paper claims the largest QNBAD effect. The benchmark infrastructure has merit and the paper could be made sound with a corrected analysis, but the central conclusion currently rests on an artifact.","major_comments":[{"comment":"The severity ranking mixes incomparable baselines. Table I reports QTrojan/FGSM/PGD amplification relative to a single-noisy clean estimate, while Table II reports QNBAD amplification relative to a ZNE-extrapolated clean estimate. On CAI-H2 the clean baseline changes from 0.0770 (Table I) to 0.0166 (Table II); on CAI-H3+ from 0.1980 to 0.0231. Re-expressing the Table II CAI absolute errors against the Table I clean baseline drops QF's GEO-Mean from 8.84× to about 1.40×, QM to about 1.32×, and QS to about 0.72×, all below QTrojan's 4.48× on the same backend. Moreover, even using the paper's own mixed baselines, QTrojan exceeds every QNBAD variant on four of the five backends (MON, GUA, ALM, AUC); the CAI reversal is entirely attributable to the inconsistent ZNE baseline. The abstract's 'up to 8.84×' vs '7.52×' also compares maxima on different backends. The claim that noise-induced attack","section":"Section V, Tables I and II, Abstract"},{"comment":"The experimental section does not report the number of shots, random seeds, or repeated execution counts. All E_abs values are point estimates of stochastic quantities on noisy backends, and ZNE extrapolation amplifies sampling noise. Without confidence intervals or at least explicit shots/seeds, the approximate equality between the top amplification factors (e.g., 8.84× vs 7.52× on different backends) is not enough to establish a robust ordering. In addition, all results use Qiskit fake backends; no real-device validation is provided, while the threat model and conclusions are phrased in terms of real VQE-as-a-service pipelines. The authors should either add real-hardware experiments or substantially qualify the claims to noise-calibrated simulation.","section":"Section IV, ZNE Setting and Evaluation Metric"},{"comment":"Three of the four attack classes that drive the ranking — QTrojan, QDoor, and the QNBAD family — are prior works of this paper's co-authors, and the paper does not identify whether independent third-party implementations were used. This is not definitional circularity, but it creates an implementation-quality confound: a benchmark whose contribution is a head-to-head severity comparison should release the code/data and document how faithfully each attack is instantiated, or include independent implementations. Otherwise the ranking may partly reflect implementation tuning rather than inherent attack effectiveness. This concern should be addressed in a revised version.","section":"Sections III and VIII; References [12,13,16]"}],"minor_comments":[{"comment":"The severity numbers should state the specific baseline and backend. The current phrasing 'up to 8.84×' vs '7.52×' is misleading because the amplification factors are computed relative to different clean baselines.","section":"Abstract and Section VIII"},{"comment":"The CAI clean values in Table II differ by factors of 4–8 from Table I (0.0166 vs 0.0770 for H2; 0.0231 vs 0.1980 for H3+). If this is a deliberate effect of the ZNE baseline, it should be explained; if it is a typographical error, it must be corrected. Either way, the discrepancy needs to be addressed explicitly.","section":"Table II, CAI rows"},{"comment":"The text says 'QTrojan, FGSM, and PGD all peak on MON and bottom out on GUA.' This is not consistent with Table I: QTrojan on H3+ peaks on ALM (0.6424) and PGD on H3+ peaks on AUC (0.6630). The statement should be restricted to the GEO-Mean or H2 columns.","section":"Section V.A"},{"comment":"Reference [16] (QNBAD) has no venue, arXiv identifier, or publication details. As a central attack family, it needs a complete citation.","section":"References"},{"comment":"The taxonomy labels QTrojan as gray-box and QDoor as white-box, but an adversary who can insert pre- and post-encoding layers into the circuit has full circuit-level access. The distinction should be justified.","section":"Section III"},{"comment":"Only one perturbation budget (ε = 0.25 rad) is reported. Since FGSM and PGD are compared across adversaries, a short sensitivity study over ε would make the relative ranking more informative.","section":"Section IV, FGSM/PGD experimental setup"}],"recommendation":"major_revision","confidential_remarks":"The benchmark protocol is a useful contribution, but the headline result is not reliable as printed. The baseline mismatch is the most serious issue: it directly determines whether QNBAD or QTrojan is reported as the top threat. The authors should re-analyze all attacks against a common baseline and, if the ordering changes, reframe the paper as a benchmark/protocol contribution rather than a severity ranking. I would also strongly encourage them to add shot/seed statistics and code release before any resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: the paper builds a genuinely useful scaffold—a unified red-teaming protocol for VQE with seven attacks and a clear taxonomy—but its central claim, that noise-induced ZNE attacks are the most damaging (up to 8.84x), does not survive a common baseline. The ranking is a comparison artifact.\n\nWhat's actually new: this is the first benchmark of its kind for VQE; the parameter-space FGSM/PGD adaptations are new; the fixed protocol across H2/H3+ and five fake backends is a step forward; and the tables are informative. The paper is also honest about scope: small molecules, shallow circuits, trusted compiler, and it explicitly says defenses are left to future work. Credit where earned—this is a solid systematization effort, not a careless one.\n\nThe big soft spot is Section V. QTrojan/FGSM/PGD are amplified relative to a single-noisy clean baseline, while QNBAD is amplified relative to a ZNE-extrapolated clean baseline. On CAI, those two 'clean' values differ by a factor of 4.6 (0.0770 vs 0.0166). Recomputing QNBAD's amplifications against the single-noisy baseline drops QF's geo-mean on CAI to about 1.40x, QM to 1.32x, and QS to 0.72x—all well below QTrojan's 4.48x on the same backend. The abstract's '8.84x vs 7.52x' also compares maxima on different backends (CAI vs MON). So the headline ordering is not robust; it is a baseline mismatch. The paper does disclose the difference, but then proceeds to present the ordering as a clear finding without flagging the incommensurability.\n\nMinor but real: no shots, seeds, or error bars reported; no code/data release; fake-backend simulations only, with no real-hardware validation; and three of the seven attack families are the authors' own implementations, so the benchmark partly tests their own code. None of these are fatal for a benchmark paper, but they compound the baseline issue.\n\nWho this is for: people working on VQE security or error-mitigation robustness. They'll get a starting point and a clear list of what needs fixing. A serious referee should see this because the protocol design and attack systematization are worth engaging with, and the flaws are correctable.\n\nRecommendation: send to peer review, with the expectation of major revision—the authors should recompute all amplifications against one shared clean baseline, report statistics, and release code.","headline":"Useful first benchmark for VQE red-teaming, but the headline severity ranking is an artifact of comparing amplifications against two different clean baselines and should be treated as provisional.","tokens_in":9286,"tokens_out":2462,"would_cite":false,"duration_ms":23284,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Lx"],"model":"deepseek-v4-flash","headline":"Under a unified red-teaming benchmark, VQE's most damaging adversarial attacks are those that corrupt the Zero-Noise Extrapolation (ZNE) error-mitigation pipeline (up to 8.84x error amplification), followed by circuit-level backdoors (7.52x","keywords":["Variational Quantum Eigensolver","Adversarial Robustness","Red-Teaming","ZNE","Backdoor Attacks","Error Mitigation","Quantum Machine Learning Security","Benchmark"],"falsifier":"Run the same seven-attack protocol on a real IBM quantum device with H2 and H3+ using the same ansatz and ZNE settings; if QNBAD FreeDrift's geometric-mean amplification does not exceed QTrojan's, or if QDoor exceeds PGD, the reported severity ordering is falsified.","tokens_in":8293,"feed_emoji":"⚛️","tokens_out":4424,"duration_ms":41606,"temperature":0.7,"pith_summary":"The paper introduces VQE-AdvBench, the first unified red-teaming benchmark for the Variational Quantum Eigensolver, and uses it to compare seven adversarial attacks under a fixed molecule–ansatz–backend–metric protocol. Its central finding is a severity ordering: attacks that manipulate the Zero-Noise Extrapolation (ZNE) error-mitigation pipeline are the most damaging, amplifying energy errors by up to 8.84x; the QTrojan circuit-level backdoor reaches 7.52x; gradient-based perturbations (FGSM/PGD) sit in between; and the QDoor parameter-level backdoor is nearly ineffective (up to 1.37x) on shallow circuits. This matters because VQE is a leading candidate for quantum chemistry and drug-discovery workloads, and the result indicates that the main security exposure in cloud-based VQE is the error-mitigation stage, not the trained parameters themselves.","feed_headline":"Noise-pipeline attacks top VQE threat list at 8.84x","feed_subtitle":"Unified red-team benchmark on H2 and H3+ shows error-mitigation, not parameters, is VQE's weak spot","key_machinery":"The benchmark itself is the key object: VQE-AdvBench fixes a molecule (H2, H3+), a hardware-efficient ansatz (efficient_su2, 24 or 48 parameters), five noise-calibrated IBM fake backends, and a metric (absolute energy error relative to a clean baseline). This fixed configuration is what makes cross-attack comparison possible. For the noise-induced attacks, the target is the Zero-Noise Extrapolation (ZNE) pipeline, which extrapolates noisy expectation values at scaling factors 1..6 to estimate the zero-noise limit; an adversary who shapes the noise trajectory can corrupt the extrapolated result.","core_discovery":"The paper's central claim is that, under a common protocol, the relative danger of known VQE attacks is measurable and reveals a clear hierarchy. The most effective attacks are the QNBAD noise-induced backdoors, which poison the parameters so that ZNE extrapolation returns far-from-correct energies without any change to the executed circuit (geometric-mean amplification up to 8.84x on H2/H3+). Next is QTrojan, a circuit-level backdoor that inserts concealed RX/RY layers and amplifies error up to 7.52x. Parameter-space gradient attacks (PGD, FGSM) achieve 2–6x, and QDoor, a parameter-level backdoor based on approximate synthesis, stays near the clean baseline (0.94–1.37x), which the authors a","pith_inferences":["Because ZNE is a widely used error-mitigation tool beyond VQE, the QNBAD attack class may transfer to other ZNE-based quantum algorithms, making error-mitigation pipelines a general security concern rather than a VQE-specific one.","The severity ordering may be an artifact of the shallow-circuit regime; QDoor's weakness could invert on deeper, more expressive circuits used for larger molecules, so the ranking should not be extrapolated without further testing.","A cheap testable mitigation is to vary the ZNE polynomial degree or noise-scaling schedule; if QNBAD effectiveness collapses under a different extrapolation fit, defenders gain a low-cost hardening option.","The lack of real-hardware validation means the amplification factors are simulation-based; real-device noise drift could change both absolute and relative values, so the ranking should be re-run on actual hardware before operational decisions."],"forward_implications":["Defenders should treat ZNE and other error-mitigation stages as a critical attack surface, since a compromised service can corrupt results even without modifying the executed circuit.","Circuit-level backdoors like QTrojan remain a top threat (7.52x), so validation checks on state-preparation components are warranted.","PGD is consistently more damaging than FGSM under the same perturbation budget, so iterative attacks should be a standard part of robustness testing for variational algorithms.","Parameter-level backdoors like QDoor appear low-risk on shallow ansatze, but this may change as circuits deepen for larger molecules.","The benchmark protocol can be extended to other variational algorithms (e.g., QAOA, VQD) and larger molecules, providing a common baseline for future red-teaming efforts."],"fun_headline_variants":["VQE noise-pipeline attacks top threat at 8.84x","Noise backdoors hit VQE hardest: 8.84x error amplification","Red-teaming VQE ranks noise injection as top attack","VQE-AdvBench: ZNE poisoning amplifies error 8.84x","QNBAD attacks dominate VQE threats, up to 8.84x"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The severity ranking assumes that noise-calibrated simulated 'fake' backends replicate real-device behavior closely enough for both ZNE extrapolation and adversarial perturbations, and that comparing each attack to its own clean baseline yields fair amplification factors.","fun_headline_variants_meta":{"raw":{"variants":["VQE noise-pipeline attacks top threat at 8.84x","Noise backdoors hit VQE hardest: 8.84x error amplification","Red-teaming VQE ranks noise injection as top attack","VQE-AdvBench: ZNE poisoning amplifies error 8.84x","QNBAD attacks dominate VQE threats, up to 8.84x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000716,"raw_usage":{"total_tokens":3149,"prompt_tokens":931,"completion_tokens":2218,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":675,"completion_tokens_details":{"reasoning_tokens":2128}},"tokens_in":675,"tokens_out":2218,"duration_ms":14058,"temperature":1.0,"reasoning_tokens":2128,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T12:45:13.025883+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same seven-attack protocol on a real IBM quantum device with H2 and H3+ using the same ansatz and ZNE settings; if QNBAD FreeDrift's geometric-mean amplification does not exceed QTrojan's, or if QDoor exceeds PGD, the reported severity ordering is falsified.","supporting_citations":[],"review_version":1}