{"id":"c2854a4b-5ee5-4a11-82dc-6d86198176ec","arxiv_id":"2509.11190","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"On toy classification tasks, variational quantum circuits show weak lottery ticket behavior: pruned circuits retain full accuracy with roughly 26% to 51% of their parameters, while strong-ticket evidence is limited to one binary dataset.","lead":"This paper tests whether the lottery ticket hypothesis, the idea that a small pruned subnetwork can match a large one, works for variational quantum circuits. It reports that on small Iris and Wine classification tasks, pruned quantum circuits keep full accuracy with as little as 26% of their parameters, and an evolutionary search found a binary circuit with 45% of weights that still reaches 100%.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Wine MVQC 'winning ticket' rests on an unverified 45% unpruned baseline; with no convergence criterion or error bars, pruning gains may be a fixed-budget artifact.","rationale":"The weakest link is not internal inconsistency but an unverified comparison. The paper's protocols (Algorithms 1–3) are clearly described, and the code uses standard libraries; however, none of that substitutes for evidence that the unpruned Wine MVQC is converged. The dramatic 45→80 result is the only place where pruning produces a large improvement, and it drives both the abstract's 26.0% claim (Iris, where baseline is 97.5%) and the BP-mitigation conclusion. I agree with the reader that the unpruned baseline is the load-bearing assumption. I would not reject outright: the Iris and simplified-Iris results are more plausible and the method itself is reasonable. But the Wine claim and the barren-plateau interpretation should be conditional on the proposed convergence control and on reporting seed-level variance. This does not move the reader's verdict, hence UNCHANGED.","tokens_in":14916,"tokens_out":6938,"duration_ms":87103,"concrete_test":"Re-run the Wine MVQC experiment under the same hyperparameters and seeds 0–9, but train the unpruned model until convergence (e.g., early stopping on validation loss with patience 20, up to 200 epochs) and record final validation accuracy. If converged unpruned accuracy reaches ~80%, then the 45% baseline in Fig. 5a is an under-training artifact and the Wine weak-LTH claim collapses. If it remains near 45%, the concern is not supported; additionally, training randomly pruned masks at 32.7% remaining weights under identical budget would clarify whether the learned mask is necessary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's abstract and §5 claim weak LTH holds for VQCs, with the strongest evidence the Wine MVQC result (§4.1.4, Fig. 5a): unpruned accuracy 45%, pruned to 32.7% weights 80%, and even 4.3% surpasses unpruned. This is a winning ticket only if the unpruned model is a fair, converged baseline. The protocol appears to stop every run at 40 epochs (Figs. 2–6) with no convergence criterion, no error bars or seed-level spread, and no loss/gradient diagnostics in §3.7/§4. A flat 45% curve is consistent with the paper's barren-plateau interpretation but equally consistent with an undertrained or poorly optimized model. If the baseline is undertrained, a smaller sparse circuit can beat it merely by converging faster within the same 40-epoch budget; any random mask at 32.7% might then also reach 80%, so the result would be a training-budget artifact, not evidence that a learned mask identifies a special subnetwork. The same baseline issue undermines the strong-LTH Wine comparison (§4.2.3) and the BP-mitigation conclusion in §5, which is asserted without measuring gradients.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper tests the weak and strong lottery ticket hypotheses for variational quantum circuits (VQCs) on Iris and Wine classification tasks. For the weak LTH, it applies iterative and one-shot magnitude pruning with reset-to-initial retraining and reports that pruned VQCs can match or exceed full-circuit accuracy, e.g., a multi-class VQC retaining 26.0% of its parameters on Iris, and a Wine MVQC reaching 80% accuracy at 32.7% remaining weights versus 45% for the unpruned model. For the strong LTH, it uses an evolutionary algorithm to learn pruning masks without training, reporting a binary VQC winning ticket at approximately 45--48% remaining weights on simplified Iris. The paper concludes that the weak LTH holds for VQCs and that pruning may mitigate barren plateaus.","tokens_in":15264,"tokens_out":7426,"duration_ms":87163,"significance":"If established, this would be one of the first systematic demonstrations of the lottery ticket hypothesis for VQCs and a concrete step toward sparse, trainable variational circuits. The paper has strengths: it documents the pruning algorithms, reports seeds 0 through 9, performs hyperparameter optimization, and includes full pruning curves in the appendix. However, the headline quantitative claims are not yet supported: there are no error bars, no formal winning-ticket criterion, no random-mask baseline for the strong LTH, and no gradient diagnostics for the barren-plateau interpretation. The contribution is potentially valuable but currently under-supported.","major_comments":[{"comment":"The Wine MVQC result is the paper's strongest evidence that pruning improves over the full model (45% vs 80% accuracy, a claimed 35% improvement). The comparison rests on the unpruned run being a converged, fairly trained baseline. The protocol fixes 40 epochs for all runs, reports only aggregate curves despite seeds 0 through 9, and provides no loss or gradient convergence diagnostics. A flat 45% curve is equally consistent with an undertrained or poorly optimized model as with a barren plateau. If the baseline is undertrained, a smaller circuit can beat it merely by converging faster within the 40-epoch budget, making the result a fixed-budget artifact rather than a winning ticket. Please report per-seed curves or mean±std, add a convergence criterion, and measure gradient variance for pruned versus unpruned models before claiming BP mitigation.","section":"§4.1.4 / Fig. 5a / §5"},{"comment":"The 'winning ticket' level is never formally defined. It appears sometimes to mean 'accuracy equals or exceeds the unpruned accuracy' (Iris, simplified Iris), sometimes 'stabilizes at 94--96%' (§4.1.5), and sometimes 'surpasses unpruned accuracy' (§4.1.4). There is no stated tolerance, no aggregation rule across the ten seeds, and no statement about which accuracy (training or validation) is being used, even though §3.7 says both were tracked. Without an explicit criterion, Table 3 cannot be reproduced and the headline '26.0%' has no precise meaning. Please give the exact condition used to select the entries in Table 3.","section":"§3.7 / Table 3"},{"comment":"The strong LTH results are not compared with random pruning masks. On simplified Iris the unpruned models already reach 100% accuracy, so an EA-evolved mask reaching 100% with about 48% of weights does not demonstrate that the mask is specially structured: a random mask at the same sparsity may do equally well. Without a random-mask baseline, the abstract's strong-LTH claim ('a binary VQC achieving 100% accuracy with only 45% of the weights') is unsupported. Please add random-pruning curves at comparable remaining-weight levels.","section":"§4.2 / Figs. 7--10"},{"comment":"The abstract and conclusion state that LTH 'may mitigate barren plateaus' and that the Wine experiment shows LTH 'can overcome' BP. This claim is not measured anywhere: the paper does not report gradient norms, gradient variance, or loss-landscape diagnostics for full or pruned circuits. The accuracy gain alone cannot distinguish BP relief from faster optimization of an easier model. If the authors wish to retain the BP contribution, they should add direct gradient diagnostics; otherwise the claim should be explicitly downgraded to a conjecture.","section":"§5 / §2.3"},{"comment":"The results for Simplified Wine are internally inconsistent. Table 3 appears to list '51.3% / 51.4%' across the VQC columns, while §4.1.5 says both BVQC and MVQC maintain accuracy only up to roughly 40% remaining weights, and §5 states that 'no winning ticket was found for the MVQC on the simplified Wine dataset' and that only the BVQC found one at 51.3%. Please clarify which model has a ticket at which level and correct the table or text.","section":"Table 3 vs §4.1.5 / §5"}],"minor_comments":[{"comment":"The reset-to-initial step is only implicit: the text says 'set the random seed' and 'create a new model' but does not explicitly state that this reproduces the original initialization. State this explicitly for clarity.","section":"§3.1.1 / §3.1.2"},{"comment":"The abstract and conclusion state '45% of the weights' for the strong-LTH BVQC winning ticket, while §4.2.2 says 'about 48%'. Make the number consistent.","section":"Abstract / §4.2.2"},{"comment":"The entry '51.3% / 51.4%' is ambiguous due to table formatting; it should be split into two unambiguous cells.","section":"Table 3"},{"comment":"The Wine dataset reference is dated '1936'; the correct UCI citation is from the early 1990s. Please correct.","section":"References"},{"comment":"The EA description does not specify population size, mutation rate, crossover operator, migration rate, or the number of generations used in the experiments. Please provide these details for reproducibility.","section":"Algorithm 3"},{"comment":"The paper says accuracy on 'training and validation sets' was tracked, but no data split is described. For small datasets, clarifying how validation accuracy was used to select tickets is important for interpreting the results.","section":"§3.7"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is likely within scope for a QML venue, but the barren-plateau claim is currently stronger than the evidence. I would be comfortable with a major revision if the authors add error bars, a formal winning-ticket definition, random-mask baselines, and either gradient diagnostics or a softened BP conclusion. The central weak-LTH claim is plausible, but the Wine MVQC baseline issue must be addressed before the quantitative headline is accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about arXiv:2509.11190. First, the core observation—that reset-to-init magnitude pruning can produce sparse VQCs matching full-circuit accuracy on Iris and simplified Iris—is real and is the paper's main contribution. Second, the much stronger claim that pruning overcomes barren plateaus on Wine, with pruned accuracy jumping to 80% from a 45% unpruned baseline, is not backed by the evidence as reported.\n\nWhat's genuinely new: transferring the Frankle–Carbin weak-LTH protocol to VQCs, testing both iterative and one-shot pruning, and adding an evolutionary search for strong tickets. The study is careful in its own terms: multiple seeds, hyperparameter search with Optuna, a classical NN baseline with matched parameter count, and clear algorithms. The weak-LTH results on Iris and simplified Iris are credible—unpruned models hit 97.5–100%, and pruning to 26–33% holds that accuracy. That is a legitimate if unsurprising transfer result.\n\nThe soft spots are concentrated in the Wine MVQC result and the strong-LTH section. The 45% unpruned accuracy is barely above chance for a 3-class problem; without a convergence criterion or seed-level error bars, the 80% from the pruned model may just be a fixed-40-epoch artifact where the smaller circuit converges faster. The paper never measures gradients, so the barren-plateau mitigation language in §4.1.4 and §5 is asserted, not demonstrated. The strong-LTH EA lacks a random-mask baseline—we can't tell if the evolutionary search is finding something special or just any mask that does well. And the 'winning ticket' threshold is never formally defined; Table 3 just lists percentages where performance 'matches,' with no tolerance.\n\nThe citation pattern is fine: the custom EA is drawn from the authors' own prior work, but it's used as a tool, not as evidence. No load-bearing circularity.\n\nWho is this for? Practitioners in quantum machine learning who want to know whether sparse VQCs can be trained cheaply, and researchers working on pruning for quantum circuits. It deserves a serious referee, but it needs revision before it can be trusted: add error bars, define the matching criterion, add a random-mask baseline, and either measure gradients or drop the BP claim. I'd send it to review with a request for major revision.","headline":"A credible weak-LTH transfer for small VQCs, but the barren-plateau win is overclaimed and the strong-LTH evidence needs a random-mask baseline.","tokens_in":15751,"tokens_out":2700,"would_cite":true,"duration_ms":29144,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sparse variational quantum circuits can preserve or improve accuracy after magnitude pruning, keeping as few as 26.0% of original parameters.","keywords":["lottery ticket hypothesis","variational quantum circuits","parameter pruning","barren plateaus","quantum machine learning","evolutionary algorithm","classification"],"falsifier":"Rerun the Wine multi-class circuit with the same architecture but longer training and several seeds while recording gradient magnitudes; if the unpruned model reaches 80% accuracy or any seed removes the pruned model's advantage, the winning-ticket claim on Wine collapses. A simpler check: repeat the simplified-Iris BVQC strong-LTH search with 20 random seeds and require all seeds to reach 100% at 45% weights; the current result does not demonstrate stability.","tokens_in":14840,"feed_emoji":"⚛️","tokens_out":7995,"duration_ms":83403,"temperature":0.7,"pith_summary":"This paper asks whether the classical lottery ticket hypothesis transfers to variational quantum circuits: within a heavily parameterized circuit, can a sparse subcircuit be trained in isolation to match or beat the full circuit? The authors report yes for the weak version: after training, pruning the smallest rotation parameters, resetting the survivors to their initial values, and retraining, winning tickets stay accurate down to about 26.0% of the original parameters. They also report one strong-LTH success: a binary circuit reaches 100% accuracy with 45% of weights using an evolutionary search over masks, but no strong tickets on the harder Wine data. The central reported result is that pruning can lift a circuit stuck at low accuracy: on Wine, a multi-class circuit reaches 80% accuracy at 32.7% weights while its unpruned version reaches only 45%. If correct, sparse pruning gives a concrete route to smaller, more trainable quantum classifiers.","feed_headline":"Sparse quantum circuits match full circuits with 26% of parameters","feed_subtitle":"Pruned quantum circuits keep full accuracy with a quarter of the parameters; on Wine a pruned model beats the full circuit by 35 points.","key_machinery":"The load-bearing mechanism is reset-to-initialization magnitude pruning. A full circuit is trained; its rotation parameters are ranked by absolute value; a fraction is set to zero; the surviving parameters are reset to their pre-training values; and the masked circuit is retrained. Iterative pruning repeats this while removing 20% of remaining weights per cycle, and one-shot pruning applies a single target ratio. The winning ticket is the sparse set of parameters plus the mask that isolates a trainable subcircuit. For the strong variant, the same parameter mask is searched by an evolutionary algorithm with crossover, mutation, and migration, with no training step, so the mask alone must enco","core_discovery":"The paper's central claim is that the weak lottery ticket hypothesis holds for variational quantum circuits on the small classification tasks tested. It identifies winning tickets at 26.0% remaining parameters for the multi-class circuit on Iris and simplified Iris, and at 33.3% for the binary circuit on simplified Iris, with accuracy matching the full circuits. It further claims that on the full Wine dataset, pruning the multi-class circuit to 32.7% of its weights raises accuracy from 45% to 80%, a 35-percentage-point improvement the authors interpret as evidence that pruning can mitigate barren-plateau-like optimization failure when the unpruned circuit is sufficiently overparameterized. F","pith_inferences":["A testable extension of the Wine result: train the unpruned multi-class circuit for substantially more epochs or with a learning-rate schedule; if it reaches 80% without pruning, the reported gain is an artifact of a weak baseline, not evidence that pruning mitigates barren plateaus.","Magnitude pruning might work by reducing the effective dimensionality of the variational search; measuring gradient variance before and after pruning would test whether the sparse circuit is genuinely easier to optimize.","If pruned gates are physically removed rather than masked, the same winning tickets could translate into shallower circuits on real hardware, where shorter depth directly reduces noise and gradient decay; the paper simulates circuits and does not test this translation.","The evolutionary search's failure on Wine suggests that strong-LTH masks may need at least a small amount of training signal; a hybrid that uses one short training run to seed the mask could close the gap."],"forward_implications":["Sparse variational circuits can match full accuracy at roughly one-quarter to one-third of the original parameters on small classification tasks, lowering simulation and hardware load.","If the Wine result generalizes, pruning can act as a rescue operation for circuits stuck at low accuracy, not just a compression technique.","The strong-LTH success on simplified Iris shows that training-free mask search can yield a functional subcircuit for easy problems, but it does not replace training on harder data.","Because iterative and one-shot pruning produced equivalent masks on these circuits, one training pass may suffice to find the ticket on small problems, skipping repeated train-prune cycles."],"fun_headline_variants":["Quantum circuits have lottery tickets: 26% parameters, same accuracy","Pruned quantum circuit beats full one by 35 points on Wine data","Weak lottery ticket hypothesis holds for variational quantum circuits","Quantum winning ticket: binary VQC hits 100% with 45% weights","Sparse quantum circuits match or beat full circuits at fraction of size"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the unpruned circuit is a fair baseline: in the Wine experiment (Section 4.1.4), the unpruned multi-class circuit sits at 45% accuracy, and the paper reports no convergence criterion, no error bars, and no gradient measurements to rule out underfitting before pruning gets credit for the jump to 80%.","fun_headline_variants_meta":{"raw":{"variants":["Quantum circuits have lottery tickets: 26% parameters, same accuracy","Pruned quantum circuit beats full one by 35 points on Wine data","Weak lottery ticket hypothesis holds for variational quantum circuits","Quantum winning ticket: binary VQC hits 100% with 45% weights","Sparse quantum circuits match or beat full circuits at fraction of size"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000206,"raw_usage":{"total_tokens":1235,"prompt_tokens":751,"completion_tokens":484,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":393}},"tokens_in":495,"tokens_out":484,"duration_ms":5803,"temperature":1.0,"reasoning_tokens":393,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T16:57:37.642637+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the Wine multi-class circuit with the same architecture but longer training and several seeds while recording gradient magnitudes; if the unpruned model reaches 80% accuracy or any seed removes the pruned model's advantage, the winning-ticket claim on Wine collapses. A simpler check: repeat the simplified-Iris BVQC strong-LTH search with 20 random seeds and require all seeds to reach 100% at 45% weights; the current result does not demonstrate stability.","supporting_citations":[],"review_version":1}