{"id":"e6f22587-2725-4e72-b1f5-0163b95ad454","arxiv_id":"2411.16530","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Distributing a quantum circuit's shots across multiple QPUs and merging with reliability-based weights reduces variance and worst-case error compared with single-QPU execution, at the cost of not beating the best QPU.","lead":"This paper proposes splitting the shots of a single quantum circuit across several quantum computers and then merging the results. The authors show this can stabilize outputs and protect against picking a bad machine, though it rarely beats the best single machine.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Calibration-transfer premise is load-bearing: weights from 5-qubit random circuits are applied to 8-qubit structured tasks; GHZ reversal (Sec. III B) shows ranking can invert, so the claimed advantage of Hellinger/MISE over uniform is not established.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing concern: the calibration transfer premise. I agree with that identification, and the paper itself provides in-text evidence of transfer failure in the GHZ results of Sec. III B, as well as an explicit caveat in Sec. II C.2. My stress-test adds two concretizations: the calibration is performed only on 5-qubit random circuits while production includes 8-qubit structured circuits, and the aggregate claims are made without confidence intervals or repeated-run statistics. Both factors make the apparent advantage of Hellinger/MISE policies fragile. Credit is due where the paper is careful: the framework is formalized, the MISE and Hellinger derivations are given, the authors explicitly state that the strategies 'generally do not exceed the best individual QPU results,' and the conclusions in the abstract and Sec. III C are more measured than the Sec. I wording. Still, the central 'more reliable' claim is supported only as a worst-case/variance-reduction statement, and that support does not depend on the calibration-informed policies. Therefore the conditional verdict stands, and the requested revision should require either matched-circuit calibration evidence or a re-scoped claim that removes the calibration advantage from the central contribution.","tokens_in":43,"tokens_out":8222,"duration_ms":208876,"concrete_test":"Re-run the production experiments with calibration performed separately for each circuit family and qubit count, e.g., calibrate unreliability weights on 10 random 8-qubit GHZ circuits with varying parameters, then compare Hellinger-MISE split/merge against uniform on held-out GHZ instances. Report paired differences and bootstrap 95% confidence intervals per circuit family and per number of QPUs. If informed policies do not beat uniform within confidence intervals on held-out structured circuits, the calibration-transfer premise fails and the paper should be re-scoped to uniform shot-wise averaging only.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that shot-wise distribution produces 'final output distributions more reliable than performing the whole computation on a single QPU' (Sec. I). The evidence for the non-uniform policies rests on the calibration ranking of Sec. III A, which is built from 10 Haar-random 5-qubit circuits (Sec. II C.2 and Sec. III A). The production evaluation, however, includes 8-qubit structured circuits such as GHZ, Grover, VQE, and QNN from MQT Bench. The authors explicitly acknowledge that 'the performance at calibration on a fixed set of random circuits does not necessarily reflect the performance observed on specific tasks' (Sec. II C.2), and Sec. III B reports exactly this failure: for GHZ, the Hellinger/MISE strategies show the opposite trend and underperform uniform. Because Figs. 4-7 report extremal summaries without confidence intervals, repeated runs, or paired comparisons, the claims that informed policies 'improve' uniform results and that the maximum error 'consistently decreases' are not statistically distinguishable from noise in a small benchmark suite. If calibration cannot be transferred to the target circuit class, the remaining support for the central claim is the variance-reduction property of uniform merging, which is a mathematical consequence of averaging and does not require the proposed calibration machinery.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Bisicchia et al. propose a 'shot-wise' framework in which the shots of a single quantum circuit are split among several heterogeneous QPUs, executed, and then merged into one output distribution. The framework includes a calibration stage that ranks QPUs by unreliability measured on random benchmark circuits, customizable split/merge policies (uniform, Hellinger, MISE), and an incremental execution/update loop. The experiments use MQT Bench circuits on IBM and IonQ simulators and compare nine split/merge combinations against single-QPU baselines. The reported findings are that split-merged results are robust and track the average baseline, that the worst-case error decreases as more QPUs are used, and that calibrated policies sometimes improve on uniform allocation, with the GHZ circuit as an acknowledged exception.","tokens_in":20910,"tokens_out":5298,"duration_ms":47364,"significance":"Shot-wise distribution is a plausible and useful robustness layer for heterogeneous NISQ backends: it reduces worst-case error and variance and is orthogonal to circuit cutting and error mitigation. The formalization of Hellinger and MISE policies with bias correction and the public dataset are assets. However, the accuracy advantage over the best single QPU claimed in the Introduction and Abstract is not supported; the empirical evidence supports only worst-case robustness and closeness to the average baseline. The calibration transfer from random 5-qubit circuits to structured 8-qubit tasks is acknowledged by the authors as unreliable, and the GHZ results illustrate a failure of the calibrated policies. Therefore the central claim as stated needs to be restricted and the statistical evidence strengthened.","major_comments":[{"comment":"The central claim of Sec. I that distributing and merging shots 'produces final output distributions more reliable than performing the whole computation on a single QPU' is broader than the evidence reported in Sec. III B. The text at the start of Sec. III B states that split-merged results 'never improve the best baseline' and are 'compatible with the average between the different baselines', and Fig. 5 shows the maximum error decreasing while the minimum error rises. That is a worst-case robustness and variance-reduction statement, not a general accuracy or reliability improvement over every single-QPU run; the abstract and conclusion should be rephrased accordingly.","section":"Sec. I and Sec. III B"},{"comment":"The calibration-transfer premise is load-bearing but not validated. The unreliability coefficients used by the Hellinger and MISE split/merge policies are computed on 10 Haar-random 5-qubit circuits (Sec. III A, Fig. 3), while the production evaluation includes structured 5- and 8-qubit circuits such as GHZ, Grover, VQE and QNN (Sec. III B). The authors explicitly write in Sec. II C.2 that 'the performance at calibration on a fixed set of random circuits does not necessarily reflect the performance observed on specific tasks', and the GHZ panels in Sec. III B show the opposite ranking, with Hellinger/MISE underperforming uniform. Consequently, the reported improvements of Hellinger/MISE over uniform in the other circuits cannot be attributed to the calibration mechanism without either per-circuit calibration or a demonstrated transferability criterion; as it stands, the evidence is compatible with calibration being unnecessary or even harmful.","section":"Sec. II C.2, Sec. III A, Sec. III B"},{"comment":"The quantitative claims about 'consistently decreases' maximum error and about policy rankings are made from extremal summaries without confidence intervals, repeated runs, or paired statistical comparisons. Figures 5-7 plot means and data ranges as a function of number of QPUs, but no standard errors, confidence intervals, or per-policy hypothesis tests are reported, and the number of independent executions per circuit/policy is not specified. On a small benchmark suite with only six circuit types, such as those in Fig. 4, the observed differences between Hellinger/MISE and uniform are not statistically distinguishable from shot noise and QPU drift; the authors should provide error bars or resampling-based intervals and pairwise tests for the key comparisons.","section":"Sec. III B, Figs. 5-7"},{"comment":"The sentence 'either splitting or merging using the Hellinger or MISE strategy improves the results of just uniformly splitting and naively merging according to the uniform strategies alone' is contradicted by the GHZ case acknowledged in the same paragraph and is not supported by any quantitative summary. A per-policy table of median and best/worst Hellinger distances across circuits, with differences relative to uniform-uniform, would make the claim checkable.","section":"Sec. III B, Fig. 4"}],"minor_comments":[{"comment":"The Abstract and Conclusion V phrase the outcome as 'often outperforming single QPU runs' and 'often superior to individual QPUs', while Sec. III B says the split-merged results 'never improve the best baseline'. Please make the claims consistent throughout.","section":"Abstract, Sec. I, Sec. V"},{"comment":"Equation (4) contains a garbled sentence: 'The variables in Eq. (4) are unbiased estimators of which is an unbiased estimator of p^{(w)}_x'. This should be rewritten for clarity.","section":"Eq. (4)"},{"comment":"Table I labels the IonQ entries as 'QPU emulators' while the surrounding text refers to them as QPUs; please clarify which entries are simulators and whether the conclusions apply to real quantum hardware.","section":"Table I and Sec. III A"},{"comment":"There is a typo in 'OPENQAMS2', which should be 'OpenQASM 2'; please also verify that the Zenodo reference [7] provides a persistent and complete dataset.","section":"Sec. III A"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about the GHZ exception and the calibration-transfer limitation in the text, which is a point in its favor. The main risk is that the abstract and conclusions overstate accuracy gains; if the authors restrict the claims to worst-case robustness and add statistical support, the contribution is publishable. The variance-reduction component is a mathematical consequence of averaging and should be presented as such rather than as evidence that the calibration machinery improves accuracy."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things. First, this is a legitimate extension of the authors' own prior shot-wise work [8,9], not a brand-new idea. What is new is the formal framework, the Hellinger and MISE split/merge policies with jackknife bias correction, and a seven-QPU experimental comparison on MQT Bench circuits. That is a real contribution, though incremental. Second, the paper is honest about its main limitation: the authors explicitly say calibration on random circuits may not transfer to specific tasks, and their GHZ results show exactly that failure. The reader's stress-test concern is on target, but the paper's core value does not collapse even if calibration transfer fails. What the paper does well: The formalization of the MISE estimator and the bias-corrected Hellinger distance is careful and correct. The experimental observation that spreading shots across QPUs reduces worst-case error and variance is robust and visible in the figures; this is a useful practical result for quantum software engineering. The writing is clear, the limitations are acknowledged in the text, and the Zenodo dataset is provided, which is more than many papers do. The GHZ reversal is reported rather than hidden, which supports the authors' credibility. Where it is soft. The abstract and Introduction claim that shot-wise outputs are 'more reliable' than single-QPU runs, but the measured claim is narrower: the method never beats the best QPU, tracks the average, and improves the worst case. That overstatement needs fixing. More seriously, the claim that Hellinger or MISE policies improve over uniform splitting depends on calibration weights learned from 5-qubit random circuits and applied to 8-qubit structured circuits. The GHZ result shows the ranking can invert, and there are no confidence intervals or paired statistical tests anywhere. So the informed-policy advantage is not established. The variance-reduction effect of uniform merging is a mathematical consequence of averaging and would occur without any calibration machinery; the paper should separate that from the added value of calibration, which is currently asserted but not demonstrated robustly. I also note the paper lacks analysis code, only the dataset, which limits reproducibility of the figures. Bottom line: this is a paper for researchers in distributed quantum computing and quantum software engineering. It deserves a serious referee—the framework and the worst-case improvement observation are worth engaging with—but the authors should be pushed to temper the 'more reliable' language, add error bars and repeated runs, and either show calibration transfer works for the target circuit class or reposition calibration as a user-tunable option rather than a reliable upgrade. I would not cite it in my own work in the next year, but I would send it to review and expect a revised version to be publishable.","headline":"Solid incremental formalization of shot-wise distribution with an honest but overclaimed experimental narrative; the calibration-transfer premise is the weak spot and the informed-policy advantage is not yet statistically supported.","tokens_in":756,"tokens_out":891,"would_cite":false,"duration_ms":35581,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68"],"pacs":[],"model":"deepseek-v4-flash","headline":"Distributing the shots of a single quantum circuit across multiple noisy QPUs and merging their outputs yields more reliable final distributions than running all shots on one QPU, without ever beating the best individual machine.","keywords":["shot-wise distribution","quantum computation","NISQ devices","shot allocation","Hellinger distance","mean integrated square error","calibration","ensemble merging"],"falsifier":"For a circuit whose ideal distribution is classically computable, calibrate on ten random circuits and then measure the Hellinger distance of the shot-wise merged output; if a calibration-informed policy yields a larger merged error than uniform splitting on the same QPU set, or if adding more QPUs increases the maximum error, the central claim fails.","tokens_in":20476,"feed_emoji":"⚛️","tokens_out":4960,"duration_ms":46606,"temperature":0.7,"pith_summary":"This paper proposes that the many repeated runs (shots) of a single quantum circuit need not all execute on one machine: they can be split across several noisy quantum processors and the resulting count distributions merged into one output. The central claim is that this shot-wise distribution makes final results more reliable than running the whole computation on one QPU, in the specific sense that worst-case error decreases as more QPUs are included while best-case error rises only slightly. The authors support this with experiments on circuits with 5 and 8 qubits across up to seven QPUs of two different hardware types, comparing uniform, Hellinger-distance-based, and mean-integrated-square-error-based split and merge policies. A sympathetic reader would care because choosing the single best QPU in advance is often impractical, and shot-wise distribution offers a fallback that tracks the average and improves the worst case.","feed_headline":"Splitting shots across QPUs tames worst-case errors","feed_subtitle":"Merging outputs from many noisy machines never beats the best one, but steadily improves the worst case.","key_machinery":"The load-bearing machinery is the calibration-and-ranking stage paired with split and merge policies. Each QPU is assigned an unreliability coefficient equal to the mean squared Hellinger distance between its output and the ideal distribution on ten random benchmark circuits, and those coefficients set the split weights for production. Merging is done either by uniform counts, by minimizing a weighted squared Hellinger distance, or by minimizing the mean integrated square error (MISE), which balances the bias of each QPU's distribution against the statistical variance from a finite number of shots. A jackknife bias-correction procedure is applied because the estimated Hellinger distance is biased upward by finite counts.","core_discovery":"The paper's empirical discovery is that merging the output distributions of many QPUs, each executing a share of the shots, produces results that are never better than the best individual QPU but consistently reduce the worst-case Hellinger distance to the ideal distribution and shrink the spread of outcomes. As the number of QPUs grows, the maximum error decreases while the minimum error increases slightly, a robustness effect rather than an accuracy gain. The paper also reports that calibration-informed split and merge policies (Hellinger and MISE) generally improve over naive uniform splitting and merging, with the GHZ circuit being a documented exception where the trend reverses.","pith_inferences":["If calibration transfer fails, as the GHZ result hints it can, informed policies may underperform uniform allocation; a safer production default could be uniform splitting with MISE merging, re-calibrated per circuit family.","The variance-reduction effect resembles ensemble averaging and is likely strongest when QPU noise is heterogeneous and only weakly correlated across machines, a prediction that can be tested by measuring per-QPU error correlations.","Combining shot-wise distribution with circuit cutting could let each fragment of a large circuit have its shots spread across machines, extending the approach beyond the current single-circuit scope.","Because the MISE policy explicitly accounts for finite-shot variance, its advantage over uniform merging should grow as the per-QPU shot budget shrinks, which is a concrete testable extension."],"forward_implications":["As the number of QPUs in the pool grows, the worst observed error drops while the best-case error rises only slightly, so users who cannot identify the best machine in advance get a dependable worst-case guarantee.","Calibration-informed Hellinger and MISE policies usually beat uniform splitting and merging, so the calibration stage has real value whenever its reliability ranking transfers to the production circuit.","The shot-wise method is orthogonal to error mitigation, error correction, and circuit cutting, and can be applied alongside all of them without modification.","The framework supports incremental execution and stopping criteria, so a user can spend shot budget adaptively and stop early when an accuracy target is reached."],"supporting_citations":[{"why":"Prior architectural proposal that this work formalizes into the general shot-wise framework.","marker":"[8]"},{"why":"Preliminary work introducing shot-wise distribution and the qualitative advantages this paper builds on.","marker":"[9]"},{"why":"Supplies the benchmark circuits used in the experimental evaluation.","marker":"[38]"},{"why":"Defines the Hellinger distance used as the unreliability metric.","marker":"[21]"},{"why":"Provides the jackknife bias-correction method for distance estimates.","marker":"[11]"},{"why":"The dataset of QPU executions that the experimental analysis is based on.","marker":"[7]"}],"fun_headline_variants":["Shot-wise QPU splits cut worst-case quantum errors","Distributing quantum shots boosts stability, not peak","Many QPUs beat single worst-case in shot distribution","Quantum shot distribution narrows error spread","Calibrated shot splitting curbs worst QPU errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Unreliability measured on ten random benchmark circuits predicts how each QPU will perform on the actual target circuit.","fun_headline_variants_meta":{"raw":{"variants":["Shot-wise QPU splits cut worst-case quantum errors","Distributing quantum shots boosts stability, not peak","Many QPUs beat single worst-case in shot distribution","Quantum shot distribution narrows error spread","Calibrated shot splitting curbs worst QPU errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1374,"prompt_tokens":886,"completion_tokens":488,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":413}},"tokens_in":502,"tokens_out":488,"duration_ms":5020,"temperature":1.0,"reasoning_tokens":413,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:00:15.048672+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a circuit whose ideal distribution is classically computable, calibrate on ten random circuits and then measure the Hellinger distance of the shot-wise merged output; if a calibration-informed policy yields a larger merged error than uniform splitting on the same QPU set, or if adding more QPUs increases the maximum error, the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior architectural proposal that this work formalizes into the general shot-wise framework."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Preliminary work introducing shot-wise distribution and the qualitative advantages this paper builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the benchmark circuits used in the experimental evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Hellinger distance used as the unreliability metric."},{"cited_title":"These distances can then be used as unreliability parameter to be associated with each QPU","cited_arxiv_id":null,"evidence_quote":"Provides the jackknife bias-correction method for distance estimates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The dataset of QPU executions that the experimental analysis is based on."}],"review_version":1}