REVIEW 3 major objections 6 minor 20 references
When you factor in price, the best quantum cloud machine is often not the one with the highest fidelity.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 03:22 UTC pith:QWQNY3QG
load-bearing objection Solid systems measurement paper: cost-aware QPU ranking really can disagree with fidelity-only ranking, and they shipped the data to prove it on a narrow probe. the 3 major comments →
Quantum Fidelity-per-Cost: A Metric for Evaluation of Quantum Computing Systems
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
At a fixed 1000-shot budget, cost-aware ranking of cloud quantum backends disagrees with purely fidelity-based ranking. The superconducting access path with the lowest mean total variation distance is not the highest-QFC path; a lower-fidelity but cheaper runtime-billed path ranks first, and the spread in QFC across paths is roughly one and a half to two orders of magnitude. That disagreement persists across moderate changes to the metric’s sensitivity weights and under an alternative billing assumption for one provider.
What carries the argument
Quantum Fidelity-per-Cost (QFC): a single score equal to exp(−α × KL divergence from the ideal output) × (shot count)^γ / (monetary cost)^β, evaluated under each provider’s documented billing model. It turns observed output quality, shot budget, and billed dollars into one comparable number so backends can be ranked by value rather than by fidelity alone.
Load-bearing premise
The claim rests on a single shallow two-qubit Bell-state circuit run with each provider’s default compiler and qubit choice; if deeper or application circuits reverse the cost-aware ranking, the practical selection advice does not carry over.
What would settle it
Re-run the same QFC pipeline on a deeper fixed circuit (for example a multi-qubit GHZ state or a small QAOA ansatz) at the same shot budgets across the same access paths; if the lowest-error device then also tops the QFC ranking under default weights, the claimed fidelity-versus-value disagreement fails for that workload class.
If this is right
- Users who can tolerate some noise should pick backends by QFC at their actual shot budget, not by fidelity tables alone, or they may overspend.
- Cost-aware rankings are only meaningful at a stated shot count, because billing model—not hardware—sets how QFC scales with shots.
- The same physical QPU reached through two clouds can land in different QFC ranks when pricing or observed output quality differs by path.
- Saved per-run counts let QFC be recomputed whenever providers change prices, so procurement monitoring does not require new hardware runs.
- QFC can serve as the objective a cost-aware quantum cloud scheduler maximizes once fidelity and runtime are predicted rather than measured.
Where Pith is reading between the lines
- Providers that bill by coarse whole-second runtime create large score noise unrelated to device quality; finer billing resolution would make value rankings more stable.
- A reliability-weighted QFC that charges failed or empty jobs would reorder paths with high backend error rates even if successful-run fidelity looks fine.
- As prices fall or new machines appear, the value leader will likely rotate faster than the fidelity leader, so published rankings need dated price snapshots to stay useful.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a cross-provider measurement study of a two-qubit Bell-state circuit (247 runs across 14 QPU access-path entries / 12 physical QPUs on AWS Braket, IBM Runtime, IQM Resonance, and OQC) and introduces Quantum Fidelity-per-Cost (QFC), a composite score that multiplies an exp(−α D_KL) fidelity factor by N^γ / C^β under documented cloud billing models. The central empirical claim is that cost-aware ranking can disagree with fidelity-only ranking: at 1000 shots IBM Marrakesh has the lowest mean TVD among superconducting entries while IQM Sirius ranks first by QFC, with the disagreement stable under moderate (α,β,γ) reweighting and an alternative IQM rounding assumption (Tables II–III, Fig. 2). A secondary result is that billing model, not hardware, fixes how QFC scales with shot count (Fig. 3 and closed-form substitutions into Eq. 4). Code, per-run counts, and analysis scripts are released for full regeneration without re-execution.
Significance. Cost-aware comparison under heterogeneous quantum-cloud billing is a genuine and underexplored practical problem; existing metrics (QV, CLOPS, application suites, Q-score) largely omit monetary cost. The paper’s contribution is concrete and usable: a documented four-class pricing taxonomy, one explicit composite score, a multi-provider dataset, and a clear ranking-disagreement observation scoped to a narrow probe. Strengths that raise the work above a pure opinion piece include the released 247-run count dataset and fully regenerable pipeline, the Kendall-τ robustness table over saved data (no refitting), and the clean derivation that billing regime—not device physics—sets QFC’s N-scaling. If the result holds as scoped, users and schedulers gain a reproducible way to trade fidelity against dollars at a stated shot budget. The Bell-only probe and price-snapshot dependence limit generality, but do not erase the value of the measurement and the metric as a starting point.
major comments (3)
- [Abstract; §I-A; §VII] §I-A and §VI correctly state that the single Bell-state probe does not support claims about deeper circuits or intrinsic hardware quality, yet the abstract and conclusion still frame the result as guidance for “which quantum computer… gives the best value for execution of their circuits” in general terms. The central claim as written (ranking can differ on this workload) is supported by Table II and Fig. 2; please tighten abstract/conclusion language so it does not over-read the probe, or add at least one deeper circuit (e.g., small GHZ or fixed-depth ansatz) on a subset of access paths to show whether disagreement persists. Without that, the practical backend-selection claim remains conditional on an untested extrapolation.
- [Table II; §IV-C; §VI] Table II reports n=3 for several AWS entries and n=5 for most others at 1000 shots; §IV-C notes the Emerald AWS-vs-IQM fidelity gap and that Garnet’s cost decomposition “does not close” with heavy-tailed Resonance costs. With such small per-cell n, and with IBM runtime quantized to whole seconds (§III-B, §VI), the precise rank order (e.g., Sirius vs. Garnet vs. Marrakesh) is noisier than the binary claim “disagreement exists.” Please either (i) report bootstrap/CI or pairwise rank uncertainty on QFC means in Table II, or (ii) state the primary finding strictly as existence of disagreement (Marrakesh best TVD, not top QFC) rather than as a stable full ordering. Table III already helps on weight sensitivity; sampling uncertainty is the missing load-bearing piece.
- [Eq. (4); §III-C; §VI; Table III] Eq. (4) and §III-C use KL with ε=10^{-12} additive smoothing and no renormalization because P_ideal(01)=P_ideal(10)=0. §VI reports that varying ε from 10^{-6} to 10^{-15} moves 1000-shot ranking by Kendall τ 0.67–0.89—comparable to the α/β rows of Table III—yet this is only in Discussion, not in the robustness table. Since the dominant error mode is leakage onto ideally zero outcomes, ε is load-bearing for absolute QFC and can affect mid-ranks. Please promote the ε sweep into Table III (or an equivalent) and, if mid-ranks move, either justify the default ε more carefully or show that the headline disagreement (lowest-TVD ≠ highest-QFC) is ε-stable over that range.
minor comments (6)
- [Figure 2] Fig. 2’s crossing-line design is effective; adding numeric TVD and QFC annotations on the two emphasized endpoints (Marrakesh, Sirius, Aria-1) would make the figure self-contained without Table II.
- [§II; §IV-D] §II lists a per-gate Azure/IonQ class “for completeness” but does not evaluate it; a one-sentence note on why Aria-1 was billed under AWS task+shot rather than Azure per-gate would avoid confusion with §IV-D.
- [Table II; Figure 1] IBM devices labeled “(ret.)” in Table II and Fig. 1—briefly state retirement status/date in the caption or §IV-A so readers know whether those rows are historical only.
- [Abstract] Typographical inconsistency: “A WS” with a space appears in the abstract; standardize to “AWS”.
- [§V] Related work §V cites Metriq for reporting execution cost; a short explicit contrast (Metriq keeps cost side-by-side; QFC folds cost into the score) would sharpen the novelty sentence already present.
- [§VI] Failed jobs are excluded and one still-billed failure is noted in §VI; flagging “reliability-weighted QFC” as future work is fine, but a footnote on how many failures occurred per provider would help readers assess selection bias.
Circularity Check
No significant circularity: QFC is an explicit constructed score; ranking disagreement is measured, not forced by definition or fit.
full rationale
The paper defines Quantum Fidelity-per-Cost as a weighted preference function QFC(N)=exp(−α D_KL) N^γ / C(N)^β under documented billing models, then scores independently collected Bell-state shot counts and provider-reported costs. The headline result—that lowest-TVD and highest-QFC devices can differ at 1000 shots, and that billing model governs shot scaling—is an empirical comparison of two different orderings of the same measured runs, not a derivation that recovers an input by construction. Closed-form substitutions after Eq. (4) explain how each billing class scales with N; they do not fit parameters to force the ranking. Sensitivity (Table III) recomputes QFC on saved per-run counts under alternate (α,β,γ) and billing assumptions without refitting to preserve disagreement. Self-citation of Szefer [20] appears only in related work on pricing exploitability and is not load-bearing for the metric or the ranking claim. No uniqueness theorem, fitted-input-as-prediction, or self-definitional loop supports the central observation. Score 0 is appropriate.
Axiom & Free-Parameter Ledger
free parameters (3)
- α, β, γ (QFC sensitivity weights) =
1 (default)
- ε (KL smoothing constant) =
1e-12
- Provider price snapshot (c_task, c_shot, c_sec) =
Table I snapshot
axioms (5)
- ad hoc to paper KL divergence with additive smoothing is an acceptable information-quality factor for composing with cost and shot count
- domain assumption Provider-reported runtime fields (IBM quantum_seconds, IQM runtime_seconds) are the correct cost basis for runtime-billed QFC even though they differ in resolution and may include classical overhead
- domain assumption Default vendor transpilation and qubit selection represent the access path an ordinary user meets
- standard math Ideal Bell distribution P(00)=P(11)=0.5, P(01)=P(10)=0 is the correct reference for fidelity error
- ad hoc to paper Failed jobs are excluded from QFC (no reliability penalty)
invented entities (1)
-
Quantum Fidelity-per-Cost (QFC) score
no independent evidence
read the original abstract
Cloud-accessible quantum computing has made hardware comparison not only a physics benchmark but also a practical purchasing decision. Cost-aware comparison of quantum computers remains underexplored and is difficult to do under the heterogeneous billing models offered by various cloud-based quantum computing providers. This paper makes two main contributions to enable price-aware comparison of quantum computers. First, this work presents a cross-provider measurement study of quantum circuit execution fidelity spanning 14 cloud QPU access-path entries (12 distinct physical QPUs) across four cloud access paths: Amazon Web Services (AWS) cloud, IBM Quantum Runtime (IBM) cloud, IQM Resonance (IQM) cloud, and Oxford Quantum Circuits (OQC) cloud. Second, this work proposes and analyzes a cost-aware score, Quantum Fidelity-per-Cost (QFC), which combines Kullback--Leibler (KL) divergence from an ideal output distribution, shot count, and monetary cost into one possible metric under a documented billing model. The main empirical observation from this work is that cost-aware ranking can differ from purely fidelity-based evaluation of quantum computers, and that users may select different quantum computing backends when they consider price in their selection, as opposed to selection based on fidelity alone. This work shows that the ranking is stable under reweighting of the metric, and that a device's billing model, not its hardware, governs how its score scales with shot count. Reported QFC values change as new machines come online or as providers revise their prices.
Figures
Reference graph
Works this paper leans on
-
[1]
Quantum computing in the NISQ era and beyond,
J. Preskill, “Quantum computing in the NISQ era and beyond,”Quan- tum, vol. 2, p. 79, 2018
2018
-
[2]
Amazon Braket Pricing,
Amazon Web Services, “Amazon Braket Pricing,” 2025, accessed: 2026-02-27. [Online]. Available: https://aws.amazon.com/braket/pricing/
2025
-
[3]
IBM Quantum Pricing,
IBM Quantum, “IBM Quantum Pricing,” 2025, accessed: 2026-02-27. [Online]. Available: https://www.ibm.com/quantum/pricing
2025
-
[4]
IQM Resonance,
IQM, “IQM Resonance,” 2025, accessed: 2026-02-27. [Online]. Available: https://meetiqm.com/products/iqm-resonance/
2025
-
[5]
Validating quantum computers using randomized model circuits,
A. W. Cross, L. S. Bishop, S. Sheldon, P. D. Nation, and J. M. Gambetta, “Validating quantum computers using randomized model circuits,”Physical Review A, vol. 100, no. 3, p. 032328, 2019
2019
-
[6]
A. Wack, H. Paik, A. Javadi-Abhari, P. Jurcevic, I. Faro, J. M. Gambetta, and B. R. Johnson, “Quality, speed, and scale: Three key attributes to measure the performance of near-term quantum computers,”arXiv preprint arXiv:2110.14108, 2021
Pith/arXiv arXiv 2021
-
[7]
Qasmbench: A low- level qasm benchmark suite for nisq evaluation and simulation,
A. Li, S. Stein, S. Krishnamoorthy, and J. Ang, “Qasmbench: A low- level qasm benchmark suite for nisq evaluation and simulation,”arXiv preprint arXiv:2005.13018, 2020
Pith/arXiv arXiv 2005
-
[8]
Supermarq: A scalable quantum benchmark suite,
T. Tomesh, P. Gokhale, V . Omole, G. S. Ravi, K. N. Smith, J. Viszlai, X.- C. Wu, N. Hardavellas, M. R. Martonosi, and F. T. Chong, “Supermarq: A scalable quantum benchmark suite,”arXiv preprint arXiv:2202.11045, 2022
Pith/arXiv arXiv 2022
-
[9]
H. Donkers, K. Mesman, Z. Al-Ars, and M. M ¨oller, “Qpack scores: Quantitative performance metrics for application-oriented quantum com- puter benchmarking,”arXiv preprint arXiv:2205.12142, 2022
Pith/arXiv arXiv 2022
-
[10]
Application-oriented performance benchmarks for quantum computing,
T. Lubinski, S. Johri, P. Varosy, J. Coleman, L. Zhao, J. Necaise, C. H. Baldwin, K. Mayer, and T. Proctor, “Application-oriented performance benchmarks for quantum computing,”arXiv preprint arXiv:2110.03137, 2021
Pith/arXiv arXiv 2021
-
[11]
QUARK: A framework for quantum computing application benchmark- ing,
J. R. Fin ˇzgar, P. Ross, L. H ¨olscher, J. Klepsch, and A. Luckow, “QUARK: A framework for quantum computing application benchmark- ing,”arXiv preprint arXiv:2202.03028, 2022
Pith/arXiv arXiv 2022
-
[12]
Benchmarking quantum copro- cessors in an application-centric, hardware-agnostic, and scalable way,
S. Martiel, T. Ayral, and C. Allouche, “Benchmarking quantum copro- cessors in an application-centric, hardware-agnostic, and scalable way,” arXiv preprint arXiv:2102.12973, 2021
Pith/arXiv arXiv 2021
-
[13]
Systematic benchmarking of quantum computers: status and recommendations,
J. M. Lorenz, T. Monz, J. Eisert, D. Reitzner, F. Schopfer, F. Barbaresco, K. Kurowski, W. van der Schoot, T. Strohm, J. Senellart, C. M. Perrault, M. Knufinke, Z. Amodjee, and M. Giardini, “Systematic benchmarking of quantum computers: status and recommendations,”arXiv preprint arXiv:2503.04905, 2025
Pith/arXiv arXiv 2025
-
[14]
Azure Quantum Pricing,
Microsoft Azure, “Azure Quantum Pricing,” 2025, accessed: 2026- 02-27. [Online]. Available: https://learn.microsoft.com/en-us/azure/ quantum/pricing
2025
-
[15]
On information and sufficiency,
S. Kullback and R. A. Leibler, “On information and sufficiency,”The Annals of Mathematical Statistics, vol. 22, no. 1, pp. 79–86, 1951
1951
-
[17]
Qonductor: A cloud orchestrator for quantum computing,
E. Giortamis, F. Rom ˜ao, N. Tornow, D. Lugovoy, and P. Bhatotia, “Qonductor: A cloud orchestrator for quantum computing,” inProceed- ings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC), 2025
2025
-
[18]
QSRA: A QPU scheduling and resource allocation approach for cloud-based quantum computing,
B. Lu, Z. Chen, and Y . Wu, “QSRA: A QPU scheduling and resource allocation approach for cloud-based quantum computing,”arXiv preprint arXiv:2411.05283, 2024
Pith/arXiv arXiv 2024
-
[19]
QuSplit: Achieving both high fidelity and throughput via job splitting on noisy quantum computers,
J. Li, Y . Song, Y . Liu, J. Pan, L. Yang, T. Humble, and W. Jiang, “QuSplit: Achieving both high fidelity and throughput via job splitting on noisy quantum computers,”Quantum Machine Intelligence, vol. 7, no. 2, 2025
2025
-
[20]
Exploiting reset operations in cloud-based quantum comput- ers to run quantum circuits for free,
J. Szefer, “Exploiting reset operations in cloud-based quantum comput- ers to run quantum circuits for free,”arXiv preprint arXiv:2512.14582, 2025
arXiv 2025
-
[2026]
Available: https://arxiv.org/abs/2603.08680
[Online]. Available: https://arxiv.org/abs/2603.08680
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.