{"id":"7c36370a-ec3f-4e20-8187-189d365368fc","arxiv_id":"2607.15426","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"A two-level runtime model decomposes hybrid quantum-classical cycles into quantum, classical, and communication time, allowing a communication-to-computation ratio to classify workflows as compute- or communication-bound.","lead":"This paper introduces a simple runtime model for hybrid quantum-classical workflows that separates application-level communication overhead from real-time feasibility constraints. It uses a communication-to-computation ratio to show which workflows tolerate remote access and which need tight integration.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4's generalization from three representatives to 'most application-level workloads' is asserted, not derived: no workflow with simultaneously HPC-scale classical compute and high communication frequency is analyzed, and Eq. 5 imposes no bound preventing such a combination.","rationale":"The reader's weakest assumption correctly identifies the representativeness of the workflow families as the main soft spot. My analysis agrees and makes the concern more concrete: the model itself cannot rule out a compute-intensive, communication-bound workflow because Rcc depends on the product F(L+V/B) and no upper bound on F is derived from large C_C. The paper's 'almost by construction' remark is an informal appeal rather than a theorem. This concern does not undermine the framework's internal validity; it only weakens the scope of the headline co-location conclusion. Since the reader already issued a CONDITIONAL verdict based on this and related issues, my stress test does not move the verdict. I would keep the conditional acceptance, with the explicit condition that Section 4 be re-scoped to the analyzed workflows or supplemented with a coverage argument that bounds F for HPC-scale workflows.","tokens_in":23618,"tokens_out":6972,"duration_ms":81075,"concrete_test":"Perform an inversion test using Eq. 5: for each integration tier, compute the minimum communication frequency F* at which Rcc = 1 for a representative HPC-scale cycle (e.g., T_Q + T_C = 10^3 s, V = 100 MB). Then scan Table B5/C7 and the cited experimental implementations ([9], [33], [36], [37], [44], [66]) for any workflow whose demonstrated or implementation-plausible F exceeds F*. If any does, the Section 4 claim is falsified for that workflow and must be re-scoped. If none does, the co-location conclusion survives the surveyed landscape.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assertion is the Section 4 generalization from 'across the application-level workflows analyzed in this work' to 'most application-level workloads in the current landscape.' The evidence consists of three representatives (SQD, GQE, QE-MCMC) plus appendix profiles. Notably, Table B5 itself lists families expected to have high Rcc (Quantum Metropolis Sampling, LAQCC), but these are not analyzed at the application level. The paper's defense is the statement that absence of a compute-intensive, communication-bound workflow is 'almost by construction.' Eq. 5 gives Rcc = F(L+V/B)/(T_Q+T_C); it imposes no upper bound on F as T_C grows. A workflow with T_C ~ 10^3 s and F ~ 10^5 blocking exchanges per cycle would have Rcc > 1 even at co-located parameters (L = 10 µs, V/B small). No scaling argument is provided for why such an F cannot coexist with large C_C. The conclusion is also sensitive to the implementation choice of batching: F is defined as implementation-dependent, so an unbated per-term measurement implementation of GQE/VQE is a legitimate workflow with F ~ 10^2-10^3, not the F=1 used in the headline calculation. The order-of-magnitude parameters (Table B6) and selected problem sizes are not accompanied by sensitivity analysis. The central claim may be true, but it is an empirical induction over an unproven sample, not a derived result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a runtime model for hybrid quantum-classical workflows that decomposes each repeated compute cycle into classical compute time T_C = C_C/τ_C, quantum compute time T_Q = C_Q/τ_Q, and communication time T_comm = F(L+V/B). It defines the communication-to-computation ratio Rcc = T_comm/(T_Q+T_C) as an application-level diagnostic, and a real-time feasibility constraint T_C,step + T_comm,step ≤ τ_phys. The model is applied to three application-level workflows (SQD, GQE, QE-MCMC) and to real-time level tasks (QEC, Pauli-based computation, Shor factoring). The paper's central claim is that, among the analyzed application workflows, none is simultaneously compute-intensive and highly communication-bound, so co-location of QPUs with HPC infrastructure offers limited performance advantage for most application-level workloads today, while tight integration remains essential for real-time tasks.","tokens_in":24099,"tokens_out":7648,"duration_ms":78723,"significance":"The model provides a useful shared vocabulary for reasoning about hybrid integration, and it is definitional rather than fitted: no constants are calibrated to data, and the SQD estimate in Section 3.1 reproduces the reported runtime within order-of-magnitude. The Rcc diagnostic cleanly separates application-level performance from real-time feasibility, and the scenario analysis for GQE (Table 5) gives concrete, falsifiable crossover conditions. The resource estimates for reaction-time effects on Shor factoring (Table 8) are based on state-of-the-art compilation and illustrate a genuine performance cascade. If the broad Section 4 conclusion is supported, the paper would be valuable for procurement and integration decisions. However, the generalization from three representative workflows to 'most application-level workloads' is an induction over an unproven sample; the manuscript itself lists high-Rcc families in Table B5 without application-level analysis.","major_comments":[{"comment":"The transition from 'across the application-level workflows analyzed in this work' to 'most application-level workloads in the current landscape' is not justified by the analysis. Eq. (5) imposes no structural bound preventing simultaneous large T_C and large F: a workflow with T_C ~ 10^3 s and F ~ 10^5 blocking exchanges per cycle would have Rcc > 1 even with negligible data volume. The statement that the absence is 'almost by construction' is therefore an empirical regularity, not a consequence of the model. Table B5 itself lists Quantum Metropolis Sampling and LAQCC as high-Rcc families, but no application-level profile is provided for them. Please either supply a concrete scaling argument bounding F relative to T_C, analyze representatives of these families, or confine the headline claim to the workflows actually profiled.","section":"Section 4, Eq. (5)"},{"comment":"There is a factor-of-ten arithmetic inconsistency. With T_Q ~ 10^-5 s, T_C negligible, F = 1, and the latency tiers in Table 3 (L = 100 ms, 10 µs, 100 ns; V/B negligible), Eq. (5) gives Rcc ≈ 10^4, 1, and 10^-2 for remote, co-located, and on-chip tiers, respectively. The text and Table 6 report ~10^3, ~10^-1, and ~10^-3. The stated values would follow from per-cycle quantum time ~10^-4 s, which is consistent with the 2–48 layer range at τ_Q ~ 10^5 layers/s. Please reconcile the reported T_Q value with the Rcc estimates.","section":"Section 3.1, QE-MCMC, Table 6"},{"comment":"Because F is acknowledged to be a workflow-implementation metric, the conclusion that workflows 'fall into two categories' is sensitive to implementation choices. For GQE, per-term measurement with F = 150 already gives Rcc ~ 10^-1 at remote parameters, within an order of magnitude of the communication-bound threshold, and faster QPU throughput or multi-QPU shot parallelism could push it higher. The paper would be strengthened by an explicit statement of which implementation version each Rcc value refers to, and by a sensitivity scan over (T_Q+T_C, F) showing how robust the compute-bound classification is across plausible parameter ranges.","section":"Section 2, Eq. (5); Section 4"}],"minor_comments":[{"comment":"Equation (7) estimates T_Q ~ 10^3 s using d ~ 10^2, s ~ 10^6, τ_Q ~ 10^5, while the subsequent reconstruction from the experimental parameters gives T_Q ≈ 2400 s. Both are order-of-magnitude consistent, but the discrepancy between the idealized d·s model and the reported shot count/circuit-layer count should be clarified in a footnote.","section":"Section 3.1, SQD, Eq. (7)"},{"comment":"The main-text feasibility condition T_C,step + T_comm,step ≤ τ_phys omits the quantum step time T_Q,step, while Eq. (C1) in the SQSP-MaF analysis explicitly includes it. Please clarify whether T_Q,step is intended to be absorbed into τ_phys or should be added to the left-hand side in the main-text statement.","section":"Section 2, Eq. (6); Appendix C.4, Eq. (C1)"},{"comment":"The description of Rcc score 2 as 'Latency has minor impact (<10% runtime)' is inconsistent with the definition Rcc = T_comm/(T_Q+T_C): the runtime overhead fraction is Rcc/(1+Rcc), so Rcc = 0.1 gives ~9% overhead, but Rcc = 0.01 gives ~1%. The score boundaries should be translated to overhead percentages consistently.","section":"Table A3"},{"comment":"The radar chart in Figure A1 and Table A4 give SQSP-MaF an application-level Rcc of 4, while the text notes that the binding constraint is at the real-time level. The caption should state more explicitly that the discrete Rcc score refers only to the idealized application-level regime and does not capture the feasibility constraint.","section":"Table B5 / C.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the modeling framework is simple and likely useful. The main risk is over-generalization from a small sample; the authors should either expand the analysis to the high-Rcc families listed in Table B5 or substantially temper the landscape-level claim. The QE-MCMC arithmetic inconsistency is easily fixed but should be corrected before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you make decisions about where QPUs should sit relative to classical HPC. The model is simple and mostly right; the headline claim about co-location being unnecessary for compute-intensive workloads is broader than the evidence.\n\nWhat's actually new: the two-level split (application vs real-time) genuinely clarifies the discussion, because people conflate 'latency kills performance' with 'latency kills correctness.' The Rcc ratio is a simple unitless diagnostic that connects algorithm structure and hardware parameters to integration choices. The authors apply it to three representative workflows (SQD, GQE, QE-MCMC) plus QEC and factoring, and the worked examples are concrete enough that a systems person can reproduce the numbers. The QE-MCMC example is useful: it shows a workflow where a single round-trip dominates, and explains why the original experiment was done offline. The real-time feasibility constraint and the reaction-time discussion (decoder latency → logical clock speed) are sensible and well grounded in the decoder literature. No parameters are fitted; the paper is transparent that the inputs are order-of-magnitude. That is a real virtue.\n\nSoft spots, in order of severity.\n\nOne: there's an arithmetic inconsistency in the QE-MCMC example. The text quotes T_Q ~1e-5 s and a remote Rcc ~1e3, but Eq. 5 with L=100 ms gives Rcc ~1e4. The co-located number (text says ~1e-1; direct calculation gives closer to 1) also doesn't line up. It's a small fix, but it's in a key example, so it should be corrected before publication.\n\nTwo: Section 4 overgeneralizes. The sentence 'Across the application-level workflows analyzed in this work, none are simultaneously compute-intensive and highly communication-bound' is supported. The next sentence, 'Co-location... offers limited performance advantage for most application-level workloads in the current landscape,' is not. The evidence is three representatives. Table B5 itself lists families (Quantum Metropolis Sampling, LAQCC) expected to have high Rcc that are not analyzed at application level. The phrase 'almost by construction' is doing too much work. Nothing in Eq. 5 prevents a workflow with T_C ~1e3 s and F ~1e5 blocking exchanges per cycle from having Rcc >1 even at co-located latency. Maybe none exists today, but the paper needs either a search argument or an explicit example from the Appendix, and definitely a sensitivity analysis on F and T_C. The fact that F is implementation-dependent (batched vs per-term) reinforces this: the headline uses F=1 for GQE, but a non-batched implementation has F~150, which changes Rcc by two orders of magnitude. That doesn't kill the framework; it means the conclusions are conditional on implementation choices in a way the paper doesn't fully acknowledge.\n\nMinor: there's no code or data beyond the tables, and no uncertainty analysis on the order-of-magnitude parameters. For a framework paper that's acceptable, but it should be flagged.\n\nThe core model holds up. This paper is for people designing QPU-HPC integration architectures, and for algorithm developers who want a common vocabulary when discussing latency. It deserves peer review. I would send it out, with the request that the authors fix the QE-MCMC arithmetic, add sensitivity bounds, and rewrite Section 4 to distinguish between 'our representative workflows' and 'most application-level workloads.' As written, it's a conditional accept with major-to-minor revisions.","headline":"A useful framework for QPU-HPC integration decisions whose headline co-location claim is broader than the evidence supports.","tokens_in":24577,"tokens_out":4273,"would_cite":true,"duration_ms":42400,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A workflow-level runtime model splits hybrid quantum-classical jobs into compute and communication, and shows most current applications are compute-bound.","keywords":["hybrid quantum-classical workflows","communication-to-computation ratio","real-time feasibility constraint","quantum error correction","co-location","reaction time","fault-tolerant quantum computation","runtime performance model"],"falsifier":"Demonstrate one real hybrid workload whose per-cycle classical compute is hours-scale or large (say T_C + T_Q > 1000 s) yet whose algorithm forces a blocking exchange frequency F >= 10^4 per cycle at current remote latency (L ~ 100 ms). Such a workflow would have T_comm ~ 1000 s and Rcc ~ 1, directly contradicting the paper's claim that compute-intensive workflows are not communication-bound; alternatively, a measurement showing that a representative SQD or GQE instance already achieves Rcc >= 0.1 under the parameters the paper assigns would falsify the numerical claim.","tokens_in":23535,"feed_emoji":"⚛️","tokens_out":5689,"duration_ms":50598,"temperature":0.7,"pith_summary":"This paper introduces a common way to reason about hybrid quantum-classical workflows by splitting each repeated quantum-classical exchange into three wall-clock pieces: quantum compute, classical compute, and communication. It then defines a single unitless diagnostic—the communication-to-computation ratio Rcc—that tells whether a workflow is communication-bound or compute-bound, and a separate feasibility constraint that tells whether real-time classical processing can keep up with the physical quantum device at all. Applied to representative workflows, the model argues that today's compute-intensive applications (subspace diagonalization, generative quantum eigensolvers) are deeply compute-bound and gain almost nothing from co-locating a QPU with an HPC facility, while latency-sensitive single-shot workflows need tight integration and fault-tolerant error correction requires it outright. The point of the framework is to turn integration decisions—remote access, co-location, or on-node co-design—into a measurable quantity rather than a matter of opinion.","feed_headline":"Co-locating QPUs with HPC barely helps today's hybrid apps","feed_subtitle":"A new runtime model splits workflows into compute-bound and latency-bound, and shows real-time error correction is the true bottleneck.","key_machinery":"The central object is the compute-cycle decomposition and its two derived diagnostics. Each compute cycle—the smallest repeated unit containing a blocking classical-quantum exchange—has time T_cycle = T_C + T_Q + T_comm, with T_comm = F(L + V/B) for exchange frequency F, latency L, volume V, and bandwidth B. The application-level diagnostic is Rcc = T_comm/(T_Q + T_C), which plays a role analogous to arithmetic intensity in classical performance analysis. The real-time-level diagnostic is the feasibility inequality T_C,step + T_comm,step <= tau_phys, where tau_phys is a device-imposed timescale such as qubit coherence time; failure means failure of execution, not mere slowdown.","core_discovery":"On the paper's terms: every hybrid workflow can be modeled as repeated compute cycles costing T_Q + T_C + T_comm. The application-level diagnostic is Rcc = T_comm/(T_Q + T_C): Rcc >> 1 means communication-bound; Rcc << 1 means compute-bound. The real-time diagnostic is whether classical step time plus communication step time fits within the physical qubit timescale; if not, execution fails outright. Surveying representative workflows, the paper finds none are simultaneously compute-intensive and highly communication-bound today, so heavy compute absorbs communication overhead, while real-time error correction remains strictly latency-limited; once the real-time constraint is met, reaction ti","pith_inferences":["The model's basic structure (T_C, T_Q, T_comm plus Rcc) is general enough that a workflow library analogous to roofline charts could be built for parameterized circuit families; the paper sketches one classification table but does not exhaust the space.","A testable consequence the paper leaves implicit: if a workflow exists with per-cycle classical compute over ~1 second and exchange frequency above ~1000 per cycle, Rcc would cross unity at current latency, and co-location would become material. The absence of such a workflow is presented as 'almost by construction' rather than demonstrated.","Quantum memory, which the paper discusses qualitatively, would restructure exactly the metrics the model tracks—reducing C_Q per cycle by amortizing state preparation and lowering F by eliminating re-initialization round trips—so the same formalism could quantify the value of quantum memory before it is built.","The model suggests a procurement rule that the community can test: for a candidate hybrid application, compute Rcc under remote and co-located parameters before choosing an integration tier; applications with Rcc below ~0.01 are safe to run remotely, while error-correction workloads must be evaluated by the feasibility inequality instead."],"forward_implications":["SQD and GQE, the representative HPC-scale workflows, have Rcc between 10^-4 and 10^-1 even at remote latency, so moving them into co-located or tightly integrated facilities will not reduce runtime; their bottleneck is quantum shot budget and classical diagonalization, not communication.","QE-MCMC, with one round-trip per accept/reject step and tiny per-step compute, has Rcc around 10^3 at remote latency, so it needs low-latency integration—but not HPC-class classical resources.","Under fault tolerance, real-time decoder reaction time becomes a continuous performance variable: doubling it increases the estimated runtime to factor a 2048-bit RSA integer by 60-100%, because slower decoding means more idle error accumulation and higher code distances.","Hardware evolution that shrinks quantum compute time—faster throughput or parallel shot distribution across multiple QPUs—raises Rcc for workflows like GQE and can move them into a communication-bound regime, so co-location decisions are not permanent.","Dynamic circuits are application-level when executed on fault-tolerant logical qubits but real-time level when executed on physical qubits: the same feed-forward can be performance-neutral or feasibility-critical depending on where it is implemented."],"fun_headline_variants":["Hybrid quantum apps: compute-bound, not communication-bound","Real-time error correction, not co-location, is the bottleneck","Model shows co-location won't speed compute-heavy quantum apps","When HPC-QPU co-location pays off: only for real-time tasks","Compute intensity trumps latency in hybrid quantum workflows"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's headline implication—that co-location offers limited performance advantage today—rests on the assumption that the workflow families and order-of-magnitude parameters it surveys adequately represent the current landscape, and in particular that no workflow is simultaneously compute-intensive and highly communication-bound; the paper says this absence is 'almost by construction' rather than proven.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid quantum apps: compute-bound, not communication-bound","Real-time error correction, not co-location, is the bottleneck","Model shows co-location won't speed compute-heavy quantum apps","When HPC-QPU co-location pays off: only for real-time tasks","Compute intensity trumps latency in hybrid quantum workflows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1251,"prompt_tokens":750,"completion_tokens":501,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":428}},"tokens_in":494,"tokens_out":501,"duration_ms":5165,"temperature":1.0,"reasoning_tokens":428,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T23:23:52.596893+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Demonstrate one real hybrid workload whose per-cycle classical compute is hours-scale or large (say T_C + T_Q > 1000 s) yet whose algorithm forces a blocking exchange frequency F >= 10^4 per cycle at current remote latency (L ~ 100 ms). Such a workflow would have T_comm ~ 1000 s and Rcc ~ 1, directly contradicting the paper's claim that compute-intensive workflows are not communication-bound; alternatively, a measurement showing that a representative SQD or GQE instance already achieves Rcc >= 0.1 under the parameters the paper assigns would falsify the numerical claim.","supporting_citations":[],"review_version":1}