{"id":"400c2f7f-3541-4b9f-964c-75a4b176dab9","arxiv_id":"2608.06875","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"QCORE is a reference architecture for QPU-side digital control that combines real-time execution, fast feedback, calibration/QEC services, and safe configuration update, evaluated only through behavioral models.","lead":"This paper introduces QCORE, a digital-control architecture for quantum processors that would handle real-time pulses, readout feedback, calibration, error correction, and safe configuration updates together on the QPU side. It reports simulated latency, drift-correction, safety, and scaling results, but no hardware or cycle-accurate validation has been done.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unreleased and under-specified behavioral simulator/quantum model underpins the P99, calibration, and scaling results; these headline numbers cannot be independently reproduced.","rationale":"The reader's weakest_assumption is exactly that the behavioral models faithfully represent a real QCORE implementation. I agree: the architecture proposal itself is coherent and the authors are transparent about the behavioral level of evidence, but the load-bearing nature of the unreleased models makes the quantitative results unverifiable. This does not change the CONDITIONAL verdict: the qualitative architectural claims (partitioning, fast sideband, safe-commit protocol) are plausibly sound, but the specific numbers should be treated as provisional pending independent reproduction. The paper deserves credit for clearly scoping the evaluation and avoiding overclaim relative to RTL/silicon. The concern is not internal inconsistency but external validity and reproducibility. A CONDITIONAL verdict (accept pending release and independent verification of the simulator) is appropriate.","tokens_in":10987,"tokens_out":17616,"duration_ms":164024,"concrete_test":"Demand the release of the simulator source, configuration files, and exact parameter definitions (Lmax, service-time distributions, arrival processes, transmon/readout parameters) as a condition of reproducibility; then independently rerun the Fig. 11–13 experiments and verify that the reported means, 95% CIs, and zero-violation rates reproduce. Additionally, sweep the background service-time distribution (e.g., uniform, exponential, deterministic, all bounded at Lmax) and the arrival process; if the P99 latency changes by more than 0.1 Lmax or the deadline-violation rate becomes nonzero, the headline latency claim is not robust to unspecified modeling details.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims—(1.984±0.004)Lmax P99 at load 0.8, 83.2%±0.8% frequency-error reduction, and 2.08× scaling ratio—are produced entirely by a transaction/event simulator and a three-level transmon/dispersive-readout model that are neither released nor fully specified. Section IV-B states the evaluation 'does not replace cycle-accurate RTL, PPA analysis, or hardware measurement,' but the abstract presents these numbers as headline results. The exact service-time distribution, arrival process, NoC arbitration granularity, SRAM bank conflict model, and parameter values (e.g., T1/T2, readout SNR, HEMT noise, Lmax definition) are not enumerated, so a different but equally plausible parameterization could change the numbers materially. In particular, the P99 latency result relies on an unstated background-traffic service-time distribution; while the zero-violation property should hold under any bounded service time with strict priority, the reported percentile value is distribution-dependent. The closed-loop calibration improvement depends on the specific drift injection and noise model. Without the simulator, these results are not auditable; this is a correctness-risk concern, not evidence of internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"QCORE is a proposed QPU-side digital control reference architecture that separates task management, shared support infrastructure, hard-real-time execution, and closed-loop services into four hardware partitions. It introduces a dual-output readout interface with a fast-result sideband and Measurement Packets, a common service-control skeleton for calibration and error correction, Tile-local QEC, and versioned Safe-Commit for configuration updates. The evaluation uses transaction/event models and a three-level transmon/dispersive-readout model, reporting P99 feedback latency of (1.984±0.004)Lmax at a background load of 0.8 under real-time priority, an 83.2%±0.8% reduction in frequency error, zero unsafe acceptances in 100,000 configuration transactions, and a 2.08× capacity-normalized scaling estimate. The paper explicitly states that the evaluation does not replace cycle-accurate RTL, PPA analysis, or hardware measurement, and it reports confidence intervals and seed counts throughout.","tokens_in":11235,"tokens_out":7154,"duration_ms":65585,"significance":"The architectural contribution is timely and useful: existing control systems provide point capabilities such as fast feedback or calibration, but QCORE targets the joint organization of deterministic control, traceable measurement, same-round feedback, calibration/error-correction services, and safe state updates within one QPU-side boundary. The proposed partitions and the separation of the short feedback path from the service path are plausible and could inform future control-electronics design. The paper is unusually careful with qualifications: it reports seed counts, confidence intervals, and explicit scope statements, and it states in Section IV-B that the evaluation is not a hardware validation. A strength is the self-identification of the behavioral-model scope and the explicit future-work requirement of cycle-accurate validation. The main limitation is that the quantitative headline results come from an unreleased and incompletely specified simulator, so the specific numbers are not auditable; the qualitative architectural claims are credible, but the numerical results should be treated as provisional.","major_comments":[{"comment":"The central quantitative claims — the P99 feedback latency of (1.984±0.004)Lmax at load 0.8, the 83.2%±0.8% frequency-error reduction, and the 5.37%±0.29% readout error — are produced by an event-driven simulator and a transmon/readout model that are neither released nor specified in sufficient detail to reproduce. The paper does not enumerate the background service-time distribution, arrival process, NoC arbitration granularity, SRAM bank-conflict model, T1/T2* values, dispersive shift, readout noise, threshold settings, drift-injection process, or NPU gating parameters. While the zero-violation property of the real-time priority policy follows from the stated bounded-transaction assumptions, the specific P99 percentile is distribution-dependent, and the calibration improvement is dependent on the injected drift and noise model. Because the abstract presents these numbers as headline results, the authors should either release the simulator and traces or provide a complete parameter table and the code or configurations needed to reproduce each figure. This is a reproducibility risk, not evidence of internal inconsistency; the manuscript's own caveat in Section IV-B ('does not replace cycle-accurate RTL, PPA analysis, or hardware measurement') does not by itself resolve the need for model transparency.","section":"Section IV-B, System-Level Quantitative Evaluation"},{"comment":"The derivation of the 2.08× scaling estimate is not auditable from the numbers given. The QEC trace in Table II is W5=(8000.00, 2000, 2167.62, 167.62, 0), but the text states that 'Tile-local QEC generates 1000 + 1000 + 0.015 + 0.041 = 2000.056 boundary transactions per tile'; the provenance of 0.015 and 0.041 is unexplained (they appear to be entries from the Calibration and RB traces rather than the QEC trace), and the 167.62 locally consumed frame actions do not appear in the sum. The ratio 4167.676/2000.056 = 2.08 therefore cannot be checked against the stated trace vector. The authors should define 'boundary transaction' precisely, show how each term is derived from the W5 columns, and either provide the trace-level calculation or remove the quantitative scaling claim.","section":"Section IV-B3, Scalability and Multiworkload Concurrency; Table II"}],"minor_comments":[{"comment":"The sentence 'The relation 𝑡̂𝜋/2 = 𝑡̂𝜋/2 is used only as an initial value' appears to contain a typo; it should likely read t̂_{π/2} = t̂_π/2 or equivalent.","section":"Section IV-A, Example 1"},{"comment":"The sentence 'A workload unit denotes a calibration shot, QEC round, or rather than an equivalent operation count' is grammatically incomplete; the intended third item (likely 'RB shot') is missing.","section":"Section IV-B3"},{"comment":"The vector notation is inconsistent: Section IV-B3 defines W as a six-tuple (N_timed, N_pkt, N_event, N_update, N_cross, N_NPU), while Table II uses W5=(T, P, E, U, X). Please state whether N_NPU is omitted by design and why.","section":"Table II and Section IV-B3"},{"comment":"The 2.08× ratio is labeled a 'capacity-normalized sustainable tile count', but the derivation only counts boundary transactions. Please clarify the capacity model (for example, equal global-service capacity at the boundary) and the relationship between transaction count and sustainable tile count.","section":"Section IV-B3"}],"recommendation":"major_revision","confidential_remarks":"The paper is an architecture proposal with a self-contained behavioral evaluation. I recommend asking the authors, as a condition of revision, to release the simulator and traces or to provide a complete parameter appendix so that the quantitative claims can be checked. The current abstract positions the simulator outputs as the main quantitative deliverables, which is risky for a journal readership expecting independent reproducibility. If the authors reposition the paper more explicitly as a reference-architecture paper with illustrative simulations, the contribution is acceptable in principle."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — This is a real architecture paper, not a measurement paper, and it should be read that way. The QCORE contribution is an integrated QPU-side digital control reference architecture: a fast-result sideband for same-round feedback, a Measurement Packet interface with versioned metadata, a common service skeleton for calibration and QEC, tile-local QEC, and safe shadow-to-active commit. Individually these ideas exist in various systems, but the combination and the explicit hardware partitioning into task management, shared resources, real-time execution, and long-timescale services is a useful organizing proposal for the field. The paper does well in carefully scoping its claims: it repeatedly states that deterministic execution is a behavioral constraint, that results are model-based, and that no RTL, PPA, or silicon validation exists. The transaction, event, and quantum-behavioral models are described with enough detail to follow the architecture's operation, and the confidence intervals and seed counts are a good sign.\n\nThe soft spots are real and concentrated in Section IV. The headline numbers—P99 at 1.984 Lmax, 83.2% frequency-error reduction, 2.08x scaling—all come from a simulator that has not been released and whose parameter values, service-time distributions, arrival processes, and quantum model constants are not fully enumerated. The stress-test note is right: a different but equally plausible parameterization could move those numbers materially. That is a correctness risk for the quantitative claims, and the abstract presents them as headline results even though the body says they do not replace RTL or hardware measurement. The architectural claims, however, do not depend on those exact numbers; the deadline-isolation property and the safe-commit consistency results are more structural and hold within the stated model.\n\nWhat is missing is reproducibility. There is no code, no artifact, no full parameter table. For a systems architecture paper in this space, that is a significant deficiency. The authors are honest about it, but it still limits what a referee can verify.\n\nWho is this for? Hardware architects working on quantum control ASICs/FPGAs and people designing control stacks for scaled superconducting processors. It deserves a serious referee because the architecture is coherent, the problem is important, and the paper is honest about its current evidence level. I would ask the authors to release the simulator and enumerate parameters, or at minimum to reframe the quantitative results as illustrative model outcomes. Fix that and the paper becomes a solid reference for the subfield.","headline":"A thoughtful, honestly scoped control architecture whose model-based headline numbers need artifact release or reframing before they can be used.","tokens_in":11784,"tokens_out":1890,"would_cite":true,"duration_ms":19184,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Quantum-control architecture caps feedback delay at 1.98x ideal.","keywords":["quantum control electronics","QPU-side digital control","real-time execution","closed-loop calibration","quantum error correction","measurement packet","safe configuration update","tile-local QEC"],"falsifier":"Build or simulate the QCORE partitions at cycle-accurate RTL level with the same $40L_{\\max}$ transaction period and 0.8 background load, and measure the P99 of the Measurement Packet/Event feedback path; if P99 rises materially above $(1.984\\pm0.004)L_{\\max}$ or deadline violations become nonzero, the real-time-isolation claim fails.","tokens_in":10776,"feed_emoji":"⚛️","tokens_out":10907,"duration_ms":92740,"temperature":0.7,"pith_summary":"QCORE is a proposed digital control architecture that sits between the host computer and a quantum processor's analog front end, with the goal of letting deterministic pulse control, same-round feedback, calibration, error correction, and safe configuration updates coexist inside one hardware boundary. The paper's central claim is that this coexistence is achieved by splitting the system into four hardware partitions — task management, shared resources, hard-real-time execution, and long-timescale services — and by giving readout two outputs: a fast-result sideband for same-round feedback and a Measurement Packet with timestamps, resource identifiers, and configuration versions for traceable services. In transaction/event/behavioral models, the shared feedback path reaches $P_{99} = (1.984\\pm0.004)L_{\\max}$ at background load 0.8 with zero deadline violations, and closed-loop calibration reduces mean frequency error by $83.2\\%\\pm0.8\\%$. The paper states that this evaluation is at the architectural level and does not replace cycle-accurate RTL, power/performance/area analysis, or hardware measurement.","feed_headline":"Quantum-control architecture caps feedback delay at 1.98x ideal","feed_subtitle":"Four-way hardware split keeps calibration, error correction, and deterministic readout together; frequency error falls 83 percent.","key_machinery":"The load-bearing mechanism is the dual-output readout interface and the safe-commit configuration path that surrounds it. The Measurement Packet is a traceable record carrying timestamps, resource identifiers, features, and the configuration version used during acquisition; the fast-result sideband is a compact classifier output that bypasses packet assembly and drives the Fast Feedback Unit directly. Around this pair, four hardware partitions isolate task management, shared resources, hard-real-time execution, and long-timescale services, while the Safe-Commit Controller switches a shadow configuration to active only after version, dependency, validity, and safe-point checks pass. Tile-local closure complements this: each Control/Readout Tile runs local gates, readout, active reset, and round-critical QEC, with only cross-tile and long-timescale traffic entering the global interconnect.","core_discovery":"The central discovery is that the tension between hard-real-time control and long-timescale closed-loop services can be managed by a boundary rather than by a single optimized datapath. On the QCORE boundary, a fast-result sideband closes same-round actions from the classifier, while a complete Measurement Packet carries traceable service data; a result may enter the feedback path only after deadline, freshness, and version checks, and a long-term parameter may become active only through a versioned shadow-to-active safe commit. Round-critical QEC actions close locally within replicable Control/Readout Tiles, so global service pressure drops. Under the paper's modeled system this produces zero deadline violations at 0.8 background load, an $83.2\\%\\pm0.8\\%$ reduction in pre-round frequency error, a drop in maximum-drift state-assignment error from $10.39\\%\\pm0.54\\%$ to $5.37\\%\\pm0.29\\%$, no unsafe or mixed-version configuration acceptance in 100,000 transactions, and a $2.08\\times$ capacity-normalized tile-scaling estimate.","pith_inferences":["If cycle-accurate RTL confirms the modeled P99, feedback budgets for error-corrected machines should be set by the classifier-to-feedback path, not by the full measurement-to-host round trip.","The $2.08\\times$ scaling ratio is a trace-derived provisioning envelope; real NoC arbitration, SRAM bank conflicts, clock-domain crossing, and front-end behavior could shrink or enlarge it.","The gated-NPU design suggests AI-assisted readout is affordable only while low-confidence requests stay rare; persistent drift would turn the shared NPU into a contention point and should be stress-tested.","A direct experimental extension would be an FPGA prototype of one tile with the safe-commit controller, measuring P99 feedback latency under synthetic background traffic and comparing it with the $1.984 L_{\\max}$ prediction."],"forward_implications":["A single quantum-processor-side digital layer can carry deterministic control and closed-loop services without dedicating a separate real-time path to each service.","Same-round conditional branch, active reset, and frame update can be driven from the fast-result sideband without waiting for full Measurement Packet assembly.","New closed-loop workloads can be added as profiles, feature pipelines, or kernels because calibration and error correction share a common service-control skeleton.","Configuration updates can be applied atomically, so long unattended runs can change calibration parameters without corrupting the active real-time state.","If tile-local closure holds, scaling to more qubits mostly means replicating tiles and local buffers rather than duplicating global management and service hardware."],"supporting_citations":[{"why":"supplies the pulse-level programming context above the QPU-side boundary that QCORE is positioned to complement.","marker":"[1]"},{"why":"supplies the integrated control/readout platform that serves as the integration baseline.","marker":"[4]"},{"why":"supplies the precisely timed control microarchitecture that serves as the timing baseline.","marker":"[5]"},{"why":"demonstrates same-round conditional feedback via projective measurement, the capability QCORE systematizes.","marker":"[10]"},{"why":"supplies the sub-microsecond decoding-feedback QEC baseline against which tile-local QEC is positioned.","marker":"[14]"},{"why":"supplies the FPGA neural-network real-time decoder baseline for QEC feedback.","marker":"[15]"}],"fun_headline_variants":["Quantum control architecture hits 1.98x ideal feedback latency","QCORE splits control to cut frequency error by 83%","Boundary-based quantum control slashes feedback delay","Four-way hardware split keeps quantum control real-time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the transaction/event simulator and the three-level transmon dispersive-readout model capture the timing, resource contention, and physical behavior of a real QCORE implementation; the paper itself says the evaluation does not replace cycle-accurate RTL, PPA analysis, or hardware measurement.","fun_headline_variants_meta":{"raw":{"variants":["Quantum control architecture hits 1.98x ideal feedback latency","QCORE splits control to cut frequency error by 83%","Boundary-based quantum control slashes feedback delay","Four-way hardware split keeps quantum control real-time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000631,"raw_usage":{"total_tokens":2974,"prompt_tokens":1068,"completion_tokens":1906,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":684,"completion_tokens_details":{"reasoning_tokens":1841}},"tokens_in":684,"tokens_out":1906,"duration_ms":14218,"temperature":1.0,"reasoning_tokens":1841,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:08:30.061999+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build or simulate the QCORE partitions at cycle-accurate RTL level with the same $40L_{\\max}$ transaction period and 0.8 background load, and measure the P99 of the Measurement Packet/Event feedback path; if P99 rises materially above $(1.984\\pm0.004)L_{\\max}$ or deadline violations become nonzero, the real-time-isolation claim fails.","supporting_citations":[{"cited_title":"Qiskit pulse: programming quantum computers through the cloud with pulses[J]","cited_arxiv_id":null,"evidence_quote":"supplies the pulse-level programming context above the QPU-side boundary that QCORE is positioned to complement."},{"cited_title":"The QICK (Quantum Instrumentation Control Kit): Readout and control for qubits and detectors[J]","cited_arxiv_id":null,"evidence_quote":"supplies the integrated control/readout platform that serves as the integration baseline."},{"cited_title":"An experimental microarchitecture for a superconducting quantum processor[C]//Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture","cited_arxiv_id":null,"evidence_quote":"supplies the precisely timed control microarchitecture that serves as the timing baseline."},{"cited_title":"Feedback control of a solid-state qubit using high-fidelity projective measurement[J]","cited_arxiv_id":null,"evidence_quote":"demonstrates same-round conditional feedback via projective measurement, the capability QCORE systematizes."},{"cited_title":"Real-time Surface-Code Error Correction Using an FPGA-based Neural-Network Decoder","cited_arxiv_id":"2605.04892","evidence_quote":"supplies the FPGA neural-network real-time decoder baseline for QEC feedback."}],"review_version":1}