{"id":"579378b6-5003-4439-a09b-173f088d6ee1","arxiv_id":"2508.21231","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Daily calibration data from a 20-qubit NISQ device cluster into stable and noisy qubit groups, and GHZ experiments confirm the stable group runs more reliable circuits.","lead":"A dense analysis of 250 days of calibration data from a 20-qubit superconducting processor shows qubits split into stable and noisy groups, and that circuits on the stable group run more reliably. The framework targets operators scheduling quantum jobs inside HPC centers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GHZ 'validation' in Section IV is confounded: overlapping circuits, no stated out-of-sample split, no readout-error correction, and no significance test, so the central predictive claim is not established.","rationale":"The reader's weakest assumption points to procedural drift in the 250-day log. That is a fair data-quality caveat, but it is not the most directly load-bearing point: even a perfectly clean log would not rescue the central claim if the GHZ validation cannot be interpreted. The more decisive weakness is that Section IV's confirmation of the clusters is not a clean, independent test. The two reported 5-qubit circuits overlap at qubit 8, the clustering features and GHZ readout may share the readout-fidelity channel, the GHZ measurement window is not shown to be out-of-sample relative to the clustering period, and no statistical test is given. These are correctable empirical artifacts, not mathematical contradictions, so the right disposition remains conditional acceptance with a concrete validation requirement. The paper does have real, long-duration operational data and shows agreement among several clustering algorithms, which supports internal reproducibility; however, algorithmic agreement within the same dataset does not establish that the clusters predict circuit performance. No code or data release is noted, which further limits independent verification but is secondary to the validation-design flaw.","tokens_in":11145,"tokens_out":7028,"duration_ms":74390,"concrete_test":"Run an out-of-sample, readout-controlled validation: derive clusters from the first 200 days only; on the remaining 50 days, measure 2-qubit GHZ fidelity for every coupled pair and 5-qubit GHZ on several disjoint qubit sets drawn from the predicted stable and noisy clusters, using readout-error-mitigated fidelities (or reporting raw and corrected). Compute the mean GHZ-fidelity difference between predicted families and its permutation p-value across all pairs/sets. If the difference is not significant at p<0.05, or if it disappears after readout-error correction, the Section IV validation does not support the clustering claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on Section IV's GHZ experiment, but the experiment as reported cannot distinguish genuine cluster predictive power from confounds. (1) The two 5-qubit circuits overlap: 'good' {13,8,12,17,14} and 'bad' {3,0,2,8,4} both contain qubit 8, so the comparison is not between disjoint stable/noisy families. (2) The Section III.C clustering features include readout and gate fidelities; GHZ fidelity is extracted from measured populations, and the paper does not state that readout errors were mitigated. If the 'bad' set simply has worse readout, the GHZ gap partly re-detects a feature already used to build clusters. (3) No temporal split is described: clusters are built from the 250-day calibration record, while the GHZ 'measurement window' is not stated to be outside that record, so agreement may be in-sample descriptive correlation rather than prediction. (4) Only one pair of circuits is reported, with no significance test; 0.74±0.05 vs 0.63±0.14 overlap within the reported spread. The Section IV sentence that circuits mapped to robust clusters yield more reliable outcomes therefore goes beyond what the evidence shows.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 250 days of calibration data from a 20-qubit IQM superconducting processor in an HPC environment. It studies temporal autocorrelation and cross-metric correlations of T1, T2, readout fidelity, and single-/two-qubit gate fidelities. The authors then apply unsupervised clustering (KMeans, GMM, Spectral, Node2Vec+KMeans) to group qubits into stable and noisy families, and report GHZ-state experiments that they interpret as validating the clusters. The paper concludes with recommendations for recalibration scheduling and circuit mapping in HPC-integrated quantum systems.","tokens_in":11454,"tokens_out":4654,"duration_ms":49049,"significance":"If the central claim were fully established, this would be a practically valuable contribution: daily calibration logs alone could identify reliability families of qubits that predict real circuit performance, informing HPC workload scheduling. The paper's strengths are its real long-term operational dataset (250 days, 20 qubits), the systematic comparison of four correlation methods, the use of multiple clustering algorithms, and the direct hardware validation attempt. These are concrete, reproducible-in-principle empirical contributions to a field where such longitudinal calibration studies are rare. The paper does not ship code or data, and the validation as reported is too weak to support the predictive claim, but the underlying framework is plausible and worth strengthening.","major_comments":[{"comment":"The GHZ validation is not sufficient to support the claim that 'circuits mapped to robust clusters indeed yield more reliable experimental outcomes.' (1) The two 5-qubit circuits overlap: {13,8,12,17,14} and {3,0,2,8,4} both contain qubit 8, so the comparison is not between disjoint stable and noisy qubit sets. (2) The reported fidelities 0.74±0.05 and 0.63±0.14 overlap within uncertainty, and no significance test, number of repetitions, or shot counts are given. (3) No temporal split is described; clusters are built from the 250-day calibration record, while the GHZ measurement window is not stated to be outside that record, so the agreement could be an in-sample descriptive correlation rather than out-of-sample prediction. (4) Readout error is not mitigated or corrected, and readout fidelity is one of the clustering features; GHZ fidelity extracted from measured populations can therefo","section":"Section IV, Fig. 9(b)"},{"comment":"The clustering setup is under-specified in ways that affect the central claim of robust stable/noisy families. The paper states that the optimal number of clusters is chosen by maximizing the silhouette score, but it does not report silhouette values or the selected k. Node2Vec hyperparameters (embedding dimension, walk length, number of walks, context window, p, q) are not given, nor is it explained how the 6-dimensional metric feature vector and the graph embedding are concatenated and normalized. The claim that KMeans and Spectral 'give the same as' Node2Vec+KMeans is only qualitative; no cluster-agreement metric (e.g., adjusted Rand index) is provided. Without these details, it is hard to assess whether the stable/noisy split is intrinsic or an artifact of particular hyperparameters. Please report the full clustering pipeline and quantitative agreement across algorithms.","section":"Section III.C"},{"comment":"The analysis relies on several time windows that are not defined. Section II.D discusses probability distributions computed over 'a representative time window' without specifying its start, end, or selection criterion. Section III.B computes metric correlations 'over 80 days after cool-down' but does not state which cool-down event (the paper mentions warm-ups at Day 130 and Day 180) or why 80 days was chosen. If these windows were chosen after inspecting the data, the reported correlations and cluster features could be influenced by selection bias. Please define the windows a priori or show that the clustering and correlation conclusions are stable across multiple sub-periods.","section":"Sections II.D and III.B"},{"comment":"The paper assumes 'no procedural drift' because the calibration schedule and pulse/readout configurations are fixed, yet the same section and Section II.C describe two cryogenic warm-up/cool-down cycles and incremental recalibration relying on prior parameters. These are procedural changes that visibly affect gate fidelities (the 'two pronounced boundaries' in Fig. 4). The clustering is performed over the entire 250-day record, so the stable/noisy classification may partly reflect qubit responses to these global events rather than intrinsic qubit health. I recommend showing that the clusters are stable when computed separately for the intervals before and after each warm-up, or that the clustering features account for the regime changes.","section":"Section II, first paragraph after dataset description"}],"minor_comments":[{"comment":"Hadamard circuit results are mentioned as part of the hourly benchmarks but no Hadamard fidelity data are shown; either include them or remove the reference.","section":"Section IV"},{"comment":"The notation for T2 is inconsistent: T2*, T2^echo, and T2 are used without a single consolidated definition. Please define all variants at first use and use consistent subscripts/superscripts throughout.","section":"Section II.A"},{"comment":"The text says 'qubit 6 has a sharpest drop' and 'qubit 3 exhibits persistent variability'; please adjust grammar and make clear whether these are the same qubits flagged by the standard-deviation bar plots.","section":"Section II.C"},{"comment":"The captions state that KMeans and Spectral give the same result as Node2Vec+KMeans, but the figure does not show those results. A quantitative cluster-comparison metric (e.g., adjusted Rand index) would be more informative than the verbal claim.","section":"Figure 8"},{"comment":"The phrase 'over 80 days after cool-down' is ambiguous: after which cool-down, and why 80 days? Please specify the exact date range and justify the choice.","section":"Section III.B"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable empirical systems study, but the validation section is currently the weakest link. The GHZ comparison has overlapping qubit sets, overlapping error bars, no significance testing, and no out-of-sample temporal split. These issues are fixable within the manuscript's scope, so I recommend major revision rather than rejection. Given the paper's reliance on proprietary operational data and closed-source analysis, I would also encourage the authors to release at least anonymized aggregate data or code to make the claimed robustness checkable. The topic fits an HPC/quantum-systems venue, but the predictive claims need to match the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful operational case study with real 250-day calibration data from a 20-qubit IQM chip, but the headline claim—that clustering predicts circuit performance—is not backed by the validation as reported. The GHZ experiment has multiple confounds, so the paper overstates what the data show.\n\nWhat's actually new and good: a 250-day longitudinal dataset from a production HPC-integrated device is genuinely rare. The descriptive analysis (heatmaps, distributions, heavy tails) is careful and reads as honest engineering. The autocorrelation analysis is a practical touch and the comparison of four correlation methods is clean. The consistency of clustering across KMeans, GMM, Spectral, and Node2Vec is a solid empirical result. The authors also flag some temporal quirks (e.g., T2 echo recorded later), which is a good sign of careful handling.\n\nSoft spots, in order of severity:\n\n1. The GHZ validation does not validate prediction. The good and bad circuits share qubit 8, so the comparison is not between disjoint stable/noisy families. The clustering features include readout and gate fidelities, and GHZ fidelity is not readout-error corrected, so the gap may simply re-detect inputs used to build the clusters. No temporal split is described: clusters come from the 250-day record, but the GHZ measurement window is not shown to be out-of-sample. Only two 5-qubit circuits are reported, with overlapping standard deviations (0.74±0.05 vs 0.63±0.14), and there is no significance test. The sentence claiming circuits mapped to robust clusters yield more reliable outcomes goes beyond the evidence.\n\n2. Clustering configuration is underreported: k, Node2Vec hyperparameters, silhouette scores, and data windows are largely unspecified. This makes the result hard to reproduce or judge.\n\n3. No code or data are released, which is a shame for a data-driven paper.\n\n4. The novelty is incremental: each individual technique is standard, and variability-aware qubit selection was already in the literature (e.g., Tannu and Qureshi). The value is in the integration and the rare dataset, not in a new method.\n\nThe circularity concern is real but not fatal for the clustering itself—unsupervised clustering is a legitimate way to identify structure. What is not legitimate is calling it a predictive model when the validating metric shares the same error sources as the features.\n\nOverall: this deserves a serious referee because real device data over 250 days is valuable and the framework is operationally motivated. But the abstract overstates the evidence. I would recommend major revision: use disjoint circuit sets, apply readout-error correction, run significance tests over more circuits, define a clear out-of-sample split, report all hyperparameters, and release code and data. If the authors do that, this becomes a credible reference for HPC centers operating NISQ devices.","headline":"Useful operational case study with rare 250-day calibration data, but the GHZ validation is confounded and the predictive claim is not established as reported.","tokens_in":11923,"tokens_out":1843,"would_cite":false,"duration_ms":20208,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that routine calibration logs alone—without extra benchmarks—can be clustered to separate reliable qubits from noisy ones, and that GHZ state experiments confirm circuits run better on the reliable group.","keywords":["quantum calibration metrics","unsupervised clustering","qubit health monitoring","NISQ devices","GHZ state validation","high-performance computing","coherence times","circuit mapping"],"falsifier":"Take the first 125 days of the calibration record, cluster qubits from that half alone, then run the same five-qubit GHZ circuits on the predicted stable and noisy clusters during the following 125 days; if the 'noisy' cluster matches or beats the 'stable' cluster in mean fidelity—or if GHZ fidelity differences vanish after controlling for the day of measurement—the central claim that clusters predict circuit reliability is falsified.","tokens_in":11027,"feed_emoji":"⚛️","tokens_out":6115,"duration_ms":56705,"temperature":0.7,"pith_summary":"The paper's claim is that the daily calibration telemetry already collected for a superconducting processor—T1 and T2 coherence times, readout fidelity, and single- and two-qubit gate fidelities, logged once per day for over 250 days on a 20-qubit IQM device—contains enough structure to split the qubits into stable and noisy families. Applying unsupervised clustering that combines each qubit's metric history with its position in the chip's connectivity graph, the authors find the same split under four different algorithms: most qubits sit in one stable group, while a smaller set is persistently noisy. The paper validates this split with real circuits: five-qubit GHZ states run on the stable cluster keep higher and steadier fidelity (0.74 ± 0.05) than the same states run on the noisy cluster (0.63 ± 0.14), and a heatmap of two-qubit GHZ fidelities matches the cluster map. If this is right, an HPC operator can turn an ordinary calibration log into a qubit health indicator that steers circuit mapping, recalibration priority, and maintenance scheduling without extra measurement overhead.","feed_headline":"Calibration logs alone predict which qubits run circuits reliably","feed_subtitle":"250 days of metrics on a 20-qubit processor split qubits into stable and noisy sets; GHZ tests confirm the split.","key_machinery":"Unsupervised clustering is the load-bearing mechanism. Each qubit gets a feature vector from the daily means of six calibration metrics, extended by a Node2Vec graph embedding of the 20-qubit connectivity graph so that both behavior and position count; the number of clusters is chosen by maximizing silhouette score. Before clustering, the paper establishes structure with autocorrelation functions (up to 30-day lags) and four cross-metric correlation measures (Pearson, Spearman, distance correlation, mutual information), which show that fidelity metrics move together, coherence times move together, and T1 has multi-day memory while fidelities forget within a day. The cluster assignments are t","core_discovery":"Using 250 days of once-daily calibration data from a 20-qubit NISQ processor, the paper shows that fidelity metrics—readout, single-qubit, and two-qubit gate fidelity—form one tightly correlated block in all four correlation measures tested, while coherence times form a second block, with weak cross-block coupling. When qubits are clustered from their six-dimensional metric histories plus a Node2Vec embedding of the chip topology, KMeans, GMM, Spectral, and Node2Vec+KMeans all separate the qubits into a dominant stable cluster and a noisy minority. The experimental validation is direct: hourly runs of GHZ circuits show per-pair fidelity patterns that mirror the cluster map, and 5-qubit GHZ s","pith_inferences":["A testable extension the authors leave implicit: cluster labels from the first half of the record should predict GHZ fidelity in the second half; if they do, the method is a genuine predictive model rather than a description of one window.","The same clustering pipeline likely transfers to the 54- and 150-qubit chips mentioned in the conclusion, since it only needs daily calibration tables and a connectivity graph; the main risk is whether cluster structure remains stable as device scale grows.","The strong cross-metric correlation block could be exploited for anomaly detection: a simultaneous drop in T1 and T2echo may be a leading indicator that a full recalibration is needed before gate fidelities visibly degrade.","Beyond scheduling, cluster membership drift over time could itself be a health signal: a qubit moving from the stable to the noisy cluster may be developing a TLS or coupling defect before standard thresholds trip."],"forward_implications":["Qubit placement for jobs can be chosen from daily calibration logs alone, steering error-sensitive circuits to the stable cluster without running extra benchmarks.","Recalibration and diagnostic effort can be prioritized on the noisy cluster, such as qubits 3, 5, 10 and their coupled pairs, which show persistent variability.","Because fidelity metrics are strongly correlated across the chip, monitoring a single representative fidelity (with coherence times as a second signal) is enough to trigger broad health alerts.","The two warm-up events visible in gate-fidelity heatmaps mean global events can degrade large regions at once; cluster-aware scheduling should re-derive clusters after each cryogenic cycle.","Autocorrelation results suggest T1 has multi-day memory while fidelity metrics lose memory after a day, so recalibration intervals can be set per metric class."],"supporting_citations":[{"why":"Supplies the calibration protocols and fit formulas (T1/T2 decay, randomized benchmarking) from which every metric in the dataset is derived.","marker":"[5]"},{"why":"Establishes the variability-aware premise that qubits differ substantially in error rates, motivating the cluster-based scheduling claims.","marker":"[13]"},{"why":"Prior noise-adaptive compiler mapping that this paper's calibration-clustering approach extends and simplifies.","marker":"[11]"},{"why":"Introduces graph-based modeling of calibration dependencies, the conceptual basis for adding topology (via Node2Vec) to metric features.","marker":"[24]"},{"why":"Earlier calibration and performance evaluation of the same Q-Exa device, providing the operational context for the 250-day dataset.","marker":"[34]"},{"why":"Supplies the Node2Vec graph embedding that encodes chip connectivity into each qubit's clustering feature vector.","marker":"[40]"},{"why":"Provides the spectral clustering algorithm used as one of the four clustering methods that agree on the stable/noisy split.","marker":"[41]"}],"fun_headline_variants":["250 days of calibration logs split qubits into stable and noisy sets","Clustering reveals qubit health from 250 days of calibration metrics","Calibration data alone flag noisy qubits, GHZ tests confirm","GHZ experiments validate clustering of qubit health from calibration","Qubit health forecast from calibration logs alone"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The entire analysis assumes the 250-day calibration record is a clean mirror of qubit health—that the daily calibration procedure, pulse shapes, readout settings, and environment stayed consistent enough that every fluctuation in the metrics is due to the device itself, not to how the data was logged.","fun_headline_variants_meta":{"raw":{"variants":["250 days of calibration logs split qubits into stable and noisy sets","Clustering reveals qubit health from 250 days of calibration metrics","Calibration data alone flag noisy qubits, GHZ tests confirm","GHZ experiments validate clustering of qubit health from calibration","Qubit health forecast from calibration logs alone"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000969,"raw_usage":{"total_tokens":3911,"prompt_tokens":653,"completion_tokens":3258,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":397,"completion_tokens_details":{"reasoning_tokens":3173}},"tokens_in":397,"tokens_out":3258,"duration_ms":20498,"temperature":1.0,"reasoning_tokens":3173,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:27:07.741942+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the first 125 days of the calibration record, cluster qubits from that half alone, then run the same five-qubit GHZ circuits on the predicted stable and noisy clusters during the following 125 days; if the 'noisy' cluster matches or beats the 'stable' cluster in mean fidelity—or if GHZ fidelity differences vanish after controlling for the day of measurement—the central claim that clusters predict circuit reliability is falsified.","supporting_citations":[{"cited_title":"Calibration and performance evaluation of a superconducting quantum processor in an hpc center,","cited_arxiv_id":null,"evidence_quote":"Earlier calibration and performance evaluation of the same Q-Exa device, providing the operational context for the 250-day dataset."}],"review_version":1}