{"id":"6bdff3f7-ec51-445b-9188-4188a37f5c55","arxiv_id":"2607.26911","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Near-collinearity of noise-source and load design-matrix columns makes REACH’s five-parameter Bayesian calibration non-reproducible; a four-parameter Chebyshev fit with hot-load T_NS(ν) recovery and spike masking stabilizes it on mock data.","lead":"REACH’s Bayesian noise-wave calibration is numerically unstable: posterior condition numbers hit 10^9–10^11, so identical data give different answers on different computers. The authors pin the cause on a built-in X_NS–X_L degeneracy and show a Chebyshev plus fixed-T_NS (then hot-load recovery) fix that restores reproducibility on mock data.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The load-bearing soft spot is mock-only validation of hot-load T_NS(ν) recovery, which the paper’s own bad-calibrator test already shows can bias the antenna.","rationale":"The reader correctly separates a strong, falsifiable numerical diagnosis (ill-conditioning from X_NS/X_L collinearity; Chebyshev + remove T_NS → κ~60 and bit-stable evidence) from a weaker generalization claim about hot-load T_NS(ν) accuracy. I find the same soft spot: the paper itself flags the failure mode in §3.5/§4/Fig. 17, yet the abstract and conclusions still present the two-step method as achieving comparable accuracy on the strength of one favourable mock. That does not overturn the instability finding or the scalar 4-NWP fix, so the verdict stays CONDITIONAL rather than REJECT—exactly as the reader has it. No stronger internal inconsistency appears in the linear algebra or the conjugate-update formulae. A real-data rerun of the hot-load path is the single check that would settle whether the packaged mitigation is ready for other 21-cm experiments.","tokens_in":22641,"tokens_out":877,"duration_ms":15383,"concrete_test":"Run the full Chebyshev 4-NWP + iterative hot-load pipeline (Fig. 11) on a real REACH lab or on-sky calibration set that includes the hot load, the same cable suite, and an independent validator (e.g. c2r91), with and without the §3.5 spike mask. Report κ(V*), NumPy 1.x vs 2.x residual agreement, and validator/antenna residuals against physical load temperatures. If antenna residuals systematically exceed the scalar-T_NS 4-NWP run by ≳ the mock noise floor (~0.26 K) or reintroduce backend dependence, the hot-load half of the claim does not transfer.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the X_NS–X_L near-collinearity (Eqs. 7, 10; §3.3) drives κ(V*)~10^9–10^11, and that a Chebyshev 4-NWP fit with T_NS fixed (scalar ENR or iterative hot-load recovery, Eqs. 22–25 / Fig. 11) restores κ~60, cross-backend reproducibility, and calibration accuracy comparable to the unstable 5-NWP pipeline. The diagnosis and the scalar 4-NWP conditioning gain are tightly supported on the mocks (SVD, Figs. 8–10; Table 1; Fig. 16). The weaker, load-bearing step is the second half of the mitigation: that iterative hot-load T_NS(ν) recovery remains unbiased and yields “comparable accuracy” once transferred beyond this mock family. Recovery divides the hot-load residual by X_NS^hot (Eq. 24); when that weight is small, or when cable-spike masking removes high-leverage edge channels (§3.5), residual NWP and edge-extrapolation errors are absorbed into T_NS(ν) and amplified on the antenna. The paper’s own stress test already exhibits this: on the deliberately poor 90–130 MHz calibrator set, hot-load+mask raises antenna RMSE from 0.140 K (scalar masked) to 0.284 K (Fig. 17, case d). All quantitative accuracy claims (Table 1; Figs. 13–15) use one REACH-like mock (lab S-parameters, simulated noise parameters, power-law antenna, no cable on the antenna). Without real receiver data, the packaged “stable, data-driven” procedure is only half-validated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper diagnoses a numerical instability in the REACH Bayesian noise-wave calibration pipeline: the posterior covariance reaches κ(V*)∼10^9–10^11, so identical code and mock data yield environment-dependent solutions (NumPy 1.x vs 2.x). SVD and a controlled synthetic experiment attribute the ill-conditioning to monomial collinearity and, more critically, near-collinearity of the design-matrix columns X_NS and X_L (Eqs. 7, 10). Mitigations are a stacked Chebyshev basis, a 4-NWP model that fixes T_NS (scalar ENR or iterative hot-load recovery of T_NS(ν)), and masking of cable standing-wave spikes in κ(X). On REACH-like mock data these steps reduce κ(V*) to ∼60, restore cross-backend reproducibility, and (for the full-band hot-load route) bring antenna residuals to a level comparable to the unstable 5-NWP fit.","tokens_in":23110,"tokens_out":1198,"duration_ms":19895,"significance":"If the diagnosis and mitigations hold on real receiver data, the work identifies a previously under-appreciated reproducibility failure that is inherent to the noise-wave formalism whenever (P_cal−P_L)/(P_NS−P_L)≈1, and supplies a practical, physically motivated fix (4-NWP + optional hot-load T_NS(ν) + spike masking) that other global 21-cm experiments can adopt. Strengths include clear ablations (Table 1), SVD/condition-number diagnostics, a controlled synthetic isolation of the X_NS–X_L driver (Fig. 7), and an explicit cross-backend reproducibility test (Figs. 2, 16). The paper correctly elevates numerical stability to a first-class calibration requirement alongside accuracy.","major_comments":[{"comment":"All quantitative accuracy claims (Table 1; Figs. 13–15) and the packaged “stable, data-driven” procedure rest on one REACH-like mock family (lab S-parameters, simulated noise parameters, power-law antenna, no antenna cable). The central conditioning diagnosis is tightly supported on these mocks, but transfer of the iterative hot-load T_NS(ν) route to real data is not demonstrated. A real-receiver or multi-mock validation (or an explicit limitation statement that accuracy claims are mock-only) is needed before the abstract’s broader claim is warranted.","section":"§4, Table 1, Abstract"},{"comment":"The paper’s own bad-calibrator subband test already shows that hot-load recovery can bias the antenna: with masking, antenna RMSE rises from 0.140 K (scalar 4-NWP) to 0.284 K (hot-load T_NS(ν); Fig. 17 case d). Recovery divides the hot-load residual by X_NS^hot (Eq. 24); when that weight is small or high-leverage edge channels are masked (§3.5), residual NWP/edge errors are absorbed into T_NS(ν) and amplified on the antenna. The abstract and §5 still present hot-load recovery as achieving “comparable calibration accuracy” without stating the regime of validity or a decision rule (when to prefer scalar ENR vs hot-load). That caveat should be quantitative and prominent.","section":"§3.4.3, §3.5, §4, Fig. 17, Abstract"}],"minor_comments":[{"comment":"Inconsistent notation for the design-matrix columns (X_NS vs XNS, bold/unbold V*) and occasional missing spaces in compound words (e.g. “design-matrixcolumns”) should be cleaned for production.","section":"§2.2–3.4"},{"comment":"Figure 10 caption reports κ≈5.94×10^1 while the text sometimes rounds to ∼60; keep a single reported value.","section":"Fig. 10, §3.4.2"},{"comment":"The conjugate-prior hyperparameters (μ0, V0, a0, b0) and the precise Chebyshev degree choices used for the main Table 1 runs should be stated explicitly so the reproducibility test can be repeated.","section":"§2.3, Table 1"},{"comment":"Typos: “Decompostition” (Fig. 4 discussion), “interation” (Fig. 11 caption), and “afactor” / similar run-ons in §3.4.","section":"§3.2, Fig. 11"},{"comment":"Related EDGES iterative calibration (Monsalve et al. 2017) is cited; a slightly sharper statement of what is identical vs analogous (C1/C2 vs T_NS/T_L) would help readers already familiar with that pipeline.","section":"§3.4.3"}],"recommendation":"minor_revision","confidential_remarks":"The conditioning diagnosis and scalar 4-NWP fix are solid and worth publishing; the hot-load accuracy claim is the only overreach and is already contradicted by the authors’ own Fig. 17. Minor revision with tighter claims and an explicit mock-only caveat should be sufficient—no need for major revision or rejection. Scope is appropriate for MNRAS."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that the REACH five-NWP Bayesian pipeline is genuinely ill-conditioned: κ(V*) hits 10^9–10^11, and identical code/data diverge across NumPy backends at the ~0.3 K level. SVD and a controlled synthetic X_NS experiment pin the driver on near-collinearity of the X_NS and X_L columns (inherent when the power ratio is near unity), plus monomial collinearity. That diagnosis is clean and portable.\n\nWhat they do well: they treat numerical stability as a first-class requirement, show the failure mode with condition numbers and cross-backend residuals, switch to a joint Chebyshev update, and demonstrate that fixing T_NS (scalar ENR or held fixed during each fit) drops κ to ~60 and restores bit-level reproducibility. Table 1 and the validator residual plots make the ablation easy to follow. The cable-spike masking section is a useful third diagnostic that other cable-calibrator experiments will recognize. Citations to Monsalve et al. and the REACH conjugate-prior papers are honest about what is borrowed.\n\nSoft spot, in proportion: everything quantitative is one REACH-like mock family (lab S-parameters, simulated noise, power-law antenna). The scalar 4-NWP conditioning gain is solid on that mock. The iterative hot-load T_NS(ν) recovery is the load-bearing second half of the “data-driven” claim, and the paper’s own bad-calibrator subband already shows it can raise antenna RMSE (0.140 K scalar-masked → 0.284 K hot-load). Division by small X_NS^hot plus edge masking is exactly where bias would enter. No real receiver data, no public code. Free parameters (orders, smoother, mask width, priors) are standard but not exhaustively stress-tested outside the mock.\n\nThis is for people who actually run global 21-cm receiver calibration, not for cosmologists hunting a detection paper. The math and the linear-algebra story hold up on what is shown. I would send it to referees; they should demand real-data validation or a clearer scope limit on the hot-load step, not a desk reject. Worth engaging if you care about REACH or noise-wave pipelines.","headline":"Solid diagnosis of a real X_NS–X_L ill-conditioning problem in noise-wave calibration, with a clean 4-NWP fix; the hot-load T_NS(ν) half is only half-validated on mocks.","tokens_in":23810,"tokens_out":585,"would_cite":true,"duration_ms":11467,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A hidden degeneracy in noise-wave calibration makes global 21-cm solutions non-reproducible across computers, and fixing the noise-source temperature restores stability.","keywords":["global 21-cm","noise wave calibration","numerical stability","condition number","Bayesian conjugate priors","Cosmic Dawn","REACH","design matrix degeneracy"],"falsifier":"Run the same mock calibrator suite through the Chebyshev 4-NWP and original monomial 5-NWP pipelines under two different BLAS backends: if κ(V*) stays near 10^10 and residuals still differ by tenths of a kelvin after the 4-NWP fix, or if real lab S-parameter data reverse the mock residual improvement, the claim fails.","tokens_in":23490,"feed_emoji":"📡","tokens_out":1104,"duration_ms":21328,"temperature":0.7,"pith_summary":"Global 21-cm experiments need receiver calibration far more accurate than the bright foregrounds that swamp the faint Cosmic Dawn signal. This paper shows that the standard five-parameter Bayesian noise-wave fit used for REACH is numerically unstable: the posterior covariance condition number reaches 10^9–10^11, so identical data and code give different millikelvin residuals under different linear-algebra backends. The root cause is near-collinearity between the design-matrix columns for excess noise-source temperature and load temperature, which is built into the noise-wave model whenever the Dicke power ratio is near unity. Switching to a Chebyshev basis, fixing the noise-source temperature (first as a scalar, then recovered frequency-by-frequency from the hot load), and masking narrow cable standing-wave spikes drops the condition number to about 60, makes solutions bit-reproducible across environments, and on mock data keeps calibration accuracy comparable to the unstable five-parameter fit. Because the degeneracy is inherent to the formalism, the same fix matters for other global 21-cm receivers.","feed_headline":"Noise-wave calibration was unstable across computers","feed_subtitle":"Fixing one degenerate temperature drops the condition number from 10^11 to ~60 and restores reproducible 21-cm residuals","key_machinery":"The X_NS–X_L degeneracy in the noise-wave design matrix (X_NS = X_L × (P_cal−P_L)/(P_NS−P_L)), diagnosed by SVD and condition number of the posterior covariance V*, and removed by a Chebyshev 4-NWP fit with T_NS held fixed (scalar or hot-load-recovered).","core_discovery":"The REACH Bayesian noise-wave posterior is driven to condition numbers κ(V*) ~ 10^9–10^11 by near-collinearity of the X_NS and X_L design-matrix columns. Fixing T_NS—either to a manufacturer scalar or to a smooth curve recovered iteratively from the hot-load residual—in a Chebyshev four-parameter fit reduces κ(V*) to ~60, restores cross-backend reproducibility, and on mock data yields calibration residuals comparable to the unstable five-parameter pipeline; masking narrow cable standing-wave channels further removes local design-matrix artefacts.","pith_inferences":["If unaddressed, environment-dependent millikelvin residuals could be absorbed into claimed 21-cm absorption features or into foreground model choices, mimicking the kinds of systematics already debated in existing global-signal claims.","The hot-load iteration is essentially a constrained scale/offset self-calibration; experiments without a well-characterized hot load may need an external ENR standard or a different absolute temperature anchor.","Sub-band calibration plus edge-aware masking may become necessary whenever high-order global polynomials are fit across cable-induced spikes.","Publishing κ(V*) and cross-backend residual differences could become a minimal reproducibility checklist for precision radio calibration papers."],"forward_implications":["Global 21-cm pipelines that use noise-wave calibration should treat condition number of V* as a standard diagnostic alongside residual RMSE.","A four-parameter fit with T_NS fixed (scalar ENR or hot-load-recovered) can replace the five-parameter joint fit without sacrificing mock accuracy while making solutions environment-independent.","Cable-connected calibrator suites need automated spike detection and narrow-channel masking where standing waves make X-matrix columns locally collinear.","Other experiments using the same noise-wave formalism (not only REACH) inherit the X_NS–X_L degeneracy and can adopt the same reduction.","Bayesian evidence used for polynomial-order selection becomes trustworthy only after the posterior is well-conditioned."],"fun_headline_variants":["Fixing T_NS drops noise-wave condition number from 10^11 to ~60","X_NS–X_L collinearity made REACH Bayesian calibration non-reproducible","Chebyshev two-step fix stabilizes noise-wave posterior for 21-cm experiments","Hot-load recovery of T_NS(ν) keeps accuracy while ending cross-backend drift","Masking cable standing-wave channels removes local design-matrix artefacts"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That results from one family of REACH-like mock datasets transfer to real on-sky receiver data, including that hot-load recovery of T_NS stays unbiased when the hot-load noise-source weight is small or edge channels are masked.","fun_headline_variants_meta":{"raw":{"variants":["Fixing T_NS drops noise-wave condition number from 10^11 to ~60","X_NS–X_L collinearity made REACH Bayesian calibration non-reproducible","Chebyshev two-step fix stabilizes noise-wave posterior for 21-cm experiments","Hot-load recovery of T_NS(ν) keeps accuracy while ending cross-backend drift","Masking cable standing-wave channels removes local design-matrix artefacts"]},"model":"grok-4.5","effort":"low","cost_usd":0.004507,"raw_usage":{"total_tokens":1425,"prompt_tokens":902,"num_sources_used":0,"completion_tokens":96,"cost_in_usd_ticks":45068000,"prompt_tokens_details":{"text_tokens":902,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":427,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":902,"tokens_out":96,"duration_ms":8039,"temperature":1.0,"reasoning_tokens":427,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T17:17:36.945365+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same mock calibrator suite through the Chebyshev 4-NWP and original monomial 5-NWP pipelines under two different BLAS backends: if κ(V*) stays near 10^10 and residuals still differ by tenths of a kelvin after the 4-NWP fix, or if real lab S-parameter data reverse the mock residual improvement, the claim fails.","supporting_citations":[],"review_version":1}