{"id":"407bc35a-a62f-4e2d-a04d-2b84677c2658","arxiv_id":"2608.10553","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"RCCP retrieves similar past residuals to set an asymmetric interval shape and calibrates that interval with a scalar conformal correction.","lead":"This paper introduces a way to build prediction intervals for time series forecasts by retrieving past errors from similar situations and then adjusting the interval width with a single correction factor. The method keeps coverage close to the target level while producing narrower intervals than existing local conformal methods.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's whole guarantee rides on Assumption 1 (|F_test(c*) − F_cal(c*)| ≤ ρ_n), never estimated in the paper, and structurally suspect: calibration B_j uses a small growing knowledge base while test B_t uses the full calibration set, so F_test and F_cal differ even under stationary residuals.","rationale":"The paper's goal is a retrieval-locality plus conformal-correction recipe whose intervals are both locally adaptive and coverage-calibrated, with Theorem 1 as the theoretical warrant and four benchmarks as evidence. For the central claim to hold, F_test(B) and F_cal(B) must be close at the operating multiplier; that is precisely Assumption 1. I agree with the Reader that this is the weakest load-bearing premise, and I sharpen the reason: the online schedule of Algorithm 1 makes the B process non-identically distributed across calibration and test by construction (growing, then much larger, knowledge base), so Assumption 1 is not merely unverified but structurally suspect even under stationarity. What is at stake: without a measured ρ_n, the theorem's bound is a conditional statement with an unknown constant; the empirical coverage could be consistent with a large ρ_n and a small realized gap, so 'attains target coverage' is not corroboration of the theory. The proposed check directly measures the object of Assumption 1 on the paper's own data, and the synthetic control isolates the KB-growth mechanism from all time-series effects. Credit: the component ablations are a genuine internal falsification exercise — removing correction induces undercoverage (−1.35 to −3.01 pp) and adding it restores coverage across α ∈ {0.05, 0.10, 0.15}, which supports the design's corrective role. The released code, frozen forecasts, and held-out test evaluation strengthen the empirical claims. The paper's Limitations section honestly flags retrieval-representation dependence, though it does not acknowledge the calibration-test knowledge-base asymmetry flagged here. Verdict: CONDITIONAL remains the right disposition. The method shows consistent empirical value; the theoretical contribution is a valid conditional bound whose key premise must either be verified on data as proposed, or the claim explicitly demoted to 'conditional on unmeasured ρ_n'. My read does not move the Reader's verdict, hence UNCHANGED, with the condition made concrete.","tokens_in":19267,"tokens_out":15126,"duration_ms":149118,"concrete_test":"With the released code, on Air and Solar with both backbones at α=0.1, compute the online calibration scores {B_j : j ∈ I_cal} exactly as in Algorithm 1 and the test scores {B_t : t ∈ I_test}, then report: (i) the two-sample Kolmogorov–Smirnov statistic and p-value between the two distributions; (ii) the empirical value |F̂_test(ĉ) − F̂_cal(ĉ)| at the operating correction factor, i.e., the realized ρ_n. If KS rejects at 0.01, or the empirical ρ_n exceeds about 1 percentage point of coverage, Assumption 1 is materially violated on the paper's own data and Theorem 1 does not certify the reported coverage; the theory should be presented as purely conditional and the coverage claim as empirical only. As a control, rerun on a synthetic stationary AR(1) with i.i.d.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Theorem 1, every non-exchangeability effect is absorbed by Assumption 1: |F_test(c*) − F_cal(c*)| ≤ ρ_n, where c* is the oracle multiplier. ρ_n is never bounded, estimated, or diagnosed. The mechanism that makes it non-negligible is built into Algorithm 1. During calibration, B_j in Eq. (6) is computed with the knowledge base containing only the training split plus earlier calibration indices; during testing, B_t is computed with the full calibration set plus earlier test indices. Retrieval at test time therefore draws from a larger, denser candidate pool, so the retrieved scales R̃⁺_t, R̃⁻_t in Eq. (4) are quantiles of a different neighborhood distribution even when the raw residual process is exactly stationary and exchangeable. Hence the normalized retrieval error B is non-stationary by construction across calibration and test, and Assumption 1 is not a benign regularity condition: it is the substantive claim the paper needs and does not verify. The proof also requires B to be a well-defined score: Lemma 1 (Appendix A) asserts R̃⁺_t, R̃⁻_t > 0, but a retrieved block of zero residuals (common on flat stretches such as Electricity) makes Eq. (6) undefined. The Table 2 coverage gaps of +0.2 to +1.8 pp for RCCP are exactly the magnitude ρ_n could absorb, so the reported 'attains target coverage' result cannot be attributed to the bound. The paper's own Limitations section flags retrieval-representation dependence but not this calibration-test knowledge-base asymmetry. Because the bound is conditional and ρ_n is unchecked, the central theoretical claim is currently unfalsifiable by the paper's own experiments.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-12T21:59:03.342101+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}