{"id":"22f5b7f1-622a-4559-bd50-cd556fab6c29","arxiv_id":"2607.26887","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"CoRAS adaptively upper-bounds each image’s reconstruction stopping time from its early residual path, with finite-sample marginal coverage and lower average sampling than fixed-rate conformal rules.","lead":"CoRAS picks how many measurements each image needs by reading its early reconstruction path, then conformal-calibrates a stopping time so error stays under a target with high probability. It matters for MRI and other costly sensors that want fewer samples without sacrificing reliability.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Error-control claim rests on unvalidated loss monotonicity; experiments only report stopping-time coverage Û≥T, not L(Û)≤c.","rationale":"The marginal conformal construction (Theorem 2) is standard and correctly specialized under exchangeability, symmetric bandwidth selection, and the stated score residualization; I see no internal proof break. The efficiency and complexity-adaptive empirical patterns in §5 are real on the reported metrics. The single point where the strongest claim is least secure is exactly the reader’s weakest assumption: monotonicity is load-bearing for turning Û≥T into L(Û)≤c, is not validated, and is silently baked into the experimental coverage definition. The MRI residual observability issue is a genuine prospective-sensing limitation but secondary to the theorem–experiment bridge for the claim as worded. No verdict shift is needed: CONDITIONAL already matches—accept-shaped method with those caveats. Concrete monotonicity and direct L(Û)≤c checks would either clear the concern or quantify how much the error-control headline overstates the evidence.","tokens_in":29497,"tokens_out":727,"duration_ms":58793,"concrete_test":"On the released Fashion-MNIST and M4Raw evaluation splits, (i) count the fraction of images for which L_i(t) increases for any t (and the mean number of up-steps), and (ii) recompute coverage as the direct fraction with L_i(Û_i)≤c alongside the reported Û_i≥T_i rate. If direct error coverage falls below 1−α by more than sampling noise, or if >5–10% of paths are non-monotone, the strongest claim’s error-control half is not empirically supported as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 2 converts stopping-time coverage into reconstruction-error control only via the assumption that L_i(t) in (4) is nonincreasing in t with probability one (also used in Theorems 1 and 3). The paper never checks this on the fitted models. Deep reconstructors need not produce monotone per-pixel MSE along a nested acquisition path; the M4Raw training recipe even adds explicit monotonicity and zero-filled-anchor penalties (Appendix A.2), which is evidence the authors treat non-monotonicity as a live risk. Critically, §5 defines empirical “coverage” as the fraction of test images with Û_i ≥ T_i and states this is “equivalent” to L≤c only under monotonicity—so the headline experimental support for P{L(Û)≤c}≥1−α is not a direct measurement of the claimed quantity. If violations are non-negligible, both the theorem-to-practice bridge and the reported coverage numbers weaken, while efficiency comparisons remain well-defined. (Separately, as the reader notes, M4Raw next-band residuals I_t are computed from the fully sampled RSS reference and are not online-observable, so that experiment is retrospective for the horizontal step.)","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes Conformalized Rate-Adaptive Sensing (CoRAS), which chooses an image-specific acquisition stopping time so that reconstruction error stays below a user target c with probability at least 1−α. After a fixed early decision time t0, a horizontal step fits a monotone polynomial to the log next-band residual history and extrapolates a plug-in stopping time; a vertical step then applies full conformal prediction with state-dependent weights on (horizontal prediction, decision-time entropy), yielding a data-dependent upper bound Û^HV. Theorem 2 gives finite-sample marginal validity under exchangeability, symmetric bandwidth selection, and nonincreasing loss; Theorem 3 gives an approximate conditional guarantee under an RKHS bias model. Experiments on Fashion-MNIST and M4Raw report target stopping-time coverage, lower average/excess sampling than fixed-rate conformal rules, and higher rates on high-entropy images.","tokens_in":29900,"tokens_out":1419,"duration_ms":33076,"significance":"The problem—when to stop acquiring measurements under a reconstruction-error guarantee—is practically important in imaging and fits a broader class of pathwise cost–quality trade-offs. The technical contribution is genuine: path-based prediction of an unobserved stopping time, embedded in full conformal calibration with a carefully symmetrized bandwidth selector, plus a model-based approximate conditional theory. Marginal validity (Theorem 2) is standard but correctly specialized; the online O(Gn + G N_cand + n N_cand) bandwidth algorithm and released code support reproducibility. If the error-control claim holds in prospective settings with verified monotone loss, CoRAS is a useful addition to conformal risk/decision methods and adaptive sensing.","major_comments":[{"comment":"Theorems 1–3 convert stopping-time coverage {T ≤ Û} into reconstruction-error control {L(Û) ≤ c} only through the assumption that L_i(t) in (4) is nonincreasing in t with probability one. Section 5 defines empirical coverage as the fraction with Û_i ≥ T_i and states this is “equivalent” to L ≤ c under that assumption, but the paper never reports the fraction of paths with non-monotone MSE, nor the direct rate P{L_{n+1}(Û) ≤ c}. Appendix A.2 even trains the M4Raw U-Net with monotonicity and zero-filled-anchor penalties, which indicates non-monotonicity is a live modeling risk. Without these checks, the headline experimental support for the abstract’s error-control claim is indirect. Please report (i) empirical monotonicity-violation rates on both datasets/models and (ii) direct error-control coverage L(Û) ≤ c alongside Û ≥ T.","section":"§5; Theorems 1–3; Eq. (4); Appendix A.2"},{"comment":"For multi-coil M4Raw, Appendix A.2 states that next-band residuals I^MRI_{i,t} are computed from the fully sampled RSS reference because coil combination is nonlinear, so the horizontal step is not online-observable and the evaluation is retrospective. The main text (§5.2, abstract, introduction) still presents M4Raw as evidence for adaptive sensing alongside Fashion-MNIST. This overstates what the MRI experiment demonstrates: calibration of a stopping rule given oracle residual histories, not prospective rate-adaptive acquisition. Either restrict M4Raw claims to retrospective validation, or supply an online-observable surrogate for I_t and re-run the adaptive rule with that surrogate.","section":"§5.2; Appendix A.2; Abstract"},{"comment":"Figures 2(a) and 4(a) show coverage falling below 1−α on the highest true-entropy bins for all rules, including CoRAS, despite increased sampling on those bins. Theorem 3’s approximate conditional bound depends on residual RKHS bias Λ_{κ,n} and neighborhood MMD D_{κ,n}; the paper does not diagnose whether the high-entropy gap is large horizontal bias, poor state matching, discrete-grid effects, or insufficient t0. A short ablation (e.g., larger t0, state without entropy, calibration-only vs exact bandwidth) and explicit discussion of when approximate conditional validity fails would make the conditional claims falsifiable rather than only marginally reassuring.","section":"§5.1–5.2 Figures 2–4; Theorem 3; Assumption 2"}],"minor_comments":[{"comment":"Assumption 1 and Propositions 1–2 motivate the log-linear/polynomial residual model on a log-spaced grid, but Fashion-MNIST uses uniform column increments θ_t = t/32. A sentence in §4.1.2 already notes the polynomial remains a local approximation on general grids; briefly quantify fit quality (e.g., CV residual or calibration of T̂^H vs T) so readers can judge misspecification severity.","section":"§4.1.2; §5.1"},{"comment":"The decision time t0 is a free design choice with large efficiency impact. Sensitivity of coverage and excess sampling to t0 (beyond the single values t0=6 and t0=9) would strengthen the experimental section.","section":"§5"},{"comment":"Clarify in the main text (not only Appendix C) whether experiments use exact candidate-specific bandwidth selection or the calibration-only approximation; Theorem 2 covers only the symmetric exact procedure.","section":"§5; Appendix C.3"},{"comment":"Minor notation: raw vs model losses L^raw vs L, and floored T^⋆ vs T, are clear once introduced but dense in §4.2; a small notation table would help.","section":"§3–4"},{"comment":"Related work on active MRI and early time-series classification is appropriate; a brief pointer to anytime/sequential conformal testing literature would better situate the Discussion’s sequential-extension limitation.","section":"§2; §6"}],"recommendation":"major_revision","confidential_remarks":"The contribution is real and the marginal conformal construction is carefully done; I would not reject on novelty or theory alone. The monotonicity gap and retrospective MRI setup are fixable with additional experiments and clearer claim scoping, which is why I recommend major rather than minor revision. Fit for a serious stat.ML / methods venue is good if those points are addressed."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a real method paper, not a thin wrapper. They take the practical question “when have I acquired enough for this image?” and turn it into a conformal upper bound on an image-specific stopping time, built from an early reconstruction path.\n\nWhat is new is the object and the two-step construction. Horizontal step: ordered-tail residual model plus image-wise monotone polynomial extrapolation of next-band energy. Vertical step: full conformal correction on a state (horizontal prediction, decision-time entropy), with a careful symmetry condition so bandwidth selection does not break exchangeability. Theorem 2 is standard rank conformal, but specialized cleanly. Theorem 3’s approximate conditional story (RKHS bias + MMD of neighborhoods) is a genuine extra, even if it is stated for fixed bandwidth rather than the data-driven selector they actually run. Experiments on Fashion-MNIST and M4Raw hit target stopping-time coverage, beat fixed-rate conformal baselines on average/excess rate, and spend more on high-entropy images. Code is linked.\n\nSoft spots, in proportion. (1) Error control is stopping-time coverage plus “L nonincreasing w.p. 1.” They never plot or tabulate whether the fitted reconstructors actually have monotone MSE along the path; M4Raw training even adds monotonicity and anchor penalties, which is a tell. Section 5 reports Û ≥ T and calls it equivalent to L ≤ c under that assumption—so the headline empirical claim for P{L(Û)≤c} is one step removed from what they measure. Efficiency comparisons still stand. (2) Multi-coil next-band residuals are computed from the fully sampled RSS reference, so the MRI horizontal step is retrospective, not online. Fashion-MNIST is cleaner on that point. (3) Conditional coverage still dips on the hardest entropy bin; they own that. Free parameters (t0, c, α, bandwidth grid, poly order) are normal for this genre.\n\nCitations look appropriate (conformal, risk control, synthetic control, adaptive MRI). Math is careful rather than ornamental. Who it is for: people in conformal methods, accelerated MRI / adaptive sensing, and anyone calibrating path-wise stopping under a loss budget.\n\nI would send it to peer review. Ask referees to require a direct L(Û)≤c check (and a monotone-violation rate) and a clearer prospective-vs-retrospective statement on MRI. Engage with it; cite if you work on conformal stopping or rate-adaptive imaging.","headline":"Solid conformal method for image-specific stopping times; main caveat is that error-control experiments lean on unchecked loss monotonicity and a retrospective MRI residual.","tokens_in":30431,"tokens_out":617,"would_cite":true,"duration_ms":19725,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"CoRAS picks a per-image acquisition rate from the early reconstruction path and still keeps reconstruction error below a target level with high probability.","keywords":["conformal prediction","adaptive sensing","image reconstruction","stopping times","rate-adaptive acquisition","MRI","distribution-free coverage"],"falsifier":"On a held-out exchangeable test set with a nonincreasing reconstruction loss, check whether the fraction of images whose CoRAS stop still has loss above the target exceeds α, or whether average measurements no longer beat the fixed-rate conformal baseline while coverage holds.","tokens_in":30385,"feed_emoji":"📷","tokens_out":802,"duration_ms":19883,"temperature":0.7,"pith_summary":"High-resolution imaging often cannot afford to collect every measurement for every image, yet the reconstructed image still has to meet a quality bar. This paper develops CoRAS, a stopping rule that watches how a reconstruction model improves as more measurements arrive, predicts when the error will first fall below a chosen target, and then calibrates that prediction against similar past images. The result is an image-specific acquisition rate with a finite-sample guarantee that the reconstruction loss stays at or under the target with probability at least 1−α. Fixed-rate conformal rules meet the same guarantee only by oversampling easy images; CoRAS spends fewer measurements on average and gives harder images more samples. The method is aimed at sensing and compression settings where cost is paid along a path and the true image is not yet fully observed when the stop decision must be made.","feed_headline":"Stop scanning sooner—and still hit the error target","feed_subtitle":"CoRAS sets a per-image acquisition rate from the early reconstruction path with a finite-sample error guarantee.","key_machinery":"Horizontal–vertical conformal stopping: a monotone polynomial fit to next-band residual history predicts the unresolved error tail; a state of that prediction plus decision-time entropy matches calibration images; full conformal residual ranks turn the corrected estimate into an upper bound on the target stopping time.","core_discovery":"From the early part of an image’s reconstruction path, CoRAS builds a horizontal estimate of the first time the reconstruction error falls below a target level, then applies a vertical conformal correction using calibration images with similar early states, yielding a stopping time that controls reconstruction error with marginal validity and approximate conditional validity across image complexity.","pith_inferences":["If reconstruction models with non-monotone error paths become common, the paper’s error-control step would need a different link from stopping time to loss, or a different score.","Prospective multi-coil MRI would require an online observable stand-in for next-band residuals; the current MRI results are retrospective on that point.","Sequential re-decisions along the path, flagged as future work, would need a new validity argument because naive reuse of conformal ranks can break exchangeability."],"forward_implications":["Imaging systems can replace a single global sampling rate with per-image rates that still meet a stated error probability.","Harder images automatically receive more measurements; easier images stop earlier, lowering average acquisition cost.","The same path-plus-conformal template applies wherever quality improves step by step and the decision must be made before the outcome is fully observed.","Approximate conditional coverage is stated given the early reconstruction state the rule can actually see, not given unobserved true-image covariates."],"fun_headline_variants":["CoRAS stops acquisition early with error guarantees","Early reconstruction path sets per-image scan rates","Conformal bound cuts measurements while hitting error targets","Adaptive sensing from early path with coverage control","Fewer scans on easy images, more on hard ones"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Reconstruction error is assumed to never get worse as more measurements are added, so stopping after the true first good time still keeps error under the target.","fun_headline_variants_meta":{"raw":{"variants":["CoRAS stops acquisition early with error guarantees","Early reconstruction path sets per-image scan rates","Conformal bound cuts measurements while hitting error targets","Adaptive sensing from early path with coverage control","Fewer scans on easy images, more on hard ones"]},"model":"grok-4.5","effort":"low","cost_usd":0.003322,"raw_usage":{"total_tokens":1054,"prompt_tokens":696,"num_sources_used":0,"completion_tokens":53,"cost_in_usd_ticks":33224000,"prompt_tokens_details":{"text_tokens":696,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":305,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":696,"tokens_out":53,"duration_ms":5796,"temperature":1.0,"reasoning_tokens":305,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T18:04:01.618911+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a held-out exchangeable test set with a nonincreasing reconstruction loss, check whether the fraction of images whose CoRAS stop still has loss above the target exceeds α, or whether average measurements no longer beat the fixed-rate conformal baseline while coverage holds.","supporting_citations":[],"review_version":1}