{"id":"ce62b99c-07d5-42f0-ad56-25ab72972705","arxiv_id":"2506.20173","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A stability-based randomized selection rule, MinSE, allows picking among multiple valid conformal prediction sets, such as the smallest, while retaining a coverage guarantee after selection, and a recalibration procedure removes the coverage inflation in split conformal settings.","lead":"Given several conformal prediction sets that each carry a coverage guarantee, this paper shows how to select one, for instance the smallest, through a randomized stability-based mechanism that preserves a coverage guarantee. The work also gives a recalibration trick for split conformal prediction that avoids the inflation penalty of the stability bounds.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2's proof replaces \\hat{R}^+_{(i)} by \\hat{R}_{(i)} using an inequality that opposes the true order-statistic ordering, so the Recal coverage guarantee is not established.","rationale":"The reader's weakest_assumption concerns the unstated proxy quantile level \\tilde{\\alpha} and the conditional-independence requirement, which are practical reproducibility issues. My stress-test identifies a more fundamental correctness concern: the proof of Theorem 2 contains an order-statistic inequality that runs counter to the definitions. This matters because Theorem 2 is the basis for the Recal method, which the paper presents as the best-performing approach in both synthetic and real experiments. If the theorem is false, the central practical claim collapses; if it is true, the proof as written still fails to establish it. The experimental results in Section 6 show Recal achieving roughly 0.9 coverage, which suggests the theorem may be salvageable, but a theorem whose proof has a sign error in a key inequality is not a finished contribution. The independent re-derivation or simulation I propose would settle whether the claim is true or merely plausible. I therefore agree with the CONDITIONAL verdict, but for a different, more load-bearing reason than the reader's reproducibility concerns.","tokens_in":24234,"tokens_out":51983,"duration_ms":468412,"concrete_test":"Independently re-derive the proof of Theorem 2 without invoking \\hat{R}^+_{(i)}\\le\\hat{R}_{(i)}; if the implication cannot be established, run a targeted simulation: use a selection rule independent of the calibration data (e.g., select predictor A for X in one half-space and predictor B otherwise, with each predictor accurate only on its own region), set m=100, K=2, \\alpha=0.1, and compute the effective-rank recalibration set over many seeds. Check whether empirical marginal coverage reaches at least 0.9. If it does not, Theorem 2 is false; if it does, the theorem may be true but a corrected proof is required.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central practical result, Theorem 2, claims that if \\hat{S} is conditionally independent of \\mathcal{D}_{cal} given X, then using the \\tau_\\alpha-th order statistic of the effective ranks recovers exact 1-\\alpha coverage. In the proof (Appendix A.3, after defining \\mathcal{R} and \\mathcal{R}^+), the authors write '\\mathcal{R}\\subset\\mathcal{R}^+ implies that \\hat{R}^+_{(m)}\\le \\hat{R}_{(m)}' and then replace \\hat{R}^+_{(i)} by \\hat{R}_{(i)} in the chain of inequalities. However, the definitions give, for each i\\le m, \\hat{R}_i \\le \\hat{R}^+_i, because adding the test score to a predictor's score set can only increase or keep the rank of a calibration point. This componentwise inequality implies \\hat{R}_{(i)} \\le \\hat{R}^+_{(i)} for every i, not the reverse. The step '\\le P{s^+_{k_{m+1},m+1}\\le s_{k_{m+1},(\\hat{R}_{(i)})}}' therefore requires a larger threshold when only a smaller one is available. The claimed coverage event is not implied by the exchangeability of \\mathcal{R}^+. Since Theorem 2 is the foundation for the Recal method, which the experiments recommend as the best performer, the paper's headline guarantee for its most useful procedure is currently unsupported by the written proof.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a stability-based framework for selecting among multiple conformal prediction sets while preserving finite-sample coverage. After introducing an (η,τ,ν)-stability notion adapted from Zrnic and Jordan, Theorem 1 gives a generic post-selection coverage bound, and Corollary 1 specializes it to conformal sets. Several stable selection mechanisms are proposed (Laplace, exponential, MinSE), with an optimality result for MinSE and extensions to adaptive, derandomized, and conditionally-valid selection. The framework is then extended to the online setting via AdaCOMA. Finally, Section 5 introduces a recalibration method (Recal) based on effective ranks in split conformal prediction, with Theorem 2 claiming exact 1−α coverage after an independence-compliant selection rule; experiments report Recal as the best-performing method. The stability-based results (Sections 3–4) appear self-contained and correct on inspection; the critical problem is that Theorem 2, the foundation of the recommended Recal method, is false as stated.","tokens_in":24470,"tokens_out":21423,"duration_ms":192686,"significance":"If the stability-based results were the whole paper, the contribution would be solid: Theorem 1 and Corollary 1 provide a clean transfer of marginal coverage under a weak stability condition, MinSE is a well-motivated and provably optimal-in-class mechanism, and the online AdaCOMA extension is a natural and useful combination with COMA. The main significance claim, however, is attached to the Recal method, which is presented as achieving tight post-selection coverage and is the best performer in the experiments. Because Theorem 2 is false, the Recal guarantee is invalid, and the empirical results for Recal in Section 6 and Appendix B are not backed by any valid theory. This substantially reduces the significance of the manuscript in its current form.","major_comments":[{"comment":"The proof of Theorem 2 is invalid and the theorem is false as stated. The proof asserts 'ℛ⊂ℛ+ implies that ^R+_{(m)}≤ ^R_{(m)}' and later uses ^R+_{(i)}≤ ^R_{(i)} to replace the threshold. However, for each i≤m, adding the test score to the calibration set can only increase (or keep) the rank of a calibration point, so pointwise R_i≤R^+_i, which implies R_{(i)}≤R^+_{(i)} for every i — the opposite of the direction used in the proof. Consequently, the step replacing ^R+_{(i)} by ^R_{(i)} is a decrease of the threshold, not an increase, and the coverage event of the final set is not implied by the exchangeability of ℛ+. This is not a minor gap: the claim is false. A concrete counterexample is obtained with m=2, α=0.4 (so τ_α=⌈0.6·3⌉=2), and a constant selection rule k_1=1, k_2=2, k_test=1, which is independent of D_cal and hence satisfies Theorem 2's assumption. If the score values s_{1,1},s_{1,2},s_{1,test},s_{2,1},s_{2,2} are iid continuous, then R_1 and R_2 are the binary ranks of the two calibration points under predictors 1 and 2, and the threshold is ^R_{(2)}=max(R_1,R_2). Conditional on max(R_1,R_2)=2, coverage occurs with probability 2/3; conditional on max(R_1,R_2)=1, it occurs with probability 1/3; and P(max(R_1,R_2)=1)=1/4. Thus the unconditional coverage is 3/4·2/3+1/4·1/3=7/12≈0.583, which is strictly less than 1−α=0.6, violating the theorem's conclusion.","section":"Theorem 2 (Section 5; proof in Appendix A.3)"},{"comment":"Because Theorem 2 is the sole theoretical justification for the Recal procedure, the claimed post-selection coverage guarantee for Recal is unsupported, and the empirical coverage reported in Figure 2 and Appendix B is purely anecdotal. The counterexample in the previous comment shows that a selection rule satisfying the theorem's independence assumption can undercover; therefore the Recal method, as described, does not provide a valid distribution-free guarantee. The authors would need to either remove Recal from the paper, replace it with a different recalibration method whose guarantee can actually be proved, or substantially revise the theory. As submitted, the paper's best-performing experimental method rests on a false theorem.","section":"Sections 5–6 (Recal and experiments)"}],"minor_comments":[{"comment":"In the display for Corollary 2, the equality 'P{Y_t ∉ C^(t)_Ŝ} = E[1{...}] = ∑_i p_i(ξ_t) 1{Y_t ∉ C^(t)_i}' omits the expectation operator around the random sum; as printed it equates a deterministic probability with a random variable. The same issue appears in Proposition 5. This should be corrected to 'E[1{...}] = E[∑_i p_i(ξ_t) 1{Y_t ∉ C^(t)_i}]'.","section":"Appendix A.2 (proofs of Corollary 2 and Proposition 5)"},{"comment":"The preliminary miscoverage rate ᾱ used to compute proxy quantiles from the auxiliary dataset D_aux is never reported in Section 6 or Appendix B, and the size of D_aux is not specified. This makes the exact Recal variant used in the experiments unclear and the results difficult to reproduce.","section":"Section 5 (Construction of an independent Ŝ)"},{"comment":"The legend in Figure 2 lists 'AdaMinSE α′=0.50', while the experimental text states 'AdaMinSE with α′=0.05'; these should be reconciled.","section":"Figure 2 caption"},{"comment":"The notation 'ℛ⊂ℛ+' is undefined and, if read as a multiset inclusion, is false: the elements of ℛ are not the same as the corresponding elements of ℛ+ because the ranks R_i and R^+_i are computed against different score sets. The proof should state the intended relationship explicitly, though the correct pointwise relationship points in the opposite direction from the one used.","section":"Appendix A.3 (proof of Theorem 2)"},{"comment":"The noise distribution 'ε∼(Lap(1/η))^{⊗K}' is nonstandard; it would be clearer to write ε_1,...,ε_K i.i.d. with distribution Lap(1/η).","section":"Lemma 1 statement"}],"recommendation":"reject","confidential_remarks":"The stress-test concern about Theorem 2 is fully confirmed, and it is more severe than a proof gap: the theorem is false, with a simple counterexample. The stability-based sections (Sections 3–4) appear sound and could in principle be salvaged, but as submitted the paper's headline practical method (Recal) is invalid, and the experimental claims attached to it are unsupported. I would advise the editor that a rejection is appropriate for this manuscript in its current form; a resubmission that removes or replaces the Recal section could be considered on the merits of the remaining stability-based contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The stability-based selection framework is the real contribution here. Framing pointwise selection among conformal sets through Zrnic-Jordan algorithmic stability is new relative to the average-performance selection in Liang et al. and Yang-Kuchibhotla, and MinSE with its optimality statement (Proposition 1) is a sensible, tightly engineered mechanism. I checked Theorem 1, Corollary 1, and the three mechanism lemmas; they are correct on my reading. The Laplace density-ratio argument works, and MinSE's feasibility argument genuinely yields (eta,tau)-stability. AdaMinSE is a nice practical wrapper. The experiments are extensive and the reported ordering Recal > AdaMinSE/MinSE > baselines is consistent across the figures. But the stress-test note is right, and it is not a minor gap. In the proof of Theorem 2 (Appendix A.3), after defining R and R+, the paper claims that R subset R+ implies R+_(i) <= R_(i) and then uses this to replace the R+ threshold by the R threshold. The componentwise truth is the opposite: adding the test score to a predictor's score set can only increase each calibration point's rank, so R_(i) <= R+_(i). With thresholds ordered the wrong way, the chain of inequalities does not establish the desired lower bound on coverage. This is not a missing epsilon in the proof; the claimed coverage event for Recal is not implied by the exchangeability of R+. The theorem may be repairable with a different argument, but as written it does not support the headline guarantee for the paper's best-performing method. Other soft spots are minor by comparison: the preliminary rate alpha-tilde used for the auxiliary-data proxy quantiles in Section 5 is never reported, so Recal is not exactly reproducible; the online guarantee in Corollary 2 is an averaged per-step probability statement, not the realized empirical coverage of (2), and the prose should be explicit about that; 4 of 50 ARMA runs are excluded without a stability criterion or sensitivity check; and the online tuning of AdaMinSE is described as selected to match COMA's coverage but no range is given. None of these affect the validity of the earlier theorems, but they matter to a reader trying to deploy the method. Who is this for: people working on conformal set aggregation or post-selection inference in split conformal settings. The stable selection part deserves to be taken seriously and cited once the recalibration is either fixed or honestly re-scoped. I would send it to peer review, but with a request for a real repair of Theorem 2 before publication.","headline":"The stable-selection framework (MinSE, Theorem 1, Corollary 1) is a genuine and sound contribution, but the paper's most practically useful result, the effective-rank recalibration (Theorem 2), has a proof with a backwards inequality, so its coverage guarantee is currently unsupported.","tokens_in":762,"tokens_out":976,"would_cite":false,"duration_ms":112345,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G15","62F25"],"pacs":[],"model":"deepseek-v4-flash","headline":"A stability condition on the selection rule transfers the marginal coverage of each conformal predictor to the chosen one, so a user can pick the smallest set per input $X$ with a certified $1-\\alpha$ guarantee.","keywords":["conformal prediction","valid selection","algorithmic stability","post-selection coverage","smallest set selection","split conformal","online conformal prediction","effective ranks"],"falsifier":"Fix the split-conformal setup with two predictors and arrange the auxiliary dataset so that its proxy set sizes rank the predictors in the opposite order to the sizes obtained from the calibration quantiles (for instance, by using different score normalizations in the two datasets). Run the effective-rank recalibration on fresh test points: if the empirical coverage of the selected set falls below $1-\\alpha$, the independence premise is not merely technical but load-bearing, and Theorem 2 as a practical recipe fails. A second decisive check is to rerun the reported real-data experiments while varying the unreported preliminary rate $\\tilde\\alpha$; if any choice breaks the nominal $1-\\alpha$ coverage, the practical version of the claim fails.","tokens_in":23959,"feed_emoji":"🎯","tokens_out":8735,"duration_ms":77165,"temperature":0.7,"pith_summary":"The paper takes up a practical annoyance of conformal prediction: when several valid prediction sets are available, selecting the most attractive one, typically the smallest, breaks the marginal coverage guarantee because the selection is informed by the same data that certified the sets. The authors claim the loss can be repaired by making the selection rule stable in a precise sense: a randomized rule whose output distribution, conditional on the observed set sizes, is within an $(e^\\eta,\\tau)$-indistinguishability budget of a fixed reference distribution. Under that condition the selected set covers the true label with probability at least $1-\\alpha e^\\eta-\\tau$, so building the individual sets at the adjusted level $1-(\\alpha-\\tau)e^{-\\eta}$ restores nominal $1-\\alpha$ coverage. In the split-conformal regime they go further and prove that exact $1-\\alpha$ coverage can be recovered by recalibrating on effective ranks, provided the selection rule is independent of the calibration data. The payoff, if the paper is right, is that a user can choose among conformal predictors pointwise, per test input $X$, without sacrificing the coverage guarantee.","feed_headline":"Randomized selection preserves coverage when picking the smallest set","feed_subtitle":"A small stability budget buys back the conformal guarantee, letting users pick among predictors for each X.","key_machinery":"The load-bearing object is $(\\eta,\\tau)$-conditional indistinguishability and the stability notion built on it: a randomized selection algorithm is stable if, conditional on the feature and the vector of set sizes, its output distribution is within factor $e^\\eta$ and additive slack $\\tau$ of a fixed reference random index $S_0$. Stability lets the proof attach a 'shadow' reference output to the selection rule and push the coverage event through the indistinguishability inequality. The second machinery, for the split conformal setting, is the effective rank: the rank of the $i$-th calibration point's non-conformity score under the predictor selected for that point. These effective ranks are exchangeable with the test point's effective rank when the selector is independent of the calibration data, so forming the set at the $\\lceil(1-\\alpha)(m+1)\\rceil$-th order statistic of the effective ranks reproduces the textbook split-conformal rank argument.","core_discovery":"The central claim is a transfer principle. Corollary 1 states that if the selection rule $\\hat S$ is $(\\eta,\\tau)$-stable, then $\\mathbb{P}\\{Y \\in C^\\alpha_{\\hat S(\\xi,\\varepsilon)}(X)\\} \\ge 1-\\alpha e^\\eta-\\tau$, so the coverage of each individual conformal set passes through the selection with only a multiplicative $e^\\eta$ inflation and an additive $\\tau$ loss. The paper introduces MinSE, the Minimum Stable Expectation mechanism, a linear program that chooses the selection distribution minimizing expected selected size subject to the stability constraints, and proves it optimal among all $(\\eta,\\tau)$-stable rules. It then shows the same principle yields long-run coverage in the online setting through AdaCOMA, and that in the split conformal setting exact coverage can be restored by using the selected predictor's calibration rank as a meta-score, taking the usual quantile of these effective ranks (Theorem 2).","pith_inferences":["Editorial inference: the same transfer principle should apply to any family of data-dependent confidence intervals beyond conformal sets, suggesting a general recipe for repairing selection among valid confidence statements by paying a small multiplicative randomization budget.","Editorial inference: the independence condition behind Theorem 2 tells practitioners to spend a slice of the calibration budget on a proxy dataset for the selector; the unreported preliminary rate $\\tilde\\alpha$ for the proxy quantiles is then a hidden tuning knob, and testing the method's sensitivity to it is the most direct check of the practical claim.","Editorial inference: one could try to lift effective-rank recalibration from split conformal to cross-conformal or jackknife+ constructions, where exchangeability of the meta-scores is not automatic, and that would require new arguments rather than a direct application of Theorem 2."],"forward_implications":["A practitioner with several conformal predictors can combine them pointwise, picking the smallest set for each $X$ with a randomized stable rule, and still certify the nominal $1-\\alpha$ marginal coverage after inflating the individual levels to $1-(\\alpha-\\tau)e^{-\\eta}$.","MinSE is a near-optimal way to do this: among all rules satisfying the same stability budget it achieves the smallest expected selected size almost surely (given a suitable prior), and the worst-case bound $\\alpha e^\\eta+\\tau$ is tight, as the oracle example shows.","In the split conformal setting, effective-rank recalibration removes the inflation entirely under the independence condition, delivering exactly the standard $1-\\alpha$ guarantee; the paper's experiments report that this version (Recal) gives the shortest average intervals among the compared methods.","In the online setting, AdaCOMA inherits COMA's historically learned weights as a prior but conditions the selection on the current set sizes, gaining pointwise adaptability while keeping the long-run coverage statement."],"supporting_citations":[{"why":"Supplies the algorithmic-stability framework and the indistinguishability lemma that the main transfer theorem builds on.","marker":"[37]"},{"why":"COMA online model aggregation whose weights become the prior for AdaCOMA and which provides the comparison coverage bound used in the online analysis.","marker":"[16]"},{"why":"Differential privacy calibration mechanism that inspires the Laplace and exponential stable selection rules.","marker":"[11]"},{"why":"ModSel selection method used as a baseline and whose open-source implementation is adapted for the experiments.","marker":"[27]"},{"why":"Split-conformal selection baselines YK-Adjust and YK-Base used as experimental comparisons.","marker":"[35]"},{"why":"Majority-vote merging technique used for the derandomized confidence set proposition.","marker":"[17]"},{"why":"Provides the exchangeability and rank-statistic foundation used in the proof of the recalibration theorem.","marker":"[1]"}],"fun_headline_variants":["Stable selection keeps conformal coverage intact","Pick the smallest conformal set without losing coverage","Coverage-preserving selection for conformal sets","A stability budget buys back conformal guarantees","Choosing among conformal sets without breaking coverage"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"For the batch theorems the selection rule must satisfy Definition 2 stability: conditional on the observed set sizes there must exist a reference random index whose distribution is within the multiplicative-additive budget $(e^\\eta,\\tau)$ of the rule's own selection probabilities, and for the exact recalibration result the rule must be conditionally independent of the calibration data, which forces the selector to use only an auxiliary dataset.","fun_headline_variants_meta":{"raw":{"variants":["Stable selection keeps conformal coverage intact","Pick the smallest conformal set without losing coverage","Coverage-preserving selection for conformal sets","A stability budget buys back conformal guarantees","Choosing among conformal sets without breaking coverage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1400,"prompt_tokens":810,"completion_tokens":590,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":426,"completion_tokens_details":{"reasoning_tokens":522}},"tokens_in":426,"tokens_out":590,"duration_ms":5803,"temperature":1.0,"reasoning_tokens":522,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:23:04.139448+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fix the split-conformal setup with two predictors and arrange the auxiliary dataset so that its proxy set sizes rank the predictors in the opposite order to the sizes obtained from the calibration quantiles (for instance, by using different score normalizations in the two datasets). Run the effective-rank recalibration on fresh test points: if the empirical coverage of the selected set falls below $1-\\alpha$, the independence premise is not merely technical but load-bearing, and Theorem 2 as a practical recipe fails. A second decisive check is to rerun the reported real-data experiments while varying the unreported preliminary rate $\\tilde\\alpha$; if any choice breaks the nominal $1-\\alpha$ coverage, the practical version of the claim fails.","supporting_citations":[{"cited_title":"Zrnic and M","cited_arxiv_id":null,"evidence_quote":"Supplies the algorithmic-stability framework and the indistinguishability lemma that the main transfer theorem builds on."},{"cited_title":"Dwork, F","cited_arxiv_id":null,"evidence_quote":"Differential privacy calibration mechanism that inspires the Laplace and exponential stable selection rules."},{"cited_title":"Yang and A","cited_arxiv_id":null,"evidence_quote":"Split-conformal selection baselines YK-Adjust and YK-Base used as experimental comparisons."}],"review_version":2}