{"id":"98bac609-d7c8-446e-9ad9-bb1a99f906a4","arxiv_id":"2502.07497","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Training-conditional conformal prediction does not estimate Bernoulli probabilities and can trivially satisfy its PAC guarantee, making it unsuitable for statistical safety certification.","lead":"This paper shows that training-conditional conformal prediction, when used to certify the safety of control systems, can satisfy its mathematical guarantee by predicting the entire sample space, producing uninformative and sometimes invalid safety claims. The authors recommend using classical binomial proportion confidence intervals to estimate the probability of unsafe behavior.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The critique of continuous-score safety-certification applications rests on an unproven equivalence to the binary indicator case; the N=2 triviality demonstration may not transfer.","rationale":"The reader's weakest_assumption identifies the Bernoulli/indicator modeling assumption as the structural basis of the paper's argument. I agree that this assumption is load-bearing, but I would sharpen it: the paper does not merely assume an indicator nonconformity measure as a theoretical simplification; it explicitly reduces the continuous-score procedures of the cited safety-certification works (R_i = J(x_i)) to that indicator case in Section 4.1. That reduction is asserted, not proved, and it is not trivial because CP predictors depend on the rank ordering of scores, which changes when a continuous score is replaced by a binary indicator. The N=2 counterexample is compelling for binary scores, but it does not automatically establish the same failure mode for continuous scores. The paper's more general point that Theorem 1 guarantees coverage of a set predictor rather than a confidence interval for a Bernoulli parameter is valid independent of the score type, so the theoretical core is sound. However, the applied claim that the specific works are 'incorrect' rests on the unproven equivalence. A careful reviewer should require the authors to either prove the equivalence (or a suitable extension of the triviality result to continuous scores) or narrow their critique to applications that genuinely use binary/indicator nonconformity scores. This is a revision-worthy gap rather than a fatal flaw, hence CONDITIONAL rather than REJECT. The concrete test would settle whether the continuous-score case exhibits the same trivial-prediction failure; if it does, the original ACCEPT verdict is confirmed, and if it does not, the paper's claims about those applications need to be modified.","tokens_in":8853,"tokens_out":18998,"duration_ms":171860,"concrete_test":"Take the safety-certification setup of Section 4.1 with a continuous nonconformity score, e.g., X = [0,1], P = Uniform, unsafe set [0,b], J(x) = x (or the distance to the unsafe set). For N = 2 and a range of epsilon, compute the CP predictor Gamma_epsilon, the event S_E, and P^2(S_E) for b > E. Determine whether the only way to satisfy P^2(S_E) >= E^2 is Gamma_epsilon = Z, or whether non-trivial Gamma_epsilon (a proper subset of X) can also meet the coverage condition. If non-trivial sets can meet it, the binary equivalence in Section 4.1 fails and the blanket claim that the cited applications provide no valid safety guarantees needs a different proof.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 4.1 claims that Lin & Bansal's use of R_i = J(x_i) with continuous J is 'equivalent' to the indicator INM (10). This equivalence is not proved, and it is not obvious: training-conditional CP set predictors depend on the ordering of scores, and replacing continuous J by 1_{x in X_A} changes the p-values and the resulting Gamma_epsilon. Example 1's triviality ('the confidence level is attained by predicting Z') is proven only for binary scores with N=2. For continuous scores, Gamma_epsilon is generally a superlevel set of J, not just Q or Z; when b > E, the coverage event S_E may be satisfied by non-trivial sets, depending on the score distribution. The paper asserts 'the same applies for any different choice' (Section 4.1) without proof. If the equivalence fails, the central claim that those applications provide no valid safety guarantees is not established for their actual continuous-score procedures. The general scope mismatch (CP guarantees P(Z in Gamma), not a confidence interval for b) still stands, but the striking trivial-prediction refutation does not transfer automatically.","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the N=2 example is the real contribution and it's correct. The paper cleanly shows that for a binary indicator nonconformity measure, the PAC bound from Vovk's training-conditional CP can be satisfied by predicting the whole space with probability b^2 (when two calibration points both fall in Q), and otherwise predicting Q. That means the confidence level is not a confidence interval for the Bernoulli parameter b. The distinction between 'the set contains the next point with probability >=1-E' and 'we have a CI for b' is well drawn. The appendix's simulation confirms the accounting. This is a useful corrective note, and the cited safety-certification papers do appear to blur that distinction.\n\nThe soft spot is Section 4.1. The authors claim that Lin & Bansal's continuous cost J is 'equivalent' to the binary indicator A(x)=1_{x in X_A}. That's not proven and not obviously true. With continuous scores, ties have probability zero, p-values are driven by rank ordering, and Gamma_epsilon is generally a superlevel set of J, not just Q or Z. The N=2 triviality argument only works for binary scores; for continuous scores, when b > E the coverage event might be satisfied by nontrivial sets, and the 'confidence earned by trivial predictions' failure mode could look different. The stress-test note has this right. The blanket sentence 'the same applies for any different choice' is doing too much work.\n\nEven so, the central argument survives. The general point is that training-conditional CP's guarantee is about the coverage of the predicted set, averaged over calibration draws, and that is not the same as a binomial confidence interval for b. That point holds regardless of whether the scores are continuous. So the critique of the safety-certification applications is directionally correct, just not as tightly proven as the N=2 example. The authors could strengthen this by either proving the reduction for continuous scores or explicitly limiting their claim to the binary-score case and presenting the rest as a scope-mismatch argument.\n\nWould I send to a serious referee? Yes. It's a clean, correct counterexample with an important caveat, and it deserves referee time. The main thing I'd ask the authors to fix is the overreach in Section 4.1. I'd cite it if I were working in learning-based safety certification.","headline":"A correct and sharp N=2 counterexample showing training-conditional CP doesn't yield binomial proportion confidence intervals, but the Section 4.1 leap from continuous to binary scores is asserted, not proved.","tokens_in":9567,"tokens_out":3776,"would_cite":true,"duration_ms":31060,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F25","68T05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Training-conditional conformal prediction does not yield valid binomial proportion confidence intervals or safety guarantees: its PAC bound can be satisfied by trivial whole-space predictions.","keywords":["conformal prediction","training-conditional conformal prediction","binomial proportion confidence intervals","PAC guarantees","safety certification","Bernoulli random variables","Clopper-Pearson interval","set predictors"],"falsifier":"Repeat the appendix's N=2 experiment with $b>E$ and $\\epsilon\\in[2/3,1)$, recording on each calibration draw whether the predicted set is $Z$ and whether the new score is covered. If the bound $P^2(S_E)\\ge E^2$ holds only through draws with $\\Gamma_\\epsilon=Z$, and the interval $[0,E]$ never contains $b$ when $b>E$, the paper's conclusion is confirmed; a draw with $\\Gamma_\\epsilon=Q$ that still covers the new score at the required level would refute it.","tokens_in":8680,"feed_emoji":"⚠️","tokens_out":7447,"duration_ms":62802,"temperature":0.7,"pith_summary":"Training-conditional conformal prediction is increasingly used for safety certification in control systems, but this paper argues the guarantees it provides are not the guarantees those applications claim. The target problem is binomial proportion estimation: given N independent Bernoulli trials with unknown success probability b, produce a valid confidence interval for b. By an explicit N=2 example, the paper shows that the PAC bound in Theorem 1 can be satisfied by predicting the entire sample space whenever b is larger than the target error E, giving no information about b. If the argument is right, several recent safety certificates built on this variant of conformal prediction do not certify what they claim. The paper recommends traditional binomial proportion confidence intervals, such as Clopper-Pearson, for statistical safety certification, and leaves open whether a different conformal formulation could work.","feed_headline":"Conformal prediction cannot certify safety via binomial intervals","feed_subtitle":"The PAC bound is met by trivial whole-space predictions, not by estimating the failure probability.","key_machinery":"The load-bearing pair is Theorem 1's PAC guarantee for training-conditional conformal prediction together with an indicator nonconformity measure (equation 8) that turns calibration scores into i.i.d. Bernoulli random variables. The argument works by enumerating the three possible calibration outcomes for N=2 (both points outside $Q$, one inside, both inside) and computing the resulting predicted set $\\Gamma_\\epsilon$: for $\\epsilon\\in[2/3,1)$, the only way to reach coverage $1-E$ when $b>E$ is the event $Q\\times Q$, which forces $\\Gamma_\\epsilon=Z$. That event has probability $b^2$, so the bound $P^2(S_E)\\ge E^2$ is met by trivial whole-space predictions, not by estimating the Bernoulli parameter $b$.","core_discovery":"The paper's central claim is that training-conditional conformal prediction, as used in recent safety-certification work, does not produce valid binomial proportion confidence intervals and therefore does not provide the statistical safety guarantees claimed. The demonstration is Example 1: with N=2 calibration points and an indicator nonconformity measure $A(z)=1_{\\{z\\in Q\\}}$, the calibration scores are i.i.d. Bernoulli with parameter $b=P(Q)$. For any $\\epsilon\\in[2/3,1)$ the predicted set $\\Gamma_\\epsilon$ is the whole space $Z$ exactly when both calibration points fall in $Q$, an event of probability $b^2$, and is $Q$ otherwise. The PAC statement $P^2(S_E)\\ge E^2$ then holds for $b>E$ only because $b^2\\ge E^2$ on those trivial whole-space predictions, while the non-trivial prediction $Q$ fails to meet the coverage requirement; when $b\\le E$ the coverage requirement is met automatically. Hence the confidence level does not estimate $b$, and the authors conclude that conformal prediction is unsuitable for binomial proportion problems and that direct binomial interval methods such as Clopper-Pearson should be used for statistical safety certification.","pith_inferences":["A testable extension: for N>2 with the same indicator score, the fraction of calibration sets that achieve coverage $1-E$ by predicting the whole space should still account for essentially all of the confidence when $b>E$; this could be quantified by simulation.","The same vacuous-satisfaction mechanism should appear in any discrete-label conformal setting where the nonconformity score takes finitely many values, because a loose prediction (all labels) contributes to coverage without estimating label probabilities.","If safety certification needs an upper confidence bound on the probability of entering an unsafe region, the Bernoulli calibration scores can be fed directly into a binomial proportion method; the paper's example suggests the conformal route cannot be repaired merely by choosing different $\\epsilon$ and $E$ values."],"forward_implications":["Safety certificates built on Theorem 1 in the style of Lin & Bansal (2024) and Chilakamarri et al. (2024) are not valid guarantees of the probability of safe operation.","The claimed equivalence by Vincent et al. (2024), that training-conditional conformal prediction with Bernoulli scores reduces to the Clopper-Pearson interval, is false.","For safety certification, a Clopper-Pearson interval computed directly from Bernoulli evaluations of a trajectory cost function gives a valid binomial proportion confidence interval, whereas Theorem 1 does not.","The mathematical content of Theorem 1 remains intact; what changes is its interpretation: it bounds how often the set predictor has coverage $1-E$, not how accurately any class or score probability is estimated."],"supporting_citations":[{"why":"Supplies Theorem 1, the training-conditional PAC guarantee that the paper analyzes and shows can be satisfied trivially.","marker":"Vovk (2012)"},{"why":"Provides the conservatively valid binomial proportion confidence interval used as the comparison baseline.","marker":"Clopper & Pearson (1934)"},{"why":"Survey that frames the binomial proportion estimation problem the paper says conformal prediction does not solve.","marker":"Dean & Pagano (2015)"},{"why":"The safety-verification application whose use of Theorem 1 is examined and criticized in Section 4.1.","marker":"Lin & Bansal (2024)"},{"why":"Follow-up verification work that the paper identifies as inheriting the same issue through Lin & Bansal (2024).","marker":"Chilakamarri et al. (2024)"},{"why":"Source of the claimed reduction of training-conditional CP to Clopper-Pearson for Bernoulli scores, which the example disproves.","marker":"Vincent et al. (2024)"}],"fun_headline_variants":["Conformal prediction falls short for binomial safety certification","Training-conditional conformal prediction invalid for safety bounds","CP fails to certify safety when failure is binomial","Binomial intervals not conformal for safety certification","Training-conditional CP gives no binomial safety guarantees"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the nonconformity score is a fixed 0/1 indicator of membership in a set $Q$ with a fixed training set, so each calibration score is a Bernoulli coin flip; if scores were continuous or $Q$ depended on calibration data, the trivial-prediction failure for $N=2$ would not automatically generalize and a different proof would be needed.","fun_headline_variants_meta":{"raw":{"variants":["Conformal prediction falls short for binomial safety certification","Training-conditional conformal prediction invalid for safety bounds","CP fails to certify safety when failure is binomial","Binomial intervals not conformal for safety certification","Training-conditional CP gives no binomial safety guarantees"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000759,"raw_usage":{"total_tokens":3368,"prompt_tokens":940,"completion_tokens":2428,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":2357}},"tokens_in":556,"tokens_out":2428,"duration_ms":14264,"temperature":1.0,"reasoning_tokens":2357,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T12:32:03.917548+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the appendix's N=2 experiment with $b>E$ and $\\epsilon\\in[2/3,1)$, recording on each calibration draw whether the predicted set is $Z$ and whether the new score is covered. If the bound $P^2(S_E)\\ge E^2$ holds only through draws with $\\Gamma_\\epsilon=Z$, and the interval $[0,E]$ never contains $b$ when $b>E$, the paper's conclusion is confirmed; a draw with $\\Gamma_\\epsilon=Q$ that still covers the new score at the required level would refute it.","supporting_citations":[{"cited_title":"Conditional validity of inductive conformal predictors","cited_arxiv_id":null,"evidence_quote":"Supplies Theorem 1, the training-conditional PAC guarantee that the paper analyzes and shows can be satisfied trivially."},{"cited_title":"The use of confidence or fiducial limits illustrated in the case of the binomial","cited_arxiv_id":null,"evidence_quote":"Provides the conservatively valid binomial proportion confidence interval used as the comparison baseline."},{"cited_title":"Evaluating confidence interval methods for binomial proportions in clustered surveys","cited_arxiv_id":null,"evidence_quote":"Survey that frames the binomial proportion estimation problem the paper says conformal prediction does not solve."},{"cited_title":"Verification of neural reachable tubes via scenario optimization and conformal prediction","cited_arxiv_id":null,"evidence_quote":"The safety-verification application whose use of Theorem 1 is examined and criticized in Section 4.1."},{"cited_title":"Reachability Analysis for Black-Box Dynamical Systems","cited_arxiv_id":"2410.07796","evidence_quote":"Follow-up verification work that the paper identifies as inheriting the same issue through Lin & Bansal (2024)."},{"cited_title":"Guarantees on robot system performance using stochastic simulation rollouts","cited_arxiv_id":null,"evidence_quote":"Source of the claimed reduction of training-conditional CP to Clopper-Pearson for Bernoulli scores, which the example disproves."}],"review_version":1}