{"id":"cf8d90e3-0519-4d17-b713-1cd546e92c0b","arxiv_id":"2501.11689","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In classification, every confidence predictor that is valid under IID data can be transformed into a conformal predictor with an explicit, constant-free bound on the loss of efficiency.","lead":"Vovk restates the known universality of conformal prediction in the functional theory of randomness, which avoids the large unspecified constants of algorithmic randomness. The main bound shows that in classification, an IID-valid predictor can be replaced by a conformal predictor with an explicit loss factor.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Corollary 3 depends on Theorems 4 and 6, both asserted without proof and deferred to the companion paper [50]; if either theorem carries unstated conditions, the chained universality result may fail.","rationale":"I read the paper as a self-contained functional translation of Nouretdinov et al., with Corollary 3 as the central technical result. I checked the proof chain: (20) calibration is standard; (21) is Corollary 1, which uses Theorem 3 (proved) and Theorem 4 (deferred); (22) uses Theorem 6 (deferred); (23) e-to-p calibration is standard; and setting G = sqrt(G1 G2) is valid by Cauchy-Schwarz. Thus the only unverified assumptions in the central argument are Theorems 4 and 6. The reader's conditional verdict identifies exactly this gap, and I agree with that assessment. The concern is not that the statements are false; it is that the paper asserts them without derivation and the only cited support is the author's own companion paper. If either theorem has a hidden extra condition, the gap is invisible to the reader of this manuscript. The concrete test of independent derivation is the appropriate check because it directly determines whether the stated generality of Corollary 3 holds. I do not see an internal inconsistency in the chaining or a calibration error, so the reader's CONDITIONAL verdict remains appropriate and no adjustment is needed.","tokens_in":24039,"tokens_out":6259,"duration_ms":62100,"concrete_test":"Independently re-derive Theorem 4 (inequality 14) and Theorem 6 from the definitions of IID/exchangeability e-variables, without consulting [50], and check whether the constants 1/(e(|Y|-1)) and 1/(|Y|-1) hold for all n≥1, all measurable object spaces X, and all finite label spaces Y with |Y|≥2. If the derivation requires an extra condition (e.g., n large, |X|≥|Y|, or a particular sigma-algebra on the bag space), that condition must be added to Corollary 3 and its absence would falsify the current statement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Corollary 3, Eq. 19) is obtained by chaining calibration, Corollary 1, Theorem 6, and e-to-p calibration. The calibration and e-to-p steps are standard and their use in the proof is sound; Theorem 3 is proved in full. The two genuinely load-bearing links are Theorem 4 (Section 6.2, inequality (14)), which transfers an invariant IID e-value to a false-label setting with constant 1/(e(|Y|-1)), and Theorem 6 (Section 6.3), the train-invariance step with constant 1/(|Y|-1). Both are stated without proofs: the text says 'A formal proof is given in [50]' (Theorem 4) and 'For a simple proof, see [50]' (Theorem 6), where [50] is the author's own companion preprint. If these proofs exploit conditions not present in the statement — e.g., restrictions on n, on the object space X beyond measurability, or extra structural assumptions on the e-variables — then Corollaries 1–4 inherit those restrictions and the advertised reduction of IID-valid predictors to conformal predictors is not established in the generality claimed. The optimality result Theorem 5 is also deferred to [50] but is not needed for Corollary 3.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reviews the relationship between the IID and exchangeability assumptions, with special attention to conformal prediction, and presents a translation of Nouretdinov, V'yugin, and Gammerman's universality result into the 'functional theory of randomness.' In the functional setting, the central result is Corollary 3: for every IID p-predictor P there exist a conformal (train-invariant exchangeability) predictor P' and an IID e-variable G such that inequality (19) holds for every false label, with the explicit constant e(|Y|-1)^2/delta and the factor P^{1-delta}. The paper contains a full proof of Theorem 3, and the proof of Corollary 3 is a transparent chaining of calibration, Kolmogorov's step, the train-invariance step, and e-to-p calibration. However, the two load-bearing steps, Theorem 4 (the 1/(e(|Y|-1)) transfer from invariant IID e-values to false labels) and Theorem 6 (the 1/(|Y|-1) train-invariance step), are stated without proofs and deferred to the author's companion preprint [50].","tokens_in":24231,"tokens_out":9868,"duration_ms":107419,"significance":"If the deferred theorems are correct in the claimed generality, the paper is a valuable contribution: it removes unspecified additive constants from the algorithmic theory, provides explicit constants, gives a self-contained proof of the key factorization theorem, and clearly identifies which steps incur the finite-label-space penalties. The paper is also honest about what is deferred, which is a methodological strength. The included proof of Theorem 3 and the chaining argument show that the overall strategy is sound, but the manuscript as submitted does not fully establish the advertised universality result because its two key technical links are not proved in the text.","major_comments":[{"comment":"Theorem 4 is essential for Corollaries 1, 2, 3, and 4, but no proof is given; the text only states 'A formal proof is given in [50].' Since [50] is the author's own companion preprint, the manuscript is not self-contained at the exact point where the constant 1/(e(|Y|-1)) is introduced. Please include the proof in this paper, or state precisely the hypotheses under which the result holds (including any restrictions on n, on the object space X, and on the e-variables) and verify that Corollary 3 remains valid under those hypotheses.","section":"6.2, Theorem 4 (Eq. (14))"},{"comment":"The train-invariance step with factor 1/(|Y|-1) is likewise deferred, with the text saying 'For a simple proof, see [50].' This theorem is used in the proof of Corollary 3 at inequality (22), so without it the chained universality statement cannot be verified from this manuscript alone. Please provide the proof or make the result a clearly marked import with a precise statement of its conditions.","section":"6.3, Theorem 6"},{"comment":"The claim that when E is train-invariant, 'the resulting predictor E' will also be train-invariant' is asserted without proof. Corollaries 2 and 4 rely on this preservation to obtain a conformal predictor, so this is another unproved link in the main chain. Please prove the preservation or give the exact argument in [50].","section":"6.2, paragraph after Corollary 1"},{"comment":"The paper states 'Very few proofs will be given, and most of them can be found in [50].' This is an explicit limitation, but it conflicts with the paper's status as a full research article, since the abstract's advertised translation result depends on these deferred proofs. The revision should either include the proofs or clearly mark the paper as a research announcement with the main results imported from [50].","section":"6, first paragraph"}],"minor_comments":[{"comment":"In the sentence 'Our argument will also establish the closeness of the conformal e-predictors (i.e., X/p/t predictors) to the R/e predictors', the parenthetical should presumably read 'X/e/t predictors' rather than 'X/p/t predictors'.","section":"6.1"},{"comment":"Theorem 5, the asymptotic optimality claim, is also deferred to [50]. It is not used in the proof of Corollary 3, so it is not blocking, but the paper should state explicitly that this is a citation result and, if possible, give the precise theorem number in [50].","section":"6.2, Theorem 5"},{"comment":"The line 'It is obvious that E' is an element of E_X' is terse; adding a one-sentence justification using the permutation-average argument would make the included proof of Theorem 3 easier to check.","section":"Theorem 3 proof"},{"comment":"The approximation 'e ≈ 2.72' is slightly misleading; using 'e ≈ 2.718' would be more accurate.","section":"2"},{"comment":"The abstract's phrase 'without losing much in predictive efficiency' should be read in the sense of inequality (19), where the loss is controlled up to an e-variable G and a power P^{1-delta}; making this qualification explicit in the abstract would prevent over-interpretation.","section":"Abstract and Corollary 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a festschrift contribution, and the historical and personal material is appropriate. The main technical issue is the systematic deferral of the key proof steps to the author's own companion preprint [50]. If the journal's policy permits proofs to be supplied by reference to an unreviewed preprint, the current version might be acceptable; otherwise, the revision should include those proofs or precise import statements. I recommend major revision rather than rejection because the missing pieces appear local and fixable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read Vovk's Festschrift paper. The thing to know: the functional version of the Nouretdinov–V'yugin–Gammerman universality result is real and clean, but the manuscript itself proves only part of it. Theorems 4 and 6, which carry the 1/(e(|Y|-1)) and train-invariance factors, are stated and then deferred to the author's own companion preprint [50]. If you referee this, you need [50] on the table. That is the whole ballgame for the advertised reduction of IID p-predictors to conformal predictors.\n\nWhat the paper does well: it gives a readable history of the IID/exchangeability distinction, explains why the algorithmic theory's additive constants are a genuine barrier for practice, and lays out the functional translation so the chain of reasoning is transparent. Theorem 3's proof is included and is a nice independent derivation of E_R = E_X E_iR. The chaining argument in Corollary 3 is clear once you grant Theorems 4 and 6. The paper is honest about what is deferred—it does not pretend to prove everything. Self-citation here is legitimate; [48] is the source of Theorem 3 and [50] is the companion.\n\nSoft spots: the two load-bearing lemmas are asserted without proof in this paper. That is a real limitation for a standalone manuscript, but not a fatal one if the companion holds up. The stress-test worry—that [50] may impose unstated conditions on n or X—needs checking, but nothing in the statements suggests hidden structure beyond finite-label classification. Theorem 5's optimality is also deferred, but it is not needed for Corollary 3. Novelty is modest: this is a translation and synthesis, not a brand-new theorem. For a Festschrift that is fine, but do not oversell it.\n\nBottom line: this paper is for a reader who wants an accessible map of how conformal prediction relates to IID validity and who does not mind chasing a companion preprint for the hard parts. It deserves refereeing—the result is foundational, the included proof is sound, and the missing pieces are checkable. I would accept it for review and expect revisions that either import or clearly locate the proofs from [50].","headline":"A clean functional translation of the universality result, but the two load-bearing lemmas live in the companion paper, so review needs [50] in hand.","tokens_in":24828,"tokens_out":2051,"would_cite":true,"duration_ms":22157,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60G09","68Q30","62F03"],"pacs":[],"model":"deepseek-v4-flash","headline":"Every IID-valid confidence predictor can be turned into a conformal predictor with only a bounded loss in efficiency.","keywords":["conformal prediction","exchangeability","IID assumption","functional theory of randomness","e-values","p-values","algorithmic randomness","classification"],"falsifier":"Check the companion proof [50] for the precise conditions under which Theorem 4 and Theorem 6 are proved; if either theorem needs a restriction that the present statement does not state, such as $n$ large, $\\mathcal{X}$ finite, or $|\\mathcal{Y}|$ above a threshold, then Corollary 3 as formulated in this paper is false. Concretely, test the minimal case $|\\mathcal{Y}|=2$, $n=1$: if the claimed factors $1/(\\mathrm{e}(|\\mathcal{Y}|-1))$ and $1/(|\\mathcal{Y}|-1)$ cannot be attained by any e-variable $G$ satisfying $\\mathbb{E}_Q[G]\\le 1$ for all IID probability measures $Q$, the inequality (19) fails.","tokens_in":23720,"feed_emoji":"🎲","tokens_out":9455,"duration_ms":97523,"temperature":0.7,"pith_summary":"This paper claims that the algorithmic theory of randomness, which links IID data to exchangeability but leaves unspecified additive constants, can be replaced by a functional theory of randomness in which the same statements become explicit inequalities between classes of confidence predictors. The main result, Corollary 3, says that in classification with a finite label space $\\mathcal{Y}$, every confidence predictor valid under IID data can be transformed into a conformal predictor with only a bounded loss: for any $\\delta\\in(0,1)$ there is an IID e-variable $G$ such that $P'(z_1,\\dots,z_n,x_{n+1},y)\\le \\mathrm{e}(|\\mathcal{Y}|-1)^2 \\delta^{-1} G(z_1,\\dots,z_{n+1})^2 P(z_1,\\dots,z_n,x_{n+1},y)^{1-\\delta}$ for every false label $y$. A sympathetic reader cares because conformal prediction is distribution-free and widely implemented; the claim is that believing data are IID rather than merely exchangeable cannot buy more than this explicit factor in classification.","feed_headline":"Conformal prediction loses little versus IID methods","feed_subtitle":"The proof converts any IID-valid predictor into a conformal one, with explicit constants instead of unspecified ones.","key_machinery":"The machinery is the cube of eight function classes obtained from three dichotomies: the data-generating assumption (IID versus exchangeability), the type of confidence value (p-values versus e-values), and train-invariance (optional). The two vertices that matter are $P_R$, the most general IID p-predictors, and $P_{tX}$, which coincides with conformal predictors. The argument travels the red path $P_R\\to E_R\\to E_X\\to E_{tX}\\to P_{tX}$, and the load-bearing identities are Theorem 3, $E_R=E_X E_{iR}$, which factors any IID e-variable into an exchangeability e-variable and an invariant IID e-variable; Theorem 4 and Theorem 6, which bound the loss in the false-label and train-invariance steps; and the two calibrators $f(p)=\\delta p^{\\delta-1}$ and $e\\mapsto 1/e$ that enter and leave the e-value world.","core_discovery":"The central claim is that conformal prediction is universal under the IID assumption in a quantitative, constant-free sense. Formally, Corollary 3 asserts that for every IID p-predictor $P\\in P_R$ and every $\\delta\\in(0,1)$, there exists a conformal predictor $P'\\in P_{tX}$ and an IID e-variable $G$ satisfying the inequality above for all training sequences, test objects, and false labels. The proof moves along the chain $P_R\\to E_R\\to E_X\\to E_{tX}\\to P_{tX}$: convert the p-value to an e-value by calibration, factor an IID e-variable into an exchangeability e-variable times an invariant IID e-variable, pass to false labels with the factor $1/(\\mathrm{e}(|\\mathcal{Y}|-1))$, make the predictor train-invariant with the factor $1/(|\\mathcal{Y}|-1)$, and convert back to a p-value. If the paper is right, any method that reports valid p-values under the standard IID assumption can be mimicked by a conformal method, up to the stated factors, unless the whole augmented data sequence itself looks non-IID.","pith_inferences":["The squared factor $(|\\mathcal{Y}|-1)^2$ in (19) comes from composing two independent losses, Theorem 4 and Theorem 6; since Theorem 5 shows the single $|\\mathcal{Y}|$ factor is asymptotically optimal, a direct p-to-p argument that bypasses e-values might reduce the squared factor, an improvement the paper's own route cannot give.","In practice this suggests that the IID-versus-exchangeability debate changes character: when data are IID, conformal methods should be competitive with bespoke IID methods up to these factors, so improving the nonconformity measure likely matters more than weakening the validity assumption.","For regression or infinite label spaces the constants have no literal meaning, since the comparison must depend on a metric between labels; the classification framing here is the favorable case, and the regression analysis in [50] may need separate constants.","The same calibrate-factor-average-calibrate template could be applied to other pairs of validity assumptions, such as covariate shift or conditional validity, suggesting that functional e-value decompositions are a reusable tool for comparing prediction settings."],"forward_implications":["If Corollary 3 is correct, the universality result of Nouretdinov, V'yugin, and Gammerman no longer depends on unspecified constants: the translation from IID p-predictors to conformal predictors is governed by explicit factors involving $\\mathrm{e}(|\\mathcal{Y}|-1)^2/\\delta$ and the IID e-variable $G$.","Because conformal predictors are exactly the train-invariant exchangeability p-predictors, the result identifies which IID methods can be replaced: any method whose p-values are valid under IID can be converted into a conformal one, so the weaker exchangeability assumption supports the same level of confidence up to these factors.","The fundamental limitation of conformal prediction, that its p-values cannot go below $1/(n+1)$, is shown to be a limitation of any IID-valid method in classification, giving $D_R \\le \\log(n+1)+O(\\log\\log(n+1))$ in the prediction-proper regime.","In the train-invariant case, Corollary 4 improves the bound to $e(|\\mathcal{Y}|-1)/\\delta \\cdot G \\cdot P^{1-\\delta}$, making the loss especially small for the most natural class of IID predictors."],"supporting_citations":[{"why":"The original universality theorem for conformal prediction under IID; it also supplies Proposition 1, which identifies train-invariant exchangeability p-tests with conformal predictors.","marker":"[27]"},{"why":"Introduces the functional (non-algorithmic) theory of randomness and provides the product decomposition $E_R=E_X E_{iR}$ used here as Theorem 3.","marker":"[48]"},{"why":"Companion preprint that contains the formal proofs of Theorems 4, 5, and 6, the load-bearing steps stated without proof in this paper.","marker":"[50]"},{"why":"The standard reference for conformal prediction; it supplies the definitions of confidence predictors and p-variables, the bag sigma-algebra apparatus used in the proof of Theorem 3, and the equivalence identifying $P_{tX}$ with conformal predictors.","marker":"[54]"},{"why":"Provides the calibration theory used in Corollary 3: the calibrator $f(p)=\\delta p^{\\delta-1}$ and the e-to-p calibration step.","marker":"[60]"},{"why":"The paper that introduced conformal prediction proper and its ideal picture; its remark on the fundamental limitation $\\log N$ is extended here to all IID-valid methods in classification.","marker":"[53]"}],"fun_headline_variants":["Any IID predictor converts to conformal with little loss","Conformal mimics any IID predictor at near-zero efficiency cost","IID-valid predicts conformal: explicit constants, small loss","Conformal prediction universal for IID: no unspecified addends"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The two steps that carry the whole argument, Theorem 4 and Theorem 6, are stated here without proof and deferred to the companion preprint, so the paper stands or falls on their validity as stated for all $n$ and all measurable object spaces.","fun_headline_variants_meta":{"raw":{"variants":["Any IID predictor converts to conformal with little loss","Conformal mimics any IID predictor at near-zero efficiency cost","IID-valid predicts conformal: explicit constants, small loss","Conformal prediction universal for IID: no unspecified addends"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000333,"raw_usage":{"total_tokens":1870,"prompt_tokens":983,"completion_tokens":887,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":816}},"tokens_in":599,"tokens_out":887,"duration_ms":9269,"temperature":1.0,"reasoning_tokens":816,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:59:29.654539+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the companion proof [50] for the precise conditions under which Theorem 4 and Theorem 6 are proved; if either theorem needs a restriction that the present statement does not state, such as $n$ large, $\\mathcal{X}$ finite, or $|\\mathcal{Y}|$ above a threshold, then Corollary 3 as formulated in this paper is false. Concretely, test the minimal case $|\\mathcal{Y}|=2$, $n=1$: if the claimed factors $1/(\\mathrm{e}(|\\mathcal{Y}|-1))$ and $1/(|\\mathcal{Y}|-1)$ cannot be attained by any e-variable $G$ satisfying $\\mathbb{E}_Q[G]\\le 1$ for all IID probability measures $Q$, the inequality (19) fails.","supporting_citations":[{"cited_title":"Transductive Confidence Machine is universal","cited_arxiv_id":null,"evidence_quote":"The original universality theorem for conformal prediction under IID; it also supplies Proposition 1, which identifies train-invariant exchangeability p-tests with conformal predictors."},{"cited_title":"Non-algorithmic theory of randomness","cited_arxiv_id":null,"evidence_quote":"Introduces the functional (non-algorithmic) theory of randomness and provides the product decomposition $E_R=E_X E_{iR}$ used here as Theorem 3."},{"cited_title":"Universality of conformal prediction under the assumption of randomness","cited_arxiv_id":"2502.19254","evidence_quote":"Companion preprint that contains the formal proofs of Theorems 4, 5, and 6, the load-bearing steps stated without proof in this paper."},{"cited_title":"Conformal e- testing","cited_arxiv_id":null,"evidence_quote":"Provides the calibration theory used in Corollary 3: the calibrator $f(p)=\\delta p^{\\delta-1}$ and the e-to-p calibration step."},{"cited_title":"Machine-learning applications of algorithmic randomness","cited_arxiv_id":null,"evidence_quote":"The paper that introduced conformal prediction proper and its ideal picture; its remark on the fundamental limitation $\\log N$ is extended here to all IID-valid methods in classification."}],"review_version":1}