{"id":"af25411b-30a5-404c-81d2-dc1c5d5bcdce","arxiv_id":"2505.22764","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Applying learned test-time augmentation before conformal scoring reduces prediction set sizes by 10-14% with no loss of nominal coverage.","lead":"Test-time augmentation, averaging a model's predictions over transformed copies of an image, shrinks the class sets produced by conformal prediction by 10-14% on average while preserving the coverage guarantee. The technique needs no retraining of the base classifier and works with standard conformal scores such as APS and RAPS.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Coverage claim depends on where RAPS hyperparameters are tuned; if tuned on the calibration split, exchangeability is broken and the 10-14% reduction may be a tuning artifact.","rationale":"The reader identified the exchangeability of D_TTA as the weakest assumption. I partially agree, but the D_TTA construction is actually standard and valid: θ is a function of a split disjoint from D_cal and test, so conditional on D_TTA the transformed calibration and test scores are exchangeable. The more concrete threat to the same assumption is the undisclosed hyperparameter selection for RAPS. The phrase 'automatically select hyperparameters kreg and λ to minimize set size' is exactly the kind of calibration-set-dependent choice that invalidates split conformal if done on D_cal. This is not a criticism of the authors' honesty; it is an ambiguity in the manuscript that determines whether the nominal coverage statement is true. The central claim 'reduces set sizes by 10-14% while maintaining coverage' is load-bearing on this point because if the threshold is tuned to minimize set size on calibration data, smaller sets and nominal-looking coverage can be obtained even when the exchangeability guarantee is broken. The empirical coverage tables are not dispositive: with 20,000 calibration points, a small tuning bias may be within the reported standard errors while still exceeding the allowed error rate in a formal sense. I therefore recommend keeping the CONDITIONAL verdict, with the explicit condition that the authors disclose and, if needed, fix the data split used for RAPS hyperparameter selection and rerun the key comparison with hyperparameters fixed on a separate split. This is a single, checkable condition that settles the concern.","tokens_in":21288,"tokens_out":11946,"duration_ms":133176,"concrete_test":"Rerun the ImageNet RAPS α=0.05 comparison twice, over the same 10 splits: (a) with k_reg and λ selected on D_cal, as the current wording may imply; (b) with k_reg and λ fixed to values selected only on D_TTA (or a third hold-out). Report empirical coverage and average set size for RAPS and RAPS+TTA-Learned. If coverage in (b) is below nominal when it was at nominal in (a), or if the roughly 13% set-size reduction disappears, the coverage-maintenance claim is unsupported; if results and coverage are unchanged, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The coverage half of the central claim rests on exchangeability of the transformed scores. The paper's separate D_TTA split for θ (Section 4) does preserve this: conditional on D_TTA, D_cal and the test point remain exchangeable, so the standard split-conformal argument applies. The unresolved risk is the RAPS hyperparameter selection. Section 5 states 'automatically select hyperparameters kreg and λ to minimize set size' without saying which data are used. If k_reg and λ are chosen on D_cal, the score function is not fixed before calibration, the calibration scores are not exchangeable with the test score in the required sense, and the nominal coverage guarantee is void. This is load-bearing because the abstract claims both smaller sets and maintained coverage; tuning the threshold on calibration data can reduce set size while making coverage appear nominal in tables even when the formal guarantee is broken. The empirical coverage tables (e.g., Table S8) cannot settle this because with n≈20,000 a modest tuning bias would be within the reported standard errors. The paper needs to disclose the split used for hyperparameter selection, or fix the hyperparameters on a split disjoint from D_cal. This is not a challenge to the D_TTA construction, which is standard and sound.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes test-time-augmented conformal prediction (TTA-CP): for a fixed augmentation policy A, a set of aggregation weights θ is learned on a labeled split D_TTA, and the classifier's output probabilities in Eq. (4) are replaced by the TTA-aggregated probabilities g(x; f, A, θ). These transformed probabilities are then fed into standard split-conformal classifiers (APS and RAPS). The authors claim that this reduces average prediction-set size by 10%–14% while maintaining nominal coverage, and they support the claim with experiments on ImageNet, iNaturalist, and CUB-Birds using ResNet-50/101/152 and MobileNetV2, two conformal scores, three α levels, and ten calibration/test splits, including four corruption-based distribution shifts. The key theoretical safeguard is the division of the validation set into D_TTA (used to learn θ) and D_cal (used to compute the conformal threshold), which preserves conditional exchangeability.","tokens_in":21490,"tokens_out":6769,"duration_ms":83884,"significance":"If the central claim holds, the paper makes a useful and practical contribution: it reduces conformal set sizes without retraining the base classifier, works with any conformal score, and is computationally cheap. The paper has real strengths: the D_TTA split is a standard and sound way to preserve exchangeability; the experimental evaluation is broad and includes significance testing; the authors honestly report that gains are small or absent on CUB-Birds; and the analysis of why TTA helps (promoting the rank of the true class) is informative. However, two load-bearing issues must be resolved before the coverage guarantee and the headline reductions can be accepted: the manuscript does not disclose which split is used to select the RAPS hyperparameters k_reg and λ, and the supplementary text contains a direct contradiction of the D_TTA safeguard. These issues affect the formal validity of the coverage claim and the interpretation of the empirical reductions.","major_comments":[{"comment":"The text says 'automatically select hyperparameters kreg and λ to minimize set size' but does not state which data are used for this selection. If k_reg and λ are chosen on D_cal, then the score function is not fixed before calibration, the calibration scores are not exchangeable with the test score in the sense required by the split-conformal argument, and the nominal coverage guarantee is void; in that case the reported set-size reductions could be a tuning artifact. The empirical coverage tables (e.g., Table S8) cannot rule this out, because with n on the order of 20,000 a modest tuning bias would be within the reported standard errors. Please disclose the exact split used for hyperparameter selection; if necessary, fix k_reg and λ on a split disjoint from D_cal, and confirm that the baseline RAPS hyperparameters are selected in the same way.","section":"Section 5, Baselines"},{"comment":"Supplementary S1.2 ('Learning aggregation function') states: 'We learn ˆg by minimizing the cross-entropy loss with respect to the true labels on the calibration set.' This directly contradicts Section 4 ('Preserving exchangeability'), which motivates the D_TTA split precisely to avoid using calibration labels when learning the TTA transformation. If θ was in fact fit on D_cal, the coverage guarantee and the claimed 10%–14% reductions are unsupported. This wording must be corrected to specify the disjoint split, or the experiments must be rerun with the correct split.","section":"Supplementary S1.2"}],"minor_comments":[{"comment":"The caption contains an unfinished placeholder reading 'FILL IN THE REST, explain how TTA’s improvement to Top-1 accuracy alone is small...' This must be completed before the manuscript is publishable.","section":"Table S5 caption"},{"comment":"Equation numbering is inconsistent: the text says APS is 'described in Eqn. 4', but Eq. (4) defines the TTA-transformed probabilities; the APS score is defined in Eqs. (1)–(3). Please renumber or correct the cross-reference.","section":"Section 5, Baselines"},{"comment":"The claim 'TTA-Learned never decreases the coverage achieved by RAPS alone' is not literally supported by the table; for example, at α=0.10 with the simple augmentation policy on CUB-Birds, RAPS+TTA-Learned reports 0.913±0.011 versus 0.919±0.014 for RAPS. If the intended claim is that there is no statistically significant decrease, please state it that way.","section":"Section 6.1 and Table S8"},{"comment":"The main-text tables should state explicitly that the baseline RAPS uses the full validation set for calibration, while TTA variants use only the 80% remaining after setting aside D_TTA; the trade-off is studied in Figure S3, but the reader should not have to infer it from the supplement.","section":"Tables 1 and 2"},{"comment":"The text says 'Code to reproduce all experiments will be made publicly available' but no repository or release is provided. Please include a link or state how the code can be obtained for review.","section":"Section 5, Evaluation"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for cs.LG and the core idea is sound if the two flagged issues are resolved. The main risk is that the RAPS hyperparameters were tuned on the calibration split, which would invalidate the coverage guarantee; the contradictory sentence in S1.2 makes this risk concrete. If the authors can disclose a disjoint hyperparameter-selection procedure and correct the supplementary wording, the paper is likely acceptable. The unfinished placeholder in Table S5 suggests the supplement was not fully polished; the authors should check for other incomplete text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid, honestly-reported empirical paper. It shows that test-time augmentation with learned per-augmentation weights shrinks conformal prediction sets by 10-14% on average without losing coverage, across ImageNet, iNaturalist, and CUB-Birds, with ResNet-50/101/152 and MobileNetV2 in the supplement, APS and RAPS scores, three alpha levels, ten splits, and ImageNet-C shifts. The core idea is simple and new: fit the aggregation weights on a separate split so exchangeability is preserved. The mechanism analysis—TTA helps by promoting the true class's rank, not by fixing top-1 errors—is a genuine contribution and worth reading on its own.\n\nWhat the paper does well: the evaluation is wide, the CUB-Birds case where gains are small or absent is reported rather than hidden, and the coverage tables are there. The D_TTA construction is standard and sound; conditional on that split, the conformal argument goes through for a fixed score function.\n\nThe soft spots are real but mostly fixable. First, the RAPS hyperparameters k_reg and lambda are \"automatically selected to minimize set size\" without saying which data are used. If they are tuned on the calibration split, the score function is not fixed in advance, the calibration scores are not exchangeable with the test score, and the coverage guarantee is void. The tables cannot rule this out because n is large and the bias could be small. This needs a clear statement, and ideally the tuning should be on a split disjoint from calibration. Second, the supplement's Section S1.2 says the aggregation weights are learned \"on the calibration set,\" which directly contradicts Section 4's separate D_TTA split. I assume it is a typo, but together with the RAPS ambiguity it makes the validity argument harder to verify than it should be. Third, the supplement contains an unfinished caption (\"FILL IN THE REST\") and a typo in a table reference; sloppy for camera-ready but not a scientific problem. No code is released, so the numbers cannot be independently checked.\n\nThe central claim is not circular—theta is fit on separate data and set sizes are measured on held-out test data—but the coverage half of the abstract hinges on the hyperparameter disclosure. If the authors fix that, I would be comfortable with the paper.\n\nRecommendation: send it to a serious referee. This is a useful, honest contribution to conformal prediction practice, and the main issue is a missing method detail rather than a wrong result.","headline":"A solid, honestly reported empirical result on TTA for conformal prediction; the coverage guarantee needs one missing detail—where RAPS hyperparameters are tuned—resolved before it's fully clean.","tokens_in":22066,"tokens_out":3865,"would_cite":true,"duration_ms":36972,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Test-time augmentation shrinks conformal prediction sets by 10-14% while preserving coverage.","keywords":["conformal prediction","test-time augmentation","prediction set efficiency","exchangeability","distribution shift","image classification","adaptive prediction sets"],"falsifier":"Learn the TTA weights on the calibration set itself and check whether empirical coverage falls below $1-\\alpha$ on a held-out test set; or apply the expanded policy to a dataset like MNIST, where inverting colors changes digit labels, and check whether set sizes grow and coverage degrades.","tokens_in":21045,"feed_emoji":"📉","tokens_out":5348,"duration_ms":61374,"temperature":0.7,"pith_summary":"The paper tries to establish that test-time augmentation (TTA), which aggregates a classifier's predictions over transformed copies of an input, can be inserted into split conformal prediction to produce smaller prediction sets at the same nominal coverage. It claims this is a general efficiency boost: it works with two conformal scores (APS and RAPS), three architectures, three image datasets, and four corruption shifts, with average set-size reductions of about 10% in-distribution and 14% under shift, and no loss of coverage. This matters because conformal classifiers often produce uninformatively large sets, and the proposed fix requires no retraining of the base model and no extra base classifiers, only a learned aggregation of augmentations plus a fraction of the already-labeled data.","feed_headline":"Test-time augmentation shrinks conformal prediction sets by 10-14%","feed_subtitle":"A learned mix of image transformations gives smaller prediction sets with no coverage loss, no retraining.","key_machinery":"The load-bearing object is the learned TTA aggregation $g(x; f, A, \\theta) = \\sigma(\\theta^\\top A(f, A, x))$, in which $A(f, A, x)$ is the $M \\times K$ matrix of logits the pretrained classifier produces on $M$ augmented views of $x$, and $\\theta$ is a vector of augmentation weights learned by cross-entropy on a dedicated split $\\mathcal{D}_{\\mathrm{TTA}}$. It does two jobs: it induces useful invariances into the probability vector that feeds the conformal score, and, because $\\theta$ is learned on examples disjoint from the calibration set, it remains a deterministic transformation applied equally to calibration and test points, preserving the exchangeability on which the coverage guarantee rests.","core_discovery":"The central claim is that replacing the classifier's raw probability vector $f(x)$ in the conformal score with a TTA-aggregated vector $g(x; f, A, \\theta) = \\sigma(\\theta^\\top A(f, A, x))$, where $A$ stacks predicted logits over $m$ augmentations and $\\theta$ is learned to maximize cross-entropy on a separate labeled split, reduces the size of split-conformal prediction sets at level $1-\\alpha$ while preserving coverage. The paper shows that the effect is not mainly about fixing top-1 errors (only about 3% of shrinking sets come from corrected top-1 classification) but about promoting the true class's rank in the predicted-probability ordering, so the cumulative-probability conformal scores (APS and RAPS) need to include fewer wrong classes. It reports average set-size reductions of about 10% in-distribution and 14% under distribution shift, with no coverage loss, and finds that classes with the largest sets and hardest classes benefit most.","pith_inferences":["The rank-promotion mechanism suggests TTA will help most in high-cardinality label spaces and at low $\\alpha$; a direct test would vary the number of classes and coverage level while holding base accuracy fixed.","Because the aggregation is a single learned linear combination of logits, per-class or per-example weighting of augmentations is a natural extension that might improve on the reported 10-14%.","The top-$k$ diagnostic the paper uses could serve as a cheap pre-screening test: estimate the $k$ required for top-$k$ coverage on the validation split before committing to the extra forward passes.","The dependence on label-preserving augmentations implies a clear boundary: on domains where the transformation menu contains few label-preserving members, learned weights cannot manufacture invariances that are not there."],"forward_implications":["Across three image datasets and three ResNet architectures, learned TTA shrinks average prediction sets by about 10% in-distribution and 14% under ImageNet-C corruptions at nominal coverage levels of 90%, 95%, and 99%.","The gains appear with both APS and RAPS scoring, so the method sits on top of the conformal score rather than replacing it.","TTA can narrow the gap between base classifiers: ResNet-101 with learned TTA yields smaller sets than ResNet-152 without it at $\\alpha = 0.01$.","The extra data cost is modest: learning the aggregation on 20% of the validation set already gives most of the benefit, and the approach needs no retraining of the base model.","Under four corruption shifts, TTA-Learned keeps coverage at least as high as plain RAPS while producing smaller sets."],"supporting_citations":[{"why":"Supplies the APS conformal score, the cumulative-probability score that the paper modifies and benchmarks against.","marker":"[34]"},{"why":"Provides the RAPS score, its implementation, and the adaptivity metric used in the main comparisons.","marker":"[1]"},{"why":"Defines the split-conformal framework and the coverage guarantee that the paper builds on.","marker":"[36]"},{"why":"Establishes that deterministic transformations applied to calibration and test points preserve exchangeability, which the paper relies on for validity.","marker":"[24]"},{"why":"Provides the learned test-time augmentation aggregation approach that the paper adapts to conformal prediction.","marker":"[37]"},{"why":"Supplies the ImageNet-C corruption benchmark used for the distribution-shift experiments.","marker":"[19]"},{"why":"Provides the ImageNet dataset used in the main experiments.","marker":"[12]"},{"why":"Basis for the expanded augmentation policy used in the paper.","marker":"[11]"}],"fun_headline_variants":["TTA trims conformal prediction sets by 10-14% without coverage loss","Test-time augmentation cuts conformal set size by up to 14% under shift","TTA reorders confidences to shrink conformal sets by 10-14%","TTA shrinks conformal sets by 14% without retraining"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The coverage guarantee rests on the assumption that the labeled split used to learn the augmentation weights comes from the same distribution as future test examples, and that the augmentations preserve each image's true class; if either fails, the smaller sets may come without the advertised coverage.","fun_headline_variants_meta":{"raw":{"variants":["TTA trims conformal prediction sets by 10-14% without coverage loss","Test-time augmentation cuts conformal set size by up to 14% under shift","TTA reorders confidences to shrink conformal sets by 10-14%","TTA shrinks conformal sets by 14% without retraining"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001216,"raw_usage":{"total_tokens":4975,"prompt_tokens":889,"completion_tokens":4086,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":3999}},"tokens_in":505,"tokens_out":4086,"duration_ms":29296,"temperature":1.0,"reasoning_tokens":3999,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:00:42.992039+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Learn the TTA weights on the calibration set itself and check whether empirical coverage falls below $1-\\alpha$ on a held-out test set; or apply the expanded policy to a dataset like MNIST, where inverting colors changes digit labels, and check whether set sizes grow and coverage degrades.","supporting_citations":[{"cited_title":"A Tutorial on Conformal Prediction","cited_arxiv_id":null,"evidence_quote":"Defines the split-conformal framework and the coverage guarantee that the paper builds on."},{"cited_title":"Better Aggregation in Test-Time Augmen- tation","cited_arxiv_id":null,"evidence_quote":"Provides the learned test-time augmentation aggregation approach that the paper adapts to conformal prediction."}],"review_version":1}