Pith. sign in

REVIEW 1 cited by

Better Aggregation in Test-Time Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.11156 v2 pith:2N2K5QWV submitted 2020-11-23 cs.CV

classification cs.CV
keywords test-timeaugmentationpredictionsmethodacrossaggregationaugmentationsaverage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Test-time augmentation -- the aggregation of predictions across transformed versions of a test input -- is a common practice in image classification. Traditionally, predictions are combined using a simple average. In this paper, we present 1) experimental analyses that shed light on cases in which the simple average is suboptimal and 2) a method to address these shortcomings. A key finding is that even when test-time augmentation produces a net improvement in accuracy, it can change many correct predictions into incorrect predictions. We delve into when and why test-time augmentation changes a prediction from being correct to incorrect and vice versa. Building on these insights, we present a learning-based method for aggregating test-time augmentations. Experiments across a diverse set of models, datasets, and augmentations show that our method delivers consistent improvements over existing approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CLaC at SemEval-2025 Task 6: A Multi-Architecture Approach for Corporate Environmental Promise Verification

    cs.CL 2025-05 conditional novelty 3.0 of 10

    A shared-task system report in which a multitask DeBERTa-v3 model with attention pooling and test-time augmentation scores 0.5268 on SemEval-2025 Task 6, a 0.4 percent relative gain over the baseline.

Pith tools