Pith. sign in

REVIEW 1 cited by

Semi-supervised Vision Transformers at Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.05688 v1 pith:INN23RVR submitted 2022-08-11 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords semi-supervisedfine-tuninglabelstransformersvisionaccuracyachievescomparable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study semi-supervised learning (SSL) for vision transformers (ViT), an under-explored topic despite the wide adoption of the ViT architectures to different tasks. To tackle this problem, we propose a new SSL pipeline, consisting of first un/self-supervised pre-training, followed by supervised fine-tuning, and finally semi-supervised fine-tuning. At the semi-supervised fine-tuning stage, we adopt an exponential moving average (EMA)-Teacher framework instead of the popular FixMatch, since the former is more stable and delivers higher accuracy for semi-supervised vision transformers. In addition, we propose a probabilistic pseudo mixup mechanism to interpolate unlabeled samples and their pseudo labels for improved regularization, which is important for training ViTs with weak inductive bias. Our proposed method, dubbed Semi-ViT, achieves comparable or better performance than the CNN counterparts in the semi-supervised classification setting. Semi-ViT also enjoys the scalability benefits of ViTs that can be readily scaled up to large-size models with increasing accuracies. For example, Semi-ViT-Huge achieves an impressive 80% top-1 accuracy on ImageNet using only 1% labels, which is comparable with Inception-v4 using 100% ImageNet labels.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Info-Coevolution: An Efficient Framework for Data Model Coevolution

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A data-model coevolution framework that fuses model and nearest-neighbor predictions to select labels, reaching ImageNet-1K accuracy with 68% of annotations and 50% under semi-supervised training.

Pith tools