Pith. sign in

REVIEW 1 cited by

CrossTransformers: spatially-aware few-shot transfer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.11498 v5 pith:4APT4SRV submitted 2020-07-22 cs.CV

classification cs.CV
keywords transfervisioncrosstransformersdomainfeaturesimagesinformationlabeled
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Given new tasks with very little data$-$such as new classes in a classification problem or a domain shift in the input$-$performance of modern vision systems degrades remarkably quickly. In this work, we illustrate how the neural network representations which underpin modern vision systems are subject to supervision collapse, whereby they lose any information that is not necessary for performing the training task, including information that may be necessary for transfer to new tasks or domains. We then propose two methods to mitigate this problem. First, we employ self-supervised learning to encourage general-purpose features that transfer better. Second, we propose a novel Transformer based neural network architecture called CrossTransformers, which can take a small number of labeled images and an unlabeled query, find coarse spatial correspondence between the query and the labeled images, and then infer class membership by computing distances between spatially-corresponding features. The result is a classifier that is more robust to task and domain shift, which we demonstrate via state-of-the-art performance on Meta-Dataset, a recent dataset for evaluating transfer from ImageNet to many other vision datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ViT-ProtoNet for Few-Shot Image Classification: A Multi-Benchmark Evaluation

    cs.CV 2025-07 reject novelty 2.0 of 10

    ViT-ProtoNet, a Prototypical Network with a ViT-Small encoder, is reported to reach 95-97% 5-shot accuracy on three benchmarks and 81.88% on FC100, but the evaluation lacks critical baselines.

Pith tools