Pith. sign in

REVIEW 1 cited by

SparseFormer: Sparse Visual Recognition via Limited Latent Tokens

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03768 v1 pith:C5GU3BJR submitted 2023-04-07 cs.CV

classification cs.CV
keywords sparseformervisualsparsedenserecognitionspaceclassificationcomputational
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human visual recognition is a sparse process, where only a few salient visual cues are attended to rather than traversing every detail uniformly. However, most current vision networks follow a dense paradigm, processing every single visual unit (e.g,, pixel or patch) in a uniform manner. In this paper, we challenge this dense paradigm and present a new method, coined SparseFormer, to imitate human's sparse visual recognition in an end-to-end manner. SparseFormer learns to represent images using a highly limited number of tokens (down to 49) in the latent space with sparse feature sampling procedure instead of processing dense units in the original pixel space. Therefore, SparseFormer circumvents most of dense operations on the image space and has much lower computational costs. Experiments on the ImageNet classification benchmark dataset show that SparseFormer achieves performance on par with canonical or well-established models while offering better accuracy-throughput tradeoff. Moreover, the design of our network can be easily extended to the video classification with promising performance at lower computational costs. We hope that our work can provide an alternative way for visual modeling and inspire further research on sparse neural architectures. The code will be publicly available at https://github.com/showlab/sparseformer

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A patch-selection backbone prunes redundant video patches using GOP motion and residual cues, then adds pseudo global context, matching ViT-B accuracy at roughly 26% of its FLOPs.

Pith tools