Pith. sign in

REVIEW 17 cited by

TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.11297 v4 pith:GXENAZTU submitted 2021-06-21 cs.CV cs.LG

TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

classification cs.CV cs.LG
keywords tokensvisualresultsvideoapproachattentionimageimages
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In this paper, we introduce a novel visual representation learning which relies on a handful of adaptively learned tokens, and which is applicable to both image and video understanding tasks. Instead of relying on hand-designed splitting strategies to obtain visual tokens and processing a large number of densely sampled patches for attention, our approach learns to mine important tokens in visual data. This results in efficiently and effectively finding a few important visual tokens and enables modeling of pairwise attention between such tokens, over a longer temporal horizon for videos, or the spatial content in images. Our experiments demonstrate strong performance on several challenging benchmarks for both image and video recognition tasks. Importantly, due to our tokens being adaptive, we accomplish competitive results at significantly reduced compute amount. We obtain comparable results to the state-of-the-arts on ImageNet while being computationally more efficient. We also confirm the effectiveness of the approach on multiple video datasets, including Kinetics-400, Kinetics-600, Charades, and AViD. The code is available at: https://github.com/google-research/scenic/tree/main/scenic/projects/token_learner

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

    cs.CV 2026-06 unverdicted novelty 7.0

    EgoSAT is the first benchmark unifying retrospective, online, and prospective reasoning tasks in egocentric streaming video to evaluate VLMs, revealing struggles with temporal modeling and mis-calibration.

  2. Why Training-Free Token Reduction Collapses: The Inherent Instability of Pairwise Scoring Signals

    cs.AI 2026-04 unverdicted novelty 7.0

    Pairwise scoring signals in Vision Transformer token reduction are inherently unstable due to high perturbation counts and degrade in deep layers, causing collapse, while unary signals with triage enable CATIS to reta...

  3. ATS-ToDMA: Adaptive Token Selection and Token-Domain Multiple Access for Cross-Modal Semantic Communications

    cs.IT 2026-07 conditional novelty 6.0

    ATS-ToDMA jointly selects semantic tokens, schedules them under a similarity-based SSINR interference model, and allocates power, yielding higher simulated semantic throughput and accuracy than OMA and Semantic NOMA.

  4. Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

    cs.CV 2026-06 unverdicted novelty 6.0

    STORM is a training-free spatial-aware token reduction framework that reformulates compression on spatial units to preserve grid topology and neighborhood coherence in visual state space models.

  5. See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs

    cs.RO 2026-06 unverdicted novelty 6.0

    S2 improves generalization in vision-language-action models by using goal-preserving refined language guidance and explicit visual evidence budgets, raising mean subtask success from 54.2% to 79.0% on eight real-robot...

  6. FlowNar: Scalable Streaming Narration for Long-Form Videos

    cs.CV 2026-05 unverdicted novelty 6.0

    FlowNar achieves bounded memory and 3x higher throughput for streaming narration on Ego4D, EgoExo4D, and EpicKitchens100 by combining dynamic historical context removal with a Cross Linear Attentive Memory module.

  7. Head Similarity: Modeling Structured Whole-Head Appearance Beyond Face Recognition

    cs.CV 2026-05 unverdicted novelty 6.0

    Head Similarity extends identity recognition to structured whole-head similarity by capturing intra-identity appearance variations via hierarchical supervision on a weakly-labeled video benchmark.

  8. VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning

    cs.RO 2026-03 conditional novelty 6.0

    An RGB-only diffusion policy that builds an explicit volumetric representation, distills it into spatial tokens, and conditions a multi-token decoder reaches 88.8% average success on LIBERO.

  9. Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

    cs.RO 2026-02 conditional novelty 6.0

    Using tokenized proprioception plus instruction to select ~15% of visual patches matches or beats full-token VLA baselines and cuts latency by ~58%.

  10. A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models

    cs.CV 2025-11 conditional novelty 6.0

    Visual token pruning causes severe loss in discrete diffusion MLLMs; only from-scratch models on long-answer tasks recover via late denoising, so redundancy is recoverability, not dispensability.

  11. One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory

    cs.CV 2025-05 unverdicted novelty 6.0

    TrajViT tokenizes videos via panoptic sub-object trajectories, achieving 10x token reduction and outperforming ViT3D by 6% on retrieval and 5.2% on VideoQA tasks with faster training and inference.

  12. PaLM-E: An Embodied Multimodal Language Model

    cs.LG 2023-03 conditional novelty 6.0

    PaLM-E is a single 562B-parameter multimodal model that performs embodied reasoning tasks like robotic manipulation planning and visual question answering by interleaving vision, state, and text inputs with positive t...

  13. Florence: A New Foundation Model for Computer Vision

    cs.CV 2021-11 unverdicted novelty 6.0

    Florence is a new vision foundation model that learns universal visual-language representations from web-scale data and reports state-of-the-art results on 44 benchmarks including 83.74% zero-shot ImageNet top-1 accuracy.

  14. Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection

    cs.CV 2026-07 conditional novelty 5.0

    A split-and-fuse video transformer with sparse winner-takes-all token selection reports 82.55% top-1 on Kinetics-400, Pareto-efficient inference, and peak EEG RSA of 0.18 (about 78% of the noise ceiling).

  15. DinoLink: A Token-Centric Representation Compression Framework for Bandwidth-Constrained Collaborative V2X Perception

    cs.CV 2026-06 unverdicted novelty 5.0

    DinoLink uses saliency-aware token pruning and residual vector quantization to cut V2X bitrate by 139x while retaining 32.8% mAP on nuScenes.

  16. Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance

    cs.CV 2025-09 conditional novelty 5.0

    IGG, an attention-based reweighting of classifier-free guidance, concentrates guidance on important tokens and modestly improves FID/IS in scale-wise autoregressive image generation.

  17. DinoLink: A Token-Centric Representation Compression Framework for Bandwidth-Constrained Collaborative V2X Perception

    cs.CV 2026-06 unverdicted novelty 4.0

    DinoLink uses saliency-aware token pruning plus residual vector quantization to cut V2X bitrate by 139x while reporting 32.8% mAP on nuScenes.