Pith. sign in

REVIEW 2 cited by

Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00989 v1 pith:X54ZTFUO submitted 2023-06-01 cs.CV cs.LG

classification cs.CVcs.LG
keywords hieravisionhierarchicaltransformeraddedbells-and-whistlescomponentsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance. While these components lead to effective accuracies and attractive FLOP counts, the added complexity actually makes these transformers slower than their vanilla ViT counterparts. In this paper, we argue that this additional bulk is unnecessary. By pretraining with a strong visual pretext task (MAE), we can strip out all the bells-and-whistles from a state-of-the-art multi-stage vision transformer without losing accuracy. In the process, we create Hiera, an extremely simple hierarchical vision transformer that is more accurate than previous models while being significantly faster both at inference and during training. We evaluate Hiera on a variety of tasks for image and video recognition. Our code and models are available at https://github.com/facebookresearch/hiera.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Vision foundation model features improve pixel and object classification in microscopy over hand-crafted features, and object-guided attentive probing (ObAP) can match or beat supervised baselines with very few labels.

  2. SAM2RL: Towards Reinforcement Learning Memory Control in Segment Anything Model 2

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A reinforcement learning agent that controls SAM 2 memory bank updates achieves a +4.91% tracking quality gain over SAM 2 when overfitted per video, indicating untapped potential in memory control.

Pith tools