REVIEW 4 major objections 5 minor 54 references
Principles of Visual Tokens for Efficient Video Understanding
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that a cheap MLP imitating a gradient oracle selects video tokens better than all existing methods, cutting compute roughly in half with little accuracy loss.
desk verdict LITE's efficiency gains look real, but the paper's claim that it works by imitating a privileged-label oracle is undercut by its own tables. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the oracle: for a target class $c$, the gradient of the pre-softmax score $y^c$ with respect to the backbone's last-block feature maps $A^d_{thw}$ is spatially averaged (Eq. 1) to give per-feature importance weights, and each token's score is the ReLU-weighted combination of its activations (Eq. 2), normalized to $[0,1]$. LITE is a three-layer MLP, applied token-wise to the patch embeddings, trained with binary cross-entropy to reproduce these oracle scores; at inference the frozen VideoMAE backbone processes only the top-budget fraction of tokens. An optional adaptive budget (LITE++) uses a fast MoviNet confidence estimate to assign fewer tokens to easy videos.
What would settle it
Measure the rank correlation between LITE's predicted token scores and the Grad-CAM oracle's scores on held-out videos; a near-zero correlation while accuracy is maintained would show the selector is not actually learning the oracle, and at matched GFLOPs on a novel dataset a tie with random token selection would falsify the claimed generalization.
Extended reading notes
Core claim
The central claim is that a Grad-CAM-style gradient score, computed from the true class label, accurately measures how much each visual token contributes to a frozen video transformer's decision, and that this score distribution is strongly Pareto: most tokens carry almost no information while a handful are essential. Keeping only the high-value tokens can outperform the full model by up to 9% accuracy, so low-value tokens are not merely redundant but actively harmful. From this the paper distills five principles: random token sampling is a stronger baseline than most learned methods; good tokens do not coincide with attention, motion, or saliency cues; low-value tokens hurt; token values follow a Pareto distribution; and easy videos need fewer tokens. LITE operationalizes the principles with a three-layer MLP that predicts the oracle scores from patch embeddings, selecting the top tokens for the unchanged backbone; on Kinetics-400 and Something-Something-V2 it sits on the Pareto front of the accuracy-versus-GFLOPs trade-off. The paper further reports that a selector trained on Kinetics-400 alone transfers to UCF101, Something-Something-V2, and AVA action detection with minimal loss.
Load-bearing premise
The whole method rests on the assumption that the gradient of the frozen backbone's true-class score is a faithful, transferable measure of a token's value, and that a small MLP trained on initial embeddings can reproduce that measure well enough to transfer across datasets and tasks.
Editorial extensions
If this is right
- LITE cuts GFLOPs by more than half with under one percentage point of accuracy loss on both Kinetics-400 and Something-Something-V2, outperforming ToMe, STA, STTS, LookupViT, and ObjectViViT at comparable compute.
- Keeping only the oracle's top tokens can beat the full model by up to 9% accuracy, which means many tokens are not merely redundant but actively hurt classification.
- The Pareto-like distribution of token values explains why random token dropping is such a strong baseline: random sampling rarely removes the few tail tokens that carry the signal.
- A LITE selector trained on Kinetics-400 alone transfers to Something-Something-V2, UCF101, and AVA action detection with minimal accuracy or mAP loss, indicating that token importance is consistent across domains.
- The adaptive budget variant, LITE++, reduces computation by up to 34% at the cost of less than one percentage point of accuracy by giving easy videos fewer tokens.
Reading between the lines
- The paper leaves implicit that the same gradient oracle could be used to denoise training itself by dropping low-value tokens during fine-tuning, turning pruning from an inference trick into a regularizer.
- A testable extension would be replacing the external MoviNet confidence estimator in LITE++ with the backbone's own softmax confidence, checking whether the adaptive budget can be made self-contained without losing savings.
- Because the oracle uses the true label, its offline scores can be computed once per training video; a natural next step is distilling the selector entirely from these offline scores without any additional labels at inference time.
- The paper's comparison suggests that future token-reduction papers should always report a random-token baseline at matched GFLOPs, since most published methods fail to beat it.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the value of visual tokens in video transformers, proposes five qualitative principles about token importance, and introduces LITE, a lightweight MLP selector trained to imitate a Grad-CAM-based oracle that scores tokens using true-label gradients. At inference, LITE keeps the highest-scoring tokens before the transformer blocks; LITE++ further adapts the token budget per video using a confidence estimate. Experiments on Something-Something-V2 and Kinetics-400 compare LITE against random token selection, token merging, and prior token-selection methods, and additional experiments report zero-shot transfer of a K400-trained selector to UCF101, SS-V2, and AVA action detection. The central claim is that LITE achieves a better accuracy-versus-GFLOPs trade-off than existing baselines, including random dropping, and that the learned selector generalizes across datasets and tasks without retraining.
Significance. If the empirical trade-off and the generalization results hold, LITE is a simple and practical plug-and-play token selector for frozen video transformers, and the five principles provide useful organizing observations for future work on token reduction. The paper contains a broad set of comparisons on two standard datasets, a cross-backbone experiment, and a cross-task experiment, which are valuable even if the mechanism is only partially understood. The main limitations are that the reported efficiency omits the selector's own compute, the oracle-approximation mechanism is not supported by the paper's own numbers, and the zero-shot claim lacks an appropriate matched baseline. The contribution is therefore significant but conditional on resolving these measurement and conceptual issues.
major comments (4)
- [Sec. 4.2, Tables 1 and 4] The paper states that the LITE selector is trained to reproduce the true-label Grad-CAM oracle of Eqs. (1)-(2), but the reported numbers contradict this mechanism. On SS-V2 at P-Ratio 0.5, the true-label oracle reaches 78.52 Top-1, the predicted-label oracle reaches 70.00, and VideoMAE-LITE50 reaches 69.91. LITE is therefore 8.6 points below its training target and essentially at the level of a label-free oracle. Because the paper credits LITE's gains to oracle approximation and uses this to argue that important tokens transfer across domains, this is a load-bearing issue. Please report the rank correlation or top-K overlap between LITE scores and oracle scores, and add the predicted-label oracle to the main Pareto plots (Fig. 2 and Fig. 1), so the reader can see whether LITE is actually learning the true-label oracle or merely a label-free proxy.
- [Sec. 5.1, Tables 1-2, Fig. 1] The reported GFLOPs for LITE do not appear to include the cost of the selector MLP or, for LITE++, the MoviNet confidence model. Since the central claim is an efficiency trade-off, the end-to-end compute including all added components must be reported; otherwise the GFLOPs savings may be overstated. In addition, the accuracy numbers are presented without error bars or multiple-seed results, and the claim that LITE forms the "optimal Pareto front" is a qualitative visual judgment. Please provide variance estimates and a quantitative statement of the accuracy loss at matched GFLOPs, including the selector overhead in the reported numbers.
- [Sec. 5.2, Table 3] The zero-shot generalization claim is not supported by the current experimental design. The VideoMAE baseline rows are described as trained and tested on the same dataset, whereas the LITE rows use a selector trained on Kinetics-400; there is no VideoMAE baseline trained on K400 and evaluated on UCF101/SS-V2/AVA without a selector, and no selector trained on the target dataset for comparison. Without these controls, the observed accuracy retention could be due to the backbone being shared or to dataset-specific training effects, rather than to the selector generalizing. Please add matched baselines (same backbone training setup, with and without a target-trained selector) to separate selector transfer from backbone behavior.
- [Sec. 3, Principle 4 and Fig. 3] The claim that token values "closely follow a Pareto distribution" is supported only by a histogram and a visual resemblance. Since the five principles are presented as a main contribution, this should be quantified with a distribution fit, an estimated tail exponent, or at least a goodness-of-fit or error-bar analysis. The term "Pareto" is used loosely; a heavy-tailed distribution with many near-zero values could be many other families. Please provide quantitative evidence for the distributional claim, or soften the statement to a qualitative heavy-tail observation.
minor comments (5)
- [Eq. (1)] The normalization constant N in Eq. (1) is not defined; the summation ranges over t, h, w should be stated explicitly, along with the layer at which A_d_thw is taken.
- [Sec. 5.1 and Fig. 2] The predicted-label oracle from Table 4 is a natural reference point for token selection and should appear in the main Pareto plot, not only in the analysis table.
- [Sec. 4.3, Table 5] The adaptive-budget thresholds tau1 = 0.1 and tau2 = 0.5 and the hand-set reductions (30% / 20%) are not accompanied by a sensitivity analysis; a short ablation would clarify how robust LITE++ is to these choices.
- [Sec. 5.4, Table 6] For the TimeSformer experiment, please state whether the selector was trained on TimeSformer features or transferred from VideoMAE, and whether the oracle was recomputed for TimeSformer; this affects the interpretation of the cross-backbone result.
- [Throughout] There are several typos and formatting issues, including "cutting over of GFLOPs" in the introduction, the missing space in "ClassificationK400" in Table 3, and inconsistent capitalization of "Pareto" and "state-of-the-art". Please proofread the final version.
Circularity Check
No significant circularity: LITE is validated against external baselines; oracle-relative token values and the 'oracle is hard to predict' admission are correctness caveats, not circular reductions.
full rationale
The paper's only internal target is the Grad-CAM oracle (Eqs. 1-2): token scores S_c are gradients of the frozen VideoMAE class score w.r.t. last-block activations, using the true label. LITE's MLP is trained with BCE against those scores (Eq. 4, Sec. 4.2), which is a normal supervised fitting step, not a renamed prediction: the selector's outputs are evaluated on held-out test videos and against independent methods (ToMe, STA, LookupViViT, STTS, etc.). The oracle is model-relative because it measures influence on the same VideoMAE backbone that LITE prunes, so the 'principles' describe this model's gradient geometry rather than an absolute perceptual quantity; however, LITE's central efficiency claim is not derived from the oracle by construction—it is an empirical Pareto-front comparison in Tables 1-2 and zero-shot transfers in Table 3. The paper's self-citations ([17],[18],[19]) appear in related work and peripheral energy/context discussions and carry no load-bearing uniqueness or ansatz. The largest internal tension is the mechanism story: Table 4 shows the true-label oracle reaches 78.52 on SS-V2 at P-Ratio 0.5 while the predicted-label oracle reaches 70.00 and VideoMAE-LITE50 reaches 69.91, and Sec. 4.4 concedes 'the oracle is extremely hard to predict'. That contradiction questions whether LITE actually learns the oracle, but it is an empirical/interpretation failure, not a circular reduction; no equation is equal to its input by construction. Score 0.
Assumptions & free parameters
free parameters (3)
- adaptive budget thresholds tau1, tau2 =
tau1=0.1, tau2=0.5
- easy-case budget reductions =
top 20%-30% tokens depending on P-Ratio
- selector MLP hidden size =
not specified
assumptions (3)
- domain assumption Grad-CAM gradients of the target class score with respect to deep activations are a valid measure of token value.
- domain assumption A frozen VideoMAE backbone remains accurate when tokens are discarded without fine-tuning.
- domain assumption Selector predictions trained on K400 transfer to unseen datasets (SS-V2, UCF101) and tasks (AVA detection).
Cite this review
Pith. "Pith review of Principles of Visual Tokens for Efficient Video Understanding." pith.science (2026). https://pith.science/paper/VXVO4RCZ
@misc{pith2026241113626,
author = {Pith},
title = {Pith review of: Principles of Visual Tokens for Efficient Video Understanding},
year = {2026},
howpublished = {\url{https://pith.science/paper/VXVO4RCZ}},
note = {Machine review of arXiv:2411.13626}
}
read the original abstract
Video understanding has made huge strides in recent years, relying largely on the power of transformers. As this architecture is notoriously expensive and video data is highly redundant, research into improving efficiency has become particularly relevant. Some creative solutions include token selection and merging. While most methods succeed in reducing the cost of the model and maintaining accuracy, an interesting pattern arises: most methods do not outperform the baseline of randomly discarding tokens. In this paper we take a closer look at this phenomenon and observe 5 principles of the nature of visual tokens. For example, we observe that the value of tokens follows a clear Pareto-distribution where most tokens have remarkably low value, and just a few carry most of the perceptual information. We build on these and further insights to propose a lightweight video model, LITE, that can select a small number of tokens effectively, outperforming state-of-the-art and existing baselines across datasets (Kinetics-400 and Something-Something-V2) in the challenging trade-off of computation (GFLOPs) vs accuracy. Experiments also show that LITE generalizes across datasets and even other tasks without the need for retraining.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Vivit: A video vi- sion transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lucic, and Cordelia Schmid. Vivit: A video vi- sion transformer. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 6816–6826, 2021. 1, 3, 7
work page 2021
-
[2]
Is space-time attention all you need for video understanding? ArXiv, abs/2102.05095, 2021
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? ArXiv, abs/2102.05095, 2021. 7, 8
arXiv 2021
-
[3]
Is space-time attention all you need for video understanding? In ICML, page 4, 2021
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, page 4, 2021. 3
2021
-
[4]
Token merging: Your ViT but faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token merging: Your ViT but faster. In International Conference on Learning Representations, 2023. 2, 3, 4, 7
work page 2023
-
[5]
Revisiting the” video” in video-language understanding
Shyamal Buch, Crist ´obal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles. Revisiting the” video” in video-language understanding. In CVPR, 2022. 3
work page 2022
-
[6]
Space-time mixing attention for video transformer
Adrian Bulat, Juan-Manuel P ´erez-R´ua, Swathikiran Sud- hakaran, Brais Mart ´ınez, and Georgios Tzimiropoulos. Space-time mixing attention for video transformer. InNeural Information Processing Systems, 2021. 3
work page 2021
-
[7]
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceed- ings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015. 1
2015
-
[8]
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In pro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017. 1, 3
work page 2017
Show all 54 references
-
[9]
An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024
Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024. 3
2024
-
[10]
Joonmyung Choi, Sanghyeok Lee, Jaewon Chu, Minhyuk Choi, and Hyunwoo J. Kim. vid-tldr: Training free token merging for light-weight video transformer. In Conference on Computer Vision and Pattern Recognition, 2024. 2, 3
2024
-
[11]
Prune spatio-temporal tokens by semantic-aware temporal accumulation
Shuangrui Ding, Peisen Zhao, Xiaopeng Zhang, Rui Qian, Hongkai Xiong, and Qi Tian. Prune spatio-temporal tokens by semantic-aware temporal accumulation. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 16899–16910, 2023. 2, 3, 4, 7
2023
-
[12]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. InInternational ...
2020
-
[13]
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. 2021 IEEE/CVF Interna- tional Conference on Computer Vision (ICCV), pages 6804– 6815, 2021. 3
2021
-
[14]
X3d: Expanding architectures for efficient video recognition
Christoph Feichtenhofer. X3d: Expanding architectures for efficient video recognition. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 200–210, 2020. 3
2020
-
[15]
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 6201–6210, 2018. 3
2019
-
[16]
Efficient video transformers via spatial-temporal token merging for action recognition
Zhanzhou Feng, Jiaming Xu, Lei Ma, and Shiliang Zhang. Efficient video transformers via spatial-temporal token merging for action recognition. ACM Transactions on Mul- timedia Computing, Communications and Applications , 20: 1 – 21, 2023. 2, 3, 7
2023
-
[17]
Smart frame selection for action recognition
Shreyank N Gowda, Marcus Rohrbach, and Laura Sevilla- Lara. Smart frame selection for action recognition. In Pro- ceedings of the AAAI Conference on Artificial Intelligence , pages 1451–1459, 2021. 3
2021
-
[18]
Watt for what: Rethinking deep learn- ing’s energy-performance relationship
Shreyank N Gowda, Xinyue Hao, Gen Li, Shashank Narayana Gowda, Xiaobo Jin, and Laura Sevilla-Lara. Watt for what: Rethinking deep learn- ing’s energy-performance relationship. arXiv preprint arXiv:2310.06522, 2023. 1
2023 arXiv
-
[19]
Optimizing factorized encoder models: Time and memory reduction for scalable and efficient action recognition
Shreyank N Gowda, Anurag Arnab, and Jonathan Huang. Optimizing factorized encoder models: Time and memory reduction for scalable and efficient action recognition. In European Conference on Computer Vision, pages 457–474. Springer, 2025. 3
2025
-
[20]
something something
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michal- ski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fr ¨und, Peter N. Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic. The “something something” video...
2017
-
[21]
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A Ross, Carl V ondrick, Car- oline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. Ava: A video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the IEEE conf...
-
[22]
D. I. Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew A. Brown, and Boqing Gong. Movinets: Mobile video networks for efficient video recog- nition. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16015–16025, 2021. 3, 6
2021
-
[23]
Lookupvit: Compressing visual information to a limited number of tokens
Rajat Koner, Gagan Jain, Prateek Jain, V olker Tresp, and Su- joy Paul. Lookupvit: Compressing visual information to a limited number of tokens. ArXiv, abs/2407.12753, 2024. 3, 7
2024 arXiv
-
[24]
Revisiting token pruning for object detection and instance segmentation
Yifei Liu, Mathias Gehrig, Nico Messikommer, Marco Can- nici, and Davide Scaramuzza. Revisiting token pruning for object detection and instance segmentation. In Proceed- ings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2658–2668, 2024. 3
2024
-
[25]
Swin transformer: 9 Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: 9 Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021. 3
2021
-
[26]
Video swin transformer
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3192–3201, 2021. 7
2022
-
[27]
Video transformer network
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Assel- mann. Video transformer network. 2021 IEEE/CVF Interna- tional Conference on Computer Vision Workshops (ICCVW), pages 3156–3165, 2021. 3
2021
-
[28]
Expanding language-image pretrained models for gen- eral video recognition
Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling. Expanding language-image pretrained models for gen- eral video recognition. In European Conference on Com- puter Vision, pages 1–18. Springer, 2022. 3
2022
-
[29]
St-adapter: Parameter-efficient image-to-video transfer learning
Junting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao, and Hong- sheng Li. St-adapter: Parameter-efficient image-to-video transfer learning. Advances in Neural Information Process- ing Systems, 35:26462–26477, 2022. 3
2022
-
[30]
K-centered patch sampling for efficient video recognition
Seong Hyeon Park, Jihoon Tack, Byeongho Heo, Jung-Woo Ha, and Jinwoo Shin. K-centered patch sampling for efficient video recognition. In European Conference on Computer Vision, 2022. 3
2022
-
[31]
Asano, Is- han Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Jo ˜ao F
Mandela Patrick, Dylan Campbell, Yuki M. Asano, Is- han Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Jo ˜ao F. Henriques. Keeping your eye on the ball: Trajectory attention in video transformers. In Neural Information Processing Systems, 2021. 2, 3, 7
2021
-
[32]
So, Maud Texier, and Jeff Dean
David Patterson, Joseph Gonzalez, Urs H ¨olzle, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David R. So, Maud Texier, and Jeff Dean. The carbon foot- print of machine learning training will plateau, then shrink. Computer, 55(7):18–28, 2022. 1
2022
-
[33]
DiCarlo, Todd M
Benjamin Peters, James J. DiCarlo, Todd M. Gureckis, Ralf Haefner, Leyla Isik, Joshua B. Tenenbaum, Talia Konkle, Thomas Naselaris, Kimberly L. Stachenfeld, Zenna Tavares, Doris Tsao, Ilker Yildirim, and Nikolaus Kriegeskorte. How does the primate brain combine generative and ...
2024 arXiv
-
[34]
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. Dynamicvit: Efficient vision transformers with dynamic token sparsification. Advances in neural information processing systems, 34:13937–13949,
-
[35]
Tokenlearner: What can 8 learned tokens do for images and videos? arXiv preprint arXiv:2106.11297, 2021
Michael S Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova. Tokenlearner: What can 8 learned tokens do for images and videos? arXiv preprint arXiv:2106.11297, 2021. 3
2021 arXiv
-
[36]
Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Ba- tra
Ramprasaath R. Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Ba- tra. Grad-cam: Visual explanations from deep networks via gradient-based localization. International Journal of Com- puter Vision, 128:336 – 359, 2016. 2, 4
2016
-
[37]
Only time can tell: Discovering temporal data for temporal modeling
Laura Sevilla-Lara, Shengxin Zha, Zhicheng Yan, Vedanuj Goswami, Matt Feiszli, and Lorenzo Torresani. Only time can tell: Discovering temporal data for temporal modeling. 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 535–544, 2019. 4
2021
-
[38]
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012. 4, 7
2012 arXiv
-
[39]
Videomae: Masked autoencoders are data-efficient learn- ers for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learn- ers for self-supervised video pre-training. ArXiv, abs/2203.12602, 2022. 4, 7, 8
2022 arXiv
-
[40]
Training data-efficient image transformers & distillation through at- tention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv ´e J´egou. Training data-efficient image transformers & distillation through at- tention. In International conference on machine learning , pages 10347–10357. PMLR, 2021. 3
2021
-
[41]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 1, 3
2017
-
[42]
Efficient video transformers with spatial- temporal token selection
Junke Wang, Xitong Yang, Hengduo Li, Zuxuan Wu, and Yu-Gang Jiang. Efficient video transformers with spatial- temporal token selection. In European Conference on Com- puter Vision, 2021. 3, 7
2021
-
[43]
Actionclip: Adapting language-image pretrained models for video action recognition
Mengmeng Wang, Jiazheng Xing, Jianbiao Mei, Yong Liu, and Yunliang Jiang. Actionclip: Adapting language-image pretrained models for video action recognition. IEEE Trans- actions on Neural Networks and Learning Systems, 2023. 3
2023
-
[44]
Vila: Efficient video-language alignment for video question answering
Xijun Wang, Junbang Liang, Chun-Kai Wang, Kenan Deng, Yu (Michael) Lou, Ming Lin, and Shan Yang. Vila: Efficient video-language alignment for video question answering. In ECCV, 2024. 3
2024
-
[45]
Video-focalnets: Spatio-temporal focal modu- lation for video action recognition
Syed Talal Wasim, Muhammad Uzair Khattak, Muzammal Naseer, Salman Khan, Mubarak Shah, and Fahad Shah- baz Khan. Video-focalnets: Spatio-temporal focal modu- lation for video action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages ...
2023
-
[46]
Manmatha, Alex Smola, and Philipp Kr¨ahenb¨uhl
Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R. Manmatha, Alex Smola, and Philipp Kr¨ahenb¨uhl. Compressed video ac- tion recognition. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6026–6035, 2017. 3
2018
-
[47]
Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognitio...
2022
-
[48]
Can i trust your answer? visually grounded video question answering
Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua. Can i trust your answer? visually grounded video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13204– 13214, 2024. 3
2024
-
[49]
Aim: Adapting image models for effi- cient video action recognition
Taojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang, Chen Chen, and Mu Li. Aim: Adapting image models for effi- cient video action recognition. In The Eleventh International Conference on Learning Representations, 2023. 3 10
2023
-
[50]
A-vit: Adap- tive tokens for efficient vision transformer
Hongxu Yin, Arash Vahdat, Jos ´e Manuel ´Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov. A-vit: Adap- tive tokens for efficient vision transformer. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10799–10808, 2021. 3
2022
-
[51]
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal. Self-chained image-language model for video localization and question answering. In NeurIPS, 2024. 3
2024
-
[52]
Pyramid feature attention net- work for saliency detection
Ting Zhao and Xiangqian Wu. Pyramid feature attention net- work for saliency detection. In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), 2019. 4
2019
-
[53]
How can objects help action recognition? 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2353–2362, 2023
Xingyi Zhou, Anurag Arnab, Chen Sun, and Cordelia Schmid. How can objects help action recognition? 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2353–2362, 2023. 7
2023
-
[54]
Eco: Efficient convolutional network for online video understanding
Mohammadreza Zolfaghari, Kamaljeet Singh, and Thomas Brox. Eco: Efficient convolutional network for online video understanding. ArXiv, abs/1804.09066, 2018. 3 11
2018 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.