REVIEW 4 major objections 4 minor 127 references
Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Uni-AdaFocus claims that video recognition can be made much cheaper by letting a learned policy crop one informative patch per frame, skip uninformative frames, and exit early on easy videos, all trained end-to-end.
desk verdict A serious, well-written AdaFocus unification with plausible training innovations, but the excerpt ends before any experiments, leaving the efficiency claim unverified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the policy network $\pi$, which outputs the patch center $(\tilde{x}^t_c, \tilde{y}^t_c)$ and, in the deformable variant, the patch height and width $(H^t_p, W^t_p)$ for each selected frame. It is trained through a differentiable interpolation-based crop: the selected patch is reconstructed from the four neighboring pixels (or from a patch of the global feature map) so that gradients can flow back to $\pi$ in both pixel space and deep-feature space. The same policy also generates frame weights for weighted sampling without replacement, giving the temporal and sample-wise dynamic computation a single common source of decisions.
What would settle it
Take a video where the action is defined by two far-apart objects, such as a person waving while another person enters from the opposite side. If Uni-AdaFocus with a single deformable patch loses accuracy relative to full-frame processing at matched compute, while a two-patch variant recovers that accuracy, the single-patch assumption is the bottleneck.
Extended reading notes
Core claim
Uni-AdaFocus unifies three kinds of dynamic computation inside one framework: spatial (which patch to attend to), temporal (which frames to process), and sample-wise (which videos need more compute). The central mechanism is a lightweight global encoder that produces a cheap overview of each frame, a policy network $\pi$ that outputs a patch center and, in the deformable version, a patch height and width, and a high-capacity local encoder that only sees the resized patch. The patch selection is made trainable end-to-end by replacing the hard crop with differentiable interpolation, so gradients flow from the recognition loss back into the policy network. Temporal frame selection is formulated as weighted sampling without replacement and optimized with a Monte Carlo estimate of the expected loss, avoiding reinforcement learning and multi-stage training. The authors claim this combination is considerably more efficient than competitive baselines while remaining hardware-friendly because all selected patches are resized to a common shape and processed in parallel.
Load-bearing premise
The central premise is that the most informative content in each frame is captured by a single rectangular patch whose location changes smoothly, so a policy can find it; if task-relevant cues are spread over multiple disjoint regions, cropping to one patch discards important information.
Editorial extensions
If this is right
- Uni-AdaFocus can wrap off-the-shelf backbones like TSM and X3D, so the efficiency gains could transfer to existing video models without redesigning their core architecture.
- The end-to-end training recipe removes the need for reinforcement learning and multi-stage procedures, making dynamic spatial cropping practical for ordinary training pipelines.
- Spatial, temporal, and sample-wise savings compose inside one framework, so the computational reductions multiply rather than merely add.
- The model's inference cost can be adjusted online by changing the exit criterion, which suits applications with fluctuating compute budgets such as mobile video search.
- The selected patches and sampled frames concentrate computation on the task-relevant content, which is claimed to preserve accuracy at a fraction of the full-frame, full-video cost.
Reading between the lines
- The differentiable cropping idea is not inherently video-specific: the same policy-plus-interpolation recipe could be applied to high-resolution images, medical volumes, or multi-view data, although the paper does not test those settings.
- The deformable patch mechanism behaves like a learned multi-scale attention, so one could diagnose whether the patch size tracks object scale or motion magnitude; that diagnostic is not reported in the paper.
- The sample-wise exit mechanism suggests a natural accuracy-compute calibration curve, which a deployment system could use to choose a per-query exit threshold; the paper mentions online adjustment but does not develop a scheduling policy around it.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Uni-AdaFocus, a video-recognition framework designed to combine spatial-, temporal-, and sample-wise dynamic computation. It formalizes an AdaFocus architecture with a lightweight global encoder, a policy network that localizes a single task-relevant patch per frame, a high-capacity local encoder, and a classifier. It then develops an end-to-end training scheme via interpolation-based patch selection, a deep-feature-based policy supervision method, a deformable-patch extension, and a unified treatment of dynamic frame sampling and conditional early exit. The authors claim substantial computational-efficiency gains on seven benchmark datasets and three application scenarios, but the provided manuscript text does not include the corresponding experimental tables or ablations.
Significance. If the empirical claims hold, Uni-AdaFocus would be a useful consolidation of several dynamic-computation ideas: the derivations in Sections 3.3 and 4.1 are detailed, the training recipe is concrete and reproducible in spirit, and the compatibility with off-the-shelf backbones such as TSM and X3D is a practical strength. The paper is also self-aware about one limitation, fixed-size patches, and proposes a deformable remedy. However, the central efficiency-accuracy claim rests on the single-rectangle-per-frame assumption and on experimental evidence that is not present in the provided text; these gaps are load-bearing rather than cosmetic.
major comments (4)
- [Sections 3.1 and 4.1.2] The system processes exactly one rectangular patch per selected frame; the deformable variant still outputs one (H_t^p, W_t^p) rectangle resized to P×P. This is load-bearing for the efficiency claim: if task-relevant information is spread over multiple disjoint regions, a single rectangle either discards important content or expands to include irrelevant pixels, degrading accuracy or erasing the efficiency gain. The manuscript acknowledges that fixed-size patches are suboptimal, but it does not analyze multi-modal spatial distributions or provide experiments demonstrating that one rectangle per frame preserves accuracy across the seven claimed datasets. Please add a matched-FLOPs comparison against multi-patch or full-frame baselines and include qualitative patch-coverage analysis on representative benchmark frames.
- [Section 4.1.1, Eqs. (13)-(14)] The deep-feature interpolation method treats feature-space cropping as a faithful proxy for pixel-space cropping, justified by the 'location-preserving nature' of deep features. The provided text gives no analysis of the approximation error due to the spatial downsampling of e_G^t (size H_G × W_G) or to the receptive field of the global encoder, and it does not include an ablation comparing the policy trained with Eq. (14) against the pixel-space-gradient baseline of AdaFocusV2. Because this is the main claimed novelty over AdaFocusV2, please provide such an ablation and, ideally, a quantitative error measurement or bound for the location-preservation assumption.
- [Abstract and Section 1] The paper's headline claim — 'considerably more efficient than the competitive baselines' — is presented as established, but the provided manuscript text contains no experimental results, ablations, or comparison tables. The only evidence cited is internal references to Tables 6, 7, 15, and 16, none of which appear in the excerpt. Without these data, the central claim cannot be assessed. Please ensure the submitted version includes the full experimental section and that every abstract-level claim is directly supported by the included tables and figures.
- [Section 4.2, Eq. (16)] The dynamic frame sampling formulation is stated as weighted sampling without replacement, and the text promises a differentiable Monte Carlo objective obtained by decomposing the expected loss, but the derivation is not present in the provided text. The correctness and training stability of the temporal selection mechanism depend on this derivation, including the gradient estimator and its bias/variance behavior. Please present the full derivation in the main text or in the appendix and state explicitly which parts of the objective are differentiable and which use a straight-through or score-function estimator.
minor comments (4)
- [Section 3.3.2] The paragraph claims that the proposed training techniques 'do not introduce additional tunable hyper-parameters,' but Section 4.1.2 introduces alpha in Eq. (15). Please clarify that the no-new-hyperparameter claim applies only to the three techniques in Section 3.3.2, and state how alpha is set.
- [Eqs. (1)-(3)] The notation e_G^t and e_L^t is used for both feature maps and pooled feature vectors. Please use distinct symbols (for example, bold for vectors) to avoid ambiguity in the classifier input in Eq. (3).
- [Section 4.1.1] AdaFocusV3 is referenced as a prior work, but it has not been introduced earlier in this manuscript. Please give a one-sentence definition or clearly point to reference [30] so the comparison in Eq. (14) is self-contained.
- [Figure 6 caption] The caption states that deformable patches yield 'significant accuracy improvements across diverse scenarios,' but no numerical evidence is shown in the provided text. Please cite the corresponding table when discussing this figure.
Circularity Check
No significant circularity: the paper's claims rest on explicit new formulations and external benchmarks, with self-citations serving as transparent lineage rather than load-bearing derivations.
full rationale
The derivation chain is not circular in the sense required by the review. The paper explicitly frames itself as an extension: "This paper extends previous conference papers that introduced the basic AdaFocus framework [28] and preliminarily discussed its end-to-end training [29]. Moreover, our deep-feature-based approach for training the patch selection policy (Section 4.1.1) is conceptually relevant to, and improved upon [30]." These self-citations are lineage statements, not load-bearing evidence: the end-to-end training mechanism is derived from the interpolation equations (6)-(9), the deep-feature policy loss is defined in Eqs. (13)-(14) and then modified in Eq. (15), and the temporal and sample-wise components are introduced as new formulations in Section 4.2 onward. The "small-patch" premise is an assumption about video content ("the most informative region in each video frame usually corresponds to a small image patch"), not a conclusion obtained from the method; any resulting accuracy-efficiency tradeoff under distributed task-relevant content is an empirical robustness concern, not circularity. The paper reports comparisons on "seven widely-used benchmark datasets" and competitive baselines, so the efficiency claim is externally checkable rather than forced by a fitted parameter renamed as a prediction. No equation in the excerpt equates a predicted quantity with a fitted input by construction, and no uniqueness or forbidden-alternative argument is imported from the authors' prior work. Therefore the paper receives a non-circular finding.
Assumptions & free parameters
free parameters (3)
- alpha (patch size regularization coefficient) =
unknown, pre-defined per dataset
- Patch size P =
not specified
- Number of frames TG and TL =
not specified
assumptions (3)
- domain assumption Informative content in each video frame is concentrated in a small patch whose shape, size, and location vary smoothly across frames.
- ad hoc to paper Deep feature maps preserve the spatial location of input contents, so that interpolation in feature space approximates interpolation in pixel space.
- domain assumption Weighted sampling without replacement with a differentiable expected loss can be trained end-to-end via Monte Carlo estimation.
Cite this review
Pith. "Pith review of Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition." pith.science (2026). https://pith.science/paper/JAS7KWHE
@misc{pith2026241211228,
author = {Pith},
title = {Pith review of: Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/JAS7KWHE}},
note = {Machine review of arXiv:2412.11228}
}
read the original abstract
This paper presents a comprehensive exploration of the phenomenon of data redundancy in video understanding, with the aim to improve computational efficiency. Our investigation commences with an examination of spatial redundancy, which refers to the observation that the most informative region in each video frame usually corresponds to a small image patch, whose shape, size and location shift smoothly across frames. Motivated by this phenomenon, we formulate the patch localization problem as a dynamic decision task, and introduce a spatially adaptive video recognition approach, termed AdaFocus. In specific, a lightweight encoder is first employed to quickly process the full video sequence, whose features are then utilized by a policy network to identify the most task-relevant regions. Subsequently, the selected patches are inferred by a high-capacity deep network for the final prediction. The full model can be trained in end-to-end conveniently. Furthermore, AdaFocus can be extended by further considering temporal and sample-wise redundancies, i.e., allocating the majority of computation to the most task-relevant frames, and minimizing the computation spent on relatively "easier" videos. Our resulting approach, Uni-AdaFocus, establishes a comprehensive framework that seamlessly integrates spatial, temporal, and sample-wise dynamic computation, while it preserves the merits of AdaFocus in terms of efficient end-to-end training and hardware friendliness. In addition, Uni-AdaFocus is general and flexible as it is compatible with off-the-shelf efficient backbones (e.g., TSM and X3D), which can be readily deployed as our feature extractor, yielding a significantly improved computational efficiency. Empirically, extensive experiments based on seven benchmark datasets and three application scenarios substantiate that Uni-AdaFocus is considerably more efficient than the competitive baselines.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
The youtube video recommendation system,
J. Davidson, B. Liebald, J. Liu, P . Nandy, T. Van Vleet, U. Gargi, S. Gupta, Y. He, M. Lambert, B. Livingston et al., “The youtube video recommendation system,” in Proceedings of the fourth ACM conference on Recommender systems, 2010, pp. 293–296
2010
-
[2]
Content-based video recommendation system based on stylistic visual features,
Y. Deldjoo, M. Elahi, P . Cremonesi, F. Garzotto, P . Piazzolla, and M. Quadrana, “Content-based video recommendation system based on stylistic visual features,” Journal on Data Semantics, vol. 5, no. 2, pp. 99–113, 2016
2016
-
[3]
A unified personalized video recommendation via dynamic recurrent neural networks,
J. Gao, T. Zhang, and C. Xu, “A unified personalized video recommendation via dynamic recurrent neural networks,” in ACM MM, 2017, pp. 127–135
2017
-
[4]
A system for video surveillance and monitoring,
R. T. Collins, A. J. Lipton, T. Kanade, H. Fujiyoshi, D. Duggins, Y. Tsin, D. Tolliver, N. Enomoto, O. Hasegawa, P . Burtet al., “A system for video surveillance and monitoring,” VSAM final report, vol. 2000, no. 1-68, p. 1, 2000
2000
-
[5]
Distributed deep learning model for intelligent video surveillance systems with edge computing,
J. Chen, K. Li, Q. Deng, K. Li, and S. Y. Philip, “Distributed deep learning model for intelligent video surveillance systems with edge computing,” IEEE Transactions on Industrial Informatics, 2019
2019
-
[6]
Searching video for complex activities with finite state models,
N. Ikizler and D. Forsyth, “Searching video for complex activities with finite state models,” in CVPR, 2007, pp. 1–8
2007
-
[7]
Slowfast networks for video recognition,
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in ICCV, 2019, pp. 6202–6211
2019
-
[8]
Deep feature flow for video recognition,
X. Zhu, Y. Xiong, J. Dai, L. Yuan, and Y. Wei, “Deep feature flow for video recognition,” in CVPR, 2017, pp. 2349–2358
2017
Show all 127 references
-
[9]
Convolutional two- stream network fusion for video action recognition,
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two- stream network fusion for video action recognition,” in CVPR, 2016, pp. 1933–1941
2016
-
[10]
Quo vadis, action recognition? a new model and the kinetics dataset,
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in CVPR, 2017, pp. 6299– 6308
2017
-
[11]
Learn- ing spatiotemporal features with 3d convolutional networks,
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learn- ing spatiotemporal features with 3d convolutional networks,” in ICCV, 2015, pp. 4489–4497
2015
-
[12]
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
K. Hara, H. Kataoka, and Y. Satoh, “Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?” in CVPR, 2018, pp. 6546–6555
2018
-
[13]
Liteeval: A coarse- to-fine framework for resource efficient video recognition,
Z. Wu, C. Xiong, Y.-G. Jiang, and L. S. Davis, “Liteeval: A coarse- to-fine framework for resource efficient video recognition,” in NeurIPS, 2019b
-
[14]
Multi-agent rein- forcement learning based frame sampling for effective untrimmed video recognition,
W. Wu, D. He, X. Tan, S. Chen, and S. Wen, “Multi-agent rein- forcement learning based frame sampling for effective untrimmed video recognition,” in ICCV, 2019a, pp. 6222–6231
-
[15]
A dynamic frame selection framework for fast video recognition,
Z. Wu, H. Li, C. Xiong, Y.-G. Jiang, and L. S. Davis, “A dynamic frame selection framework for fast video recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020b
-
[16]
Scsampler: Sampling salient clips from video for efficient action recognition,
B. Korbar, D. Tran, and L. Torresani, “Scsampler: Sampling salient clips from video for efficient action recognition,” in ICCV, 2019, pp. 6232–6242
2019
-
[17]
Listen to look: Action recognition by previewing audio,
R. Gao, T.-H. Oh, K. Grauman, and L. Torresani, “Listen to look: Action recognition by previewing audio,” in CVPR, 2020, pp. 10 457–10 467
2020
-
[18]
Ar-net: Adaptive frame resolution for efficient action recognition,
Y. Meng, C.-C. Lin, R. Panda, P . Sattigeri, L. Karlinsky, A. Oliva, K. Saenko, and R. Feris, “Ar-net: Adaptive frame resolution for efficient action recognition,” in ECCV, 2020, pp. 86–104
2020
-
[19]
Ocsampler: Compressing videos to one clip with single-step sampling,
J. Lin, H. Duan, K. Chen, D. Lin, and L. Wang, “Ocsampler: Compressing videos to one clip with single-step sampling,” in CVPR, 2022
2022
-
[20]
Recurrent models of visual attention,
V . Mnih, N. Heess, A. Graveset al., “Recurrent models of visual attention,” in NeurIPS, 2014, pp. 2204–2212
2014
-
[21]
Look closer to see better: Recurrent attention convolutional neural network for fine-grained image recognition,
J. Fu, H. Zheng, and T. Mei, “Look closer to see better: Recurrent attention convolutional neural network for fine-grained image recognition,” in CVPR, 2017, pp. 4438–4446
2017
-
[22]
Spatially adaptive inference with stochastic feature sampling and interpolation,
Z. Xie, Z. Zhang, X. Zhu, G. Huang, and S. Lin, “Spatially adaptive inference with stochastic feature sampling and interpolation,” in ECCV, 2020, pp. 531–548
2020
-
[23]
Dynamic neural networks: A survey,
Y. Han, G. Huang, S. Song, L. Yang, H. Wang, and Y. Wang, “Dynamic neural networks: A survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 11, pp. 7436–7456, 2022
2022
-
[24]
Glance and focus networks for dynamic visual recognition,
G. Huang, Y. Wang, K. Lv, H. Jiang, W. Huang, P . Qi, and S. Song, “Glance and focus networks for dynamic visual recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 4, pp. 4605–4621, 2023
2023
-
[25]
Dynamic spatial sparsification for efficient vision transformers and convolutional neural networks,
Y. Rao, Z. Liu, W. Zhao, J. Zhou, and J. Lu, “Dynamic spatial sparsification for efficient vision transformers and convolutional neural networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–14, 2023
2023
-
[26]
Tsm: Temporal shift module for efficient video understanding,
J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” in ICCV, 2019, pp. 7083–7093
2019
-
[27]
X3d: Expanding architectures for efficient video recognition,
C. Feichtenhofer, “X3d: Expanding architectures for efficient video recognition,” in CVPR, 2020, pp. 203–213
2020
-
[28]
Adaptive focus for efficient video recognition,
Y. Wang, Z. Chen, H. Jiang, S. Song, Y. Han, and G. Huang, “Adaptive focus for efficient video recognition,” in ICCV, 2021
2021
-
[29]
Adafocus v2: End-to-end training of spatial dynamic networks for video recognition,
Y. Wang, Y. Yue, Y. Lin, H. Jiang, Z. Lai, V . Kulikov, N. Orlov, H. Shi, and G. Huang, “Adafocus v2: End-to-end training of spatial dynamic networks for video recognition,” in CVPR, 2022
2022
-
[30]
Adafocusv3: On unified spatial-temporal dynamic video recognition,
Y. Wang, Y. Yue, X. Xu, A. Hassani, V . Kulikov, N. Orlov, S. Song, H. Shi, and G. Huang, “Adafocusv3: On unified spatial-temporal dynamic video recognition,” in ECCV, 2022, pp. 226–243
2022
-
[31]
Activitynet: A large-scale video benchmark for human activity understanding,
F. Caba Heilbron, V . Escorcia, B. Ghanem, and J. Carlos Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in CVPR, 2015, pp. 961–970
2015
-
[32]
Exploiting feature and class relationships in video categorization with regularized deep neural networks,
Y.-G. Jiang, Z. Wu, J. Wang, X. Xue, and S.-F. Chang, “Exploiting feature and class relationships in video categorization with regularized deep neural networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 2, pp. 352–364, 2018
2018
-
[33]
The kinetics human action video dataset,
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijaya- narasimhan, F. Viola, T. Green, T. Back, P . Natsev et al. , “The kinetics human action video dataset,” arXiv:1705.06950, 2017
2017 arXiv
-
[34]
The "something something
R. Goyal, S. Ebrahimi Kahou, V . Michalski, J. Materzynska, S. Westphal, H. Kim, V . Haenel, I. Fruend, P . Yianilos, M. Mueller- Freitag et al. , “The "something something" video database for learning and evaluating visual common sense,” in ICCV, 2017, pp. 5842–5850
2017
-
[35]
The jester dataset: A large-scale video dataset of human gestures,
J. Materzynska, G. Berger, I. Bax, and R. Memisevic, “The jester dataset: A large-scale video dataset of human gestures,” inICCVW, 2019
2019
-
[36]
Temporal segment networks: Towards good prac- tices for deep action recognition,
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks: Towards good prac- tices for deep action recognition,” in ECCV, 2016, pp. 20–36
2016
-
[37]
Long-term recurrent convolutional networks for visual recognition and description,
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR, 2015, pp. 2625–2634
2015
-
[38]
Recurrent tubelet proposal and recognition networks for action detection,
D. Li, Z. Qiu, Q. Dai, T. Yao, and T. Mei, “Recurrent tubelet proposal and recognition networks for action detection,” in ECCV, 2018, pp. 303–318
2018
-
[39]
Beyond short snippets: Deep net- works for video classification,
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici, “Beyond short snippets: Deep net- works for video classification,” in CVPR, 2015, pp. 4694–4702
2015
-
[40]
Gate-shift networks for video action recognition,
S. Sudhakaran, S. Escalera, and O. Lanz, “Gate-shift networks for video action recognition,” in CVPR, 2020, pp. 1102–1111
2020
-
[41]
Adafuse: Adaptive temporal fusion network for efficient action recognition,
Y. Meng, R. Panda, C.-C. Lin, P . Sattigeri, L. Karlinsky, K. Saenko, A. Oliva, and R. Feris, “Adafuse: Adaptive temporal fusion network for efficient action recognition,” in ICLR, 2021
2021
-
[42]
Spatiotemporal multiplier networks for video action recognition,
C. Feichtenhofer, A. Pinz, and R. P . Wildes, “Spatiotemporal multiplier networks for video action recognition,” in CVPR, 2017, pp. 4768–4777
2017
-
[43]
Searching for two-stream models in multivariate space for video recognition,
X. Gong, H. Wang, M. Z. Shou, M. Feiszli, Z. Wang, and Z. Yan, “Searching for two-stream models in multivariate space for video recognition,” in ICCV, 2021, pp. 8033–8042
2021
-
[44]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR, 2021
2021
-
[45]
Vivit: A video vision transformer,
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Luˇ ci´ c, and C. Schmid, “Vivit: A video vision transformer,” in ICCV, 2021
2021
-
[46]
Video swin transformer,
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in CVPR, 2022, pp. 3202–3211
2022
-
[47]
Is space-time attention all you need for video understanding?
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in ICML, 2021
2021
-
[48]
Video transformer network,
D. Neimark, O. Bar, M. Zohar, and D. Asselmann, “Video transformer network,” in ICCV, 2021, pp. 3163–3172
2021
-
[49]
A closer look at spatiotemporal convolutions for action recognition,
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in CVPR, 2018, pp. 6450–6459. UNI-ADAFOCUS: SPATIAL-TEMPORAL DYNAMIC COMPUTATION FOR VIDEO RECOGNITION 17
2018
-
[50]
Eco: Efficient convolutional network for online video understanding,
M. Zolfaghari, K. Singh, and T. Brox, “Eco: Efficient convolutional network for online video understanding,” in ECCV, 2018, pp. 695–712
2018
-
[51]
Video classification with channel-separated convolutional networks,
D. Tran, H. Wang, L. Torresani, and M. Feiszli, “Video classification with channel-separated convolutional networks,” in ICCV, 2019, pp. 5552–5561
2019
-
[52]
Teinet: Towards an efficient architecture for video recognition,
Z. Liu, D. Luo, Y. Wang, L. Wang, Y. Tai, C. Wang, J. Li, F. Huang, and T. Lu, “Teinet: Towards an efficient architecture for video recognition,” in AAAI, vol. 34, no. 07, 2020, pp. 11 669–11 676
2020
-
[53]
Tam: Temporal adaptive module for video recognition,
Z. Liu, L. Wang, W. Wu, C. Qian, and T. Lu, “Tam: Temporal adaptive module for video recognition,” in ICCV, 2021, pp. 13 708– 13 718
2021
-
[54]
End-to-end learning of action detection from frame glimpses in videos,
S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei, “End-to-end learning of action detection from frame glimpses in videos,” in CVPR, 2016, pp. 2678–2687
2016
-
[55]
Frameexit: Condi- tional early exiting for efficient video recognition,
A. Ghodrati, B. E. Bejnordi, and A. Habibian, “Frameexit: Condi- tional early exiting for efficient video recognition,” in CVPR, 2021, pp. 15 608–15 618
2021
-
[56]
Efficient action recognition via dynamic knowledge propagation,
H. Kim, M. Jain, J.-T. Lee, S. Yun, and F. Porikli, “Efficient action recognition via dynamic knowledge propagation,” in ICCV, 2021, pp. 13 719–13 728
2021
-
[57]
Dynamic network quantization for efficient video inference,
X. Sun, R. Panda, C.-F. R. Chen, A. Oliva, R. Feris, and K. Saenko, “Dynamic network quantization for efficient video inference,” in ICCV, 2021, pp. 7375–7385
2021
-
[58]
Nsnet: Non-saliency suppression sampler for effi- cient video recognition,
B. Xia, W. Wu, H. Wang, R. Su, D. He, H. Yang, X. Fan, and W. Ouyang, “Nsnet: Non-saliency suppression sampler for effi- cient video recognition,” in ECCV, 2022, pp. 705–723
2022
-
[59]
Temporal saliency query network for efficient video recognition,
B. Xia, Z. Wang, W. Wu, H. Wang, and J. Han, “Temporal saliency query network for efficient video recognition,” in ECCV, 2022, pp. 741–759
2022
-
[60]
Spatial trans- former networks,
M. Jaderberg, K. Simonyan, A. Zisserman et al., “Spatial trans- former networks,” in NeurIPS, 2015, pp. 2017–2025
2015
-
[61]
Sbnet: Sparse blocks network for fast inference,
M. Ren, A. Pokrovsky, B. Yang, and R. Urtasun, “Sbnet: Sparse blocks network for fast inference,” in CVPR, 2018, pp. 8711–8720
2018
-
[62]
Resolution adaptive networks for efficient inference,
L. Yang, Y. Han, X. Chen, S. Song, J. Dai, and G. Huang, “Resolution adaptive networks for efficient inference,” in CVPR, 2020, pp. 2369–2378
2020
-
[63]
Adaptively connected neural networks,
G. Wang, K. Wang, and L. Lin, “Adaptively connected neural networks,” in CVPR, 2019, pp. 1781–1790
2019
-
[64]
Dynamic region- aware convolution,
J. Chen, X. Wang, Z. Guo, X. Zhang, and J. Sun, “Dynamic region- aware convolution,” in CVPR, 2021, pp. 8064–8073
2021
-
[65]
Spatially adaptive computation time for residual networks,
M. Figurnov, M. D. Collins, Y. Zhu, L. Zhang, J. Huang, D. Vetrov, and R. Salakhutdinov, “Spatially adaptive computation time for residual networks,” in CVPR, 2017, pp. 1039–1048
2017
-
[66]
Dynamic convolutions: Exploiting spatial sparsity for faster inference,
T. Verelst and T. Tuytelaars, “Dynamic convolutions: Exploiting spatial sparsity for faster inference,” in CVPR, 2020, pp. 2320–2329
2020
-
[67]
Interpretable spatio-temporal attention for video action recognition,
L. Meng, B. Zhao, B. Chang, G. Huang, W. Sun, F. Tung, and L. Sigal, “Interpretable spatio-temporal attention for video action recognition,” in ICCVW, 2019
2019
-
[68]
Generalized domain conditioned adaptation network,
S. Li, B. Xie, Q. Lin, C. H. Liu, G. Huang, and G. Wang, “Generalized domain conditioned adaptation network,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 8, pp. 4093–4109, 2021
2021
-
[69]
Learning deep features for discriminative localization,
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in CVPR, 2016, pp. 2921–2929
2016
-
[70]
Grad-cam: Visual explanations from deep networks via gradient-based localization,
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in ICCV, 2017, pp. 618–626
2017
-
[71]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR, 2021
2021
-
[72]
Delving into details: Synopsis-to-detail networks for video recognition,
S. Liang, X. Shen, J. Huang, and X.-S. Hua, “Delving into details: Synopsis-to-detail networks for video recognition,” in ECCV, 2022, pp. 262–278
2022
-
[73]
Long short-term memory,
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997
1997
-
[74]
Learning phrase rep- resentations using RNN encoder–decoder for statistical machine translation,
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase rep- resentations using RNN encoder–decoder for statistical machine translation,” in EMNLP, 2014, pp. 1724–1734
2014
-
[75]
Playing atari with deep reinforce- ment learning,
V . Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing atari with deep reinforce- ment learning,” arXiv:1312.5602, 2013
2013 arXiv
-
[76]
Proximal policy optimization algorithms,
J. Schulman, F. Wolski, P . Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv:1707.06347, 2017
2017 arXiv
-
[77]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016, pp. 770–778
2016
-
[78]
Convolutional networks with dense connectivity,
G. Huang, Z. Liu, G. Pleiss, L. Van Der Maaten, and K. Wein- berger, “Convolutional networks with dense connectivity,” IEEE transactions on pattern analysis and machine intelligence , 2019
2019
-
[79]
Better mixing via deep representations,
Y. Bengio, G. Mesnil, Y. Dauphin, and S. Rifai, “Better mixing via deep representations,” in ICML, 2013, pp. 552–560
2013
-
[80]
Implicit semantic data augmentation for deep networks,
Y. Wang, X. Pan, S. Song, H. Zhang, G. Huang, and C. Wu, “Implicit semantic data augmentation for deep networks,” in NeurIPS, 2019
2019
-
[81]
Deep residual correction network for partial domain adaptation,
S. Li, C. H. Liu, Q. Lin, Q. Wen, L. Su, G. Huang, and Z. Ding, “Deep residual correction network for partial domain adaptation,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 7, pp. 2329–2344, 2020
2020
-
[82]
Sepico: Semantic-guided pixel contrast for domain adaptive semantic segmentation,
B. Xie, S. Li, M. Li, C. H. Liu, G. Huang, and G. Wang, “Sepico: Semantic-guided pixel contrast for domain adaptive semantic segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 7, pp. 9004–9021, 2023
2023
-
[83]
Adapting across domains via target-oriented transferable semantic augmentation under prototype constraint,
M. Xie, S. Li, K. Gong, Y. Wang, and G. Huang, “Adapting across domains via target-oriented transferable semantic augmentation under prototype constraint,” International Journal of Computer Vision, vol. 132, no. 4, pp. 1417–1441, 2024
2024
-
[84]
Multi-scale dense networks for resource efficient image classification,
G. Huang, D. Chen, T. Li, F. Wu, L. van der Maaten, and K. Q. Weinberger, “Multi-scale dense networks for resource efficient image classification,” in ICLR, 2018
2018
-
[85]
Glance and focus: a dynamic approach to reducing spatial redundancy in image classification,
Y. Wang, K. Lv, R. Huang, S. Song, L. Yang, and G. Huang, “Glance and focus: a dynamic approach to reducing spatial redundancy in image classification,” in NeurIPS, 2020
2020
-
[86]
Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition,
Y. Wang, R. Huang, S. Song, Z. Huang, and G. Huang, “Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition,” in NeurIPS, 2021
2021
-
[87]
Watching a small portion could be as good as watching all: Towards efficient video classification,
H. Fan, Z. Xu, L. Zhu, C. Yan, J. Ge, and Y. Yang, “Watching a small portion could be as good as watching all: Towards efficient video classification,” in IJCAI, 2018
2018
-
[88]
Adamml: Adaptive multi-modal learning for efficient video recognition,
R. Panda, C.-F. R. Chen, Q. Fan, X. Sun, K. Saenko, A. Oliva, and R. Feris, “Adamml: Adaptive multi-modal learning for efficient video recognition,” in ICCV, 2021, pp. 7576–7585
2021
-
[89]
Smart frame selection for action recognition,
S. N. Gowda, M. Rohrbach, and L. Sevilla-Lara, “Smart frame selection for action recognition,” in AAAI, vol. 35, no. 2, 2021, pp. 1451–1459
2021
-
[90]
Look more but care less in video recognition,
Y. Zhang, Y. Bai, H. Wang, Y. Xu, and Y. Fu, “Look more but care less in video recognition,” in NeurIPS, 2022
2022
-
[91]
Resound: Towards action recognition without representation bias,
Y. Li, Y. Li, and N. Vasconcelos, “Resound: Towards action recognition without representation bias,” in ECCV, 2018, pp. 513– 528
2018
-
[92]
https://adni.loni.usc.edu/
-
[93]
https://www.oasis-brains.org/
-
[94]
https://www.ppmi-info.org/
-
[95]
Violence recognition from videos using deep learning techniques,
M. M. Soliman, M. H. Kamal, M. A. E.-M. Nashed, Y. M. Mostafa, B. S. Chawky, and D. Khattab, “Violence recognition from videos using deep learning techniques,” in 2019 Ninth International Conference on Intelligent Computing and Information Systems (ICICIS), 2019, pp. 80–85
2019
-
[96]
Mobilenetv2: Inverted residuals and linear bottlenecks,
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in CVPR, 2018, pp. 4510–4520
2018
-
[97]
Videos as space-time region graphs,
X. Wang and A. Gupta, “Videos as space-time region graphs,” in ECCV, 2018, pp. 399–417
2018
-
[98]
Temporal relational reasoning in videos,
B. Zhou, A. Andonian, A. Oliva, and A. Torralba, “Temporal relational reasoning in videos,” in ECCV, 2018, pp. 803–818
2018
-
[99]
Smallbignet: Integrating core and contextual views for video classification,
X. Li, Y. Wang, Z. Zhou, and Y. Qiao, “Smallbignet: Integrating core and contextual views for video classification,” in CVPR, 2020, pp. 1092–1101
2020
-
[100]
More is less: Learning efficient video representations by big-little network and depthwise temporal aggregation,
Q. Fan, C.-F. R. Chen, H. Kuehne, M. Pistoia, and D. Cox, “More is less: Learning efficient video representations by big-little network and depthwise temporal aggregation,” in NeurIPS, 2019
2019
-
[101]
Stm: Spatiotempo- ral and motion encoding for action recognition,
B. Jiang, M. Wang, W. Gan, W. Wu, and J. Yan, “Stm: Spatiotempo- ral and motion encoding for action recognition,” in ICCV, 2019, pp. 2000–2009
2019
-
[102]
Tea: Temporal excitation and aggregation for action recognition,
Y. Li, B. Ji, X. Shi, J. Zhang, B. Kang, and L. Wang, “Tea: Temporal excitation and aggregation for action recognition,” in CVPR, 2020, pp. 909–918
2020
-
[103]
Appearance-and-relation networks for video classification,
L. Wang, W. Li, W. Li, and L. Van Gool, “Appearance-and-relation networks for video classification,” in CVPR, 2018, pp. 1430–1439
2018
-
[104]
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,” in ECCV, 2018, pp. 305–321. UNI-ADAFOCUS: SPATIAL-TEMPORAL DYNAMIC COMPUTATION FOR VIDEO RECOGNITION 18
2018
-
[105]
Deep analysis of cnn-based spatio-temporal representations for action recognition,
C.-F. R. Chen, R. Panda, K. Ramakrishnan, R. Feris, J. Cohn, A. Oliva, and Q. Fan, “Deep analysis of cnn-based spatio-temporal representations for action recognition,” in CVPR, 2021, pp. 6165– 6175
2021
-
[106]
Non-local neural networks,
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR, 2018, pp. 7794–7803
2018
-
[107]
Tdn: Temporal difference networks for efficient action recognition,
L. Wang, Z. Tong, B. Ji, and G. Wu, “Tdn: Temporal difference networks for efficient action recognition,” in CVPR, 2021, pp. 1895–1904
2021
-
[108]
Efficientnet: Rethinking model scaling for convolutional neural networks,
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in ICML, 2019, pp. 6105–6114
2019
-
[109]
Movinets: Mobile video networks for efficient video recognition,
D. Kondratyuk, L. Yuan, Y. Li, L. Zhang, M. Tan, M. Brown, and B. Gong, “Movinets: Mobile video networks for efficient video recognition,” in CVPR, 2021, pp. 16 020–16 030
2021
-
[110]
Mgsampler: An explainable sampling strategy for video action recognition,
Y. Zhi, Z. Tong, L. Wang, and G. Wu, “Mgsampler: An explainable sampling strategy for video action recognition,” in ICCV, 2021, pp. 1513–1522
2021
-
[111]
Alzheimer’s disease,
P . Scheltens, B. De Strooper, M. Kivipelto, H. Holstege, G. Chéte- lat, C. E. Teunissen, J. Cummings, and W. M. van der Flier, “Alzheimer’s disease,” The Lancet, vol. 397, no. 10284, pp. 1577– 1590, 2021
2021
-
[112]
Parkinson’s disease,
B. R. Bloem, M. S. Okun, and C. Klein, “Parkinson’s disease,” The Lancet, vol. 397, no. 10291, pp. 2284–2303, 2021
2021
-
[113]
Parkinson’s disease: mechanisms and models,
W. Dauer and S. Przedborski, “Parkinson’s disease: mechanisms and models,” Neuron, vol. 39, no. 6, pp. 889–909, 2003
2003
-
[114]
Relational self- attention: What’s missing in attention for video understanding,
M. Kim, H. Kwon, C. Wang, S. Kwak, and M. Cho, “Relational self- attention: What’s missing in attention for video understanding,” in NeurIPS, 2021, pp. 8046–8059
2021
-
[115]
Hierarchical fully con- volutional network for joint atrophy localization and alzheimer’s disease diagnosis using structural mri,
C. Lian, M. Liu, J. Zhang, and D. Shen, “Hierarchical fully con- volutional network for joint atrophy localization and alzheimer’s disease diagnosis using structural mri,”IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 4, pp. 880–893, 2018
2018
-
[116]
End-to-end automatic pathology localization for alzheimer’s disease diagnosis using structural mri,
G. Cao, M. Zhang, Y. Wang, J. Zhang, Y. Han, X. Xu, J. Huang, and G. Kang, “End-to-end automatic pathology localization for alzheimer’s disease diagnosis using structural mri,” Computers in Biology and Medicine, p. 107110, 2023
2023
-
[117]
Deep neural networks with broad views for parkinson’s disease screening,
X. Zhang, Y. Yang, H. Wang, S. Ning, and H. Wang, “Deep neural networks with broad views for parkinson’s disease screening,” in 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2019, pp. 1018–1022
2019
-
[118]
Med3d: Transfer learning for 3d medical image analysis,
S. Chen, K. Ma, and Y. Zheng, “Med3d: Transfer learning for 3d medical image analysis,” arXiv:1904.00625, 2019
1904 arXiv
-
[119]
M3t: three-dimensional medical image classifier using multi-plane and multi-slice transformer,
J. Jang and D. Hwang, “M3t: three-dimensional medical image classifier using multi-plane and multi-slice transformer,” in CVPR, 2022, pp. 20 718–20 729
2022
-
[120]
Learn-explain-reinforce: Coun- terfactual reasoning and its guidance to reinforce an alzheimer’s disease diagnosis model,
K. Oh, J. S. Yoon, and H.-I. Suk, “Learn-explain-reinforce: Coun- terfactual reasoning and its guidance to reinforce an alzheimer’s disease diagnosis model,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 4, pp. 4843–4857, 2023
2023
-
[121]
2d bidirectional gated recurrent unit convolutional neural networks for end-to-end violence detec- tion in videos,
A. Traoré and M. A. Akhloufi, “2d bidirectional gated recurrent unit convolutional neural networks for end-to-end violence detec- tion in videos,” in Image Analysis and Recognition: 17th International Conference, ICIAR 2020, Póvoa de Varzim, Portugal, June 24–26, 2020, Proceed...
2020
-
[122]
A temporal fusion approach for video classification with convolutional and lstm neural networks applied to violence detection,
J. P . de Oliveira Lima and C. M. S. Figueiredo, “A temporal fusion approach for video classification with convolutional and lstm neural networks applied to violence detection,” Inteligencia Artificial, vol. 24, no. 67, pp. 40–50, 2021
2021
-
[123]
Data efficient video transformer for violence detection,
A. R. Abdali, “Data efficient video transformer for violence detection,” in 2021 IEEE International Conference on Communication, Networks and Satellite (COMNETSAT), 2021, pp. 195–199
2021
-
[124]
Pytorch: An imperative style, high-performance deep learning library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al., “Pytorch: An imperative style, high-performance deep learning library,” in NeurIPS, 2019, pp. 8026–8037. UNI-ADAFOCUS: SPATIAL-TEMPORAL DYNAMIC COMPUTATION FOR...
2019
-
[125]
Moreover, we conduct more experiments to evaluate our method on top of the efficient X3D [27] backbone networks on the large-scale Kinetics-400 benchmark (Table 5)
(Figures 10 and 11, Table 3). Moreover, we conduct more experiments to evaluate our method on top of the efficient X3D [27] backbone networks on the large-scale Kinetics-400 benchmark (Table 5). • We further apply our approach to three real-world application scenarios (i.e., f...
-
[126]
Ground Truth
have demonstrated that the vision backbones learned for recognition generally excel at localizing task-relevant regions with their deep representations. Ideally, the reward rt is expected to measure the value of selecting ˜vt in terms of video recognition. With this aim, we de...
-
[127]
easy” and “hard
and ResNet-50 [77] as the global encoder fG and local encoder fL in Uni-AdaFocus. The temporal accumulated max- pooling module proposed in [55] is deployed as the classifier fC. The policy network π adopts an efficient two-branch architecture, corresponding to producing the te...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.