Pith. sign in

REVIEW 4 major objections 4 minor 127 references

Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition

T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Uni-AdaFocus claims that video recognition can be made much cheaper by letting a learned policy crop one informative patch per frame, skip uninformative frames, and exit early on easy videos, all trained end-to-end.

desk verdict A serious, well-written AdaFocus unification with plausible training innovations, but the excerpt ends before any experiments, leaving the efficiency claim unverified. read the letter →

arxiv 2412.11228 v1 pith:JAS7KWHE submitted 2024-12-15 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords dynamicneuralnetworksefficientvideorecognitionspatialredundancyadaptivepatchselectionframesamplingsample-adaptiveinferenceend-to-endtraining
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that video recognition wastes computation on redundant pixels, because the most informative content in each frame is usually a small patch whose location shifts smoothly across frames. It proposes Uni-AdaFocus, a network that first scans every frame cheaply, then uses a learned policy to crop the most informative patch per frame, processes only that patch with a strong encoder, and additionally spends less computation on uninformative frames and easier videos. The whole model trains end-to-end by making patch selection differentiable through bilinear interpolation, and it can wrap standard backbones such as TSM and X3D. If the central claim is right, video recognition can become substantially more efficient without giving up accuracy across action recognition, medical imaging, and content moderation.

What carries the argument

The load-bearing object is the policy network $\pi$, which outputs the patch center $(\tilde{x}^t_c, \tilde{y}^t_c)$ and, in the deformable variant, the patch height and width $(H^t_p, W^t_p)$ for each selected frame. It is trained through a differentiable interpolation-based crop: the selected patch is reconstructed from the four neighboring pixels (or from a patch of the global feature map) so that gradients can flow back to $\pi$ in both pixel space and deep-feature space. The same policy also generates frame weights for weighted sampling without replacement, giving the temporal and sample-wise dynamic computation a single common source of decisions.

What would settle it

Take a video where the action is defined by two far-apart objects, such as a person waving while another person enters from the opposite side. If Uni-AdaFocus with a single deformable patch loses accuracy relative to full-frame processing at matched compute, while a two-patch variant recovers that accuracy, the single-patch assumption is the bottleneck.

Watch

Extended reading notes

Core claim

Uni-AdaFocus unifies three kinds of dynamic computation inside one framework: spatial (which patch to attend to), temporal (which frames to process), and sample-wise (which videos need more compute). The central mechanism is a lightweight global encoder that produces a cheap overview of each frame, a policy network $\pi$ that outputs a patch center and, in the deformable version, a patch height and width, and a high-capacity local encoder that only sees the resized patch. The patch selection is made trainable end-to-end by replacing the hard crop with differentiable interpolation, so gradients flow from the recognition loss back into the policy network. Temporal frame selection is formulated as weighted sampling without replacement and optimized with a Monte Carlo estimate of the expected loss, avoiding reinforcement learning and multi-stage training. The authors claim this combination is considerably more efficient than competitive baselines while remaining hardware-friendly because all selected patches are resized to a common shape and processed in parallel.

Load-bearing premise

The central premise is that the most informative content in each frame is captured by a single rectangular patch whose location changes smoothly, so a policy can find it; if task-relevant cues are spread over multiple disjoint regions, cropping to one patch discards important information.

Editorial extensions

If this is right

  • Uni-AdaFocus can wrap off-the-shelf backbones like TSM and X3D, so the efficiency gains could transfer to existing video models without redesigning their core architecture.
  • The end-to-end training recipe removes the need for reinforcement learning and multi-stage procedures, making dynamic spatial cropping practical for ordinary training pipelines.
  • Spatial, temporal, and sample-wise savings compose inside one framework, so the computational reductions multiply rather than merely add.
  • The model's inference cost can be adjusted online by changing the exit criterion, which suits applications with fluctuating compute budgets such as mobile video search.
  • The selected patches and sampled frames concentrate computation on the task-relevant content, which is claimed to preserve accuracy at a fraction of the full-frame, full-video cost.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The differentiable cropping idea is not inherently video-specific: the same policy-plus-interpolation recipe could be applied to high-resolution images, medical volumes, or multi-view data, although the paper does not test those settings.
  • The deformable patch mechanism behaves like a learned multi-scale attention, so one could diagnose whether the patch size tracks object scale or motion magnitude; that diagnostic is not reported in the paper.
  • The sample-wise exit mechanism suggests a natural accuracy-compute calibration curve, which a deployment system could use to choose a per-query exit threshold; the paper mentions online adjustment but does not develop a scheduling policy around it.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes Uni-AdaFocus, a video-recognition framework designed to combine spatial-, temporal-, and sample-wise dynamic computation. It formalizes an AdaFocus architecture with a lightweight global encoder, a policy network that localizes a single task-relevant patch per frame, a high-capacity local encoder, and a classifier. It then develops an end-to-end training scheme via interpolation-based patch selection, a deep-feature-based policy supervision method, a deformable-patch extension, and a unified treatment of dynamic frame sampling and conditional early exit. The authors claim substantial computational-efficiency gains on seven benchmark datasets and three application scenarios, but the provided manuscript text does not include the corresponding experimental tables or ablations.

Significance. If the empirical claims hold, Uni-AdaFocus would be a useful consolidation of several dynamic-computation ideas: the derivations in Sections 3.3 and 4.1 are detailed, the training recipe is concrete and reproducible in spirit, and the compatibility with off-the-shelf backbones such as TSM and X3D is a practical strength. The paper is also self-aware about one limitation, fixed-size patches, and proposes a deformable remedy. However, the central efficiency-accuracy claim rests on the single-rectangle-per-frame assumption and on experimental evidence that is not present in the provided text; these gaps are load-bearing rather than cosmetic.

major comments (4)
  1. [Sections 3.1 and 4.1.2] The system processes exactly one rectangular patch per selected frame; the deformable variant still outputs one (H_t^p, W_t^p) rectangle resized to P×P. This is load-bearing for the efficiency claim: if task-relevant information is spread over multiple disjoint regions, a single rectangle either discards important content or expands to include irrelevant pixels, degrading accuracy or erasing the efficiency gain. The manuscript acknowledges that fixed-size patches are suboptimal, but it does not analyze multi-modal spatial distributions or provide experiments demonstrating that one rectangle per frame preserves accuracy across the seven claimed datasets. Please add a matched-FLOPs comparison against multi-patch or full-frame baselines and include qualitative patch-coverage analysis on representative benchmark frames.
  2. [Section 4.1.1, Eqs. (13)-(14)] The deep-feature interpolation method treats feature-space cropping as a faithful proxy for pixel-space cropping, justified by the 'location-preserving nature' of deep features. The provided text gives no analysis of the approximation error due to the spatial downsampling of e_G^t (size H_G × W_G) or to the receptive field of the global encoder, and it does not include an ablation comparing the policy trained with Eq. (14) against the pixel-space-gradient baseline of AdaFocusV2. Because this is the main claimed novelty over AdaFocusV2, please provide such an ablation and, ideally, a quantitative error measurement or bound for the location-preservation assumption.
  3. [Abstract and Section 1] The paper's headline claim — 'considerably more efficient than the competitive baselines' — is presented as established, but the provided manuscript text contains no experimental results, ablations, or comparison tables. The only evidence cited is internal references to Tables 6, 7, 15, and 16, none of which appear in the excerpt. Without these data, the central claim cannot be assessed. Please ensure the submitted version includes the full experimental section and that every abstract-level claim is directly supported by the included tables and figures.
  4. [Section 4.2, Eq. (16)] The dynamic frame sampling formulation is stated as weighted sampling without replacement, and the text promises a differentiable Monte Carlo objective obtained by decomposing the expected loss, but the derivation is not present in the provided text. The correctness and training stability of the temporal selection mechanism depend on this derivation, including the gradient estimator and its bias/variance behavior. Please present the full derivation in the main text or in the appendix and state explicitly which parts of the objective are differentiable and which use a straight-through or score-function estimator.
minor comments (4)
  1. [Section 3.3.2] The paragraph claims that the proposed training techniques 'do not introduce additional tunable hyper-parameters,' but Section 4.1.2 introduces alpha in Eq. (15). Please clarify that the no-new-hyperparameter claim applies only to the three techniques in Section 3.3.2, and state how alpha is set.
  2. [Eqs. (1)-(3)] The notation e_G^t and e_L^t is used for both feature maps and pooled feature vectors. Please use distinct symbols (for example, bold for vectors) to avoid ambiguity in the classifier input in Eq. (3).
  3. [Section 4.1.1] AdaFocusV3 is referenced as a prior work, but it has not been introduced earlier in this manuscript. Please give a one-sentence definition or clearly point to reference [30] so the comparison in Eq. (14) is self-contained.
  4. [Figure 6 caption] The caption states that deformable patches yield 'significant accuracy improvements across diverse scenarios,' but no numerical evidence is shown in the provided text. Please cite the corresponding table when discussing this figure.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims rest on explicit new formulations and external benchmarks, with self-citations serving as transparent lineage rather than load-bearing derivations.

full rationale

The derivation chain is not circular in the sense required by the review. The paper explicitly frames itself as an extension: "This paper extends previous conference papers that introduced the basic AdaFocus framework [28] and preliminarily discussed its end-to-end training [29]. Moreover, our deep-feature-based approach for training the patch selection policy (Section 4.1.1) is conceptually relevant to, and improved upon [30]." These self-citations are lineage statements, not load-bearing evidence: the end-to-end training mechanism is derived from the interpolation equations (6)-(9), the deep-feature policy loss is defined in Eqs. (13)-(14) and then modified in Eq. (15), and the temporal and sample-wise components are introduced as new formulations in Section 4.2 onward. The "small-patch" premise is an assumption about video content ("the most informative region in each video frame usually corresponds to a small image patch"), not a conclusion obtained from the method; any resulting accuracy-efficiency tradeoff under distributed task-relevant content is an empirical robustness concern, not circularity. The paper reports comparisons on "seven widely-used benchmark datasets" and competitive baselines, so the efficiency claim is externally checkable rather than forced by a fitted parameter renamed as a prediction. No equation in the excerpt equates a predicted quantity with a fitted input by construction, and no uniqueness or forbidden-alternative argument is imported from the authors' prior work. Therefore the paper receives a non-circular finding.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The method introduces a few hyperparameters (alpha, patch size, frame budgets) but no new physical or mathematical entities. The key assumptions are the single-patch spatial redundancy hypothesis and the location-preserving property of deep features, both of which are plausible but not rigorously established in the excerpt.

free parameters (3)
  • alpha (patch size regularization coefficient) = unknown, pre-defined per dataset
    Introduced in Eq. (15) to prevent degenerate patch sizes; its value likely tuned on validation data, affecting learned patch shapes and thus the efficiency-accuracy tradeoff.
  • Patch size P = not specified
    Fixed resize size for local encoder input; a design choice that determines the spatial budget and accuracy.
  • Number of frames TG and TL = not specified
    Uniformly sampled global frames TG and dynamically selected local frames TL control the temporal budget; chosen by the user.
assumptions (3)
  • domain assumption Informative content in each video frame is concentrated in a small patch whose shape, size, and location vary smoothly across frames.
    Stated in abstract and Section 1; this is the core spatial redundancy assumption that justifies patch-based processing.
  • ad hoc to paper Deep feature maps preserve the spatial location of input contents, so that interpolation in feature space approximates interpolation in pixel space.
    Invoked in Section 4.1.1 to justify training the policy network via feature interpolation (Eq. 13,14).
  • domain assumption Weighted sampling without replacement with a differentiable expected loss can be trained end-to-end via Monte Carlo estimation.
    Used for dynamic frame sampling in Section 4.2; the full derivation is partially described and relies on standard policy-gradient or reparameterization-type estimates.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition." pith.science (2026). https://pith.science/paper/JAS7KWHE

@misc{pith2026241211228,
  author       = {Pith},
  title        = {Pith review of: Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JAS7KWHE}},
  note         = {Machine review of arXiv:2412.11228}
}
read the original abstract

This paper presents a comprehensive exploration of the phenomenon of data redundancy in video understanding, with the aim to improve computational efficiency. Our investigation commences with an examination of spatial redundancy, which refers to the observation that the most informative region in each video frame usually corresponds to a small image patch, whose shape, size and location shift smoothly across frames. Motivated by this phenomenon, we formulate the patch localization problem as a dynamic decision task, and introduce a spatially adaptive video recognition approach, termed AdaFocus. In specific, a lightweight encoder is first employed to quickly process the full video sequence, whose features are then utilized by a policy network to identify the most task-relevant regions. Subsequently, the selected patches are inferred by a high-capacity deep network for the final prediction. The full model can be trained in end-to-end conveniently. Furthermore, AdaFocus can be extended by further considering temporal and sample-wise redundancies, i.e., allocating the majority of computation to the most task-relevant frames, and minimizing the computation spent on relatively "easier" videos. Our resulting approach, Uni-AdaFocus, establishes a comprehensive framework that seamlessly integrates spatial, temporal, and sample-wise dynamic computation, while it preserves the merits of AdaFocus in terms of efficient end-to-end training and hardware friendliness. In addition, Uni-AdaFocus is general and flexible as it is compatible with off-the-shelf efficient backbones (e.g., TSM and X3D), which can be readily deployed as our feature extractor, yielding a significantly improved computational efficiency. Empirically, extensive experiments based on seven benchmark datasets and three application scenarios substantiate that Uni-AdaFocus is considerably more efficient than the competitive baselines.

Figures

Figures reproduced from arXiv: 2412.11228 by the authors.

Figure 1
Figure 1. Comparisons between existing temporal-based methods and our proposed approaches. Most existing works aim to reduce computational costs by selecting a few informative frames to process. Orthogonal to them, AdaFocus reveals that a superior computational efficiency can be achieved by reducing the spatial redundancy. Built upon this finding, we further demonstrate that it is feasible to formulate flexible and highly eff… view at source ↗
Figure 2
Figure 2. Overview of AdaFocus. It first takes a quick glance at each frame vt using a lightweight global encoder fG. Then a policy network π is built on top of fG to select the most important image region v˜t in terms of recognition. A high-capacity local encoder fL is adopted to extract features from v˜t. Finally, a classifier aggregates the features across frames to obtain the prediction pt. 3 ADAPTIVE FOCUS NETWORK (ADAFO… view at source ↗
Figure 3
Figure 3. Illustration of the policy network π in AdaFocusV1. The outputs of π parameterize a categorical distribution π(·|e G 1 , . . . , e G t ) on multiple patch candidates (here we take 25 as an example). During training, we sample v˜t from π(·|e G 1 , . . . , e G t ), while at test time, we directly select the patch with the largest softmax probability. facilitating efficient feature reuse. This design is natural since i… view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Illustration of interpolation-based patch selection. This operation is differentiable, i.e., the gradients can be directly back￾propagated into the policy network π through the selected image patch v˜t. Consequently, integrating the learning of π into a unified end-to-…
Figure 5
Figure 5. Figure 5: Guiding the training of π with deep features. The gradients for π are obtained by minimizing the deep-feature-based loss L spatial π , instead of being back propagated from the pixel space. With this design, the supervision signals for π contain much more semantic-leve…
Figure 6
Figure 6. Figure 6: Deformable patches, which enable the patch selection policy to adapt flexibly to the task-relevant regions in various shapes, scales, and locations. The selected patches are resized to a common size P 2 to be processed efficiently by fL on hardwares (e.g., GPUs). Notab…
Figure 7
Figure 7. Figure 7: Overview of Uni-AdaFocus, a unified framework that models spatial-temporal dynamic computation concurrently, leading to a considerably improved computational efficiency for inference. Compared with [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Illustration of the conditional-exit algorithm. to other samples. We propose to model this sample-wise redundancy through a straightforward early-exit algorithm. This approach can be seamlessly integrated into the models trained using the previously outlined paradigm, …
Figure 9
Figure 9. Figure 9: Comparisons of Uni-AdaFocus and state-of-the-art efficient video understanding approaches on ActivityNet in terms of inference efficiency. Our method adopts P 2 ∈ {962 , 1282 , 1602 }, corresponding to the three black curves. Notably, the inference cost of Uni-AdaFocus…
Figure 10
Figure 10. Figure 10: Comparisons between Uni-AdaFocus and the preliminary versions of AdaFocus on ActivityNet in terms of inference efficiency [PITH_FULL_IMAGE:figures/full_fig_p011_10.png]
Figure 11
Figure 11. Figure 11: Uni-AdaFocus v.s. the preliminary versions of AdaFocus on Sth-Sth V1 (left) and V2 (right) in terms of inference efficiency. For fair comparison, we do not perform sample-wise dynamic computation for Uni-AdaFocus. mAP by 1.6-1.9% on top of AdaFocusV2. Moreover, it can…
Figure 12
Figure 12. Figure 12: Ablation studies of the conditional-exit algorithm. The results on ActivityNet are provided. For a fair comparison, different variants are deployed on top of the same base model. Our method significantly reduces the computational cost when achieving the same mAP. TABL…
Figure 13
Figure 13. Figure 13: Examples of the task-relevant patches and informative videos frames selected by Uni-AdaFocus (zoom in for details). We present a variety of representative input videos, where “easy/hard” videos refer to the samples to which Uni-AdaFocus allocates a relatively smaller …
Figure 14
Figure 14. Figure 14: Representative examples of the failure cases of our model (zoom in for details). Blue boxes indicate the patches selected by Uni￾AdaFocus. “Ground Truth” and “Prediction” denote the ground truth labels of the videos and the predictions of Uni-AdaFocus, respectively. D…
Figure 15
Figure 15. Figure 15: Additional visualization results focusing primarily on videos that incorporate multi-person/object interactions (zoom in for details). Blue boxes indicate the patches selected by Uni-AdaFocus. for the model. Moreover, it appears to be more challenging for the model to…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

127 extracted references · 70 canonical work pages

  1. [1]

    The youtube video recommendation system,

    J. Davidson, B. Liebald, J. Liu, P . Nandy, T. Van Vleet, U. Gargi, S. Gupta, Y. He, M. Lambert, B. Livingston et al., “The youtube video recommendation system,” in Proceedings of the fourth ACM conference on Recommender systems, 2010, pp. 293–296

  2. [2]

    Content-based video recommendation system based on stylistic visual features,

    Y. Deldjoo, M. Elahi, P . Cremonesi, F. Garzotto, P . Piazzolla, and M. Quadrana, “Content-based video recommendation system based on stylistic visual features,” Journal on Data Semantics, vol. 5, no. 2, pp. 99–113, 2016

  3. [3]

    A unified personalized video recommendation via dynamic recurrent neural networks,

    J. Gao, T. Zhang, and C. Xu, “A unified personalized video recommendation via dynamic recurrent neural networks,” in ACM MM, 2017, pp. 127–135

  4. [4]

    A system for video surveillance and monitoring,

    R. T. Collins, A. J. Lipton, T. Kanade, H. Fujiyoshi, D. Duggins, Y. Tsin, D. Tolliver, N. Enomoto, O. Hasegawa, P . Burtet al., “A system for video surveillance and monitoring,” VSAM final report, vol. 2000, no. 1-68, p. 1, 2000

  5. [5]

    Distributed deep learning model for intelligent video surveillance systems with edge computing,

    J. Chen, K. Li, Q. Deng, K. Li, and S. Y. Philip, “Distributed deep learning model for intelligent video surveillance systems with edge computing,” IEEE Transactions on Industrial Informatics, 2019

  6. [6]

    Searching video for complex activities with finite state models,

    N. Ikizler and D. Forsyth, “Searching video for complex activities with finite state models,” in CVPR, 2007, pp. 1–8

  7. [7]

    Slowfast networks for video recognition,

    C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in ICCV, 2019, pp. 6202–6211

  8. [8]

    Deep feature flow for video recognition,

    X. Zhu, Y. Xiong, J. Dai, L. Yuan, and Y. Wei, “Deep feature flow for video recognition,” in CVPR, 2017, pp. 2349–2358

Show all 127 references
  1. [9]

    Convolutional two- stream network fusion for video action recognition,

    C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two- stream network fusion for video action recognition,” in CVPR, 2016, pp. 1933–1941

  2. [10]

    Quo vadis, action recognition? a new model and the kinetics dataset,

    J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in CVPR, 2017, pp. 6299– 6308

  3. [11]

    Learn- ing spatiotemporal features with 3d convolutional networks,

    D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learn- ing spatiotemporal features with 3d convolutional networks,” in ICCV, 2015, pp. 4489–4497

  4. [12]

    Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?

    K. Hara, H. Kataoka, and Y. Satoh, “Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?” in CVPR, 2018, pp. 6546–6555

  5. [13]

    Liteeval: A coarse- to-fine framework for resource efficient video recognition,

    Z. Wu, C. Xiong, Y.-G. Jiang, and L. S. Davis, “Liteeval: A coarse- to-fine framework for resource efficient video recognition,” in NeurIPS, 2019b

  6. [14]

    Multi-agent rein- forcement learning based frame sampling for effective untrimmed video recognition,

    W. Wu, D. He, X. Tan, S. Chen, and S. Wen, “Multi-agent rein- forcement learning based frame sampling for effective untrimmed video recognition,” in ICCV, 2019a, pp. 6222–6231

  7. [15]

    A dynamic frame selection framework for fast video recognition,

    Z. Wu, H. Li, C. Xiong, Y.-G. Jiang, and L. S. Davis, “A dynamic frame selection framework for fast video recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020b

  8. [16]

    Scsampler: Sampling salient clips from video for efficient action recognition,

    B. Korbar, D. Tran, and L. Torresani, “Scsampler: Sampling salient clips from video for efficient action recognition,” in ICCV, 2019, pp. 6232–6242

  9. [17]

    Listen to look: Action recognition by previewing audio,

    R. Gao, T.-H. Oh, K. Grauman, and L. Torresani, “Listen to look: Action recognition by previewing audio,” in CVPR, 2020, pp. 10 457–10 467

  10. [18]

    Ar-net: Adaptive frame resolution for efficient action recognition,

    Y. Meng, C.-C. Lin, R. Panda, P . Sattigeri, L. Karlinsky, A. Oliva, K. Saenko, and R. Feris, “Ar-net: Adaptive frame resolution for efficient action recognition,” in ECCV, 2020, pp. 86–104

  11. [19]

    Ocsampler: Compressing videos to one clip with single-step sampling,

    J. Lin, H. Duan, K. Chen, D. Lin, and L. Wang, “Ocsampler: Compressing videos to one clip with single-step sampling,” in CVPR, 2022

  12. [20]

    Recurrent models of visual attention,

    V . Mnih, N. Heess, A. Graveset al., “Recurrent models of visual attention,” in NeurIPS, 2014, pp. 2204–2212

  13. [21]

    Look closer to see better: Recurrent attention convolutional neural network for fine-grained image recognition,

    J. Fu, H. Zheng, and T. Mei, “Look closer to see better: Recurrent attention convolutional neural network for fine-grained image recognition,” in CVPR, 2017, pp. 4438–4446

  14. [22]

    Spatially adaptive inference with stochastic feature sampling and interpolation,

    Z. Xie, Z. Zhang, X. Zhu, G. Huang, and S. Lin, “Spatially adaptive inference with stochastic feature sampling and interpolation,” in ECCV, 2020, pp. 531–548

  15. [23]

    Dynamic neural networks: A survey,

    Y. Han, G. Huang, S. Song, L. Yang, H. Wang, and Y. Wang, “Dynamic neural networks: A survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 11, pp. 7436–7456, 2022

  16. [24]

    Glance and focus networks for dynamic visual recognition,

    G. Huang, Y. Wang, K. Lv, H. Jiang, W. Huang, P . Qi, and S. Song, “Glance and focus networks for dynamic visual recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 4, pp. 4605–4621, 2023

  17. [25]

    Dynamic spatial sparsification for efficient vision transformers and convolutional neural networks,

    Y. Rao, Z. Liu, W. Zhao, J. Zhou, and J. Lu, “Dynamic spatial sparsification for efficient vision transformers and convolutional neural networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–14, 2023

  18. [26]

    Tsm: Temporal shift module for efficient video understanding,

    J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” in ICCV, 2019, pp. 7083–7093

  19. [27]

    X3d: Expanding architectures for efficient video recognition,

    C. Feichtenhofer, “X3d: Expanding architectures for efficient video recognition,” in CVPR, 2020, pp. 203–213

  20. [28]

    Adaptive focus for efficient video recognition,

    Y. Wang, Z. Chen, H. Jiang, S. Song, Y. Han, and G. Huang, “Adaptive focus for efficient video recognition,” in ICCV, 2021

  21. [29]

    Adafocus v2: End-to-end training of spatial dynamic networks for video recognition,

    Y. Wang, Y. Yue, Y. Lin, H. Jiang, Z. Lai, V . Kulikov, N. Orlov, H. Shi, and G. Huang, “Adafocus v2: End-to-end training of spatial dynamic networks for video recognition,” in CVPR, 2022

  22. [30]

    Adafocusv3: On unified spatial-temporal dynamic video recognition,

    Y. Wang, Y. Yue, X. Xu, A. Hassani, V . Kulikov, N. Orlov, S. Song, H. Shi, and G. Huang, “Adafocusv3: On unified spatial-temporal dynamic video recognition,” in ECCV, 2022, pp. 226–243

  23. [31]

    Activitynet: A large-scale video benchmark for human activity understanding,

    F. Caba Heilbron, V . Escorcia, B. Ghanem, and J. Carlos Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in CVPR, 2015, pp. 961–970

  24. [32]

    Exploiting feature and class relationships in video categorization with regularized deep neural networks,

    Y.-G. Jiang, Z. Wu, J. Wang, X. Xue, and S.-F. Chang, “Exploiting feature and class relationships in video categorization with regularized deep neural networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 2, pp. 352–364, 2018

  25. [33]

    The kinetics human action video dataset,

    W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijaya- narasimhan, F. Viola, T. Green, T. Back, P . Natsev et al. , “The kinetics human action video dataset,” arXiv:1705.06950, 2017

  26. [34]

    The "something something

    R. Goyal, S. Ebrahimi Kahou, V . Michalski, J. Materzynska, S. Westphal, H. Kim, V . Haenel, I. Fruend, P . Yianilos, M. Mueller- Freitag et al. , “The "something something" video database for learning and evaluating visual common sense,” in ICCV, 2017, pp. 5842–5850

  27. [35]

    The jester dataset: A large-scale video dataset of human gestures,

    J. Materzynska, G. Berger, I. Bax, and R. Memisevic, “The jester dataset: A large-scale video dataset of human gestures,” inICCVW, 2019

  28. [36]

    Temporal segment networks: Towards good prac- tices for deep action recognition,

    L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks: Towards good prac- tices for deep action recognition,” in ECCV, 2016, pp. 20–36

  29. [37]

    Long-term recurrent convolutional networks for visual recognition and description,

    J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR, 2015, pp. 2625–2634

  30. [38]

    Recurrent tubelet proposal and recognition networks for action detection,

    D. Li, Z. Qiu, Q. Dai, T. Yao, and T. Mei, “Recurrent tubelet proposal and recognition networks for action detection,” in ECCV, 2018, pp. 303–318

  31. [39]

    Beyond short snippets: Deep net- works for video classification,

    J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici, “Beyond short snippets: Deep net- works for video classification,” in CVPR, 2015, pp. 4694–4702

  32. [40]

    Gate-shift networks for video action recognition,

    S. Sudhakaran, S. Escalera, and O. Lanz, “Gate-shift networks for video action recognition,” in CVPR, 2020, pp. 1102–1111

  33. [41]

    Adafuse: Adaptive temporal fusion network for efficient action recognition,

    Y. Meng, R. Panda, C.-C. Lin, P . Sattigeri, L. Karlinsky, K. Saenko, A. Oliva, and R. Feris, “Adafuse: Adaptive temporal fusion network for efficient action recognition,” in ICLR, 2021

  34. [42]

    Spatiotemporal multiplier networks for video action recognition,

    C. Feichtenhofer, A. Pinz, and R. P . Wildes, “Spatiotemporal multiplier networks for video action recognition,” in CVPR, 2017, pp. 4768–4777

  35. [43]

    Searching for two-stream models in multivariate space for video recognition,

    X. Gong, H. Wang, M. Z. Shou, M. Feiszli, Z. Wang, and Z. Yan, “Searching for two-stream models in multivariate space for video recognition,” in ICCV, 2021, pp. 8033–8042

  36. [44]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR, 2021

  37. [45]

    Vivit: A video vision transformer,

    A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Luˇ ci´ c, and C. Schmid, “Vivit: A video vision transformer,” in ICCV, 2021

  38. [46]

    Video swin transformer,

    Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in CVPR, 2022, pp. 3202–3211

  39. [47]

    Is space-time attention all you need for video understanding?

    G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in ICML, 2021

  40. [48]

    Video transformer network,

    D. Neimark, O. Bar, M. Zohar, and D. Asselmann, “Video transformer network,” in ICCV, 2021, pp. 3163–3172

  41. [49]

    A closer look at spatiotemporal convolutions for action recognition,

    D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in CVPR, 2018, pp. 6450–6459. UNI-ADAFOCUS: SPATIAL-TEMPORAL DYNAMIC COMPUTATION FOR VIDEO RECOGNITION 17

  42. [50]

    Eco: Efficient convolutional network for online video understanding,

    M. Zolfaghari, K. Singh, and T. Brox, “Eco: Efficient convolutional network for online video understanding,” in ECCV, 2018, pp. 695–712

  43. [51]

    Video classification with channel-separated convolutional networks,

    D. Tran, H. Wang, L. Torresani, and M. Feiszli, “Video classification with channel-separated convolutional networks,” in ICCV, 2019, pp. 5552–5561

  44. [52]

    Teinet: Towards an efficient architecture for video recognition,

    Z. Liu, D. Luo, Y. Wang, L. Wang, Y. Tai, C. Wang, J. Li, F. Huang, and T. Lu, “Teinet: Towards an efficient architecture for video recognition,” in AAAI, vol. 34, no. 07, 2020, pp. 11 669–11 676

  45. [53]

    Tam: Temporal adaptive module for video recognition,

    Z. Liu, L. Wang, W. Wu, C. Qian, and T. Lu, “Tam: Temporal adaptive module for video recognition,” in ICCV, 2021, pp. 13 708– 13 718

  46. [54]

    End-to-end learning of action detection from frame glimpses in videos,

    S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei, “End-to-end learning of action detection from frame glimpses in videos,” in CVPR, 2016, pp. 2678–2687

  47. [55]

    Frameexit: Condi- tional early exiting for efficient video recognition,

    A. Ghodrati, B. E. Bejnordi, and A. Habibian, “Frameexit: Condi- tional early exiting for efficient video recognition,” in CVPR, 2021, pp. 15 608–15 618

  48. [56]

    Efficient action recognition via dynamic knowledge propagation,

    H. Kim, M. Jain, J.-T. Lee, S. Yun, and F. Porikli, “Efficient action recognition via dynamic knowledge propagation,” in ICCV, 2021, pp. 13 719–13 728

  49. [57]

    Dynamic network quantization for efficient video inference,

    X. Sun, R. Panda, C.-F. R. Chen, A. Oliva, R. Feris, and K. Saenko, “Dynamic network quantization for efficient video inference,” in ICCV, 2021, pp. 7375–7385

  50. [58]

    Nsnet: Non-saliency suppression sampler for effi- cient video recognition,

    B. Xia, W. Wu, H. Wang, R. Su, D. He, H. Yang, X. Fan, and W. Ouyang, “Nsnet: Non-saliency suppression sampler for effi- cient video recognition,” in ECCV, 2022, pp. 705–723

  51. [59]

    Temporal saliency query network for efficient video recognition,

    B. Xia, Z. Wang, W. Wu, H. Wang, and J. Han, “Temporal saliency query network for efficient video recognition,” in ECCV, 2022, pp. 741–759

  52. [60]

    Spatial trans- former networks,

    M. Jaderberg, K. Simonyan, A. Zisserman et al., “Spatial trans- former networks,” in NeurIPS, 2015, pp. 2017–2025

  53. [61]

    Sbnet: Sparse blocks network for fast inference,

    M. Ren, A. Pokrovsky, B. Yang, and R. Urtasun, “Sbnet: Sparse blocks network for fast inference,” in CVPR, 2018, pp. 8711–8720

  54. [62]

    Resolution adaptive networks for efficient inference,

    L. Yang, Y. Han, X. Chen, S. Song, J. Dai, and G. Huang, “Resolution adaptive networks for efficient inference,” in CVPR, 2020, pp. 2369–2378

  55. [63]

    Adaptively connected neural networks,

    G. Wang, K. Wang, and L. Lin, “Adaptively connected neural networks,” in CVPR, 2019, pp. 1781–1790

  56. [64]

    Dynamic region- aware convolution,

    J. Chen, X. Wang, Z. Guo, X. Zhang, and J. Sun, “Dynamic region- aware convolution,” in CVPR, 2021, pp. 8064–8073

  57. [65]

    Spatially adaptive computation time for residual networks,

    M. Figurnov, M. D. Collins, Y. Zhu, L. Zhang, J. Huang, D. Vetrov, and R. Salakhutdinov, “Spatially adaptive computation time for residual networks,” in CVPR, 2017, pp. 1039–1048

  58. [66]

    Dynamic convolutions: Exploiting spatial sparsity for faster inference,

    T. Verelst and T. Tuytelaars, “Dynamic convolutions: Exploiting spatial sparsity for faster inference,” in CVPR, 2020, pp. 2320–2329

  59. [67]

    Interpretable spatio-temporal attention for video action recognition,

    L. Meng, B. Zhao, B. Chang, G. Huang, W. Sun, F. Tung, and L. Sigal, “Interpretable spatio-temporal attention for video action recognition,” in ICCVW, 2019

  60. [68]

    Generalized domain conditioned adaptation network,

    S. Li, B. Xie, Q. Lin, C. H. Liu, G. Huang, and G. Wang, “Generalized domain conditioned adaptation network,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 8, pp. 4093–4109, 2021

  61. [69]

    Learning deep features for discriminative localization,

    B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in CVPR, 2016, pp. 2921–2929

  62. [70]

    Grad-cam: Visual explanations from deep networks via gradient-based localization,

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in ICCV, 2017, pp. 618–626

  63. [71]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR, 2021

  64. [72]

    Delving into details: Synopsis-to-detail networks for video recognition,

    S. Liang, X. Shen, J. Huang, and X.-S. Hua, “Delving into details: Synopsis-to-detail networks for video recognition,” in ECCV, 2022, pp. 262–278

  65. [73]

    Long short-term memory,

    S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997

  66. [74]

    Learning phrase rep- resentations using RNN encoder–decoder for statistical machine translation,

    K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase rep- resentations using RNN encoder–decoder for statistical machine translation,” in EMNLP, 2014, pp. 1724–1734

  67. [75]

    Playing atari with deep reinforce- ment learning,

    V . Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing atari with deep reinforce- ment learning,” arXiv:1312.5602, 2013

  68. [76]

    Proximal policy optimization algorithms,

    J. Schulman, F. Wolski, P . Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv:1707.06347, 2017

  69. [77]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016, pp. 770–778

  70. [78]

    Convolutional networks with dense connectivity,

    G. Huang, Z. Liu, G. Pleiss, L. Van Der Maaten, and K. Wein- berger, “Convolutional networks with dense connectivity,” IEEE transactions on pattern analysis and machine intelligence , 2019

  71. [79]

    Better mixing via deep representations,

    Y. Bengio, G. Mesnil, Y. Dauphin, and S. Rifai, “Better mixing via deep representations,” in ICML, 2013, pp. 552–560

  72. [80]

    Implicit semantic data augmentation for deep networks,

    Y. Wang, X. Pan, S. Song, H. Zhang, G. Huang, and C. Wu, “Implicit semantic data augmentation for deep networks,” in NeurIPS, 2019

  73. [81]

    Deep residual correction network for partial domain adaptation,

    S. Li, C. H. Liu, Q. Lin, Q. Wen, L. Su, G. Huang, and Z. Ding, “Deep residual correction network for partial domain adaptation,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 7, pp. 2329–2344, 2020

  74. [82]

    Sepico: Semantic-guided pixel contrast for domain adaptive semantic segmentation,

    B. Xie, S. Li, M. Li, C. H. Liu, G. Huang, and G. Wang, “Sepico: Semantic-guided pixel contrast for domain adaptive semantic segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 7, pp. 9004–9021, 2023

  75. [83]

    Adapting across domains via target-oriented transferable semantic augmentation under prototype constraint,

    M. Xie, S. Li, K. Gong, Y. Wang, and G. Huang, “Adapting across domains via target-oriented transferable semantic augmentation under prototype constraint,” International Journal of Computer Vision, vol. 132, no. 4, pp. 1417–1441, 2024

  76. [84]

    Multi-scale dense networks for resource efficient image classification,

    G. Huang, D. Chen, T. Li, F. Wu, L. van der Maaten, and K. Q. Weinberger, “Multi-scale dense networks for resource efficient image classification,” in ICLR, 2018

  77. [85]

    Glance and focus: a dynamic approach to reducing spatial redundancy in image classification,

    Y. Wang, K. Lv, R. Huang, S. Song, L. Yang, and G. Huang, “Glance and focus: a dynamic approach to reducing spatial redundancy in image classification,” in NeurIPS, 2020

  78. [86]

    Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition,

    Y. Wang, R. Huang, S. Song, Z. Huang, and G. Huang, “Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition,” in NeurIPS, 2021

  79. [87]

    Watching a small portion could be as good as watching all: Towards efficient video classification,

    H. Fan, Z. Xu, L. Zhu, C. Yan, J. Ge, and Y. Yang, “Watching a small portion could be as good as watching all: Towards efficient video classification,” in IJCAI, 2018

  80. [88]

    Adamml: Adaptive multi-modal learning for efficient video recognition,

    R. Panda, C.-F. R. Chen, Q. Fan, X. Sun, K. Saenko, A. Oliva, and R. Feris, “Adamml: Adaptive multi-modal learning for efficient video recognition,” in ICCV, 2021, pp. 7576–7585

  81. [89]

    Smart frame selection for action recognition,

    S. N. Gowda, M. Rohrbach, and L. Sevilla-Lara, “Smart frame selection for action recognition,” in AAAI, vol. 35, no. 2, 2021, pp. 1451–1459

  82. [90]

    Look more but care less in video recognition,

    Y. Zhang, Y. Bai, H. Wang, Y. Xu, and Y. Fu, “Look more but care less in video recognition,” in NeurIPS, 2022

  83. [91]

    Resound: Towards action recognition without representation bias,

    Y. Li, Y. Li, and N. Vasconcelos, “Resound: Towards action recognition without representation bias,” in ECCV, 2018, pp. 513– 528

  84. [92]

    https://adni.loni.usc.edu/

  85. [93]

    https://www.oasis-brains.org/

  86. [94]

    https://www.ppmi-info.org/

  87. [95]

    Violence recognition from videos using deep learning techniques,

    M. M. Soliman, M. H. Kamal, M. A. E.-M. Nashed, Y. M. Mostafa, B. S. Chawky, and D. Khattab, “Violence recognition from videos using deep learning techniques,” in 2019 Ninth International Conference on Intelligent Computing and Information Systems (ICICIS), 2019, pp. 80–85

  88. [96]

    Mobilenetv2: Inverted residuals and linear bottlenecks,

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in CVPR, 2018, pp. 4510–4520

  89. [97]

    Videos as space-time region graphs,

    X. Wang and A. Gupta, “Videos as space-time region graphs,” in ECCV, 2018, pp. 399–417

  90. [98]

    Temporal relational reasoning in videos,

    B. Zhou, A. Andonian, A. Oliva, and A. Torralba, “Temporal relational reasoning in videos,” in ECCV, 2018, pp. 803–818

  91. [99]

    Smallbignet: Integrating core and contextual views for video classification,

    X. Li, Y. Wang, Z. Zhou, and Y. Qiao, “Smallbignet: Integrating core and contextual views for video classification,” in CVPR, 2020, pp. 1092–1101

  92. [100]

    More is less: Learning efficient video representations by big-little network and depthwise temporal aggregation,

    Q. Fan, C.-F. R. Chen, H. Kuehne, M. Pistoia, and D. Cox, “More is less: Learning efficient video representations by big-little network and depthwise temporal aggregation,” in NeurIPS, 2019

  93. [101]

    Stm: Spatiotempo- ral and motion encoding for action recognition,

    B. Jiang, M. Wang, W. Gan, W. Wu, and J. Yan, “Stm: Spatiotempo- ral and motion encoding for action recognition,” in ICCV, 2019, pp. 2000–2009

  94. [102]

    Tea: Temporal excitation and aggregation for action recognition,

    Y. Li, B. Ji, X. Shi, J. Zhang, B. Kang, and L. Wang, “Tea: Temporal excitation and aggregation for action recognition,” in CVPR, 2020, pp. 909–918

  95. [103]

    Appearance-and-relation networks for video classification,

    L. Wang, W. Li, W. Li, and L. Van Gool, “Appearance-and-relation networks for video classification,” in CVPR, 2018, pp. 1430–1439

  96. [104]

    Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,

    S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,” in ECCV, 2018, pp. 305–321. UNI-ADAFOCUS: SPATIAL-TEMPORAL DYNAMIC COMPUTATION FOR VIDEO RECOGNITION 18

  97. [105]

    Deep analysis of cnn-based spatio-temporal representations for action recognition,

    C.-F. R. Chen, R. Panda, K. Ramakrishnan, R. Feris, J. Cohn, A. Oliva, and Q. Fan, “Deep analysis of cnn-based spatio-temporal representations for action recognition,” in CVPR, 2021, pp. 6165– 6175

  98. [106]

    Non-local neural networks,

    X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR, 2018, pp. 7794–7803

  99. [107]

    Tdn: Temporal difference networks for efficient action recognition,

    L. Wang, Z. Tong, B. Ji, and G. Wu, “Tdn: Temporal difference networks for efficient action recognition,” in CVPR, 2021, pp. 1895–1904

  100. [108]

    Efficientnet: Rethinking model scaling for convolutional neural networks,

    M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in ICML, 2019, pp. 6105–6114

  101. [109]

    Movinets: Mobile video networks for efficient video recognition,

    D. Kondratyuk, L. Yuan, Y. Li, L. Zhang, M. Tan, M. Brown, and B. Gong, “Movinets: Mobile video networks for efficient video recognition,” in CVPR, 2021, pp. 16 020–16 030

  102. [110]

    Mgsampler: An explainable sampling strategy for video action recognition,

    Y. Zhi, Z. Tong, L. Wang, and G. Wu, “Mgsampler: An explainable sampling strategy for video action recognition,” in ICCV, 2021, pp. 1513–1522

  103. [111]

    Alzheimer’s disease,

    P . Scheltens, B. De Strooper, M. Kivipelto, H. Holstege, G. Chéte- lat, C. E. Teunissen, J. Cummings, and W. M. van der Flier, “Alzheimer’s disease,” The Lancet, vol. 397, no. 10284, pp. 1577– 1590, 2021

  104. [112]

    Parkinson’s disease,

    B. R. Bloem, M. S. Okun, and C. Klein, “Parkinson’s disease,” The Lancet, vol. 397, no. 10291, pp. 2284–2303, 2021

  105. [113]

    Parkinson’s disease: mechanisms and models,

    W. Dauer and S. Przedborski, “Parkinson’s disease: mechanisms and models,” Neuron, vol. 39, no. 6, pp. 889–909, 2003

  106. [114]

    Relational self- attention: What’s missing in attention for video understanding,

    M. Kim, H. Kwon, C. Wang, S. Kwak, and M. Cho, “Relational self- attention: What’s missing in attention for video understanding,” in NeurIPS, 2021, pp. 8046–8059

  107. [115]

    Hierarchical fully con- volutional network for joint atrophy localization and alzheimer’s disease diagnosis using structural mri,

    C. Lian, M. Liu, J. Zhang, and D. Shen, “Hierarchical fully con- volutional network for joint atrophy localization and alzheimer’s disease diagnosis using structural mri,”IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 4, pp. 880–893, 2018

  108. [116]

    End-to-end automatic pathology localization for alzheimer’s disease diagnosis using structural mri,

    G. Cao, M. Zhang, Y. Wang, J. Zhang, Y. Han, X. Xu, J. Huang, and G. Kang, “End-to-end automatic pathology localization for alzheimer’s disease diagnosis using structural mri,” Computers in Biology and Medicine, p. 107110, 2023

  109. [117]

    Deep neural networks with broad views for parkinson’s disease screening,

    X. Zhang, Y. Yang, H. Wang, S. Ning, and H. Wang, “Deep neural networks with broad views for parkinson’s disease screening,” in 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2019, pp. 1018–1022

  110. [118]

    Med3d: Transfer learning for 3d medical image analysis,

    S. Chen, K. Ma, and Y. Zheng, “Med3d: Transfer learning for 3d medical image analysis,” arXiv:1904.00625, 2019

  111. [119]

    M3t: three-dimensional medical image classifier using multi-plane and multi-slice transformer,

    J. Jang and D. Hwang, “M3t: three-dimensional medical image classifier using multi-plane and multi-slice transformer,” in CVPR, 2022, pp. 20 718–20 729

  112. [120]

    Learn-explain-reinforce: Coun- terfactual reasoning and its guidance to reinforce an alzheimer’s disease diagnosis model,

    K. Oh, J. S. Yoon, and H.-I. Suk, “Learn-explain-reinforce: Coun- terfactual reasoning and its guidance to reinforce an alzheimer’s disease diagnosis model,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 4, pp. 4843–4857, 2023

  113. [121]

    2d bidirectional gated recurrent unit convolutional neural networks for end-to-end violence detec- tion in videos,

    A. Traoré and M. A. Akhloufi, “2d bidirectional gated recurrent unit convolutional neural networks for end-to-end violence detec- tion in videos,” in Image Analysis and Recognition: 17th International Conference, ICIAR 2020, Póvoa de Varzim, Portugal, June 24–26, 2020, Proceed...

  114. [122]

    A temporal fusion approach for video classification with convolutional and lstm neural networks applied to violence detection,

    J. P . de Oliveira Lima and C. M. S. Figueiredo, “A temporal fusion approach for video classification with convolutional and lstm neural networks applied to violence detection,” Inteligencia Artificial, vol. 24, no. 67, pp. 40–50, 2021

  115. [123]

    Data efficient video transformer for violence detection,

    A. R. Abdali, “Data efficient video transformer for violence detection,” in 2021 IEEE International Conference on Communication, Networks and Satellite (COMNETSAT), 2021, pp. 195–199

  116. [124]

    Pytorch: An imperative style, high-performance deep learning library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al., “Pytorch: An imperative style, high-performance deep learning library,” in NeurIPS, 2019, pp. 8026–8037. UNI-ADAFOCUS: SPATIAL-TEMPORAL DYNAMIC COMPUTATION FOR...

  117. [125]

    Moreover, we conduct more experiments to evaluate our method on top of the efficient X3D [27] backbone networks on the large-scale Kinetics-400 benchmark (Table 5)

    (Figures 10 and 11, Table 3). Moreover, we conduct more experiments to evaluate our method on top of the efficient X3D [27] backbone networks on the large-scale Kinetics-400 benchmark (Table 5). • We further apply our approach to three real-world application scenarios (i.e., f...

  118. [126]

    Ground Truth

    have demonstrated that the vision backbones learned for recognition generally excel at localizing task-relevant regions with their deep representations. Ideally, the reward rt is expected to measure the value of selecting ˜vt in terms of video recognition. With this aim, we de...

  119. [127]

    easy” and “hard

    and ResNet-50 [77] as the global encoder fG and local encoder fL in Uni-AdaFocus. The temporal accumulated max- pooling module proposed in [55] is deployed as the classifier fC. The policy network π adopts an efficient two-branch architecture, corresponding to producing the te...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.