Pith. sign in

REVIEW 4 major objections 5 minor 250 references

Video Understanding by Design: How Datasets Shape Video Models

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read This survey argues that the structure of video datasets—motion complexity, temporal span, compositionality, and multimodal richness—is the principal force shaping model architecture, making dataset design a strategic lever for the field.

desk verdict A useful survey with a serious internal contradiction: its own benchmark table refutes its central claim that early 3D CNNs dominate short-clip datasets. read the letter →

arxiv 2509.09151 v2 pith:FKHX2A4F submitted 2025-09-11 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords videounderstandingdataset-centricanalysisinductivebiasarchitectureevolutionactionrecognitiontransformersvision-languagemodelsbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish a single organizing claim: video-understanding architectures do not evolve on their own; they are responses to the structural properties of the datasets the field builds and adopts. Four properties—motion amplitude, temporal span, compositional/hierarchical structure, and multimodal richness—act as structural pressures that favor specific inductive biases, from two-stream CNNs and 3D convolutions for short motion clips, to temporal reasoning networks and transformers for long procedural sequences, to vision-language models for text-paired corpora. A sympathetic reader should care because this reframes dataset construction as an active design choice that determines which kinds of video intelligence become possible, and it offers a matching rule: pick the architecture whose inductive bias fits the dataset's structure. The paper supports the claim with a large comparative table of datasets and benchmarks spanning two decades.

What carries the argument

The carrying mechanism is the dataset-bias-architecture framework. It breaks video datasets into four structural properties—motion amplitude, temporal span, compositionality/hierarchy, and multimodal richness (plus agent density)—and treats every architecture as an enforced inductive bias that matches (or mismatches) those properties. Table II operationalizes the framework as a compact rating system (H/M/L for amplitude and agents, S/M/L for span, -/C/H for composition) across a century of datasets; Tables III and IV connect those ratings to measured performance of representative models. The framework does the explanatory work: it turns the history of video understanding into a sequence of d

What would settle it

Run a controlled study that keeps the architecture family fixed, varies a single dataset attribute (e.g., temporal span while holding motion amplitude constant), and shows no systematic performance ordering; alternatively, find two datasets with identical attribute ratings that produced very different dominant architectures. Either result would break the claimed causal link.

Watch

Extended reading notes

Core claim

The central discovery is that datasets operate as inductive-bias generators. Each dataset imposes invariances its model must internalize: coarse high-amplitude motions reward instantaneous motion capture (optical flow, shallow 3D filters); long-horizon, overlapping activities reward temporal memory and hierarchy; multi-agent scenes reward relational or graph representations; and video-text corpora reward cross-modal alignment. On this reading, the milestone trajectory—two-stream networks, 3D CNNs, temporal segment/relation networks, transformers, masked self-supervised models, and video-language foundation models—is not a random succession of fashions but a systematic accommodation of increa

Load-bearing premise

The load-bearing premise is that the four hand-selected dataset attributes are the dominant cause of architectural change—rather than compute availability, leaderboard incentives, or model-family trends—and that the paper's H/M/L ratings of each dataset are accurate and sufficient.

Editorial extensions

If this is right

  • Matching architecture to dataset structure pays off: short-clip motion datasets favor two-stream and 3D CNN models, compositional and interaction-heavy datasets favor sequential and transformer models, and text-paired corpora favor video-language pretraining.
  • Training on coarse, motion-only datasets yields fragile transfer; if robustness in nuanced real-world settings is desired, motion granularity must appear in the data.
  • Simply scaling class counts or clip counts will not yield general video intelligence; the decisive ingredient is structure—procedural hierarchies, temporal continuity, and precise cross-modal alignment.
  • Future architectures should integrate temporal precision, hierarchical composition, long-horizon attention, and multimodal grounding; future datasets should be built with sub-second audio-text alignment, multi-agent annotations, and compositional evaluation splits.
  • Dataset design should be treated as a strategic lever, not a scaling exercise, because datasets generate the invariance pressures that architectures evolve to accommodate.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: The causal arrow (datasets shape architectures) could be tested directly by controlled experiments that fix the architecture family and vary one structural attribute at a time; the survey does not run such ablations, so the claim remains an interpretation of correlated historical patterns.
  • Editorial inference: If the framework holds, it predicts that next-generation long-horizon, multi-agent, multimodal corpora will push the field toward memory-augmented and state-space models plus retrieval-augmented video understanding, since those structures directly target temporal span and compositionality.
  • Editorial inference: The authors' H/M/L ratings are assigned by hand; a community-validated or automatically computed scoring of dataset attributes would let the framework serve as a reusable diagnostic tool for predicting which architecture family suits any new benchmark.
  • Editorial inference: The same lens could be applied prospectively during dataset construction: deliberately vary the four attributes to probe whether an architecture's inductive bias is genuinely being challenged, rather than relying on leaderboard rankings alone.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This survey argues that the evolution of video understanding architectures is fundamentally shaped by dataset structure. It introduces a 'dataset-bias-architecture' framework in which four dataset properties—motion amplitude, temporal span, compositionality/hierarchy, and multimodal richness—impose inductive biases that drive architectural choices. The paper organizes major datasets into these categories (Table II), reviews milestones from two-stream CNNs and 3D CNNs through transformers and video-language models, and uses curated benchmark tables (Tables III and IV) to claim that dataset properties predict which model families succeed. It concludes with a prescriptive roadmap for aligning model design with dataset structure and for constructing future datasets.

Significance. If the central causal claim—that datasets are 'the principal structural force shaping model design'—were established, the survey would offer a useful organizing perspective for a fragmented literature. The paper has clear strengths: Table II is a broad compendium of datasets with structural annotations, the coverage of modern video-language and egocentric datasets is current, the authors provide code and dynamic visualizations, and the roadmap in Section V is actionable. However, the significance is currently undermined by the fact that the empirical support consists of selectively curated benchmark numbers and hand-assigned dataset ratings, with no controlled comparison. The central claim is plausible as a retrospective narrative but is not demonstrated at the strength asserted.

major comments (4)
  1. [Section IV.A, Table III] The text states that 'On HMDB51 and UCF101, early Two-Stream variants and 3D CNNs consistently outperform others,' citing Two-Stream'16 (69.2/93.5) and RGB-I3D (74.8/95.6). The same Table III lists VideoMAE V2 at 88.1 and 99.6 on these datasets, and InternVideo at 89.3 on HMDB51. These are 15–20 points higher, so the sentence is factually contradicted by the paper's own table. The table caption also says 'the best-performing model variant is reported,' meaning rows are not comparable under a fixed protocol: models differ in pretraining data, compute, input sampling, and evaluation settings. This contradiction undermines the inference that short-clip datasets 'strongly favor' two-stream/3D CNNs. A similar issue appears in Section IV.B, where early backbones are said to 'consistently excel' in temporal localization, while Table IV lists InternVideo2 at 72.0 mAP on THUMOS'14 versus I3D+Flow
  2. [Sections III.A and V.B] The central thesis that datasets are 'the principal structural force shaping model design' is asserted on the basis of a correlation between hand-assigned dataset attributes (Table II) and selected architecture successes (Tables III–IV). No attempt is made to hold constant model scale, pretraining data, compute budget, or evaluation protocol, so the observed alignment is equally consistent with compute-driven or pretraining-driven evolution. Moreover, Section V.A itself acknowledges that 'evaluation fragmentation' and leaderboard incentives 'shape architectural incentives'—a non-dataset confound. The causal claim is therefore not supported by the presented evidence. Please either weaken the claim to 'an important and underexamined influence' or provide a more rigorous argument, e.g., a historical timeline showing architecture transitions following dataset releases, or citations to ablati
  3. [Table II] The structural ratings (Amp/Span/Comp/Agents) are central to the framework, but they are assigned without an explicit rubric, operational definitions, inter-annotator agreement, or sensitivity analysis. For example, Kinetics-400 is rated Amp=H, Span=S, Comp=-, Agents=M, but the criteria for these levels are not given, and a different researcher could plausibly rate the same dataset differently. Because these ratings are used to support the paper's main narrative, their subjectivity is load-bearing. Please provide a coding protocol, report reliability, or explicitly relabel the ratings as informal and reduce their role in the causal argument.
  4. [Tables III and IV] The selective reporting in Tables III and IV makes it difficult to interpret 'dashes' as capabilities. For example, VideoMAE V2 has no retrieval or QA entries in Table IV, and InternVideo2 lacks QA entries, yet the text interprets such absences as evidence that certain model families are specialized or limited. A model with no reported number may simply have not been evaluated on that benchmark. The tables should include a completeness statement or a reference to the original papers' evaluation suites, and the text should avoid reading missing entries as negative evidence.
minor comments (5)
  1. [Section IV.A and References] The model 'Two-Stream'16' is cited as [184] (Feichtenhofer et al., 2016), but the original two-stream architecture is [38] (Simonyan and Zisserman, 2014). Please clarify the naming to avoid confusion between the two papers.
  2. [Table III] The row labeled 'Swin' refers to Video Swin Transformer [239]; using the unqualified name may be confused with the image Swin Transformer. Please rename to 'Video Swin'.
  3. [Table II] The UCF101-24 dataset is listed with year 2024, but UCF101-24 is a subset of UCF101 with spatio-temporal annotations and is much older. Please correct the year or clarify the provenance.
  4. [Front matter] The arXiv abstract uses the title 'Video Understanding by Design: How Datasets Shape Video Models,' while the manuscript header title is '... How Datasets Shape Architectures and Insights.' Please unify the title and abstract wording.
  5. [Figure 5] The caption says 'Images adopted from [180]'; for a journal submission please confirm that permission or license for reuse is obtained and that the source is clearly credited.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's dataset-centric synthesis is an independent reading of external benchmark results, not a derivation from its own inputs.

full rationale

This is a survey with no equations, fitted parameters, or uniqueness theorems, so the classic failure modes (self-definitional equations, fitted inputs renamed as predictions, ansatz smuggled via citation) do not apply. The central claim—that dataset structure (motion complexity, temporal span, compositionality, multimodal richness) shaped architecture evolution—is supported by Table II's historical categorization and Tables III–IV, which compile externally published benchmark numbers. Those tables are not derived from the framework; they are independent evidence. The authors' self-citations (e.g., refs [4]–[13]) are contextual and not load-bearing for the dataset-centric thesis. The paper even acknowledges in Section V.A that evaluation fragmentation and leaderboard tuning shape architectural incentives, which weakens the causal exclusivity of the central claim but is a correctness/evidential concern, not circularity. Table III's 'best-performing model variant is reported' protocol may bias the qualitative reading, but selecting published results is not the same as fitting a parameter to the data being predicted. No specific reduction of a conclusion to an input by construction can be exhibited, so no circular step is identified.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The paper introduces a conceptual entity, 'structural pressure', but it is a framing device rather than a physical or mathematical entity with independent falsifiable evidence. The load-bearing assumptions are the causal arrow from datasets to architectures and the validity of the hand-coded dataset attribute ratings and compiled benchmark numbers.

free parameters (1)
  • Hand-assigned dataset attribute ratings (motion amplitude, temporal span, compositionality, agent density) in Table II = H/M/L per dataset, assigned by authors
    These categorical ratings are chosen without an explicit measurement protocol and are used to support the dataset-driven narrative. They are not fitted numbers but are hand-set values that the framework depends on.
assumptions (3)
  • domain assumption The four structural pressures (motion complexity, temporal span, hierarchical structure, multimodal richness) are the dominant forces driving video model evolution, and they exclude class distribution, labeling schemes, and collection biases (footnote 1, Section I).
    This definition is the foundation of the framework and is asserted without empirical validation. The choice of which properties count as 'structural' is arbitrary and excludes alternative causes.
  • domain assumption Datasets induce inductive biases in architectures (Section III.A), meaning the causal arrow runs from data to model design.
    The survey assumes causality from dataset structure to architectural innovation. Alternative explanations such as compute growth, benchmark leaderboards, and theoretical insights are not controlled or refuted.
  • domain assumption Benchmark numbers compiled from different papers are accurate and mutually comparable (Tables III and IV).
    The empirical support relies on reported numbers from heterogeneous evaluation protocols. The paper does not standardize splits, input sizes, or pretraining, and it does not verify the numbers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Video Understanding by Design: How Datasets Shape Video Models." pith.science (2026). https://pith.science/paper/FKHX2A4F

@misc{pith2026250909151,
  author       = {Pith},
  title        = {Pith review of: Video Understanding by Design: How Datasets Shape Video Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FKHX2A4F}},
  note         = {Machine review of arXiv:2509.09151}
}
read the original abstract

Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures. While existing surveys typically organize progress by tasks, benchmarks, or model families, they provide limited insight into why particular architectures emerged and succeeded. In this survey, we argue that the evolution of video understanding is fundamentally shaped by dataset structure. We present a dataset-centric perspective that connects dataset structure, inductive biases, and architectural design within a unified framework. We show that different datasets require models to capture specific invariances and capabilities, such as robustness to viewpoint changes, sensitivity to temporal ordering, reasoning over long-range dependencies, relational interactions, and cross-modal alignment. These requirements naturally give rise to inductive biases, i.e., architectural assumptions that favor particular patterns of reasoning and generalization. From this perspective, milestone architectures, including two-stream networks, 3D CNNs, temporal models, transformers, graph-based methods, and multimodal foundation models, can be understood as architectural responses to the challenges posed by evolving datasets. Building on this framework, we systematically analyze how dataset characteristics have shaped architectural innovation across video understanding tasks and discuss the representational biases induced by different data regimes. By unifying datasets, inductive biases, and architectures into a coherent perspective, this survey offers both a retrospective explanation of the field's evolution and a forward-looking roadmap toward general-purpose video understanding systems. Code and dynamic video visualizations of dataset-induced biases are available at https://time.griffith.edu.au/paper-sites/video-understanding/.

Figures

Figures reproduced from arXiv: 2509.09151 by the authors.

Figure 1
Figure 1. Datasets as structural lenses. Key attributes, motion [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Motion complexity across datasets. UCF101 (top) [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Hierarchical and compositional structures in video datasets. (a) Kinetics-400: [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Examples of egocentric and long-horizon video datasets highlighting temporal and procedural complexity. (a) [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Koala-36M illustrates the power of large-scale, [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

250 extracted references · 33 linked inside Pith

  1. [1]

    Large-scale video classification with convolutional neural networks,

    A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” inProceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2014, pp. 1725–1732

  2. [2]

    Learning spatiotemporal features with 3d convolutional networks,

    D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” inProceedings of the IEEE international conference on computer vision, 2015, pp. 4489–4497

  3. [3]

    Slowfast networks for video recognition,

    C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 6202–6211

  4. [4]

    Motion meets attention: Video motion prompts,

    Q. Chen, L. Wang, P. Koniusz, and T. Gedeon, “Motion meets attention: Video motion prompts,” inAsian Conference on Machine Learning. PMLR, 2025, pp. 591–606. 15

  5. [5]

    Taylor videos for action recognition,

    L. Wang, X. Yuan, T. Gedeon, and L. Zheng, “Taylor videos for action recognition,” inProceedings of the 41st International Conference on Machine Learning, 2024, pp. 52 117–52 133

  6. [6]

    Learnable expansion of graph operators for multi-modal feature fusion,

    D. Ding, L. Wang, L. Zhu, T. Gedeon, and P. Koniusz, “Learnable expansion of graph operators for multi-modal feature fusion,” inThe Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https://openreview.net/forum?id=SMZqIOSdlN

  7. [7]

    Meet jeanie: a similarity measure for 3d skeleton sequences via temporal-viewpoint alignment,

    L. Wang, J. Liu, L. Zheng, T. Gedeon, and P. Koniusz, “Meet jeanie: a similarity measure for 3d skeleton sequences via temporal-viewpoint alignment,”International Journal of Computer Vision, vol. 132, no. 9, pp. 4091–4122, 2024

  8. [8]

    Evolving skeletons: Motion dynamics in action recognition,

    J. Qiu and L. Wang, “Evolving skeletons: Motion dynamics in action recognition,” inCompanion Proceedings of the ACM on Web Confer- ence 2025, 2025, pp. 1916–1937

Show all 250 references
  1. [9]

    Feature hallucination for self-supervised action recognition,

    L. Wang and P. Koniusz, “Feature hallucination for self-supervised action recognition,”International Journal of Computer Vision, 2025

  2. [10]

    Do language models understand time?

    X. Ding and L. Wang, “Do language models understand time?” in Companion Proceedings of the ACM on Web Conference 2025, 2025, pp. 1855–1868

  3. [11]

    Quo vadis, anomaly detection? llms and vlms in the spotlight,

    ——, “Quo vadis, anomaly detection? llms and vlms in the spotlight,” arXiv preprint arXiv:2412.18298, 2024

  4. [12]

    The journey of action recognition,

    ——, “The journey of action recognition,” inCompanion Proceedings of the ACM on Web Conference 2025, 2025, pp. 1869–1884

  5. [13]

    Representation-centric survey of skeletal action recognition and the anubis benchmark,

    Y. Liu, J. Yang, M. Perera, P. Ji, D. Kim, M. Xu, T. Wang, S. Anwar, T. Gedeon, L. Wanget al., “Representation-centric survey of skeletal action recognition and the anubis benchmark,”CoRR, 2025

  6. [14]

    Foundation models for video understanding: A survey,

    N. Madan, A. Møgelmose, R. Modi, Y. S. Rawat, and T. B. Moeslund, “Foundation models for video understanding: A survey,”arXiv preprint arXiv:2405.03770, 2024

  7. [15]

    Video understanding with large language models: A survey,

    Y. Tang, J. Bi, S. Xu, L. Song, S. Liang, T. Wang, D. Zhang, J. An, J. Lin, R. Zhuet al., “Video understanding with large language models: A survey,”IEEE Transactions on Circuits and Systems for Video Technology, 2025

  8. [16]

    The kinetics human action video dataset,

    W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijaya- narasimhan, F. Viola, T. Green, T. Back, P. Natsevet al., “The kinetics human action video dataset,”arXiv preprint arXiv:1705.06950, 2017

  9. [17]

    The” something something

    R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitaget al., “The” something something” video database for learning and evaluating visual common sense,” inProceedings of the IEEE international conf...

  10. [18]

    Activi- tynet: A large-scale video benchmark for human activity understanding,

    F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, “Activi- tynet: A large-scale video benchmark for human activity understanding,” inProceedings of the ieee conference on computer vision and pattern recognition, 2015, pp. 961–970

  11. [19]

    Hollywood in homes: Crowdsourcing data collection for activity understanding,

    G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, “Hollywood in homes: Crowdsourcing data collection for activity understanding,” inEuropean conference on computer vision. Springer, 2016, pp. 510–526

  12. [20]

    Charades-ego: A large-scale dataset of paired third and first person videos,

    G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari, “Charades-ego: A large-scale dataset of paired third and first person videos,”arXiv preprint arXiv:1804.09626, 2018

  13. [21]

    Ava: A video dataset of spatio-temporally localized atomic visual actions,

    C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankaret al., “Ava: A video dataset of spatio-temporally localized atomic visual actions,” in Proceedings of the IEEE conference on computer vision and pattern recog...

  14. [22]

    Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,

    D. Damen, H. Doughty, G. M. Farinella, A. Furnari, E. Kazakos, J. Ma, D. Moltisanti, J. Munro, T. Perrett, W. Priceet al., “Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,”International Journal of Computer Vision, vol. 130, no. 1, pp. 3...

  15. [23]

    Graph based skeleton motion representation and similarity measurement for action recognition,

    P. Wang, C. Yuan, W. Hu, B. Li, and Y. Zhang, “Graph based skeleton motion representation and similarity measurement for action recognition,” inEuropean conference on computer vision. Springer, 2016, pp. 370–385

  16. [24]

    Quo vadis, action recognition? a new model and the kinetics dataset,

    J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” inproceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 6299–6308

  17. [25]

    Non-local neural networks,

    X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7794–7803

  18. [26]

    Spatial temporal graph convolutional networks for skeleton-based action recognition,

    S. Yan, Y. Xiong, and D. Lin, “Spatial temporal graph convolutional networks for skeleton-based action recognition,” inProceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018

  19. [27]

    Vivit: A video vision transformer,

    A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lu ˇci´c, and C. Schmid, “Vivit: A video vision transformer,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 6836–6846

  20. [28]

    Is space-time attention all you need for video understanding?

    G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” inIcml, vol. 2, no. 3, 2021, p. 4

  21. [29]

    Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,

    Z. Tong, Y. Song, J. Wang, and L. Wang, “Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,”Advances in neural information processing systems, vol. 35, pp. 10 078–10 093, 2022

  22. [30]

    Internvideo: General video foundation models via gen- erative and discriminative learning,

    Y. Wang, K. Li, Y. Li, Y. He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y. Liu, Z. Wanget al., “Internvideo: General video foundation models via gen- erative and discriminative learning,”arXiv preprint arXiv:2212.03191, 2022

  23. [31]

    Videomae v2: Scaling video masked autoencoders with dual masking,

    L. Wang, B. Huang, Z. Zhao, Z. Tong, Y. He, Y. Wang, Y. Wang, and Y. Qiao, “Videomae v2: Scaling video masked autoencoders with dual masking,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 14 549–14 560

  24. [32]

    Internvideo2: Scaling foundation models for multimodal video understanding,

    Y. Wang, K. Li, X. Li, J. Yu, Y. He, G. Chen, B. Pei, R. Zheng, Z. Wang, Y. Shiet al., “Internvideo2: Scaling foundation models for multimodal video understanding,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 396–416

  25. [33]

    Ucf101: A dataset of 101 human actions classes from videos in the wild,

    K. Soomro, A. R. Zamir, and M. Shah, “Ucf101: A dataset of 101 human actions classes from videos in the wild,”arXiv preprint arXiv:1212.0402, 2012

  26. [34]

    Hmdb: a large video database for human motion recognition,

    H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “Hmdb: a large video database for human motion recognition,” in2011 International conference on computer vision. IEEE, 2011, pp. 2556–2563

  27. [35]

    A spatio-temporal descriptor based on 3d-gradients,

    A. Klaser, M. Marsza lek, and C. Schmid, “A spatio-temporal descriptor based on 3d-gradients,” inBMVC 2008-19th British machine vision conference. British Machine Vision Association, 2008, pp. 275–1

  28. [36]

    Action recognition with improved trajectories,

    H. Wang and C. Schmid, “Action recognition with improved trajectories,” inProceedings of the IEEE international conference on computer vision, 2013, pp. 3551–3558

  29. [37]

    Scaling egocentric vision: The epic-kitchens dataset,

    D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Priceet al., “Scaling egocentric vision: The epic-kitchens dataset,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 720–736

  30. [38]

    Two-stream convolutional networks for action recognition in videos,

    K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,”Advances in neural information processing systems, vol. 27, 2014

  31. [39]

    Omnivl: One foundation model for image-language and video-language tasks,

    J. Wang, D. Chen, Z. Wu, C. Luo, L. Zhou, Y. Zhao, Y. Xie, C. Liu, Y.-G. Jiang, and L. Yuan, “Omnivl: One foundation model for image-language and video-language tasks,”Advances in neural information processing systems, vol. 35, pp. 5696–5710, 2022

  32. [40]

    Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,

    J. Wu, M. Zhong, S. Xing, Z. Lai, Z. Liu, Z. Chen, W. Wang, X. Zhu, L. Lu, T. Luet al., “Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,”Advances in Neural Information Processing Systems, vol. 37, pp. 69 925–69 975, 2024

  33. [41]

    Timechat: A time-sensitive multimodal large language model for long video understanding,

    S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, “Timechat: A time-sensitive multimodal large language model for long video understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14 313–14 323

  34. [42]

    Videollama 3: Frontier multimodal foundation models for image and video understanding,

    B. Zhang, K. Li, Z. Cheng, Z. Hu, Y. Yuan, G. Chen, S. Leng, Y. Jiang, H. Zhang, X. Liet al., “Videollama 3: Frontier multimodal foundation models for image and video understanding,”arXiv preprint arXiv:2501.13106, 2025

  35. [43]

    Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,

    J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi, “Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26...

  36. [44]

    Masked feature prediction for self-supervised visual pre-training,

    C. Wei, H. Fan, S. Xie, C.-Y. Wu, A. Yuille, and C. Feichtenhofer, “Masked feature prediction for self-supervised visual pre-training,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 14 668–14 678

  37. [45]

    Transductive zero-shot action recog- nition by word-vector embedding,

    X. Xu, T. Hospedales, and S. Gong, “Transductive zero-shot action recog- nition by word-vector embedding,”International Journal of Computer Vision, vol. 123, no. 3, pp. 309–333, 2017

  38. [46]

    Out-of-distribution detection for generalized zero-shot action recognition,

    D. Mandal, S. Narayan, S. K. Dwivedi, V. Gupta, S. Ahmed, F. S. Khan, and L. Shao, “Out-of-distribution detection for generalized zero-shot action recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 9985–9993

  39. [47]

    Few-shot action recognition with permutation-invariant attention,

    H. Zhang, L. Zhang, X. Qi, H. Li, P. H. Torr, and P. Koniusz, “Few-shot action recognition with permutation-invariant attention,” inEuropean conference on computer vision. Springer, 2020, pp. 525–542. 16

  40. [48]

    Actionclip: A new paradigm for video action recognition,

    M. Wang, J. Xing, and Y. Liu, “Actionclip: A new paradigm for video action recognition,”arXiv preprint arXiv:2109.08472, 2021

  41. [49]

    Temporal-relational crosstransformers for few-shot action recognition,

    T. Perrett, A. Masullo, T. Burghardt, M. Mirmehdi, and D. Damen, “Temporal-relational crosstransformers for few-shot action recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 475–484

  42. [50]

    Temporal-viewpoint transportation plan for skeletal few-shot action recognition,

    L. Wang and P. Koniusz, “Temporal-viewpoint transportation plan for skeletal few-shot action recognition,” inProceedings of the Asian conference on computer vision, 2022, pp. 4176–4193

  43. [51]

    Uncertainty-dtw for time series and sequences,

    ——, “Uncertainty-dtw for time series and sequences,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 176–195

  44. [52]

    Reinforced video captioning with entail- ment rewards,

    R. Pasunuru and M. Bansal, “Reinforced video captioning with entail- ment rewards,”arXiv preprint arXiv:1708.02300, 2017

  45. [53]

    Embodied question answering,

    A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra, “Embodied question answering,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 1–10

  46. [54]

    Merlot: Multimodal neural script knowledge models,

    R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi, “Merlot: Multimodal neural script knowledge models,” Advances in neural information processing systems, vol. 34, pp. 23 634– 23 651, 2021

  47. [55]

    Flamingo: a visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynoldset al., “Flamingo: a visual language model for few-shot learning,”Advances in neural information processing systems, vol. 35, pp. 23 716–23 736, 2022

  48. [56]

    Human action recognition from various data modalities: A review,

    Z. Sun, Q. Ke, H. Rahmani, M. Bennamoun, G. Wang, and J. Liu, “Human action recognition from various data modalities: A review,” IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 3, pp. 3200–3225, 2022

  49. [57]

    Video-language understanding: A survey from model architecture, model training, and data perspectives,

    T. Nguyen, Y. Bin, J. Xiao, L. Qu, Y. Li, J. Z. Wu, C.-D. Nguyen, S.-K. Ng, and L. A. Tuan, “Video-language understanding: A survey from model architecture, model training, and data perspectives,”arXiv preprint arXiv:2406.05615, 2024

  50. [58]

    Video question answering: A survey of the state-of-the-art,

    J. P.J. and B. C. Kovoor, “Video question answering: A survey of the state-of-the-art,”J. Vis. Comun. Image Represent., vol. 105, no. C, Dec

  51. [59]

    A survey on generative ai and llm for video generation, understanding, and streaming,

    P. Zhou, L. Wang, Z. Liu, Y. Hao, P. Hui, S. Tarkoma, and J. Kangasharju, “A survey on generative ai and llm for video generation, understanding, and streaming,”arXiv preprint arXiv:2404.16038, 2024

  52. [60]

    Human activity analysis: A review,

    J. K. Aggarwal and M. S. Ryoo, “Human activity analysis: A review,” Acm Computing Surveys (Csur), vol. 43, no. 3, pp. 1–43, 2011

  53. [61]

    Going deeper into action recognition: A survey,

    S. Herath, M. Harandi, and F. Porikli, “Going deeper into action recognition: A survey,”Image and vision computing, vol. 60, pp. 4– 21, 2017

  54. [62]

    A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,

    P. Pareek and A. Thakkar, “A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,” Artificial Intelligence Review, vol. 54, no. 3, pp. 2259–2322, 2021

  55. [63]

    Vision transformers for action recognition: A survey,

    A. Ulhaq, N. Akhtar, G. Pogrebna, and A. Mian, “Vision transformers for action recognition: A survey,”arXiv preprint arXiv:2209.05700, 2022

  56. [64]

    Video transformers: A survey,

    J. Selva, A. S. Johansen, S. Escalera, K. Nasrollahi, T. B. Moeslund, and A. Clap ´es, “Video transformers: A survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 11, pp. 12 922– 12 943, 2023

  57. [65]

    End-to-end learning of visual representations from uncurated instruc- tional videos,

    A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman, “End-to-end learning of visual representations from uncurated instruc- tional videos,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 9879–9889

  58. [66]

    Multimodal learning with transform- ers: A survey,

    P. Xu, X. Zhu, and D. A. Clifton, “Multimodal learning with transform- ers: A survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 10, pp. 12 113–12 132, 2023

  59. [67]

    A survey on human activity recognition from videos,

    T. Subetha and S. Chitrakala, “A survey on human activity recognition from videos,” in2016 international conference on information commu- nication and embedded systems (ICICES). IEEE, 2016, pp. 1–7

  60. [68]

    A comparative review of recent kinect-based action recognition algorithms,

    L. Wang, D. Q. Huynh, and P. Koniusz, “A comparative review of recent kinect-based action recognition algorithms,”IEEE Transactions on Image Processing, vol. 29, pp. 15–28, 2019

  61. [69]

    Skeleton- based action recognition with shift graph convolutional network,

    K. Cheng, Y. Zhang, X. He, W. Chen, J. Cheng, and H. Lu, “Skeleton- based action recognition with shift graph convolutional network,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 183–192

  62. [70]

    Graph convo- lutional neural network for human action recognition: A comprehensive survey,

    T. Ahmad, L. Jin, X. Zhang, S. Lai, G. Tang, and L. Lin, “Graph convo- lutional neural network for human action recognition: A comprehensive survey,”IEEE Transactions on Artificial Intelligence, vol. 2, no. 2, pp. 128–145, 2021

  63. [71]

    A survey on deep learning for skeleton-based human animation,

    L. Mourot, L. Hoyet, F. Le Clerc, F. Schnitzler, and P. Hellier, “A survey on deep learning for skeleton-based human animation,” inComputer Graphics Forum, vol. 41, no. 1. Wiley Online Library, 2022, pp. 122– 157

  64. [72]

    A survey on 3d skeleton-based action recognition using learning method,

    B. Ren, M. Liu, R. Ding, and H. Liu, “A survey on 3d skeleton-based action recognition using learning method,”Cyborg and Bionic Systems, vol. 5, p. 0100, 2024

  65. [73]

    Self-supervised visual feature learning with deep neural networks: A survey,

    L. Jing and Y. Tian, “Self-supervised visual feature learning with deep neural networks: A survey,”IEEE transactions on pattern analysis and machine intelligence, vol. 43, no. 11, pp. 4037–4058, 2020

  66. [74]

    Self- supervised learning: Generative or contrastive,

    X. Liu, F. Zhang, Z. Hou, L. Mian, Z. Wang, J. Zhang, and J. Tang, “Self- supervised learning: Generative or contrastive,”IEEE transactions on knowledge and data engineering, vol. 35, no. 1, pp. 857–876, 2021

  67. [75]

    Self-supervised representation learning: Introduction, advances, and challenges,

    L. Ericsson, H. Gouk, C. C. Loy, and T. M. Hospedales, “Self-supervised representation learning: Introduction, advances, and challenges,”IEEE Signal Processing Magazine, vol. 39, no. 3, pp. 42–62, 2022

  68. [76]

    Deep generative models: Survey,

    A. Oussidi and A. Elhassouny, “Deep generative models: Survey,” in2018 International conference on intelligent systems and computer vision (ISCV). IEEE, 2018, pp. 1–8

  69. [77]

    A survey of multimodal deep generative models,

    M. Suzuki and Y. Matsuo, “A survey of multimodal deep generative models,”Advanced Robotics, vol. 36, no. 5-6, pp. 261–278, 2022

  70. [78]

    Sora as an agi world model? a complete survey on text-to-video generation,

    J. Cho, F. D. Puspitasari, S. Zheng, J. Zheng, L.-H. Lee, T.-H. Kim, C. S. Hong, and C. Zhang, “Sora as an agi world model? a complete survey on text-to-video generation,”arXiv preprint arXiv:2403.05131, 2024

  71. [79]

    A survey on video diffusion models,

    Z. Xing, Q. Feng, H. Chen, Q. Dai, H. Hu, H. Xu, Z. Wu, and Y.-G. Jiang, “A survey on video diffusion models,”ACM Computing Surveys, vol. 57, no. 2, pp. 1–42, 2024

  72. [80]

    Benchmarking a multimodal and multiview and interactive dataset for human action recognition,

    A.-A. Liu, N. Xu, W.-Z. Nie, Y.-T. Su, Y. Wong, and M. Kankanhalli, “Benchmarking a multimodal and multiview and interactive dataset for human action recognition,”IEEE Transactions on cybernetics, vol. 47, no. 7, pp. 1781–1794, 2016

  73. [81]

    Video benchmarks of human action datasets: a review,

    T. Singh and D. K. Vishwakarma, “Video benchmarks of human action datasets: a review,”Artificial Intelligence Review, vol. 52, no. 2, pp. 1107–1154, 2019

  74. [82]

    A review of convolutional-neural-network- based action recognition,

    G. Yao, T. Lei, and J. Zhong, “A review of convolutional-neural-network- based action recognition,”Pattern Recognition Letters, vol. 118, pp. 14–22, 2019

  75. [83]

    Benchmarking micro-action recognition: Dataset, methods, and applications,

    D. Guo, K. Li, B. Hu, Y. Zhang, and M. Wang, “Benchmarking micro-action recognition: Dataset, methods, and applications,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 7, pp. 6238–6252, 2024

  76. [84]

    An outlook into the future of egocentric vision,

    C. Plizzari, G. Goletto, A. Furnari, S. Bansal, F. Ragusa, G. M. Farinella, D. Damen, and T. Tommasi, “An outlook into the future of egocentric vision,”International Journal of Computer Vision, vol. 132, no. 11, pp. 4880–4936, 2024

  77. [85]

    A survey of content-aware video analysis for sports,

    H.-C. Shih, “A survey of content-aware video analysis for sports,”IEEE Transactions on circuits and systems for video technology, vol. 28, no. 5, pp. 1212–1231, 2017

  78. [86]

    Video transcoding: an overview of various techniques and research issues,

    I. Ahmad, X. Wei, Y. Sun, and Y.-Q. Zhang, “Video transcoding: an overview of various techniques and research issues,”IEEE Transactions on multimedia, vol. 7, no. 5, pp. 793–804, 2005

  79. [87]

    Video description: A survey of methods, datasets, and evaluation metrics,

    N. Aafaq, A. Mian, W. Liu, S. Z. Gilani, and M. Shah, “Video description: A survey of methods, datasets, and evaluation metrics,”ACM Computing Surveys (CSUR), vol. 52, no. 6, pp. 1–37, 2019

  80. [88]

    Generative multi- view human action recognition,

    L. Wang, Z. Ding, Z. Tao, Y. Liu, and Y. Fu, “Generative multi- view human action recognition,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 6212–6221

  81. [89]

    Human action recognition and prediction: A survey,

    Y. Kong and Y. Fu, “Human action recognition and prediction: A survey,” International Journal of Computer Vision, vol. 130, no. 5, pp. 1366– 1401, 2022

  82. [90]

    Video generative adversarial networks: a review,

    N. Aldausari, A. Sowmya, N. Marcus, and G. Mohammadi, “Video generative adversarial networks: a review,”ACM Computing Surveys (CSUR), vol. 55, no. 2, pp. 1–25, 2022

  83. [91]

    Search-map-search: a frame selection paradigm for action recognition,

    M. Zhao, Y. Yu, X. Wang, L. Yang, and D. Niu, “Search-map-search: a frame selection paradigm for action recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 10 627–10 636

  84. [92]

    Self-supervised learning for videos: A survey,

    M. C. Schiappa, Y. S. Rawat, and M. Shah, “Self-supervised learning for videos: A survey,”ACM Computing Surveys, vol. 55, no. 13s, pp. 1–37, 2023

  85. [93]

    On space-time interest points,

    I. Laptev, “On space-time interest points,”International journal of computer vision, vol. 64, no. 2, pp. 107–123, 2005

  86. [94]

    Histograms of oriented gradients for human detection,

    N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), vol. 1. Ieee, 2005, pp. 886–893

  87. [95]

    Dense trajectories and motion boundary descriptors for action recognition,

    H. Wang, A. Kl ¨aser, C. Schmid, and C.-L. Liu, “Dense trajectories and motion boundary descriptors for action recognition,”International journal of computer vision, vol. 103, no. 1, pp. 60–79, 2013. 17

  88. [96]

    Transformers in vision: A survey,

    S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,”ACM computing surveys (CSUR), vol. 54, no. 10s, pp. 1–41, 2022

  89. [97]

    Deep reinforcement learning: An overview,

    Y. Li, “Deep reinforcement learning: An overview,”arXiv preprint arXiv:1701.07274, 2017

  90. [98]

    Deep reinforcement learning: A brief survey,

    K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “Deep reinforcement learning: A brief survey,”IEEE signal processing magazine, vol. 34, no. 6, pp. 26–38, 2017

  91. [99]

    Continual lifelong learning with neural networks: A review,

    G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,”Neural networks, vol. 113, pp. 54–71, 2019

  92. [100]

    Federated machine learning: Concept and applications,

    Q. Yang, Y. Liu, T. Chen, and Y. Tong, “Federated machine learning: Concept and applications,”ACM Transactions on Intelligent Systems and Technology (TIST), vol. 10, no. 2, pp. 1–19, 2019

  93. [101]

    A continual learning survey: Defying forgetting in classification tasks,

    M. De Lange, R. Aljundi, M. Masana, S. Parisot, X. Jia, A. Leonardis, G. Slabaugh, and T. Tuytelaars, “A continual learning survey: Defying forgetting in classification tasks,”IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 7, pp. 3366–3385, 2021

  94. [102]

    A comprehensive survey of privacy- preserving federated learning: A taxonomy, review, and future direc- tions,

    X. Yin, Y. Zhu, and J. Hu, “A comprehensive survey of privacy- preserving federated learning: A taxonomy, review, and future direc- tions,”ACM Computing Surveys (CSUR), vol. 54, no. 6, pp. 1–36, 2021

  95. [103]

    Recognizing human actions: a local svm approach,

    C. Schuldt, I. Laptev, and B. Caputo, “Recognizing human actions: a local svm approach,” inProceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004., vol. 3. IEEE, 2004, pp. 32–36

  96. [104]

    Actions as space-time shapes,

    M. Blank, L. Gorelick, E. Shechtman, M. Irani, and R. Basri, “Actions as space-time shapes,” inTenth IEEE International Conference on Computer Vision (ICCV’05) Volume 1, vol. 2. IEEE, 2005, pp. 1395– 1402

  97. [105]

    Free viewpoint action recognition using motion history volumes,

    D. Weinland, R. Ronfard, and E. Boyer, “Free viewpoint action recognition using motion history volumes,”Computer vision and image understanding, vol. 104, no. 2-3, pp. 249–257, 2006

  98. [106]

    Learning realistic human actions from movies,

    I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld, “Learning realistic human actions from movies,” in2008 IEEE conference on computer vision and pattern recognition. IEEE, 2008, pp. 1–8

  99. [107]

    Actions in context,

    M. Marszalek, I. Laptev, and C. Schmid, “Actions in context,” in2009 IEEE conference on computer vision and pattern recognition. IEEE, 2009, pp. 2929–2936

  100. [108]

    What are they doing?: Collective activity classification using spatio-temporal relationship among people,

    W. Choi, K. Shahid, and S. Savarese, “What are they doing?: Collective activity classification using spatio-temporal relationship among people,” in2009 IEEE 12th international conference on computer vision work- shops, ICCV Workshops. IEEE, 2009, pp. 1282–1289

  101. [109]

    Modeling temporal structure of decomposable motion segments for activity classification,

    J. C. Niebles, C.-W. Chen, and L. Fei-Fei, “Modeling temporal structure of decomposable motion segments for activity classification,” inEuro- pean conference on computer vision. Springer, 2010, pp. 392–405

  102. [110]

    Action recognition based on a bag of 3d points,

    W. Li, Z. Zhang, and Z. Liu, “Action recognition based on a bag of 3d points,” in2010 IEEE computer society conference on computer vision and pattern recognition-workshops. IEEE, 2010, pp. 9–14

  103. [111]

    View invariant human action recognition using histograms of 3d joints,

    L. Xia, C.-C. Chen, and J. K. Aggarwal, “View invariant human action recognition using histograms of 3d joints,” in2012 IEEE computer soci- ety conference on computer vision and pattern recognition workshops. IEEE, 2012, pp. 20–27

  104. [112]

    G3d: A gaming action dataset and real time action recognition evaluation framework,

    V. Bloom, D. Makris, and V. Argyriou, “G3d: A gaming action dataset and real time action recognition evaluation framework,” in2012 IEEE Computer society conference on computer vision and pattern recognition workshops. IEEE, 2012, pp. 7–12

  105. [113]

    Recognizing 50 human action categories of web videos,

    K. K. Reddy and M. Shah, “Recognizing 50 human action categories of web videos,”Machine vision and applications, vol. 24, no. 5, pp. 971–981, 2013

  106. [114]

    Recognizing actions from depth cameras as weakly aligned multi-part bag-of-poses,

    L. Seidenari, V. Varano, S. Berretti, A. Bimbo, and P. Pala, “Recognizing actions from depth cameras as weakly aligned multi-part bag-of-poses,” inProceedings of the IEEE conference on computer vision and pattern recognition workshops, 2013, pp. 479–485

  107. [115]

    Towards under- standing action recognition,

    H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black, “Towards under- standing action recognition,” inProceedings of the IEEE international conference on computer vision, 2013, pp. 3192–3199

  108. [116]

    Cross-view action modeling, learning and recognition,

    J. Wang, X. Nie, Y. Xia, Y. Wu, and S.-C. Zhu, “Cross-view action modeling, learning and recognition,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 2649– 2656

  109. [117]

    Ntu rgb+ d: A large scale dataset for 3d human activity analysis,

    A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+ d: A large scale dataset for 3d human activity analysis,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1010–1019

  110. [118]

    Infar dataset: Infrared action recognition at different times,

    C. Gao, Y. Du, J. Liu, J. Lv, L. Yang, D. Meng, and A. G. Hauptmann, “Infar dataset: Infrared action recognition at different times,”Neurocom- puting, vol. 212, pp. 36–47, 2016

  111. [119]

    Thermal imaging based elderly fall detection,

    S. Vadivelu, S. Ganesan, O. R. Murthy, and A. Dhall, “Thermal imaging based elderly fall detection,” inAsian conference on computer vision. Springer, 2016, pp. 541–553

  112. [120]

    Human action localization with sparse spatial supervision,

    P. Weinzaepfel, X. Martin, and C. Schmid, “Human action localization with sparse spatial supervision,”arXiv preprint arXiv:1605.05197, 2016

  113. [121]

    Every moment counts: Dense detailed labeling of actions in complex videos,

    S. Yeung, O. Russakovsky, N. Jin, M. Andriluka, G. Mori, and L. Fei-Fei, “Every moment counts: Dense detailed labeling of actions in complex videos,”International Journal of Computer Vision, vol. 126, no. 2, pp. 375–389, 2018

  114. [122]

    A hierarchical deep temporal model for group activity recognition,

    M. S. Ibrahim, S. Muralidharan, Z. Deng, A. Vahdat, and G. Mori, “A hierarchical deep temporal model for group activity recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1971–1980

  115. [123]

    Need for speed: A benchmark for higher frame rate object tracking,

    H. Kiani Galoogahi, A. Fagg, C. Huang, D. Ramanan, and S. Lucey, “Need for speed: A benchmark for higher frame rate object tracking,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 1125–1134

  116. [124]

    Audio set: An ontology and human- labeled dataset for audio events,

    J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human- labeled dataset for audio events,” in2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2017, p...

  117. [125]

    Attentive spatio-temporal representation learning for diving classification,

    G. Kanojia, S. Kumawat, and S. Raman, “Attentive spatio-temporal representation learning for diving classification,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2019, pp. 0–0

  118. [126]

    Moments in time dataset: one million videos for event understanding,

    M. Monfort, A. Andonian, B. Zhou, K. Ramakrishnan, S. A. Bargal, T. Yan, L. Brown, Q. Fan, D. Gutfreund, C. Vondricket al., “Moments in time dataset: one million videos for event understanding,”IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 2, pp....

  119. [127]

    Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,

    J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot, “Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,”IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 10, pp. 2684–2701, 2019

  120. [128]

    Finegym: A hierarchical video dataset for fine-grained action understanding,

    D. Shao, Y. Zhao, B. Dai, and D. Lin, “Finegym: A hierarchical video dataset for fine-grained action understanding,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2616–2625

  121. [129]

    Vggsound: A large- scale audio-visual dataset,

    H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “Vggsound: A large- scale audio-visual dataset,” inICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 721–725

  122. [130]

    Ava active speaker: An audio-visual dataset for active speaker detection,

    J. Roth, S. Chaudhuri, O. Klejch, R. Marvin, A. Gallagher, L. Kaver, S. Ramaswamy, A. Stopczynski, C. Schmid, Z. Xiet al., “Ava active speaker: An audio-visual dataset for active speaker detection,” inICASSP 2020-2020 IEEE international conference on acoustics, speech and sign...

  123. [131]

    Uav-human: A large benchmark for human behavior understanding with unmanned aerial vehicles,

    T. Li, J. Liu, W. Zhang, Y. Ni, W. Wang, and Z. Li, “Uav-human: A large benchmark for human behavior understanding with unmanned aerial vehicles,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 16 266–16 275

  124. [132]

    End-to-end spatio-temporal action localisation with video transformers,

    A. A. Gritsenko, X. Xiong, J. Djolonga, M. Dehghani, C. Sun, M. Lucic, C. Schmid, and A. Arnab, “End-to-end spatio-temporal action localisation with video transformers,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 18 373–18 383

  125. [133]

    Epic- sounds: A large-scale dataset of actions that sound,

    J. Huh, J. Chalk, E. Kazakos, D. Damen, and A. Zisserman, “Epic- sounds: A large-scale dataset of actions that sound,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  126. [134]

    Berkeley mhad: A comprehensive multimodal human action database,

    F. Ofli, R. Chaudhry, G. Kurillo, R. Vidal, and R. Bajcsy, “Berkeley mhad: A comprehensive multimodal human action database,” in2013 IEEE workshop on applications of computer vision (WACV). IEEE, 2013, pp. 53–60

  127. [135]

    Unstructured human activity detection from rgbd images,

    J. Sung, C. Ponce, B. Selman, and A. Saxena, “Unstructured human activity detection from rgbd images,” in2012 IEEE international conference on robotics and automation. IEEE, 2012, pp. 842–849

  128. [136]

    Learning to recognize daily actions using gaze,

    A. Fathi, Y. Li, and J. M. Rehg, “Learning to recognize daily actions using gaze,” inEuropean Conference on Computer Vision. Springer, 2012, pp. 314–327

  129. [137]

    Learning human activities and object affordances from rgb-d videos,

    H. S. Koppula, R. Gupta, and A. Saxena, “Learning human activities and object affordances from rgb-d videos,”The International journal of robotics research, vol. 32, no. 8, pp. 951–970, 2013

  130. [138]

    Combining embedded accelerometers with computer vision for recognizing food preparation activities,

    S. Stein and S. J. McKenna, “Combining embedded accelerometers with computer vision for recognizing food preparation activities,” inProceedings of the 2013 ACM international joint conference on Pervasive and ubiquitous computing, 2013, pp. 729–738. 18

  131. [139]

    The language of actions: Recovering the syntax and semantics of goal-directed human activities,

    H. Kuehne, A. Arslan, and T. Serre, “The language of actions: Recovering the syntax and semantics of goal-directed human activities,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 780–787

  132. [140]

    Delving into egocentric actions,

    Y. Li, Z. Ye, and J. M. Rehg, “Delving into egocentric actions,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 287–295

  133. [141]

    Jointly learning hetero- geneous features for rgb-d activity recognition,

    J.-F. Hu, W.-S. Zheng, J. Lai, and J. Zhang, “Jointly learning hetero- geneous features for rgb-d activity recognition,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 5344–5352

  134. [142]

    Towards automatic learning of procedures from web instructional videos,

    L. Zhou, C. Xu, and J. Corso, “Towards automatic learning of procedures from web instructional videos,” inProceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018

  135. [143]

    In the eye of beholder: Joint learning of gaze and actions in first person video,

    Y. Li, M. Liu, and J. M. Rehg, “In the eye of beholder: Joint learning of gaze and actions in first person video,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 619–635

  136. [144]

    Weakly-supervised video object grounding from text by loss weighting and object interaction,

    L. Zhou, N. Louis, and J. J. Corso, “Weakly-supervised video object grounding from text by loss weighting and object interaction,”arXiv preprint arXiv:1805.02834, 2018

  137. [145]

    Coin: A large-scale dataset for comprehensive instructional video analysis,

    Y. Tang, D. Ding, Y. Rao, Y. Zheng, D. Zhang, L. Zhao, J. Lu, and J. Zhou, “Coin: A large-scale dataset for comprehensive instructional video analysis,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 1207–1216

  138. [146]

    Cater: A diagnostic dataset for compositional actions and temporal reasoning,

    R. Girdhar and D. Ramanan, “Cater: A diagnostic dataset for compositional actions and temporal reasoning,”arXiv preprint arXiv:1910.04744, 2019

  139. [147]

    Clevrer: Collision events for video representation and reasoning,

    K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum, “Clevrer: Collision events for video representation and reasoning,”arXiv preprint arXiv:1910.01442, 2019

  140. [148]

    Cross-task weakly supervised learning from instructional videos,

    D. Zhukov, J.-B. Alayrac, R. G. Cinbis, D. Fouhey, I. Laptev, and J. Sivic, “Cross-task weakly supervised learning from instructional videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 3537–3545

  141. [149]

    Moma: Multi-object multi-actor activity parsing,

    Z. Luo, W. Xie, S. Kapoor, Y. Liang, M. Cooper, J. C. Niebles, E. Adeli, and F.-F. Li, “Moma: Multi-object multi-actor activity parsing,” Advances in neural information processing systems, vol. 34, pp. 17 939– 17 955, 2021

  142. [150]

    Moma-lrg: Language-refined graphs for multi- object multi-actor activity parsing,

    Z. Luo, Z. Durante, L. Li, W. Xie, R. Liu, E. Jin, Z. Huang, L. Y. Li, J. Wu, J. C. Niebleset al., “Moma-lrg: Language-refined graphs for multi- object multi-actor activity parsing,”Advances in Neural Information Processing Systems, vol. 35, pp. 5282–5298, 2022

  143. [151]

    Detecting activities of daily living in first-person camera views,

    H. Pirsiavash and D. Ramanan, “Detecting activities of daily living in first-person camera views,” in2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012, pp. 2847–2854

  144. [152]

    Mining actionlet ensemble for action recognition with depth cameras,

    J. Wang, Z. Liu, Y. Wu, and J. Yuan, “Mining actionlet ensemble for action recognition with depth cameras,” in2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012, pp. 1290–1297

  145. [153]

    A database for fine grained activity detection of cooking activities,

    M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele, “A database for fine grained activity detection of cooking activities,” in2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012, pp. 1194–1201

  146. [154]

    The thumos challenge on action recognition for videos “in the wild

    H. Idrees, A. R. Zamir, Y.-G. Jiang, A. Gorban, I. Laptev, R. Sukthankar, and M. Shah, “The thumos challenge on action recognition for videos “in the wild”,”Computer Vision and Image Understanding, vol. 155, pp. 1–23, 2017

  147. [155]

    Pku-mmd: A large scale benchmark for continuous multi-modal human action understanding,

    C. Liu, Y. Hu, Y. Li, S. Song, and J. Liu, “Pku-mmd: A large scale benchmark for continuous multi-modal human action understanding,” arXiv preprint arXiv:1703.07475, 2017

  148. [156]

    Soccernet: A scalable dataset for action spotting in soccer videos,

    S. Giancola, M. Amine, T. Dghaily, and B. Ghanem, “Soccernet: A scalable dataset for action spotting in soccer videos,” inProceedings of the IEEE conference on computer vision and pattern recognition workshops, 2018, pp. 1711–1721

  149. [157]

    Hacs: Human action clips and segments dataset for recognition and temporal localization,

    H. Zhao, A. Torralba, L. Torresani, and Z. Yan, “Hacs: Human action clips and segments dataset for recognition and temporal localization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 8668–8678

  150. [158]

    A benchmark dataset and comparison study for multi-modal human action analytics,

    J. Liu, S. Song, C. Liu, Y. Li, and Y. Hu, “A benchmark dataset and comparison study for multi-modal human action analytics,”ACM Trans- actions on Multimedia Computing, Communications, and Applications (TOMM), vol. 16, no. 2, pp. 1–24, 2020

  151. [159]

    Bdd100k: A diverse driving dataset for heterogeneous multitask learning,

    F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, F. Liu, V. Madhavan, and T. Darrell, “Bdd100k: A diverse driving dataset for heterogeneous multitask learning,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2636–2645

  152. [160]

    Collecting highly parallel data for paraphrase evaluation,

    D. Chen and W. B. Dolan, “Collecting highly parallel data for paraphrase evaluation,” inProceedings of the 49th annual meeting of the associa- tion for computational linguistics: human language technologies, 2011, pp. 190–200

  153. [161]

    Msr-vtt: A large video description dataset for bridging video and language,

    J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video description dataset for bridging video and language,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 5288– 5296

  154. [162]

    Movie description,

    A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele, “Movie description,”International Journal of Computer Vision, vol. 123, no. 1, pp. 94–120, 2017

  155. [163]

    Localizing moments in video with natural language,

    L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 5803–5812

  156. [164]

    Dense- captioning events in videos,

    R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, “Dense- captioning events in videos,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 706–715

  157. [165]

    Tgif-qa: Toward spatio- temporal reasoning in visual question answering,

    Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “Tgif-qa: Toward spatio- temporal reasoning in visual question answering,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2758–2766

  158. [166]

    Video question answering via gradually refined attention over appearance and motion,

    D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, “Video question answering via gradually refined attention over appearance and motion,” inProceedings of the 25th ACM international conference on Multimedia, 2017, pp. 1645–1653

  159. [167]

    Tall: Temporal activity local- ization via language query,

    J. Gao, C. Sun, Z. Yang, and R. Nevatia, “Tall: Temporal activity local- ization via language query,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 5267–5275

  160. [168]

    Tvqa: Localized, compositional video question answering,

    J. Lei, L. Yu, M. Bansal, and T. L. Berg, “Tvqa: Localized, compositional video question answering,”arXiv preprint arXiv:1809.01696, 2018

  161. [169]

    Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,

    X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang, “Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 4581–4591

  162. [170]

    Next-qa: Next phase of question-answering to explaining temporal actions,

    J. Xiao, X. Shang, A. Yao, and T.-S. Chua, “Next-qa: Next phase of question-answering to explaining temporal actions,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 9777–9786

  163. [171]

    Agqa: A benchmark for compositional spatio-temporal reasoning,

    M. Grunde-McLaughlin, R. Krishna, and M. Agrawala, “Agqa: A benchmark for compositional spatio-temporal reasoning,” inProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 11 287–11 297

  164. [172]

    Frozen in time: A joint video and image encoder for end-to-end retrieval,

    M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 1728–1738

  165. [173]

    Advancing high-resolution video-language representation with large- scale video transcriptions,

    H. Xue, T. Hang, Y. Zeng, Y. Sun, B. Liu, H. Yang, J. Fu, and B. Guo, “Advancing high-resolution video-language representation with large- scale video transcriptions,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5036–5045

  166. [174]

    Ego4d: Around the world in 3,000 hours of egocentric video,

    K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liuet al., “Ego4d: Around the world in 3,000 hours of egocentric video,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 18 9...

  167. [175]

    Vidchapters-7m: Video chapters at scale,

    A. Yang, A. Nagrani, I. Laptev, J. Sivic, and C. Schmid, “Vidchapters-7m: Video chapters at scale,”Advances in Neural Information Processing Systems, vol. 36, pp. 49 428–49 444, 2023

  168. [176]

    Internvid: A large-scale video-text dataset for multimodal understanding and generation,

    Y. Wang, Y. He, Y. Li, K. Li, J. Yu, X. Ma, X. Li, G. Chen, X. Chen, Y. Wanget al., “Internvid: A large-scale video-text dataset for multimodal understanding and generation,”arXiv preprint arXiv:2307.06942, 2023

  169. [177]

    Panda- 70m: Captioning 70m videos with multiple cross-modality teachers,

    T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H.-w. Chao, B. E. Jeon, Y. Fang, H.-Y. Lee, J. Ren, M.-H. Yanget al., “Panda- 70m: Captioning 70m videos with multiple cross-modality teachers,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...

  170. [178]

    Miradata: A large-scale video dataset with long durations and structured captions,

    X. Ju, Y. Gao, Z. Zhang, Z. Yuan, X. Wang, A. Zeng, Y. Xiong, Q. Xu, and Y. Shan, “Miradata: A large-scale video dataset with long durations and structured captions,”Advances in Neural Information Processing Systems, vol. 37, pp. 48 955–48 970, 2024

  171. [179]

    Openvid-1m: A large-scale high-quality dataset for text-to-video generation,

    K. Nan, R. Xie, P. Zhou, T. Fan, Z. Yang, Z. Chen, X. Li, J. Yang, and Y. Tai, “Openvid-1m: A large-scale high-quality dataset for text-to-video generation,”arXiv preprint arXiv:2407.02371, 2024

  172. [180]

    Koala-36m: A large-scale video dataset 19 improving consistency between fine-grained conditions and video con- tent,

    Q. Wang, Y. Shi, J. Ou, R. Chen, K. Lin, J. Wang, B. Jiang, H. Yang, M. Zheng, X. Taoet al., “Koala-36m: A large-scale video dataset 19 improving consistency between fine-grained conditions and video con- tent,” inProceedings of the Computer Vision and Pattern Recognition Conf...

  173. [181]

    The ava-kinetics localized human actions video dataset,

    A. Li, M. Thotakuri, D. A. Ross, J. Carreira, A. Vostrikov, and A. Zisserman, “The ava-kinetics localized human actions video dataset,” arXiv preprint arXiv:2005.00214, 2020

  174. [182]

    Grounding action descriptions in videos,

    M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal, “Grounding action descriptions in videos,”Transactions of the Association for Computational Linguistics, vol. 1, pp. 25–36, 2013

  175. [183]

    Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,

    A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 2630–2640

  176. [184]

    Convolutional two-stream network fusion for video action recognition,

    C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1933– 1941

  177. [185]

    Temporal segment networks: Towards good practices for deep action recognition,

    L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks: Towards good practices for deep action recognition,” inEuropean conference on computer vision. Springer, 2016, pp. 20–36

  178. [186]

    Anticipative video transformer,

    R. Girdhar and K. Grauman, “Anticipative video transformer,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 13 505–13 515

  179. [187]

    Video modeling with correlation networks,

    H. Wang, D. Tran, L. Torresani, and M. Feiszli, “Video modeling with correlation networks,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 352–361

  180. [188]

    Gate-shift networks for video action recognition,

    S. Sudhakaran, S. Escalera, and O. Lanz, “Gate-shift networks for video action recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1102–1111

  181. [189]

    Sportscap: Monocular 3d human motion capture and fine-grained understanding in challenging sports videos,

    X. Chen, A. Pang, W. Yang, Y. Ma, L. Xu, and J. Yu, “Sportscap: Monocular 3d human motion capture and fine-grained understanding in challenging sports videos,”International Journal of Computer Vision, vol. 129, no. 10, pp. 2846–2864, 2021

  182. [190]

    Learning correlation structures for vision transformers,

    M. Kim, P. H. Seo, C. Schmid, and M. Cho, “Learning correlation structures for vision transformers,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 18 941–18 951

  183. [191]

    Ms-tcn: Multi-stage temporal convolutional network for action segmentation,

    Y. A. Farha and J. Gall, “Ms-tcn: Multi-stage temporal convolutional network for action segmentation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3575– 3584

  184. [192]

    Temporal aggregate representa- tions for long-range video understanding,

    F. Sener, D. Singhania, and A. Yao, “Temporal aggregate representa- tions for long-range video understanding,” inEuropean conference on computer vision. Springer, 2020, pp. 154–171

  185. [193]

    Selective structured state-spaces for long-form video understanding,

    J. Wang, W. Zhu, P. Wang, X. Yu, L. Liu, M. Omar, and R. Hamid, “Selective structured state-spaces for long-form video understanding,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6387–6397

  186. [194]

    Masked autoencoders as spatiotemporal learners,

    C. Feichtenhofer, Y. Li, K. Heet al., “Masked autoencoders as spatiotemporal learners,”Advances in neural information processing systems, vol. 35, pp. 35 946–35 958, 2022

  187. [195]

    Omnivore: A single model for many visual modalities,

    R. Girdhar, M. Singh, N. Ravi, L. Van Der Maaten, A. Joulin, and I. Misra, “Omnivore: A single model for many visual modalities,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16 102–16 112

  188. [196]

    Hiera: A hierarchical vision transformer without the bells-and-whistles,

    C. Ryali, Y.-T. Hu, D. Bolya, C. Wei, H. Fan, P.-Y. Huang, V. Aggarwal, A. Chowdhury, O. Poursaeed, J. Hoffmanet al., “Hiera: A hierarchical vision transformer without the bells-and-whistles,” inInternational conference on machine learning. PMLR, 2023, pp. 29 441–29 454

  189. [197]

    Beyond short snippets: Deep networks for video classification,

    J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici, “Beyond short snippets: Deep networks for video classification,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 4694–4702

  190. [198]

    Learning spatio-temporal representation with pseudo-3d residual networks,

    Z. Qiu, T. Yao, and T. Mei, “Learning spatio-temporal representation with pseudo-3d residual networks,” inproceedings of the IEEE International Conference on Computer Vision, 2017, pp. 5533–5541

  191. [199]

    A closer look at spatiotemporal convolutions for action recognition,

    D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, pp. 6450–6459

  192. [200]

    Long-term feature banks for detailed video understanding,

    C.-Y. Wu, C. Feichtenhofer, H. Fan, K. He, P. Krahenbuhl, and R. Girshick, “Long-term feature banks for detailed video understanding,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 284–293

  193. [201]

    Sequence level semantics aggregation for video object detection,

    H. Wu, Y. Chen, N. Wang, and Z. Zhang, “Sequence level semantics aggregation for video object detection,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9217–9225

  194. [202]

    Epic-fusion: Audio-visual temporal binding for egocentric action recognition,

    E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen, “Epic-fusion: Audio-visual temporal binding for egocentric action recognition,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 5492–5501

  195. [203]

    Timeception for complex action recognition,

    N. Hussein, E. Gavves, and A. W. Smeulders, “Timeception for complex action recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 254–263

  196. [204]

    Videobert: A joint model for video and language representation learning,

    C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 7464–7473

  197. [205]

    Multi-modal domain adaptation for fine- grained action recognition,

    J. Munro and D. Damen, “Multi-modal domain adaptation for fine- grained action recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 122–132

  198. [206]

    Temporal pyramid network for action recognition,

    C. Yang, Y. Xu, J. Shi, B. Dai, and B. Zhou, “Temporal pyramid network for action recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 591–600

  199. [207]

    Multi-modal transformer for video retrieval,

    V. Gabeur, C. Sun, K. Alahari, and C. Schmid, “Multi-modal transformer for video retrieval,” inEuropean Conference on Computer Vision. Springer, 2020, pp. 214–229

  200. [208]

    Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,

    H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y. Cui, and B. Gong, “Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,”Advances in neural information processing systems, vol. 34, pp. 24 206–24 221, 2021

  201. [209]

    Memvit: Memory-augmented multiscale vision trans- former for efficient long-term video recognition,

    C.-Y. Wu, Y. Li, K. Mangalam, H. Fan, B. Xiong, J. Malik, and C. Feichtenhofer, “Memvit: Memory-augmented multiscale vision trans- former for efficient long-term video recognition,” inProceedings of the ieee/cvf conference on computer vision and pattern recognition, 2022, pp. ...

  202. [210]

    Long movie clip classification with state- space video models,

    M. M. Islam and G. Bertasius, “Long movie clip classification with state- space video models,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 87–104

  203. [211]

    Language models with image descriptors are strong few-shot video-language learners,

    Z. Wang, M. Li, R. Xu, L. Zhou, J. Lei, X. Lin, S. Wang, Z. Yang, C. Zhu, D. Hoiemet al., “Language models with image descriptors are strong few-shot video-language learners,”Advances in Neural Information Processing Systems, vol. 35, pp. 8483–8497, 2022

  204. [212]

    Diffusion action segmentation,

    D. Liu, Q. Li, A.-D. Dinh, T. Jiang, M. Shah, and C. Xu, “Diffusion action segmentation,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 10 139–10 149

  205. [213]

    Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,

    A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid, “Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp....

  206. [214]

    Videomamba: State space model for efficient video understanding,

    K. Li, X. Li, Y. Wang, Y. He, Y. Wang, L. Wang, and Y. Qiao, “Videomamba: State space model for efficient video understanding,” in European conference on computer vision. Springer, 2024, pp. 237–255

  207. [215]

    Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,

    B. He, H. Li, Y. K. Jang, M. Jia, X. Cao, A. Shah, A. Shrivastava, and S.-N. Lim, “Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13 504–13 514

  208. [216]

    Mamba-nd: Selective state space modeling for multi-dimensional data,

    S. Li, H. Singh, and A. Grover, “Mamba-nd: Selective state space modeling for multi-dimensional data,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 75–92

  209. [217]

    Video action transformer network,

    R. Girdhar, J. Carreira, C. Doersch, and A. Zisserman, “Video action transformer network,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 244–253

  210. [218]

    X3d: Expanding architectures for efficient video recognition,

    C. Feichtenhofer, “X3d: Expanding architectures for efficient video recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 203–213

  211. [219]

    Spatiotemporal contrastive video representation learning,

    R. Qian, T. Meng, B. Gong, M.-H. Yang, H. Wang, S. Belongie, and Y. Cui, “Spatiotemporal contrastive video representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 6964–6974

  212. [220]

    Un- masked teacher: Towards training-efficient video foundation models,

    K. Li, Y. Wang, Y. Li, Y. Wang, Y. He, L. Wang, and Y. Qiao, “Un- masked teacher: Towards training-efficient video foundation models,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 19 948–19 960

  213. [221]

    Temporal relational reasoning in videos,

    B. Zhou, A. Andonian, A. Oliva, and A. Torralba, “Temporal relational reasoning in videos,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 803–818. 20

  214. [222]

    Videos as space-time region graphs,

    X. Wang and A. Gupta, “Videos as space-time region graphs,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 399–417

  215. [223]

    Multiscale vision transformers,

    H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 6824–6835

  216. [224]

    Videoclip: Contrastive pre-training for zero-shot video-text understanding,

    H. Xu, G. Ghosh, P.-Y. Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer, “Videoclip: Contrastive pre-training for zero-shot video-text understanding,”arXiv preprint arXiv:2109.14084, 2021

  217. [225]

    Videococa: Video-text modeling with zero-shot transfer from contrastive captioners,

    S. Yan, T. Zhu, Z. Wang, Y. Cao, M. Zhang, S. Ghosh, Y. Wu, and J. Yu, “Videococa: Video-text modeling with zero-shot transfer from contrastive captioners,”arXiv preprint arXiv:2212.04979, 2022

  218. [226]

    Violet: End-to-end video-language transformers with masked visual- token modeling,

    T.-J. Fu, L. Li, Z. Gan, K. Lin, W. Y. Wang, L. Wang, and Z. Liu, “Violet: End-to-end video-language transformers with masked visual- token modeling,”arXiv preprint arXiv:2111.12681, 2021

  219. [227]

    Align and prompt: Video-and-language pre-training with entity prompts,

    D. Li, J. Li, H. Li, J. C. Niebles, and S. C. Hoi, “Align and prompt: Video-and-language pre-training with entity prompts,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 4953–4963

  220. [228]

    Video-llama: An instruction-tuned audio-visual language model for video understanding,

    H. Zhang, X. Li, and L. Bing, “Video-llama: An instruction-tuned audio-visual language model for video understanding,”arXiv preprint arXiv:2306.02858, 2023

  221. [229]

    Video-chatgpt: Towards detailed video understanding via large vision and language models,

    M. Maaz, H. Rasheed, S. Khan, and F. S. Khan, “Video-chatgpt: Towards detailed video understanding via large vision and language models,” arXiv preprint arXiv:2306.05424, 2023

  222. [230]

    Human action recognition using factorized spatio-temporal convolutional networks,

    L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi, “Human action recognition using factorized spatio-temporal convolutional networks,” inProceed- ings of the IEEE international conference on computer vision, 2015, pp. 4597–4605

  223. [231]

    Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,

    S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 305–321

  224. [232]

    Tsm: Temporal shift module for efficient video understanding,

    J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 7083–7093

  225. [233]

    Tiny video networks,

    A. Piergiovanni, A. Angelova, and M. S. Ryoo, “Tiny video networks,” Applied AI Letters, vol. 3, no. 1, p. e38, 2022

  226. [234]

    Movinets: Mobile video networks for efficient video recognition,

    D. Kondratyuk, L. Yuan, Y. Li, L. Zhang, M. Tan, M. Brown, and B. Gong, “Movinets: Mobile video networks for efficient video recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 16 020–16 030

  227. [235]

    Eco: Efficient convolutional network for online video understanding,

    M. Zolfaghari, K. Singh, and T. Brox, “Eco: Efficient convolutional network for online video understanding,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 695–712

  228. [236]

    Video transformer network,

    D. Neimark, O. Bar, M. Zohar, and D. Asselmann, “Video transformer network,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 3163–3172

  229. [237]

    Assemblenet: Searching for multi-stream neural connectivity in video architectures,

    M. S. Ryoo, A. Piergiovanni, M. Tan, and A. Angelova, “Assemblenet: Searching for multi-stream neural connectivity in video architectures,” arXiv preprint arXiv:1905.13209, 2019

  230. [238]

    Keeping your eye on the ball: Trajec- tory attention in video transformers,

    M. Patrick, D. Campbell, Y. Asano, I. Misra, F. Metze, C. Feichtenhofer, A. Vedaldi, and J. F. Henriques, “Keeping your eye on the ball: Trajec- tory attention in video transformers,”Advances in neural information processing systems, vol. 34, pp. 12 493–12 506, 2021

  231. [239]

    Video swin transformer,

    Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 3202–3211

  232. [240]

    Mvitv2: Improved multiscale vision transformers for classification and detection,

    Y. Li, C.-Y. Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, and C. Feichtenhofer, “Mvitv2: Improved multiscale vision transformers for classification and detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 4804– 4814

  233. [241]

    Heterogeneous memory enhanced multimodal attention model for video question answering,

    C. Fan, X. Zhang, S. Zhang, W. Wang, C. Zhang, and H. Huang, “Heterogeneous memory enhanced multimodal attention model for video question answering,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 1999–2007

  234. [242]

    Less is more: Clipbert for video-and-language learning via sparse sampling,

    J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu, “Less is more: Clipbert for video-and-language learning via sparse sampling,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 7331–7341

  235. [243]

    Clip4clip: An empirical study of clip for end to end video clip retrieval,

    H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li, “Clip4clip: An empirical study of clip for end to end video clip retrieval,”arXiv preprint arXiv:2104.08860, 2021

  236. [244]

    Zero-shot video question answering via frozen bidirectional language models,

    A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Zero-shot video question answering via frozen bidirectional language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 124–141, 2022

  237. [245]

    All in one: Exploring unified video-language pre- training,

    J. Wang, Y. Ge, R. Yan, Y. Ge, K. Q. Lin, S. Tsutsui, X. Lin, G. Cai, J. Wu, Y. Shanet al., “All in one: Exploring unified video-language pre- training,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6598–6608

  238. [246]

    Videochat: Chat-centric video understanding,

    K. Li, Y. He, Y. Wang, Y. Li, W. Wang, P. Luo, Y. Wang, L. Wang, and Y. Qiao, “Videochat: Chat-centric video understanding,”arXiv preprint arXiv:2305.06355, 2023

  239. [247]

    Llama-adapter: Efficient fine-tuning of language models with zero-init attention,

    R. Zhang, J. Han, C. Liu, P. Gao, A. Zhou, X. Hu, S. Yan, P. Lu, H. Li, and Y. Qiao, “Llama-adapter: Efficient fine-tuning of language models with zero-init attention,”arXiv preprint arXiv:2303.16199, 2023

  240. [248]

    Valley: Video assistant with large language model enhanced ability,

    R. Luo, Z. Zhao, M. Yang, J. Dong, D. Li, P. Lu, T. Wang, L. Hu, M. Qiu, and Z. Wei, “Valley: Video assistant with large language model enhanced ability,”arXiv preprint arXiv:2306.07207, 2023

  241. [249]

    Groundinggpt: Language enhanced multi-modal grounding model,

    Z. Li, Q. Xu, D. Zhang, H. Song, Y. Cai, Q. Qi, R. Zhou, J. Pan, Z. Li, V. T. Vuet al., “Groundinggpt: Language enhanced multi-modal grounding model,”arXiv preprint arXiv:2401.06071, 2024. Lei Wangreceived his M.E. in Software Engineering from the University of Western Austral...

  242. [2024]

    Available: https://doi.org/10.1016/j.jvcir.2024.104320

    [Online]. Available: https://doi.org/10.1016/j.jvcir.2024.104320

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.