REVIEW 4 major objections 5 minor 250 references
Video Understanding by Design: How Datasets Shape Video Models
T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read This survey argues that the structure of video datasets—motion complexity, temporal span, compositionality, and multimodal richness—is the principal force shaping model architecture, making dataset design a strategic lever for the field.
desk verdict A useful survey with a serious internal contradiction: its own benchmark table refutes its central claim that early 3D CNNs dominate short-clip datasets. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the dataset-bias-architecture framework. It breaks video datasets into four structural properties—motion amplitude, temporal span, compositionality/hierarchy, and multimodal richness (plus agent density)—and treats every architecture as an enforced inductive bias that matches (or mismatches) those properties. Table II operationalizes the framework as a compact rating system (H/M/L for amplitude and agents, S/M/L for span, -/C/H for composition) across a century of datasets; Tables III and IV connect those ratings to measured performance of representative models. The framework does the explanatory work: it turns the history of video understanding into a sequence of d
What would settle it
Run a controlled study that keeps the architecture family fixed, varies a single dataset attribute (e.g., temporal span while holding motion amplitude constant), and shows no systematic performance ordering; alternatively, find two datasets with identical attribute ratings that produced very different dominant architectures. Either result would break the claimed causal link.
Extended reading notes
Core claim
The central discovery is that datasets operate as inductive-bias generators. Each dataset imposes invariances its model must internalize: coarse high-amplitude motions reward instantaneous motion capture (optical flow, shallow 3D filters); long-horizon, overlapping activities reward temporal memory and hierarchy; multi-agent scenes reward relational or graph representations; and video-text corpora reward cross-modal alignment. On this reading, the milestone trajectory—two-stream networks, 3D CNNs, temporal segment/relation networks, transformers, masked self-supervised models, and video-language foundation models—is not a random succession of fashions but a systematic accommodation of increa
Load-bearing premise
The load-bearing premise is that the four hand-selected dataset attributes are the dominant cause of architectural change—rather than compute availability, leaderboard incentives, or model-family trends—and that the paper's H/M/L ratings of each dataset are accurate and sufficient.
Editorial extensions
If this is right
- Matching architecture to dataset structure pays off: short-clip motion datasets favor two-stream and 3D CNN models, compositional and interaction-heavy datasets favor sequential and transformer models, and text-paired corpora favor video-language pretraining.
- Training on coarse, motion-only datasets yields fragile transfer; if robustness in nuanced real-world settings is desired, motion granularity must appear in the data.
- Simply scaling class counts or clip counts will not yield general video intelligence; the decisive ingredient is structure—procedural hierarchies, temporal continuity, and precise cross-modal alignment.
- Future architectures should integrate temporal precision, hierarchical composition, long-horizon attention, and multimodal grounding; future datasets should be built with sub-second audio-text alignment, multi-agent annotations, and compositional evaluation splits.
- Dataset design should be treated as a strategic lever, not a scaling exercise, because datasets generate the invariance pressures that architectures evolve to accommodate.
Reading between the lines
- Editorial inference: The causal arrow (datasets shape architectures) could be tested directly by controlled experiments that fix the architecture family and vary one structural attribute at a time; the survey does not run such ablations, so the claim remains an interpretation of correlated historical patterns.
- Editorial inference: If the framework holds, it predicts that next-generation long-horizon, multi-agent, multimodal corpora will push the field toward memory-augmented and state-space models plus retrieval-augmented video understanding, since those structures directly target temporal span and compositionality.
- Editorial inference: The authors' H/M/L ratings are assigned by hand; a community-validated or automatically computed scoring of dataset attributes would let the framework serve as a reusable diagnostic tool for predicting which architecture family suits any new benchmark.
- Editorial inference: The same lens could be applied prospectively during dataset construction: deliberately vary the four attributes to probe whether an architecture's inductive bias is genuinely being challenged, rather than relying on leaderboard rankings alone.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey argues that the evolution of video understanding architectures is fundamentally shaped by dataset structure. It introduces a 'dataset-bias-architecture' framework in which four dataset properties—motion amplitude, temporal span, compositionality/hierarchy, and multimodal richness—impose inductive biases that drive architectural choices. The paper organizes major datasets into these categories (Table II), reviews milestones from two-stream CNNs and 3D CNNs through transformers and video-language models, and uses curated benchmark tables (Tables III and IV) to claim that dataset properties predict which model families succeed. It concludes with a prescriptive roadmap for aligning model design with dataset structure and for constructing future datasets.
Significance. If the central causal claim—that datasets are 'the principal structural force shaping model design'—were established, the survey would offer a useful organizing perspective for a fragmented literature. The paper has clear strengths: Table II is a broad compendium of datasets with structural annotations, the coverage of modern video-language and egocentric datasets is current, the authors provide code and dynamic visualizations, and the roadmap in Section V is actionable. However, the significance is currently undermined by the fact that the empirical support consists of selectively curated benchmark numbers and hand-assigned dataset ratings, with no controlled comparison. The central claim is plausible as a retrospective narrative but is not demonstrated at the strength asserted.
major comments (4)
- [Section IV.A, Table III] The text states that 'On HMDB51 and UCF101, early Two-Stream variants and 3D CNNs consistently outperform others,' citing Two-Stream'16 (69.2/93.5) and RGB-I3D (74.8/95.6). The same Table III lists VideoMAE V2 at 88.1 and 99.6 on these datasets, and InternVideo at 89.3 on HMDB51. These are 15–20 points higher, so the sentence is factually contradicted by the paper's own table. The table caption also says 'the best-performing model variant is reported,' meaning rows are not comparable under a fixed protocol: models differ in pretraining data, compute, input sampling, and evaluation settings. This contradiction undermines the inference that short-clip datasets 'strongly favor' two-stream/3D CNNs. A similar issue appears in Section IV.B, where early backbones are said to 'consistently excel' in temporal localization, while Table IV lists InternVideo2 at 72.0 mAP on THUMOS'14 versus I3D+Flow
- [Sections III.A and V.B] The central thesis that datasets are 'the principal structural force shaping model design' is asserted on the basis of a correlation between hand-assigned dataset attributes (Table II) and selected architecture successes (Tables III–IV). No attempt is made to hold constant model scale, pretraining data, compute budget, or evaluation protocol, so the observed alignment is equally consistent with compute-driven or pretraining-driven evolution. Moreover, Section V.A itself acknowledges that 'evaluation fragmentation' and leaderboard incentives 'shape architectural incentives'—a non-dataset confound. The causal claim is therefore not supported by the presented evidence. Please either weaken the claim to 'an important and underexamined influence' or provide a more rigorous argument, e.g., a historical timeline showing architecture transitions following dataset releases, or citations to ablati
- [Table II] The structural ratings (Amp/Span/Comp/Agents) are central to the framework, but they are assigned without an explicit rubric, operational definitions, inter-annotator agreement, or sensitivity analysis. For example, Kinetics-400 is rated Amp=H, Span=S, Comp=-, Agents=M, but the criteria for these levels are not given, and a different researcher could plausibly rate the same dataset differently. Because these ratings are used to support the paper's main narrative, their subjectivity is load-bearing. Please provide a coding protocol, report reliability, or explicitly relabel the ratings as informal and reduce their role in the causal argument.
- [Tables III and IV] The selective reporting in Tables III and IV makes it difficult to interpret 'dashes' as capabilities. For example, VideoMAE V2 has no retrieval or QA entries in Table IV, and InternVideo2 lacks QA entries, yet the text interprets such absences as evidence that certain model families are specialized or limited. A model with no reported number may simply have not been evaluated on that benchmark. The tables should include a completeness statement or a reference to the original papers' evaluation suites, and the text should avoid reading missing entries as negative evidence.
minor comments (5)
- [Section IV.A and References] The model 'Two-Stream'16' is cited as [184] (Feichtenhofer et al., 2016), but the original two-stream architecture is [38] (Simonyan and Zisserman, 2014). Please clarify the naming to avoid confusion between the two papers.
- [Table III] The row labeled 'Swin' refers to Video Swin Transformer [239]; using the unqualified name may be confused with the image Swin Transformer. Please rename to 'Video Swin'.
- [Table II] The UCF101-24 dataset is listed with year 2024, but UCF101-24 is a subset of UCF101 with spatio-temporal annotations and is much older. Please correct the year or clarify the provenance.
- [Front matter] The arXiv abstract uses the title 'Video Understanding by Design: How Datasets Shape Video Models,' while the manuscript header title is '... How Datasets Shape Architectures and Insights.' Please unify the title and abstract wording.
- [Figure 5] The caption says 'Images adopted from [180]'; for a journal submission please confirm that permission or license for reuse is obtained and that the source is clearly credited.
Circularity Check
No circularity: the survey's dataset-centric synthesis is an independent reading of external benchmark results, not a derivation from its own inputs.
full rationale
This is a survey with no equations, fitted parameters, or uniqueness theorems, so the classic failure modes (self-definitional equations, fitted inputs renamed as predictions, ansatz smuggled via citation) do not apply. The central claim—that dataset structure (motion complexity, temporal span, compositionality, multimodal richness) shaped architecture evolution—is supported by Table II's historical categorization and Tables III–IV, which compile externally published benchmark numbers. Those tables are not derived from the framework; they are independent evidence. The authors' self-citations (e.g., refs [4]–[13]) are contextual and not load-bearing for the dataset-centric thesis. The paper even acknowledges in Section V.A that evaluation fragmentation and leaderboard tuning shape architectural incentives, which weakens the causal exclusivity of the central claim but is a correctness/evidential concern, not circularity. Table III's 'best-performing model variant is reported' protocol may bias the qualitative reading, but selecting published results is not the same as fitting a parameter to the data being predicted. No specific reduction of a conclusion to an input by construction can be exhibited, so no circular step is identified.
Assumptions & free parameters
free parameters (1)
- Hand-assigned dataset attribute ratings (motion amplitude, temporal span, compositionality, agent density) in Table II =
H/M/L per dataset, assigned by authors
assumptions (3)
- domain assumption The four structural pressures (motion complexity, temporal span, hierarchical structure, multimodal richness) are the dominant forces driving video model evolution, and they exclude class distribution, labeling schemes, and collection biases (footnote 1, Section I).
- domain assumption Datasets induce inductive biases in architectures (Section III.A), meaning the causal arrow runs from data to model design.
- domain assumption Benchmark numbers compiled from different papers are accurate and mutually comparable (Tables III and IV).
Cite this review
Pith. "Pith review of Video Understanding by Design: How Datasets Shape Video Models." pith.science (2026). https://pith.science/paper/FKHX2A4F
@misc{pith2026250909151,
author = {Pith},
title = {Pith review of: Video Understanding by Design: How Datasets Shape Video Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/FKHX2A4F}},
note = {Machine review of arXiv:2509.09151}
}
read the original abstract
Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures. While existing surveys typically organize progress by tasks, benchmarks, or model families, they provide limited insight into why particular architectures emerged and succeeded. In this survey, we argue that the evolution of video understanding is fundamentally shaped by dataset structure. We present a dataset-centric perspective that connects dataset structure, inductive biases, and architectural design within a unified framework. We show that different datasets require models to capture specific invariances and capabilities, such as robustness to viewpoint changes, sensitivity to temporal ordering, reasoning over long-range dependencies, relational interactions, and cross-modal alignment. These requirements naturally give rise to inductive biases, i.e., architectural assumptions that favor particular patterns of reasoning and generalization. From this perspective, milestone architectures, including two-stream networks, 3D CNNs, temporal models, transformers, graph-based methods, and multimodal foundation models, can be understood as architectural responses to the challenges posed by evolving datasets. Building on this framework, we systematically analyze how dataset characteristics have shaped architectural innovation across video understanding tasks and discuss the representational biases induced by different data regimes. By unifying datasets, inductive biases, and architectures into a coherent perspective, this survey offers both a retrospective explanation of the field's evolution and a forward-looking roadmap toward general-purpose video understanding systems. Code and dynamic video visualizations of dataset-induced biases are available at https://time.griffith.edu.au/paper-sites/video-understanding/.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Large-scale video classification with convolutional neural networks,
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” inProceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2014, pp. 1725–1732
2014
-
[2]
Learning spatiotemporal features with 3d convolutional networks,
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” inProceedings of the IEEE international conference on computer vision, 2015, pp. 4489–4497
2015
-
[3]
Slowfast networks for video recognition,
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 6202–6211
2019
-
[4]
Motion meets attention: Video motion prompts,
Q. Chen, L. Wang, P. Koniusz, and T. Gedeon, “Motion meets attention: Video motion prompts,” inAsian Conference on Machine Learning. PMLR, 2025, pp. 591–606. 15
2025
-
[5]
Taylor videos for action recognition,
L. Wang, X. Yuan, T. Gedeon, and L. Zheng, “Taylor videos for action recognition,” inProceedings of the 41st International Conference on Machine Learning, 2024, pp. 52 117–52 133
2024
-
[6]
Learnable expansion of graph operators for multi-modal feature fusion,
D. Ding, L. Wang, L. Zhu, T. Gedeon, and P. Koniusz, “Learnable expansion of graph operators for multi-modal feature fusion,” inThe Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https://openreview.net/forum?id=SMZqIOSdlN
2025
-
[7]
Meet jeanie: a similarity measure for 3d skeleton sequences via temporal-viewpoint alignment,
L. Wang, J. Liu, L. Zheng, T. Gedeon, and P. Koniusz, “Meet jeanie: a similarity measure for 3d skeleton sequences via temporal-viewpoint alignment,”International Journal of Computer Vision, vol. 132, no. 9, pp. 4091–4122, 2024
2024
-
[8]
Evolving skeletons: Motion dynamics in action recognition,
J. Qiu and L. Wang, “Evolving skeletons: Motion dynamics in action recognition,” inCompanion Proceedings of the ACM on Web Confer- ence 2025, 2025, pp. 1916–1937
2025
Show all 250 references
-
[9]
Feature hallucination for self-supervised action recognition,
L. Wang and P. Koniusz, “Feature hallucination for self-supervised action recognition,”International Journal of Computer Vision, 2025
2025
-
[10]
Do language models understand time?
X. Ding and L. Wang, “Do language models understand time?” in Companion Proceedings of the ACM on Web Conference 2025, 2025, pp. 1855–1868
2025
-
[11]
Quo vadis, anomaly detection? llms and vlms in the spotlight,
——, “Quo vadis, anomaly detection? llms and vlms in the spotlight,” arXiv preprint arXiv:2412.18298, 2024
2024 arXiv
-
[12]
The journey of action recognition,
——, “The journey of action recognition,” inCompanion Proceedings of the ACM on Web Conference 2025, 2025, pp. 1869–1884
2025
-
[13]
Representation-centric survey of skeletal action recognition and the anubis benchmark,
Y. Liu, J. Yang, M. Perera, P. Ji, D. Kim, M. Xu, T. Wang, S. Anwar, T. Gedeon, L. Wanget al., “Representation-centric survey of skeletal action recognition and the anubis benchmark,”CoRR, 2025
2025
-
[14]
Foundation models for video understanding: A survey,
N. Madan, A. Møgelmose, R. Modi, Y. S. Rawat, and T. B. Moeslund, “Foundation models for video understanding: A survey,”arXiv preprint arXiv:2405.03770, 2024
2024 arXiv
-
[15]
Video understanding with large language models: A survey,
Y. Tang, J. Bi, S. Xu, L. Song, S. Liang, T. Wang, D. Zhang, J. An, J. Lin, R. Zhuet al., “Video understanding with large language models: A survey,”IEEE Transactions on Circuits and Systems for Video Technology, 2025
2025
-
[16]
The kinetics human action video dataset,
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijaya- narasimhan, F. Viola, T. Green, T. Back, P. Natsevet al., “The kinetics human action video dataset,”arXiv preprint arXiv:1705.06950, 2017
2017 arXiv
-
[17]
The” something something
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitaget al., “The” something something” video database for learning and evaluating visual common sense,” inProceedings of the IEEE international conf...
2017
-
[18]
Activi- tynet: A large-scale video benchmark for human activity understanding,
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, “Activi- tynet: A large-scale video benchmark for human activity understanding,” inProceedings of the ieee conference on computer vision and pattern recognition, 2015, pp. 961–970
2015
-
[19]
Hollywood in homes: Crowdsourcing data collection for activity understanding,
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, “Hollywood in homes: Crowdsourcing data collection for activity understanding,” inEuropean conference on computer vision. Springer, 2016, pp. 510–526
2016
-
[20]
Charades-ego: A large-scale dataset of paired third and first person videos,
G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari, “Charades-ego: A large-scale dataset of paired third and first person videos,”arXiv preprint arXiv:1804.09626, 2018
2018 arXiv
-
[21]
Ava: A video dataset of spatio-temporally localized atomic visual actions,
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankaret al., “Ava: A video dataset of spatio-temporally localized atomic visual actions,” in Proceedings of the IEEE conference on computer vision and pattern recog...
2018
-
[22]
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,
D. Damen, H. Doughty, G. M. Farinella, A. Furnari, E. Kazakos, J. Ma, D. Moltisanti, J. Munro, T. Perrett, W. Priceet al., “Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,”International Journal of Computer Vision, vol. 130, no. 1, pp. 3...
2022
-
[23]
Graph based skeleton motion representation and similarity measurement for action recognition,
P. Wang, C. Yuan, W. Hu, B. Li, and Y. Zhang, “Graph based skeleton motion representation and similarity measurement for action recognition,” inEuropean conference on computer vision. Springer, 2016, pp. 370–385
2016
-
[24]
Quo vadis, action recognition? a new model and the kinetics dataset,
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” inproceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 6299–6308
2017
-
[25]
Non-local neural networks,
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7794–7803
2018
-
[26]
Spatial temporal graph convolutional networks for skeleton-based action recognition,
S. Yan, Y. Xiong, and D. Lin, “Spatial temporal graph convolutional networks for skeleton-based action recognition,” inProceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018
2018
-
[27]
Vivit: A video vision transformer,
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lu ˇci´c, and C. Schmid, “Vivit: A video vision transformer,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 6836–6846
2021
-
[28]
Is space-time attention all you need for video understanding?
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” inIcml, vol. 2, no. 3, 2021, p. 4
2021
-
[29]
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,
Z. Tong, Y. Song, J. Wang, and L. Wang, “Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,”Advances in neural information processing systems, vol. 35, pp. 10 078–10 093, 2022
2022
-
[30]
Internvideo: General video foundation models via gen- erative and discriminative learning,
Y. Wang, K. Li, Y. Li, Y. He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y. Liu, Z. Wanget al., “Internvideo: General video foundation models via gen- erative and discriminative learning,”arXiv preprint arXiv:2212.03191, 2022
2022 arXiv
-
[31]
Videomae v2: Scaling video masked autoencoders with dual masking,
L. Wang, B. Huang, Z. Zhao, Z. Tong, Y. He, Y. Wang, Y. Wang, and Y. Qiao, “Videomae v2: Scaling video masked autoencoders with dual masking,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 14 549–14 560
2023
-
[32]
Internvideo2: Scaling foundation models for multimodal video understanding,
Y. Wang, K. Li, X. Li, J. Yu, Y. He, G. Chen, B. Pei, R. Zheng, Z. Wang, Y. Shiet al., “Internvideo2: Scaling foundation models for multimodal video understanding,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 396–416
2024
-
[33]
Ucf101: A dataset of 101 human actions classes from videos in the wild,
K. Soomro, A. R. Zamir, and M. Shah, “Ucf101: A dataset of 101 human actions classes from videos in the wild,”arXiv preprint arXiv:1212.0402, 2012
2012 arXiv
-
[34]
Hmdb: a large video database for human motion recognition,
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “Hmdb: a large video database for human motion recognition,” in2011 International conference on computer vision. IEEE, 2011, pp. 2556–2563
2011
-
[35]
A spatio-temporal descriptor based on 3d-gradients,
A. Klaser, M. Marsza lek, and C. Schmid, “A spatio-temporal descriptor based on 3d-gradients,” inBMVC 2008-19th British machine vision conference. British Machine Vision Association, 2008, pp. 275–1
2008
-
[36]
Action recognition with improved trajectories,
H. Wang and C. Schmid, “Action recognition with improved trajectories,” inProceedings of the IEEE international conference on computer vision, 2013, pp. 3551–3558
2013
-
[37]
Scaling egocentric vision: The epic-kitchens dataset,
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Priceet al., “Scaling egocentric vision: The epic-kitchens dataset,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 720–736
2018
-
[38]
Two-stream convolutional networks for action recognition in videos,
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,”Advances in neural information processing systems, vol. 27, 2014
2014
-
[39]
Omnivl: One foundation model for image-language and video-language tasks,
J. Wang, D. Chen, Z. Wu, C. Luo, L. Zhou, Y. Zhao, Y. Xie, C. Liu, Y.-G. Jiang, and L. Yuan, “Omnivl: One foundation model for image-language and video-language tasks,”Advances in neural information processing systems, vol. 35, pp. 5696–5710, 2022
2022
-
[40]
Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,
J. Wu, M. Zhong, S. Xing, Z. Lai, Z. Liu, Z. Chen, W. Wang, X. Zhu, L. Lu, T. Luet al., “Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,”Advances in Neural Information Processing Systems, vol. 37, pp. 69 925–69 975, 2024
2024
-
[41]
Timechat: A time-sensitive multimodal large language model for long video understanding,
S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, “Timechat: A time-sensitive multimodal large language model for long video understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14 313–14 323
2024
-
[42]
Videollama 3: Frontier multimodal foundation models for image and video understanding,
B. Zhang, K. Li, Z. Cheng, Z. Hu, Y. Yuan, G. Chen, S. Leng, Y. Jiang, H. Zhang, X. Liet al., “Videollama 3: Frontier multimodal foundation models for image and video understanding,”arXiv preprint arXiv:2501.13106, 2025
2025 arXiv
-
[43]
Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,
J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi, “Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26...
2024
-
[44]
Masked feature prediction for self-supervised visual pre-training,
C. Wei, H. Fan, S. Xie, C.-Y. Wu, A. Yuille, and C. Feichtenhofer, “Masked feature prediction for self-supervised visual pre-training,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 14 668–14 678
2022
-
[45]
Transductive zero-shot action recog- nition by word-vector embedding,
X. Xu, T. Hospedales, and S. Gong, “Transductive zero-shot action recog- nition by word-vector embedding,”International Journal of Computer Vision, vol. 123, no. 3, pp. 309–333, 2017
2017
-
[46]
Out-of-distribution detection for generalized zero-shot action recognition,
D. Mandal, S. Narayan, S. K. Dwivedi, V. Gupta, S. Ahmed, F. S. Khan, and L. Shao, “Out-of-distribution detection for generalized zero-shot action recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 9985–9993
2019
-
[47]
Few-shot action recognition with permutation-invariant attention,
H. Zhang, L. Zhang, X. Qi, H. Li, P. H. Torr, and P. Koniusz, “Few-shot action recognition with permutation-invariant attention,” inEuropean conference on computer vision. Springer, 2020, pp. 525–542. 16
2020
-
[48]
Actionclip: A new paradigm for video action recognition,
M. Wang, J. Xing, and Y. Liu, “Actionclip: A new paradigm for video action recognition,”arXiv preprint arXiv:2109.08472, 2021
2021 arXiv
-
[49]
Temporal-relational crosstransformers for few-shot action recognition,
T. Perrett, A. Masullo, T. Burghardt, M. Mirmehdi, and D. Damen, “Temporal-relational crosstransformers for few-shot action recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 475–484
2021
-
[50]
Temporal-viewpoint transportation plan for skeletal few-shot action recognition,
L. Wang and P. Koniusz, “Temporal-viewpoint transportation plan for skeletal few-shot action recognition,” inProceedings of the Asian conference on computer vision, 2022, pp. 4176–4193
2022
-
[51]
Uncertainty-dtw for time series and sequences,
——, “Uncertainty-dtw for time series and sequences,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 176–195
2022
-
[52]
Reinforced video captioning with entail- ment rewards,
R. Pasunuru and M. Bansal, “Reinforced video captioning with entail- ment rewards,”arXiv preprint arXiv:1708.02300, 2017
2017 arXiv
-
[53]
Embodied question answering,
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra, “Embodied question answering,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 1–10
2018
-
[54]
Merlot: Multimodal neural script knowledge models,
R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi, “Merlot: Multimodal neural script knowledge models,” Advances in neural information processing systems, vol. 34, pp. 23 634– 23 651, 2021
2021
-
[55]
Flamingo: a visual language model for few-shot learning,
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynoldset al., “Flamingo: a visual language model for few-shot learning,”Advances in neural information processing systems, vol. 35, pp. 23 716–23 736, 2022
2022
-
[56]
Human action recognition from various data modalities: A review,
Z. Sun, Q. Ke, H. Rahmani, M. Bennamoun, G. Wang, and J. Liu, “Human action recognition from various data modalities: A review,” IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 3, pp. 3200–3225, 2022
2022
-
[57]
Video-language understanding: A survey from model architecture, model training, and data perspectives,
T. Nguyen, Y. Bin, J. Xiao, L. Qu, Y. Li, J. Z. Wu, C.-D. Nguyen, S.-K. Ng, and L. A. Tuan, “Video-language understanding: A survey from model architecture, model training, and data perspectives,”arXiv preprint arXiv:2406.05615, 2024
2024 arXiv
-
[58]
Video question answering: A survey of the state-of-the-art,
J. P.J. and B. C. Kovoor, “Video question answering: A survey of the state-of-the-art,”J. Vis. Comun. Image Represent., vol. 105, no. C, Dec
-
[59]
A survey on generative ai and llm for video generation, understanding, and streaming,
P. Zhou, L. Wang, Z. Liu, Y. Hao, P. Hui, S. Tarkoma, and J. Kangasharju, “A survey on generative ai and llm for video generation, understanding, and streaming,”arXiv preprint arXiv:2404.16038, 2024
2024 arXiv
-
[60]
Human activity analysis: A review,
J. K. Aggarwal and M. S. Ryoo, “Human activity analysis: A review,” Acm Computing Surveys (Csur), vol. 43, no. 3, pp. 1–43, 2011
2011
-
[61]
Going deeper into action recognition: A survey,
S. Herath, M. Harandi, and F. Porikli, “Going deeper into action recognition: A survey,”Image and vision computing, vol. 60, pp. 4– 21, 2017
2017
-
[62]
A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,
P. Pareek and A. Thakkar, “A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,” Artificial Intelligence Review, vol. 54, no. 3, pp. 2259–2322, 2021
2021
-
[63]
Vision transformers for action recognition: A survey,
A. Ulhaq, N. Akhtar, G. Pogrebna, and A. Mian, “Vision transformers for action recognition: A survey,”arXiv preprint arXiv:2209.05700, 2022
2022 arXiv
-
[64]
Video transformers: A survey,
J. Selva, A. S. Johansen, S. Escalera, K. Nasrollahi, T. B. Moeslund, and A. Clap ´es, “Video transformers: A survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 11, pp. 12 922– 12 943, 2023
2023
-
[65]
End-to-end learning of visual representations from uncurated instruc- tional videos,
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman, “End-to-end learning of visual representations from uncurated instruc- tional videos,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 9879–9889
2020
-
[66]
Multimodal learning with transform- ers: A survey,
P. Xu, X. Zhu, and D. A. Clifton, “Multimodal learning with transform- ers: A survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 10, pp. 12 113–12 132, 2023
2023
-
[67]
A survey on human activity recognition from videos,
T. Subetha and S. Chitrakala, “A survey on human activity recognition from videos,” in2016 international conference on information commu- nication and embedded systems (ICICES). IEEE, 2016, pp. 1–7
2016
-
[68]
A comparative review of recent kinect-based action recognition algorithms,
L. Wang, D. Q. Huynh, and P. Koniusz, “A comparative review of recent kinect-based action recognition algorithms,”IEEE Transactions on Image Processing, vol. 29, pp. 15–28, 2019
2019
-
[69]
Skeleton- based action recognition with shift graph convolutional network,
K. Cheng, Y. Zhang, X. He, W. Chen, J. Cheng, and H. Lu, “Skeleton- based action recognition with shift graph convolutional network,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 183–192
2020
-
[70]
Graph convo- lutional neural network for human action recognition: A comprehensive survey,
T. Ahmad, L. Jin, X. Zhang, S. Lai, G. Tang, and L. Lin, “Graph convo- lutional neural network for human action recognition: A comprehensive survey,”IEEE Transactions on Artificial Intelligence, vol. 2, no. 2, pp. 128–145, 2021
2021
-
[71]
A survey on deep learning for skeleton-based human animation,
L. Mourot, L. Hoyet, F. Le Clerc, F. Schnitzler, and P. Hellier, “A survey on deep learning for skeleton-based human animation,” inComputer Graphics Forum, vol. 41, no. 1. Wiley Online Library, 2022, pp. 122– 157
2022
-
[72]
A survey on 3d skeleton-based action recognition using learning method,
B. Ren, M. Liu, R. Ding, and H. Liu, “A survey on 3d skeleton-based action recognition using learning method,”Cyborg and Bionic Systems, vol. 5, p. 0100, 2024
2024
-
[73]
Self-supervised visual feature learning with deep neural networks: A survey,
L. Jing and Y. Tian, “Self-supervised visual feature learning with deep neural networks: A survey,”IEEE transactions on pattern analysis and machine intelligence, vol. 43, no. 11, pp. 4037–4058, 2020
2020
-
[74]
Self- supervised learning: Generative or contrastive,
X. Liu, F. Zhang, Z. Hou, L. Mian, Z. Wang, J. Zhang, and J. Tang, “Self- supervised learning: Generative or contrastive,”IEEE transactions on knowledge and data engineering, vol. 35, no. 1, pp. 857–876, 2021
2021
-
[75]
Self-supervised representation learning: Introduction, advances, and challenges,
L. Ericsson, H. Gouk, C. C. Loy, and T. M. Hospedales, “Self-supervised representation learning: Introduction, advances, and challenges,”IEEE Signal Processing Magazine, vol. 39, no. 3, pp. 42–62, 2022
2022
-
[76]
Deep generative models: Survey,
A. Oussidi and A. Elhassouny, “Deep generative models: Survey,” in2018 International conference on intelligent systems and computer vision (ISCV). IEEE, 2018, pp. 1–8
2018
-
[77]
A survey of multimodal deep generative models,
M. Suzuki and Y. Matsuo, “A survey of multimodal deep generative models,”Advanced Robotics, vol. 36, no. 5-6, pp. 261–278, 2022
2022
-
[78]
Sora as an agi world model? a complete survey on text-to-video generation,
J. Cho, F. D. Puspitasari, S. Zheng, J. Zheng, L.-H. Lee, T.-H. Kim, C. S. Hong, and C. Zhang, “Sora as an agi world model? a complete survey on text-to-video generation,”arXiv preprint arXiv:2403.05131, 2024
2024
-
[79]
A survey on video diffusion models,
Z. Xing, Q. Feng, H. Chen, Q. Dai, H. Hu, H. Xu, Z. Wu, and Y.-G. Jiang, “A survey on video diffusion models,”ACM Computing Surveys, vol. 57, no. 2, pp. 1–42, 2024
2024
-
[80]
Benchmarking a multimodal and multiview and interactive dataset for human action recognition,
A.-A. Liu, N. Xu, W.-Z. Nie, Y.-T. Su, Y. Wong, and M. Kankanhalli, “Benchmarking a multimodal and multiview and interactive dataset for human action recognition,”IEEE Transactions on cybernetics, vol. 47, no. 7, pp. 1781–1794, 2016
2016
-
[81]
Video benchmarks of human action datasets: a review,
T. Singh and D. K. Vishwakarma, “Video benchmarks of human action datasets: a review,”Artificial Intelligence Review, vol. 52, no. 2, pp. 1107–1154, 2019
2019
-
[82]
A review of convolutional-neural-network- based action recognition,
G. Yao, T. Lei, and J. Zhong, “A review of convolutional-neural-network- based action recognition,”Pattern Recognition Letters, vol. 118, pp. 14–22, 2019
2019
-
[83]
Benchmarking micro-action recognition: Dataset, methods, and applications,
D. Guo, K. Li, B. Hu, Y. Zhang, and M. Wang, “Benchmarking micro-action recognition: Dataset, methods, and applications,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 7, pp. 6238–6252, 2024
2024
-
[84]
An outlook into the future of egocentric vision,
C. Plizzari, G. Goletto, A. Furnari, S. Bansal, F. Ragusa, G. M. Farinella, D. Damen, and T. Tommasi, “An outlook into the future of egocentric vision,”International Journal of Computer Vision, vol. 132, no. 11, pp. 4880–4936, 2024
2024
-
[85]
A survey of content-aware video analysis for sports,
H.-C. Shih, “A survey of content-aware video analysis for sports,”IEEE Transactions on circuits and systems for video technology, vol. 28, no. 5, pp. 1212–1231, 2017
2017
-
[86]
Video transcoding: an overview of various techniques and research issues,
I. Ahmad, X. Wei, Y. Sun, and Y.-Q. Zhang, “Video transcoding: an overview of various techniques and research issues,”IEEE Transactions on multimedia, vol. 7, no. 5, pp. 793–804, 2005
2005
-
[87]
Video description: A survey of methods, datasets, and evaluation metrics,
N. Aafaq, A. Mian, W. Liu, S. Z. Gilani, and M. Shah, “Video description: A survey of methods, datasets, and evaluation metrics,”ACM Computing Surveys (CSUR), vol. 52, no. 6, pp. 1–37, 2019
2019
-
[88]
Generative multi- view human action recognition,
L. Wang, Z. Ding, Z. Tao, Y. Liu, and Y. Fu, “Generative multi- view human action recognition,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 6212–6221
2019
-
[89]
Human action recognition and prediction: A survey,
Y. Kong and Y. Fu, “Human action recognition and prediction: A survey,” International Journal of Computer Vision, vol. 130, no. 5, pp. 1366– 1401, 2022
2022
-
[90]
Video generative adversarial networks: a review,
N. Aldausari, A. Sowmya, N. Marcus, and G. Mohammadi, “Video generative adversarial networks: a review,”ACM Computing Surveys (CSUR), vol. 55, no. 2, pp. 1–25, 2022
2022
-
[91]
Search-map-search: a frame selection paradigm for action recognition,
M. Zhao, Y. Yu, X. Wang, L. Yang, and D. Niu, “Search-map-search: a frame selection paradigm for action recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 10 627–10 636
2023
-
[92]
Self-supervised learning for videos: A survey,
M. C. Schiappa, Y. S. Rawat, and M. Shah, “Self-supervised learning for videos: A survey,”ACM Computing Surveys, vol. 55, no. 13s, pp. 1–37, 2023
2023
-
[93]
On space-time interest points,
I. Laptev, “On space-time interest points,”International journal of computer vision, vol. 64, no. 2, pp. 107–123, 2005
2005
-
[94]
Histograms of oriented gradients for human detection,
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), vol. 1. Ieee, 2005, pp. 886–893
2005
-
[95]
Dense trajectories and motion boundary descriptors for action recognition,
H. Wang, A. Kl ¨aser, C. Schmid, and C.-L. Liu, “Dense trajectories and motion boundary descriptors for action recognition,”International journal of computer vision, vol. 103, no. 1, pp. 60–79, 2013. 17
2013
-
[96]
Transformers in vision: A survey,
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,”ACM computing surveys (CSUR), vol. 54, no. 10s, pp. 1–41, 2022
2022
-
[97]
Deep reinforcement learning: An overview,
Y. Li, “Deep reinforcement learning: An overview,”arXiv preprint arXiv:1701.07274, 2017
2017 arXiv
-
[98]
Deep reinforcement learning: A brief survey,
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “Deep reinforcement learning: A brief survey,”IEEE signal processing magazine, vol. 34, no. 6, pp. 26–38, 2017
2017
-
[99]
Continual lifelong learning with neural networks: A review,
G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,”Neural networks, vol. 113, pp. 54–71, 2019
2019
-
[100]
Federated machine learning: Concept and applications,
Q. Yang, Y. Liu, T. Chen, and Y. Tong, “Federated machine learning: Concept and applications,”ACM Transactions on Intelligent Systems and Technology (TIST), vol. 10, no. 2, pp. 1–19, 2019
2019
-
[101]
A continual learning survey: Defying forgetting in classification tasks,
M. De Lange, R. Aljundi, M. Masana, S. Parisot, X. Jia, A. Leonardis, G. Slabaugh, and T. Tuytelaars, “A continual learning survey: Defying forgetting in classification tasks,”IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 7, pp. 3366–3385, 2021
2021
-
[102]
A comprehensive survey of privacy- preserving federated learning: A taxonomy, review, and future direc- tions,
X. Yin, Y. Zhu, and J. Hu, “A comprehensive survey of privacy- preserving federated learning: A taxonomy, review, and future direc- tions,”ACM Computing Surveys (CSUR), vol. 54, no. 6, pp. 1–36, 2021
2021
-
[103]
Recognizing human actions: a local svm approach,
C. Schuldt, I. Laptev, and B. Caputo, “Recognizing human actions: a local svm approach,” inProceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004., vol. 3. IEEE, 2004, pp. 32–36
2004
-
[104]
Actions as space-time shapes,
M. Blank, L. Gorelick, E. Shechtman, M. Irani, and R. Basri, “Actions as space-time shapes,” inTenth IEEE International Conference on Computer Vision (ICCV’05) Volume 1, vol. 2. IEEE, 2005, pp. 1395– 1402
2005
-
[105]
Free viewpoint action recognition using motion history volumes,
D. Weinland, R. Ronfard, and E. Boyer, “Free viewpoint action recognition using motion history volumes,”Computer vision and image understanding, vol. 104, no. 2-3, pp. 249–257, 2006
2006
-
[106]
Learning realistic human actions from movies,
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld, “Learning realistic human actions from movies,” in2008 IEEE conference on computer vision and pattern recognition. IEEE, 2008, pp. 1–8
2008
-
[107]
Actions in context,
M. Marszalek, I. Laptev, and C. Schmid, “Actions in context,” in2009 IEEE conference on computer vision and pattern recognition. IEEE, 2009, pp. 2929–2936
2009
-
[108]
What are they doing?: Collective activity classification using spatio-temporal relationship among people,
W. Choi, K. Shahid, and S. Savarese, “What are they doing?: Collective activity classification using spatio-temporal relationship among people,” in2009 IEEE 12th international conference on computer vision work- shops, ICCV Workshops. IEEE, 2009, pp. 1282–1289
2009
-
[109]
Modeling temporal structure of decomposable motion segments for activity classification,
J. C. Niebles, C.-W. Chen, and L. Fei-Fei, “Modeling temporal structure of decomposable motion segments for activity classification,” inEuro- pean conference on computer vision. Springer, 2010, pp. 392–405
2010
-
[110]
Action recognition based on a bag of 3d points,
W. Li, Z. Zhang, and Z. Liu, “Action recognition based on a bag of 3d points,” in2010 IEEE computer society conference on computer vision and pattern recognition-workshops. IEEE, 2010, pp. 9–14
2010
-
[111]
View invariant human action recognition using histograms of 3d joints,
L. Xia, C.-C. Chen, and J. K. Aggarwal, “View invariant human action recognition using histograms of 3d joints,” in2012 IEEE computer soci- ety conference on computer vision and pattern recognition workshops. IEEE, 2012, pp. 20–27
2012
-
[112]
G3d: A gaming action dataset and real time action recognition evaluation framework,
V. Bloom, D. Makris, and V. Argyriou, “G3d: A gaming action dataset and real time action recognition evaluation framework,” in2012 IEEE Computer society conference on computer vision and pattern recognition workshops. IEEE, 2012, pp. 7–12
2012
-
[113]
Recognizing 50 human action categories of web videos,
K. K. Reddy and M. Shah, “Recognizing 50 human action categories of web videos,”Machine vision and applications, vol. 24, no. 5, pp. 971–981, 2013
2013
-
[114]
Recognizing actions from depth cameras as weakly aligned multi-part bag-of-poses,
L. Seidenari, V. Varano, S. Berretti, A. Bimbo, and P. Pala, “Recognizing actions from depth cameras as weakly aligned multi-part bag-of-poses,” inProceedings of the IEEE conference on computer vision and pattern recognition workshops, 2013, pp. 479–485
2013
-
[115]
Towards under- standing action recognition,
H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black, “Towards under- standing action recognition,” inProceedings of the IEEE international conference on computer vision, 2013, pp. 3192–3199
2013
-
[116]
Cross-view action modeling, learning and recognition,
J. Wang, X. Nie, Y. Xia, Y. Wu, and S.-C. Zhu, “Cross-view action modeling, learning and recognition,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 2649– 2656
2014
-
[117]
Ntu rgb+ d: A large scale dataset for 3d human activity analysis,
A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+ d: A large scale dataset for 3d human activity analysis,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1010–1019
2016
-
[118]
Infar dataset: Infrared action recognition at different times,
C. Gao, Y. Du, J. Liu, J. Lv, L. Yang, D. Meng, and A. G. Hauptmann, “Infar dataset: Infrared action recognition at different times,”Neurocom- puting, vol. 212, pp. 36–47, 2016
2016
-
[119]
Thermal imaging based elderly fall detection,
S. Vadivelu, S. Ganesan, O. R. Murthy, and A. Dhall, “Thermal imaging based elderly fall detection,” inAsian conference on computer vision. Springer, 2016, pp. 541–553
2016
-
[120]
Human action localization with sparse spatial supervision,
P. Weinzaepfel, X. Martin, and C. Schmid, “Human action localization with sparse spatial supervision,”arXiv preprint arXiv:1605.05197, 2016
2016 arXiv
-
[121]
Every moment counts: Dense detailed labeling of actions in complex videos,
S. Yeung, O. Russakovsky, N. Jin, M. Andriluka, G. Mori, and L. Fei-Fei, “Every moment counts: Dense detailed labeling of actions in complex videos,”International Journal of Computer Vision, vol. 126, no. 2, pp. 375–389, 2018
2018
-
[122]
A hierarchical deep temporal model for group activity recognition,
M. S. Ibrahim, S. Muralidharan, Z. Deng, A. Vahdat, and G. Mori, “A hierarchical deep temporal model for group activity recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1971–1980
2016
-
[123]
Need for speed: A benchmark for higher frame rate object tracking,
H. Kiani Galoogahi, A. Fagg, C. Huang, D. Ramanan, and S. Lucey, “Need for speed: A benchmark for higher frame rate object tracking,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 1125–1134
2017
-
[124]
Audio set: An ontology and human- labeled dataset for audio events,
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human- labeled dataset for audio events,” in2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2017, p...
2017
-
[125]
Attentive spatio-temporal representation learning for diving classification,
G. Kanojia, S. Kumawat, and S. Raman, “Attentive spatio-temporal representation learning for diving classification,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2019, pp. 0–0
2019
-
[126]
Moments in time dataset: one million videos for event understanding,
M. Monfort, A. Andonian, B. Zhou, K. Ramakrishnan, S. A. Bargal, T. Yan, L. Brown, Q. Fan, D. Gutfreund, C. Vondricket al., “Moments in time dataset: one million videos for event understanding,”IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 2, pp....
2019
-
[127]
Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,
J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot, “Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,”IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 10, pp. 2684–2701, 2019
2019
-
[128]
Finegym: A hierarchical video dataset for fine-grained action understanding,
D. Shao, Y. Zhao, B. Dai, and D. Lin, “Finegym: A hierarchical video dataset for fine-grained action understanding,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2616–2625
2020
-
[129]
Vggsound: A large- scale audio-visual dataset,
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “Vggsound: A large- scale audio-visual dataset,” inICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 721–725
2020
-
[130]
Ava active speaker: An audio-visual dataset for active speaker detection,
J. Roth, S. Chaudhuri, O. Klejch, R. Marvin, A. Gallagher, L. Kaver, S. Ramaswamy, A. Stopczynski, C. Schmid, Z. Xiet al., “Ava active speaker: An audio-visual dataset for active speaker detection,” inICASSP 2020-2020 IEEE international conference on acoustics, speech and sign...
2020
-
[131]
Uav-human: A large benchmark for human behavior understanding with unmanned aerial vehicles,
T. Li, J. Liu, W. Zhang, Y. Ni, W. Wang, and Z. Li, “Uav-human: A large benchmark for human behavior understanding with unmanned aerial vehicles,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 16 266–16 275
2021
-
[132]
End-to-end spatio-temporal action localisation with video transformers,
A. A. Gritsenko, X. Xiong, J. Djolonga, M. Dehghani, C. Sun, M. Lucic, C. Schmid, and A. Arnab, “End-to-end spatio-temporal action localisation with video transformers,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 18 373–18 383
2024
-
[133]
Epic- sounds: A large-scale dataset of actions that sound,
J. Huh, J. Chalk, E. Kazakos, D. Damen, and A. Zisserman, “Epic- sounds: A large-scale dataset of actions that sound,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
2025
-
[134]
Berkeley mhad: A comprehensive multimodal human action database,
F. Ofli, R. Chaudhry, G. Kurillo, R. Vidal, and R. Bajcsy, “Berkeley mhad: A comprehensive multimodal human action database,” in2013 IEEE workshop on applications of computer vision (WACV). IEEE, 2013, pp. 53–60
2013
-
[135]
Unstructured human activity detection from rgbd images,
J. Sung, C. Ponce, B. Selman, and A. Saxena, “Unstructured human activity detection from rgbd images,” in2012 IEEE international conference on robotics and automation. IEEE, 2012, pp. 842–849
2012
-
[136]
Learning to recognize daily actions using gaze,
A. Fathi, Y. Li, and J. M. Rehg, “Learning to recognize daily actions using gaze,” inEuropean Conference on Computer Vision. Springer, 2012, pp. 314–327
2012
-
[137]
Learning human activities and object affordances from rgb-d videos,
H. S. Koppula, R. Gupta, and A. Saxena, “Learning human activities and object affordances from rgb-d videos,”The International journal of robotics research, vol. 32, no. 8, pp. 951–970, 2013
2013
-
[138]
Combining embedded accelerometers with computer vision for recognizing food preparation activities,
S. Stein and S. J. McKenna, “Combining embedded accelerometers with computer vision for recognizing food preparation activities,” inProceedings of the 2013 ACM international joint conference on Pervasive and ubiquitous computing, 2013, pp. 729–738. 18
2013
-
[139]
The language of actions: Recovering the syntax and semantics of goal-directed human activities,
H. Kuehne, A. Arslan, and T. Serre, “The language of actions: Recovering the syntax and semantics of goal-directed human activities,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 780–787
2014
-
[140]
Delving into egocentric actions,
Y. Li, Z. Ye, and J. M. Rehg, “Delving into egocentric actions,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 287–295
2015
-
[141]
Jointly learning hetero- geneous features for rgb-d activity recognition,
J.-F. Hu, W.-S. Zheng, J. Lai, and J. Zhang, “Jointly learning hetero- geneous features for rgb-d activity recognition,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 5344–5352
2015
-
[142]
Towards automatic learning of procedures from web instructional videos,
L. Zhou, C. Xu, and J. Corso, “Towards automatic learning of procedures from web instructional videos,” inProceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018
2018
-
[143]
In the eye of beholder: Joint learning of gaze and actions in first person video,
Y. Li, M. Liu, and J. M. Rehg, “In the eye of beholder: Joint learning of gaze and actions in first person video,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 619–635
2018
-
[144]
Weakly-supervised video object grounding from text by loss weighting and object interaction,
L. Zhou, N. Louis, and J. J. Corso, “Weakly-supervised video object grounding from text by loss weighting and object interaction,”arXiv preprint arXiv:1805.02834, 2018
2018 arXiv
-
[145]
Coin: A large-scale dataset for comprehensive instructional video analysis,
Y. Tang, D. Ding, Y. Rao, Y. Zheng, D. Zhang, L. Zhao, J. Lu, and J. Zhou, “Coin: A large-scale dataset for comprehensive instructional video analysis,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 1207–1216
2019
-
[146]
Cater: A diagnostic dataset for compositional actions and temporal reasoning,
R. Girdhar and D. Ramanan, “Cater: A diagnostic dataset for compositional actions and temporal reasoning,”arXiv preprint arXiv:1910.04744, 2019
1910 arXiv
-
[147]
Clevrer: Collision events for video representation and reasoning,
K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum, “Clevrer: Collision events for video representation and reasoning,”arXiv preprint arXiv:1910.01442, 2019
1910 arXiv
-
[148]
Cross-task weakly supervised learning from instructional videos,
D. Zhukov, J.-B. Alayrac, R. G. Cinbis, D. Fouhey, I. Laptev, and J. Sivic, “Cross-task weakly supervised learning from instructional videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 3537–3545
2019
-
[149]
Moma: Multi-object multi-actor activity parsing,
Z. Luo, W. Xie, S. Kapoor, Y. Liang, M. Cooper, J. C. Niebles, E. Adeli, and F.-F. Li, “Moma: Multi-object multi-actor activity parsing,” Advances in neural information processing systems, vol. 34, pp. 17 939– 17 955, 2021
2021
-
[150]
Moma-lrg: Language-refined graphs for multi- object multi-actor activity parsing,
Z. Luo, Z. Durante, L. Li, W. Xie, R. Liu, E. Jin, Z. Huang, L. Y. Li, J. Wu, J. C. Niebleset al., “Moma-lrg: Language-refined graphs for multi- object multi-actor activity parsing,”Advances in Neural Information Processing Systems, vol. 35, pp. 5282–5298, 2022
2022
-
[151]
Detecting activities of daily living in first-person camera views,
H. Pirsiavash and D. Ramanan, “Detecting activities of daily living in first-person camera views,” in2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012, pp. 2847–2854
2012
-
[152]
Mining actionlet ensemble for action recognition with depth cameras,
J. Wang, Z. Liu, Y. Wu, and J. Yuan, “Mining actionlet ensemble for action recognition with depth cameras,” in2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012, pp. 1290–1297
2012
-
[153]
A database for fine grained activity detection of cooking activities,
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele, “A database for fine grained activity detection of cooking activities,” in2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012, pp. 1194–1201
2012
-
[154]
The thumos challenge on action recognition for videos “in the wild
H. Idrees, A. R. Zamir, Y.-G. Jiang, A. Gorban, I. Laptev, R. Sukthankar, and M. Shah, “The thumos challenge on action recognition for videos “in the wild”,”Computer Vision and Image Understanding, vol. 155, pp. 1–23, 2017
2017
-
[155]
Pku-mmd: A large scale benchmark for continuous multi-modal human action understanding,
C. Liu, Y. Hu, Y. Li, S. Song, and J. Liu, “Pku-mmd: A large scale benchmark for continuous multi-modal human action understanding,” arXiv preprint arXiv:1703.07475, 2017
2017 arXiv
-
[156]
Soccernet: A scalable dataset for action spotting in soccer videos,
S. Giancola, M. Amine, T. Dghaily, and B. Ghanem, “Soccernet: A scalable dataset for action spotting in soccer videos,” inProceedings of the IEEE conference on computer vision and pattern recognition workshops, 2018, pp. 1711–1721
2018
-
[157]
Hacs: Human action clips and segments dataset for recognition and temporal localization,
H. Zhao, A. Torralba, L. Torresani, and Z. Yan, “Hacs: Human action clips and segments dataset for recognition and temporal localization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 8668–8678
2019
-
[158]
A benchmark dataset and comparison study for multi-modal human action analytics,
J. Liu, S. Song, C. Liu, Y. Li, and Y. Hu, “A benchmark dataset and comparison study for multi-modal human action analytics,”ACM Trans- actions on Multimedia Computing, Communications, and Applications (TOMM), vol. 16, no. 2, pp. 1–24, 2020
2020
-
[159]
Bdd100k: A diverse driving dataset for heterogeneous multitask learning,
F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, F. Liu, V. Madhavan, and T. Darrell, “Bdd100k: A diverse driving dataset for heterogeneous multitask learning,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2636–2645
2020
-
[160]
Collecting highly parallel data for paraphrase evaluation,
D. Chen and W. B. Dolan, “Collecting highly parallel data for paraphrase evaluation,” inProceedings of the 49th annual meeting of the associa- tion for computational linguistics: human language technologies, 2011, pp. 190–200
2011
-
[161]
Msr-vtt: A large video description dataset for bridging video and language,
J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video description dataset for bridging video and language,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 5288– 5296
2016
-
[162]
Movie description,
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele, “Movie description,”International Journal of Computer Vision, vol. 123, no. 1, pp. 94–120, 2017
2017
-
[163]
Localizing moments in video with natural language,
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 5803–5812
2017
-
[164]
Dense- captioning events in videos,
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, “Dense- captioning events in videos,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 706–715
2017
-
[165]
Tgif-qa: Toward spatio- temporal reasoning in visual question answering,
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “Tgif-qa: Toward spatio- temporal reasoning in visual question answering,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2758–2766
2017
-
[166]
Video question answering via gradually refined attention over appearance and motion,
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, “Video question answering via gradually refined attention over appearance and motion,” inProceedings of the 25th ACM international conference on Multimedia, 2017, pp. 1645–1653
2017
-
[167]
Tall: Temporal activity local- ization via language query,
J. Gao, C. Sun, Z. Yang, and R. Nevatia, “Tall: Temporal activity local- ization via language query,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 5267–5275
2017
-
[168]
Tvqa: Localized, compositional video question answering,
J. Lei, L. Yu, M. Bansal, and T. L. Berg, “Tvqa: Localized, compositional video question answering,”arXiv preprint arXiv:1809.01696, 2018
2018 arXiv
-
[169]
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,
X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang, “Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 4581–4591
2019
-
[170]
Next-qa: Next phase of question-answering to explaining temporal actions,
J. Xiao, X. Shang, A. Yao, and T.-S. Chua, “Next-qa: Next phase of question-answering to explaining temporal actions,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 9777–9786
2021
-
[171]
Agqa: A benchmark for compositional spatio-temporal reasoning,
M. Grunde-McLaughlin, R. Krishna, and M. Agrawala, “Agqa: A benchmark for compositional spatio-temporal reasoning,” inProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 11 287–11 297
2021
-
[172]
Frozen in time: A joint video and image encoder for end-to-end retrieval,
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 1728–1738
2021
-
[173]
Advancing high-resolution video-language representation with large- scale video transcriptions,
H. Xue, T. Hang, Y. Zeng, Y. Sun, B. Liu, H. Yang, J. Fu, and B. Guo, “Advancing high-resolution video-language representation with large- scale video transcriptions,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5036–5045
2022
-
[174]
Ego4d: Around the world in 3,000 hours of egocentric video,
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liuet al., “Ego4d: Around the world in 3,000 hours of egocentric video,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 18 9...
2022
-
[175]
Vidchapters-7m: Video chapters at scale,
A. Yang, A. Nagrani, I. Laptev, J. Sivic, and C. Schmid, “Vidchapters-7m: Video chapters at scale,”Advances in Neural Information Processing Systems, vol. 36, pp. 49 428–49 444, 2023
2023
-
[176]
Internvid: A large-scale video-text dataset for multimodal understanding and generation,
Y. Wang, Y. He, Y. Li, K. Li, J. Yu, X. Ma, X. Li, G. Chen, X. Chen, Y. Wanget al., “Internvid: A large-scale video-text dataset for multimodal understanding and generation,”arXiv preprint arXiv:2307.06942, 2023
2023 arXiv
-
[177]
Panda- 70m: Captioning 70m videos with multiple cross-modality teachers,
T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H.-w. Chao, B. E. Jeon, Y. Fang, H.-Y. Lee, J. Ren, M.-H. Yanget al., “Panda- 70m: Captioning 70m videos with multiple cross-modality teachers,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...
2024
-
[178]
Miradata: A large-scale video dataset with long durations and structured captions,
X. Ju, Y. Gao, Z. Zhang, Z. Yuan, X. Wang, A. Zeng, Y. Xiong, Q. Xu, and Y. Shan, “Miradata: A large-scale video dataset with long durations and structured captions,”Advances in Neural Information Processing Systems, vol. 37, pp. 48 955–48 970, 2024
2024
-
[179]
Openvid-1m: A large-scale high-quality dataset for text-to-video generation,
K. Nan, R. Xie, P. Zhou, T. Fan, Z. Yang, Z. Chen, X. Li, J. Yang, and Y. Tai, “Openvid-1m: A large-scale high-quality dataset for text-to-video generation,”arXiv preprint arXiv:2407.02371, 2024
2024 arXiv
-
[180]
Koala-36m: A large-scale video dataset 19 improving consistency between fine-grained conditions and video con- tent,
Q. Wang, Y. Shi, J. Ou, R. Chen, K. Lin, J. Wang, B. Jiang, H. Yang, M. Zheng, X. Taoet al., “Koala-36m: A large-scale video dataset 19 improving consistency between fine-grained conditions and video con- tent,” inProceedings of the Computer Vision and Pattern Recognition Conf...
2025
-
[181]
The ava-kinetics localized human actions video dataset,
A. Li, M. Thotakuri, D. A. Ross, J. Carreira, A. Vostrikov, and A. Zisserman, “The ava-kinetics localized human actions video dataset,” arXiv preprint arXiv:2005.00214, 2020
2005 arXiv
-
[182]
Grounding action descriptions in videos,
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal, “Grounding action descriptions in videos,”Transactions of the Association for Computational Linguistics, vol. 1, pp. 25–36, 2013
2013
-
[183]
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 2630–2640
2019
-
[184]
Convolutional two-stream network fusion for video action recognition,
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1933– 1941
2016
-
[185]
Temporal segment networks: Towards good practices for deep action recognition,
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks: Towards good practices for deep action recognition,” inEuropean conference on computer vision. Springer, 2016, pp. 20–36
2016
-
[186]
Anticipative video transformer,
R. Girdhar and K. Grauman, “Anticipative video transformer,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 13 505–13 515
2021
-
[187]
Video modeling with correlation networks,
H. Wang, D. Tran, L. Torresani, and M. Feiszli, “Video modeling with correlation networks,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 352–361
2020
-
[188]
Gate-shift networks for video action recognition,
S. Sudhakaran, S. Escalera, and O. Lanz, “Gate-shift networks for video action recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1102–1111
2020
-
[189]
Sportscap: Monocular 3d human motion capture and fine-grained understanding in challenging sports videos,
X. Chen, A. Pang, W. Yang, Y. Ma, L. Xu, and J. Yu, “Sportscap: Monocular 3d human motion capture and fine-grained understanding in challenging sports videos,”International Journal of Computer Vision, vol. 129, no. 10, pp. 2846–2864, 2021
2021
-
[190]
Learning correlation structures for vision transformers,
M. Kim, P. H. Seo, C. Schmid, and M. Cho, “Learning correlation structures for vision transformers,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 18 941–18 951
2024
-
[191]
Ms-tcn: Multi-stage temporal convolutional network for action segmentation,
Y. A. Farha and J. Gall, “Ms-tcn: Multi-stage temporal convolutional network for action segmentation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3575– 3584
2019
-
[192]
Temporal aggregate representa- tions for long-range video understanding,
F. Sener, D. Singhania, and A. Yao, “Temporal aggregate representa- tions for long-range video understanding,” inEuropean conference on computer vision. Springer, 2020, pp. 154–171
2020
-
[193]
Selective structured state-spaces for long-form video understanding,
J. Wang, W. Zhu, P. Wang, X. Yu, L. Liu, M. Omar, and R. Hamid, “Selective structured state-spaces for long-form video understanding,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6387–6397
2023
-
[194]
Masked autoencoders as spatiotemporal learners,
C. Feichtenhofer, Y. Li, K. Heet al., “Masked autoencoders as spatiotemporal learners,”Advances in neural information processing systems, vol. 35, pp. 35 946–35 958, 2022
2022
-
[195]
Omnivore: A single model for many visual modalities,
R. Girdhar, M. Singh, N. Ravi, L. Van Der Maaten, A. Joulin, and I. Misra, “Omnivore: A single model for many visual modalities,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16 102–16 112
2022
-
[196]
Hiera: A hierarchical vision transformer without the bells-and-whistles,
C. Ryali, Y.-T. Hu, D. Bolya, C. Wei, H. Fan, P.-Y. Huang, V. Aggarwal, A. Chowdhury, O. Poursaeed, J. Hoffmanet al., “Hiera: A hierarchical vision transformer without the bells-and-whistles,” inInternational conference on machine learning. PMLR, 2023, pp. 29 441–29 454
2023
-
[197]
Beyond short snippets: Deep networks for video classification,
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici, “Beyond short snippets: Deep networks for video classification,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 4694–4702
2015
-
[198]
Learning spatio-temporal representation with pseudo-3d residual networks,
Z. Qiu, T. Yao, and T. Mei, “Learning spatio-temporal representation with pseudo-3d residual networks,” inproceedings of the IEEE International Conference on Computer Vision, 2017, pp. 5533–5541
2017
-
[199]
A closer look at spatiotemporal convolutions for action recognition,
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, pp. 6450–6459
2018
-
[200]
Long-term feature banks for detailed video understanding,
C.-Y. Wu, C. Feichtenhofer, H. Fan, K. He, P. Krahenbuhl, and R. Girshick, “Long-term feature banks for detailed video understanding,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 284–293
2019
-
[201]
Sequence level semantics aggregation for video object detection,
H. Wu, Y. Chen, N. Wang, and Z. Zhang, “Sequence level semantics aggregation for video object detection,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9217–9225
2019
-
[202]
Epic-fusion: Audio-visual temporal binding for egocentric action recognition,
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen, “Epic-fusion: Audio-visual temporal binding for egocentric action recognition,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 5492–5501
2019
-
[203]
Timeception for complex action recognition,
N. Hussein, E. Gavves, and A. W. Smeulders, “Timeception for complex action recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 254–263
2019
-
[204]
Videobert: A joint model for video and language representation learning,
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 7464–7473
2019
-
[205]
Multi-modal domain adaptation for fine- grained action recognition,
J. Munro and D. Damen, “Multi-modal domain adaptation for fine- grained action recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 122–132
2020
-
[206]
Temporal pyramid network for action recognition,
C. Yang, Y. Xu, J. Shi, B. Dai, and B. Zhou, “Temporal pyramid network for action recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 591–600
2020
-
[207]
Multi-modal transformer for video retrieval,
V. Gabeur, C. Sun, K. Alahari, and C. Schmid, “Multi-modal transformer for video retrieval,” inEuropean Conference on Computer Vision. Springer, 2020, pp. 214–229
2020
-
[208]
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,
H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y. Cui, and B. Gong, “Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,”Advances in neural information processing systems, vol. 34, pp. 24 206–24 221, 2021
2021
-
[209]
Memvit: Memory-augmented multiscale vision trans- former for efficient long-term video recognition,
C.-Y. Wu, Y. Li, K. Mangalam, H. Fan, B. Xiong, J. Malik, and C. Feichtenhofer, “Memvit: Memory-augmented multiscale vision trans- former for efficient long-term video recognition,” inProceedings of the ieee/cvf conference on computer vision and pattern recognition, 2022, pp. ...
2022
-
[210]
Long movie clip classification with state- space video models,
M. M. Islam and G. Bertasius, “Long movie clip classification with state- space video models,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 87–104
2022
-
[211]
Language models with image descriptors are strong few-shot video-language learners,
Z. Wang, M. Li, R. Xu, L. Zhou, J. Lei, X. Lin, S. Wang, Z. Yang, C. Zhu, D. Hoiemet al., “Language models with image descriptors are strong few-shot video-language learners,”Advances in Neural Information Processing Systems, vol. 35, pp. 8483–8497, 2022
2022
-
[212]
Diffusion action segmentation,
D. Liu, Q. Li, A.-D. Dinh, T. Jiang, M. Shah, and C. Xu, “Diffusion action segmentation,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 10 139–10 149
2023
-
[213]
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,
A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid, “Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp....
2023
-
[214]
Videomamba: State space model for efficient video understanding,
K. Li, X. Li, Y. Wang, Y. He, Y. Wang, L. Wang, and Y. Qiao, “Videomamba: State space model for efficient video understanding,” in European conference on computer vision. Springer, 2024, pp. 237–255
2024
-
[215]
Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,
B. He, H. Li, Y. K. Jang, M. Jia, X. Cao, A. Shah, A. Shrivastava, and S.-N. Lim, “Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13 504–13 514
2024
-
[216]
Mamba-nd: Selective state space modeling for multi-dimensional data,
S. Li, H. Singh, and A. Grover, “Mamba-nd: Selective state space modeling for multi-dimensional data,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 75–92
2024
-
[217]
Video action transformer network,
R. Girdhar, J. Carreira, C. Doersch, and A. Zisserman, “Video action transformer network,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 244–253
2019
-
[218]
X3d: Expanding architectures for efficient video recognition,
C. Feichtenhofer, “X3d: Expanding architectures for efficient video recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 203–213
2020
-
[219]
Spatiotemporal contrastive video representation learning,
R. Qian, T. Meng, B. Gong, M.-H. Yang, H. Wang, S. Belongie, and Y. Cui, “Spatiotemporal contrastive video representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 6964–6974
2021
-
[220]
Un- masked teacher: Towards training-efficient video foundation models,
K. Li, Y. Wang, Y. Li, Y. Wang, Y. He, L. Wang, and Y. Qiao, “Un- masked teacher: Towards training-efficient video foundation models,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 19 948–19 960
2023
-
[221]
Temporal relational reasoning in videos,
B. Zhou, A. Andonian, A. Oliva, and A. Torralba, “Temporal relational reasoning in videos,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 803–818. 20
2018
-
[222]
Videos as space-time region graphs,
X. Wang and A. Gupta, “Videos as space-time region graphs,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 399–417
2018
-
[223]
Multiscale vision transformers,
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 6824–6835
2021
-
[224]
Videoclip: Contrastive pre-training for zero-shot video-text understanding,
H. Xu, G. Ghosh, P.-Y. Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer, “Videoclip: Contrastive pre-training for zero-shot video-text understanding,”arXiv preprint arXiv:2109.14084, 2021
2021 arXiv
-
[225]
Videococa: Video-text modeling with zero-shot transfer from contrastive captioners,
S. Yan, T. Zhu, Z. Wang, Y. Cao, M. Zhang, S. Ghosh, Y. Wu, and J. Yu, “Videococa: Video-text modeling with zero-shot transfer from contrastive captioners,”arXiv preprint arXiv:2212.04979, 2022
2022 arXiv
-
[226]
Violet: End-to-end video-language transformers with masked visual- token modeling,
T.-J. Fu, L. Li, Z. Gan, K. Lin, W. Y. Wang, L. Wang, and Z. Liu, “Violet: End-to-end video-language transformers with masked visual- token modeling,”arXiv preprint arXiv:2111.12681, 2021
2021 arXiv
-
[227]
Align and prompt: Video-and-language pre-training with entity prompts,
D. Li, J. Li, H. Li, J. C. Niebles, and S. C. Hoi, “Align and prompt: Video-and-language pre-training with entity prompts,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 4953–4963
2022
-
[228]
Video-llama: An instruction-tuned audio-visual language model for video understanding,
H. Zhang, X. Li, and L. Bing, “Video-llama: An instruction-tuned audio-visual language model for video understanding,”arXiv preprint arXiv:2306.02858, 2023
2023 arXiv
-
[229]
Video-chatgpt: Towards detailed video understanding via large vision and language models,
M. Maaz, H. Rasheed, S. Khan, and F. S. Khan, “Video-chatgpt: Towards detailed video understanding via large vision and language models,” arXiv preprint arXiv:2306.05424, 2023
2023 arXiv
-
[230]
Human action recognition using factorized spatio-temporal convolutional networks,
L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi, “Human action recognition using factorized spatio-temporal convolutional networks,” inProceed- ings of the IEEE international conference on computer vision, 2015, pp. 4597–4605
2015
-
[231]
Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 305–321
2018
-
[232]
Tsm: Temporal shift module for efficient video understanding,
J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 7083–7093
2019
-
[233]
Tiny video networks,
A. Piergiovanni, A. Angelova, and M. S. Ryoo, “Tiny video networks,” Applied AI Letters, vol. 3, no. 1, p. e38, 2022
2022
-
[234]
Movinets: Mobile video networks for efficient video recognition,
D. Kondratyuk, L. Yuan, Y. Li, L. Zhang, M. Tan, M. Brown, and B. Gong, “Movinets: Mobile video networks for efficient video recognition,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 16 020–16 030
2021
-
[235]
Eco: Efficient convolutional network for online video understanding,
M. Zolfaghari, K. Singh, and T. Brox, “Eco: Efficient convolutional network for online video understanding,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 695–712
2018
-
[236]
Video transformer network,
D. Neimark, O. Bar, M. Zohar, and D. Asselmann, “Video transformer network,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 3163–3172
2021
-
[237]
Assemblenet: Searching for multi-stream neural connectivity in video architectures,
M. S. Ryoo, A. Piergiovanni, M. Tan, and A. Angelova, “Assemblenet: Searching for multi-stream neural connectivity in video architectures,” arXiv preprint arXiv:1905.13209, 2019
1905 arXiv
-
[238]
Keeping your eye on the ball: Trajec- tory attention in video transformers,
M. Patrick, D. Campbell, Y. Asano, I. Misra, F. Metze, C. Feichtenhofer, A. Vedaldi, and J. F. Henriques, “Keeping your eye on the ball: Trajec- tory attention in video transformers,”Advances in neural information processing systems, vol. 34, pp. 12 493–12 506, 2021
2021
-
[239]
Video swin transformer,
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 3202–3211
2022
-
[240]
Mvitv2: Improved multiscale vision transformers for classification and detection,
Y. Li, C.-Y. Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, and C. Feichtenhofer, “Mvitv2: Improved multiscale vision transformers for classification and detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 4804– 4814
2022
-
[241]
Heterogeneous memory enhanced multimodal attention model for video question answering,
C. Fan, X. Zhang, S. Zhang, W. Wang, C. Zhang, and H. Huang, “Heterogeneous memory enhanced multimodal attention model for video question answering,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 1999–2007
2019
-
[242]
Less is more: Clipbert for video-and-language learning via sparse sampling,
J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu, “Less is more: Clipbert for video-and-language learning via sparse sampling,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 7331–7341
2021
-
[243]
Clip4clip: An empirical study of clip for end to end video clip retrieval,
H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li, “Clip4clip: An empirical study of clip for end to end video clip retrieval,”arXiv preprint arXiv:2104.08860, 2021
2021 arXiv
-
[244]
Zero-shot video question answering via frozen bidirectional language models,
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Zero-shot video question answering via frozen bidirectional language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 124–141, 2022
2022
-
[245]
All in one: Exploring unified video-language pre- training,
J. Wang, Y. Ge, R. Yan, Y. Ge, K. Q. Lin, S. Tsutsui, X. Lin, G. Cai, J. Wu, Y. Shanet al., “All in one: Exploring unified video-language pre- training,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6598–6608
2023
-
[246]
Videochat: Chat-centric video understanding,
K. Li, Y. He, Y. Wang, Y. Li, W. Wang, P. Luo, Y. Wang, L. Wang, and Y. Qiao, “Videochat: Chat-centric video understanding,”arXiv preprint arXiv:2305.06355, 2023
2023 arXiv
-
[247]
Llama-adapter: Efficient fine-tuning of language models with zero-init attention,
R. Zhang, J. Han, C. Liu, P. Gao, A. Zhou, X. Hu, S. Yan, P. Lu, H. Li, and Y. Qiao, “Llama-adapter: Efficient fine-tuning of language models with zero-init attention,”arXiv preprint arXiv:2303.16199, 2023
2023 arXiv
-
[248]
Valley: Video assistant with large language model enhanced ability,
R. Luo, Z. Zhao, M. Yang, J. Dong, D. Li, P. Lu, T. Wang, L. Hu, M. Qiu, and Z. Wei, “Valley: Video assistant with large language model enhanced ability,”arXiv preprint arXiv:2306.07207, 2023
2023 arXiv
-
[249]
Groundinggpt: Language enhanced multi-modal grounding model,
Z. Li, Q. Xu, D. Zhang, H. Song, Y. Cai, Q. Qi, R. Zhou, J. Pan, Z. Li, V. T. Vuet al., “Groundinggpt: Language enhanced multi-modal grounding model,”arXiv preprint arXiv:2401.06071, 2024. Lei Wangreceived his M.E. in Software Engineering from the University of Western Austral...
2024 arXiv
-
[2024]
Available: https://doi.org/10.1016/j.jvcir.2024.104320
[Online]. Available: https://doi.org/10.1016/j.jvcir.2024.104320
2024
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.