REVIEW 3 major objections 6 minor 2 cited by
Do Language Models Understand Time?
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper argues that today's video-LLMs achieve benchmark success without genuinely understanding time, because temporal structure is supplied by pretrained encoders and datasets, not learned by the LLM.
desk verdict A useful survey of video-LLMs and datasets whose central claim about temporal awareness is asserted rather than tested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the interaction/fusion mechanism between a pretrained visual encoder and an LLM: projection layers, cross-attention modules, Q-Formers, and temporal-specific connectors such as temporal samplers and scene-level sequential alignment. The paper's argument is that this coupling is where temporal information is filtered or lost. Encoders such as CLIP, I3D, TimeSformer, and Video Swin supply frame-level tokens and short-range motion features; the LLM then aligns those with language without any inherent model of sequence order, so any true temporal structure must come from outside. The distribution of fusion mechanisms across recent models, and the benchmark plots built on this coupling, are the evidence the review leans on.
What would settle it
Run one controlled experiment: take a current video-LLM and ask the same set of order-sensitive questions on the same videos in correct and in temporally shuffled frame order, with subtitles off and a fixed prompt. If accuracy on the shuffled version stays at the same level, the model is not using order and the claim of lacking temporal awareness is supported; if accuracy collapses on the shuffled version, the claim that LLMs lack direct temporal awareness would be falsified.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a negative result about current architectures: state-of-the-art video-LLMs achieve competitive task performance through the interaction of pretrained visual encoders and LLMs, but this performance does not amount to temporal comprehension. The paper claims that LLMs lack direct temporal awareness, that encoders focus on short-term patterns and fragmented cues, and that datasets rarely carry the temporal annotations—event order, duration, causality—that would let either component learn long-term dependencies. The conclusion is that no current video-LLM excels across video tasks and that the apparent progress is partly an artifact of inconsistent evaluation.
Load-bearing premise
The argument depends on the benchmark numbers plotted across Figures 4-6 being comparable; if those models were evaluated under different prompts, subtitles settings, sampling, or metric versions, the performance gaps and the conclusion that no video-LLM excels across tasks do not follow.
Editorial extensions
If this is right
- Benchmark numbers on Video-MME, MSVD-QA, MSRVTT-QA, ActivityNet-QA, and retrieval or captioning sets should not be read as measures of temporal reasoning, because they mix encoder quality, language priors, and short-term cues.
- Improving temporal understanding requires training encoders and LLMs jointly on temporally annotated data, not just scaling model size or context length.
- New benchmarks must include explicit temporal labels such as event order, duration, and causality, and must use long-form videos; otherwise progress in temporal reasoning will be mismeasured.
- Most current models use projection layers and cross-attention rather than temporal-specific fusion mechanisms, so temporal structure is imposed only weakly and inconsistently.
- Comparisons of video-LLMs against traditional video models such as I3D, SlowFast, and Video Swin are unfair; within-paradigm benchmarking against other video-LLMs is needed to reveal real progress.
Reading between the lines
- If the paper is right, a cheap diagnostic follows: shuffling frame order or reversing event order in existing test videos should produce a large accuracy drop on order-sensitive questions, and this could be measured today without new training.
- The same short-term bias would likely carry over to video generation and long-context models, because a longer context window does not by itself supply training signal about causality or event progression.
- One way to test the dataset bottleneck directly is to take a fixed video-LLM and fine-tune it on a small set of temporally annotated long videos; if temporal reasoning improves substantially, the bottleneck is data rather than architecture.
- The paper's critique generalizes to time-series and event-forecasting uses of LLMs, where surface task success may again come from textual priors rather than from an internal model of time.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a survey/critical review of video-based large language models (video-LLMs) and their temporal reasoning capabilities. It reviews recent video-LLM architectures and their interaction mechanisms with pretrained visual encoders (Tables 1–2), catalogs video datasets across tasks such as action recognition, video QA, captioning, retrieval, and anomaly detection (Table 3), and presents benchmark comparisons of video-LLMs (Figures 4–6). The central claim is that LLMs, despite achieving strong task performance when paired with pretrained video encoders, lack direct temporal awareness and fall short in understanding long-term temporal dependencies. The paper identifies dataset limitations (lack of temporal annotations, short-term bias, low diversity) and proposes future directions including joint encoder-LLM training, richer temporal annotations, and new fusion architectures.
Significance. The paper is a broad and timely survey that usefully organizes a fast-moving literature on video-LLMs, fusion mechanisms, and video datasets. Its strength is the critical framing: it explicitly argues that benchmark comparisons are often inconsistent and that dataset design shapes temporal reasoning, which is a valuable message for the community. The GitHub repository and the extensive tables (Tables 1–3) provide a convenient reference. However, the paper's central claim about the absence of direct temporal awareness is asserted rather than demonstrated, and the benchmark figures are presented without a documented extraction protocol. If the central claim were supported by controlled evidence or a systematic benchmark analysis, the paper would be a significant position piece; as it stands, its conclusions outrun its evidence. The paper contains no new experiments, code, or machine-checked derivations, so its value depends on the accuracy and representativeness of the plotted numbers and the force of its conceptual argument.
major comments (3)
- [Section 3, 'Analysis and Discussion'] The central thesis that LLMs 'lack direct temporal awareness' is load-bearing for the paper, but it is neither operationalized nor tested. The same section states that LLMs 'can infer temporal relationships through contextual cues such as "first", "then", and "after"', which is itself a form of temporal inference, and the paper cites Ref. [64] (Gurnee and Tegmark) and Ref. [106] (TempCompass) without reconciling their evidence of temporal representation in LLMs with the negative claim. Please define what would count as 'direct' temporal awareness (e.g., whether positional embeddings, frame-order tokens, or learned temporal projections count), state the claim as a falsifiable hypothesis, and support it with either a controlled experiment or a systematic analysis of existing temporal-reasoning benchmarks such as TempCompass, TGIF-QA, or MVBench.
- [Figures 4–6] The benchmark figures are presented without the extraction protocol needed to support the claim that no video-LLM excels across tasks. For each plotted number, the paper should report the exact model checkpoint, prompt template, subtitles-on/off condition, frame sampling, metric version, and original source. As written, the numbers mix models, tasks, and possibly evaluation settings, so the performance gaps in Figures 4–6 do not logically follow. This is a particular problem because the paper itself (Section 'Fair evaluation is needed') argues that inconsistent evaluations produce misleading conclusions; the charts should not reproduce the very inconsistency they criticize.
- [Section 5, Conclusion] The conclusion states as a finding that LLMs 'fall short in understanding long-term temporal dependencies', but the evidence consists of architectural observations and aggregate benchmark charts rather than a comparison that isolates temporal reasoning from other factors. A model could fail on long videos because of context-window limits, token subsampling, or dataset bias rather than because it lacks temporal awareness. The causal attribution to 'the encoders' focus on short-term patterns' is underdetermined. Please either add an analysis that isolates the temporal component (e.g., comparing shuffled vs. ordered frames, short vs. long clips, or temporal-order questions vs. content questions) or explicitly soften the conclusion to a research hypothesis.
minor comments (6)
- [Keywords] The keyword 'Language language models' contains a typo; it should be 'Large language models'.
- [Table 3] Some dataset facts appear questionable: MSVD-QA is listed as having 'Start and end timestamps provided', but MSVD-QA questions are not typically timestamped; please verify and clarify the annotation type. Also, the average length listed for EPIC-KITCHENS (~458 s) refers to full raw videos rather than the annotated segments used in most benchmarks; please clarify.
- [References] References [26], [170], and [187] lack complete publication data (venue/year); please complete these entries.
- [Figure 4] The caption says 'Performance (accuracy) comparison', but the exact metric (overall accuracy vs. per-category accuracy, with or without subtitles) should be specified in the caption or text, and the definition should match the source benchmark.
- [Section 3, 'State-of-the-Art video LLMs'] The description of ActionFormer as part of 'video LLMs' is imprecise; ActionFormer is a temporal action localization transformer, not an LLM-based video model. Please tighten terminology to avoid conflating video transformers with video LLMs.
- [Section 2 and Section 3] The discussion of dataset limitations appears twice with overlapping content (Section 2 'Datasets for video understanding' and Section 3 'Video datasets: an enabler or bottleneck?'); consider consolidating to reduce redundancy.
Circularity Check
No significant circularity: the paper is a critical survey whose central claim is asserted, not derived, so there is no input-output equivalence to collapse.
full rationale
This is a position/survey paper, not a derivation or fitting exercise. It contains no equations, fitted parameters, or benchmark experiments through which a prediction could reduce by construction to an input. The load-bearing negative claim that LLMs "lack direct temporal awareness" and "fall short in understanding long-term temporal dependencies" is asserted from architectural observation and literature citations, not derived from any quantity defined in terms of itself. Unsupported or under-evidenced claims are an epistemic weakness, but unsupportedness is not circularity. The paper does cite the authors' own prior work (e.g., refs. [26], [125], [157], [212]), but those citations support auxiliary points about motion prompts, tracking, skeleton action recognition, and anomaly-detection datasets; they are not used to force the central conclusion about LLM temporal awareness, and no uniqueness theorem or fitted value is imported from them. The Figures 4-6 benchmark comparisons are drawn from external model reports, and the paper itself warns that evaluations are inconsistent; that weakens the empirical grounding of the survey but is a correctness/comparability concern, not a circular one. Accordingly, no circular step meets the quoted-reduction standard, and the appropriate score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Standard LLMs do not inherently model the flow of time unless explicitly trained on sequential video data.
- domain assumption Pretrained visual encoders generally emphasize short-term motion and spatial content over long-term temporal structure.
- domain assumption Existing video datasets largely lack temporal annotations and are biased toward short clips.
Cite this review
Pith. "Pith review of Do Language Models Understand Time?." pith.science (2026). https://pith.science/paper/G3V2S6PG
@misc{pith2026241213845,
author = {Pith},
title = {Pith review of: Do Language Models Understand Time?},
year = {2026},
howpublished = {\url{https://pith.science/paper/G3V2S6PG}},
note = {Machine review of arXiv:2412.13845}
}
read the original abstract
Large language models (LLMs) have revolutionized video-based computer vision applications, including action recognition, anomaly detection, and video summarization. Videos inherently pose unique challenges, combining spatial complexity with temporal dynamics that are absent in static images or textual data. Current approaches to video understanding with LLMs often rely on pretrained video encoders to extract spatiotemporal features and text encoders to capture semantic meaning. These representations are integrated within LLM frameworks, enabling multimodal reasoning across diverse video tasks. However, the critical question persists: Can LLMs truly understand the concept of time, and how effectively can they reason about temporal relationships in videos? This work critically examines the role of LLMs in video processing, with a specific focus on their temporal reasoning capabilities. We identify key limitations in the interaction between LLMs and pretrained encoders, revealing gaps in their ability to model long-term dependencies and abstract temporal concepts such as causality and event progression. Furthermore, we analyze challenges posed by existing video datasets, including biases, lack of temporal annotations, and domain-specific limitations that constrain the temporal understanding of LLMs. To address these gaps, we explore promising future directions, including the co-evolution of LLMs and encoders, the development of enriched datasets with explicit temporal labels, and innovative architectures for integrating spatial, temporal, and semantic reasoning. By addressing these challenges, we aim to advance the temporal comprehension of LLMs, unlocking their full potential in video analysis and beyond. Our paper's GitHub repository can be found at https://github.com/Darcyddx/Video-LLM.
Figures
Forward citations
Cited by 2 Pith papers
-
Evolving Skeletons: Motion Dynamics in Action Recognition
Taylor-transformed skeletons improve ST-GCN accuracy but reduce Hyperformer accuracy on NTU-60/120, indicating that motion-injected inputs do not universally benefit skeleton-based action recognition models.
-
Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight
A survey of 13 recent LLM/VLM-based video anomaly detection methods, organized by interpretability, temporal modeling, few-shot learning, and open-world detection.
Reference graph
Works this paper leans on
-
[64]
Wes Gurnee and Max Tegmark. 2024. Language Models Represent Space and Time. In The Twelfth International Conference on Learning Representations. https: //openreview.net/forum?id=jE8xbmvFin
2024
-
[106]
Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou. 2024. Tempcompass: Do video llms really understand videos? arXiv preprint arXiv:2403.00476 (2024)
arXiv 2024
-
[1]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Weizhu Chen, Yen- Chun Chen, Yi-Ling Chen, Hao Cheng, Parul Chopra, Xiyang Dai, Mat...
arXiv 2024
-
[2]
Amit Adam, Ehud Rivlin, Ilan Shimshoni, and Daviv Reinitz. 2008. Robust Real- Time Unusual Event Detection using Multiple Fixed-Location Monitors. IEEE Transactions on Pattern Analysis and Machine Intelligence 30, 3 (2008), 555–560. https://doi.org/10.1109/TPAMI.2007.70825
arXiv 2008
-
[3]
01. AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and...
arXiv 2024
-
[4]
Meta AI. 2024. Introducing Meta Llama 3: The most capable openly available LLM to date. https://ai.meta.com/blog/meta-llama-3/
2024
-
[5]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35 (2022), 23716–23736
2022
-
[6]
Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Deyao Zhu, Jian Ding, and Mohamed Elhoseiny. 2024. Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens. arXiv preprint arXiv:2404.03413 (2024)
arXiv 2024
Show all 229 references
-
[7]
Piyush Bagad, Makarand Tapaswi, and Cees GM Snoek. 2023. Test of time: Instilling video-language models with a sense of time. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2503–2516
2023
-
[8]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609 (2023)
2023 arXiv
-
[9]
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time atten- tion all you need for video understanding?. In ICML, Vol. 2. 4
2021
-
[10]
Andy Brock, Soham De, Samuel L Smith, and Karen Simonyan. 2021. High- performance large-scale image recognition without normalization. In Interna- tional conference on machine learning . PMLR, 1059–1071
2021
-
[11]
Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, and Desmond Elliott
-
[12]
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. 2022. COYO-700M: Image-Text Pair Dataset. https://github. com/kakaobrain/coyo-dataset
2022
-
[13]
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Car- los Niebles. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pat- tern recognition. 961–970
2015
-
[14]
Yuxuan Cai, Yizhuang Zhou, Qi Han, Jianjian Sun, Xiangwen Kong, Jun Li, and Xiangyu Zhang. 2022. Reversible column networks. arXiv preprint arXiv:2212.11696 (2022)
2022 arXiv
-
[15]
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiaoyi Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting Huang, Tao Jiang,...
2024 arXiv
-
[16]
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. 2018. A short note about kinetics-600.arXiv preprint arXiv:1808.01340 (2018)
2018 arXiv
-
[17]
Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. 2019. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987 (2019)
2019 arXiv
-
[18]
Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6299–6308
2017
-
[19]
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021. Con- ceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3558–3568
2021
-
[20]
Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, et al. 2023. Videollm: Modeling video sequence with large language models. arXiv preprint arXiv:2305.13292 (2023)
2023 arXiv
-
[21]
Huilin Chen, Lei Wang, Yifan Chen, Tom Gedeon, and Piotr Koniusz. 2024. When Spatial meets Temporal in Action Recognition. arXiv preprint arXiv:2411.15284 (2024)
2024 arXiv
-
[22]
Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, and Mike Zheng Shou. 2024. VideoLLM-online: Online Video Large Language Model for Streaming Video. In Proceedings of the IEEE/CVF Conference on Computer Vision...
2024
-
[23]
Jin Chen, Xinxiao Wu, Yao Hu, and Jiebo Luo. 2021. Spatial-temporal causal inference for partial image-to-video adaptation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 1027–1035
2021
-
[24]
Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, et al . 2024. Sharegpt4video: Improving video understanding and generation with better captions. arXiv preprint arXiv:2406.04325 (2024)
2024 arXiv
-
[25]
Ling-Hao Chen, Shunlin Lu, Ailing Zeng, Hao Zhang, Benyou Wang, Ruimao Zhang, and Lei Zhang. 2024. MotionLLM: Understanding Human Behaviors from Human Motions and Videos. arXiv preprint arXiv:2405.20340 (2024)
2024 arXiv
-
[26]
Qixiang Chen, Lei Wang, Piotr Koniusz, and Tom Gedeon. [n. d.]. Motion meets attention: Video motion prompts. In The 16th Asian Conference on Machine Learning (Conference Track)
-
[27]
Sihan Chen, Handong Li, Qunbo Wang, Zijia Zhao, Mingzhen Sun, Xinxin Zhu, and Jing Liu. 2023. Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset. Advances in Neural Information Processing Systems 36 (2023), 72842–72866
2023
-
[28]
Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, and Furu Wei. 2022. Beats: Audio pre-training with acoustic tokenizers. arXiv preprint arXiv:2212.09058 (2022)
2022 arXiv
-
[29]
Wenshuo Chen, Hongru Xiao, Erhang Zhang, Lijie Hu, Lei Wang, Mengyuan Liu, and Chen Chen. 2024. SATO: Stable Text-to-Motion Framework. In Proceedings of the 32nd ACM International Conference on Multimedia . 6989–6997
2024
-
[30]
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. 2022. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794 (2022)
2022 arXiv
-
[31]
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al . 2024. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling. arXiv preprint arXiv:2412.05271 (2024)
2024 arXiv
-
[32]
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al . 2024. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites. arXiv preprint arXiv:2404.16821 (2024)
2024 arXiv
-
[33]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Pro- ceedings of the IEEE/CVF Conference on Comp...
2024
-
[34]
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. 2024. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs. arXiv preprint arXiv:2406.07476 (2024)
2024 arXiv
-
[35]
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al
-
[36]
Clément Christophe, Praveen K Kanithi, Prateek Munjal, Tathagata Raha, Nasir Hayat, Ronnie Rajan, Ahmed Al-Mahrooqi, Avani Gupta, Muhammad Umar Salman, Gurpreet Gosal, et al. 2024. Med42–Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient...
2024 arXiv
-
[37]
StableLM contributors. 2023. StableLM: Stability AI language models. https: //github.com/stability-AI/stableLM
2023
-
[38]
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2018. Scaling Egocentric Vision: The EPIC- KITCHENS Dataset. ArXiv abs/1804.02748 (2018). http...
2018 arXiv
-
[39]
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al . 2022. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International J...
2022
-
[40]
Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso. 2013. A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2634–2641
2013
-
[41]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition . Ieee, 248–255
2009
-
[42]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Hum...
2019
-
[43]
Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W
Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2022. Time-Aware Language Models as Temporal Knowledge Bases. Transactions of the Association for Computational Linguistics 10 (03 2022), 257–273. https://doi.org...
2022 doi
-
[44]
Dexuan Ding, Lei Wang, Liyun Zhu, Tom Gedeon, and Piotr Koniusz. 2024. Lego: Learnable expansion of graph operators for multi-modal feature fusion. arXiv preprint arXiv:2410.01506 (2024)
2024 arXiv
-
[45]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognit...
2021
-
[46]
Hang Du, Sicheng Zhang, Binzhu Xie, Guoshun Nan, Jiayang Zhang, Junrui Xu, Hangyu Liu, Sicong Leng, Jiangming Liu, Hehe Fan, Dajiu Huang, Jing Feng, Linli Chen, Can Zhang, Xuhuan Li, Hao Zhang, Jianhang Chen, Qimei Cui, and Xiaofeng Tao. 2024. Uncovering What, Why and How: A C...
2024
-
[47]
Jiajun Fei, Dian Li, Zhidong Deng, Zekun Wang, Gang Liu, and Hui Wang
-
[48]
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019. Slow- fast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision . 6202–6211
2019
-
[49]
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. 2024. Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075 (2024)
2024 arXiv
-
[50]
Chaoyou Fu, Haojia Lin, Zuwei Long, Yunhang Shen, Meng Zhao, Yifan Zhang, Shaoqi Dong, Xiong Wang, Di Yin, Long Ma, et al. 2024. Vita: Towards open- source interactive omni multimodal llm. arXiv preprint arXiv:2408.05211 (2024)
2024 arXiv
-
[51]
Zhangwei Gao, Zhe Chen, Erfei Cui, Yiming Ren, Weiyun Wang, Jinguo Zhu, Hao Tian, Shenglong Ye, Junjun He, Xizhou Zhu, et al. 2024. Mini-internvl: A flexible- transfer pocket multimodal model with 5% parameters and 90% performance. arXiv preprint arXiv:2410.16261 (2024)
2024 arXiv
-
[52]
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Au- dio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal proc...
2017
-
[53]
Akash Ghosh, Arkadeep Acharya, Sriparna Saha, Vinija Jain, and Aman Chadha
-
[54]
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15180–15190
2023
-
[55]
Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh, Kalyan Vasudev Alwala, Ar- mand Joulin, and Ishan Misra. 2023. Omnimae: Single model masked pretraining on images and videos. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10406–10417
2023
-
[56]
Rohit Girdhar and Deva Ramanan. 2019. CATER: A diagnostic dataset for Com- positional Actions and TEmporal Reasoning. arXiv preprint arXiv:1910.04744 (2019)
2019 arXiv
-
[57]
arXiv preprint arXiv:2404.07214 (2024)
Exploring the frontier of vision-language models: A survey of current methodologies and future directions. arXiv preprint arXiv:2404.07214 (2024)
2024
-
[58]
something something
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. 2017. The" something something" video database for learning and evaluating visual common sense....
2017
-
[59]
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. 2022. Ego4d: Around the world in 3,000 hours of egocentric video. In Pro- ceedings of the IEEE/CVF Conference on Computer...
2022
-
[60]
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. 2018. Ava: A video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the IEEE c...
2018
-
[61]
Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens Van Der Maaten, Armand Joulin, and Ishan Misra. 2022. Omnivore: A single model for many visual modalities. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 16102–16112
2022
-
[62]
Yongxin Guo, Jingyu Liu, Mingda Li, Xiaoying Tang, Xi Chen, and Bo Zhao. 2024. VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding. arXiv preprint arXiv:2405.13382 (2024)
2024 arXiv
-
[63]
Yongxin Guo, Jingyu Liu, Mingda Li, Xiaoying Tang, Qingbin Liu, and Xi Chen
-
[65]
Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou, Niu Minzhe, Xiaodan Liang, Lewei Yao, Runhui Huang, Wei Zhang, Xin Jiang, et al. 2022. Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark. Advances in Neural Information Processing Systems 35 (2022), 26418–26431
2022
-
[66]
Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, and Andrew Zisserman. 2024. AutoAD III: The Prequel-Back to the Pixels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18164– 18174
2024
-
[67]
Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim. 2024. Ma-lmm: Memory-augmented large multimodal model for long-term video understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...
2024
-
[68]
arXiv preprint arXiv:2410.05643 (2024)
Trace: Temporal grounding video llm via causal event modeling. arXiv preprint arXiv:2410.05643 (2024)
2024 arXiv
-
[69]
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017. Localizing Moments in Video with Natural Language. arXiv:1708.01641 [cs.CV] https://arxiv.org/abs/1708.01641
2017 arXiv
-
[70]
Tengda Han, Max Bain, Arsha Nagrani, Gul Varol, Weidi Xie, and Andrew Zisserman. 2023. Autoad ii: The sequel-who, when, and what in movie audio description. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 13645–13655
2023
-
[71]
Hang Hua, Yunlong Tang, Chenliang Xu, and Jiebo Luo. 2024. V2xum-llm: Cross-modal video summarization with temporal prompt instruction tuning. arXiv preprint arXiv:2404.12353 (2024)
2024
-
[72]
Bin Huang, Xin Wang, Hong Chen, Zihan Song, and Wenwu Zhu. 2024. Vtimellm: Empower llm to grasp video moments. InProceedings of the IEEE/CVF Conference Do Language Models Understand Time? WWW Companion ’25, April 28-May 2, 2025, Sydney, NSW, Australia on Computer Vision and Pa...
2024
-
[73]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition . 770–778
2016
-
[74]
Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, and Gedas Bertasius. 2024. Video ReCap: Recursive Captioning of Hour-Long Videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18198–18208
2024
-
[75]
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556 (2022)
2022 arXiv
-
[76]
Raghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha, Adam Jatowt, and Sandipan Dandapat. 2023. Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language Models. In The 2023 Conference on Empirical Met...
2023
-
[77]
Yunseok Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2019. Video question answering with spatio-temporal reasoning. International Journal of Computer Vision 127 (2019), 1385–1412
2019
-
[78]
Thomas Hummel, Shyamgopal Karthik, Mariana-Iuliana Georgescu, and Zeynep Akata. 2024. EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval. arXiv preprint arXiv:2407.16658 (2024)
2024 arXiv
-
[79]
Cheonsu Jeong. 2024. Fine-tuning and utilization methods of domain-specific llms. arXiv preprint arXiv:2401.02981 (2024)
2024 arXiv
-
[80]
Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Daniel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, et al
-
[81]
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al . 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088 (2024)
2024 arXiv
-
[82]
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheen- dra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al
-
[83]
Byoungjip Kim, Dasol Hwang, Sungjun Cho, Youngsoo Jang, Honglak Lee, and Moontae Lee. 2024. Show Think and Tell: Thought-Augmented Fine-Tuning of Large Language Models for Video Captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . ...
2024
-
[84]
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2758–2766
2017
-
[85]
Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, and Ioannis Kompatsiaris. 2019. FIVR: Fine-grained incident video retrieval. IEEE Transac- tions on Multimedia 21, 10 (2019), 2638–2652
2019
-
[86]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al . 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)
2023 arXiv
-
[87]
Kuehne, H
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. 2011. HMDB: a large video database for human motion recognition. InProceedings of the International Conference on Computer Vision (ICCV)
2011
-
[88]
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre. 2011. HMDB: a large video database for human motion recogni- tion. In 2011 International conference on computer vision . IEEE, 2556–2563
2011
-
[89]
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. 2018. Tvqa: Localized, compositional video question answering. arXiv preprint arXiv:1809.01696 (2018)
2018 arXiv
-
[90]
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. 2020. Tvr: A large-scale dataset for video-subtitle moment retrieval. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI 16. Springer, 447–463
2020
-
[91]
Piotr Koniusz, Lei Wang, and Anoop Cherian. 2021. Tensor representations for action recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 2 (2021), 648–665
2021
-
[92]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning . PMLR, 19730– 19742
2023
-
[93]
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles
-
[94]
In Proceedings of the IEEE international conference on computer vision
Dense-captioning events in videos. In Proceedings of the IEEE international conference on computer vision . 706–715
-
[95]
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu
-
[96]
Yanwei Li, Chengyao Wang, and Jiaya Jia. 2025. Llama-vid: An image is worth 2 tokens in large language models. In European Conference on Computer Vision . Springer, 323–340
2025
-
[97]
Ruotong Liao, Max Erler, Huiyu Wang, Guangyao Zhai, Gengyuan Zhang, Yunpu Ma, and Volker Tresp. 2024. VideoINSTA: Zero-shot Long Video Understand- ing via Informative Spatial-Temporal Reasoning with LLMs. arXiv preprint arXiv:2409.20365 (2024)
2024 arXiv
-
[98]
Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. 2023. Video-llava: Learning united visual representation by alignment before projec- tion. arXiv preprint arXiv:2311.10122 (2023)
2023 arXiv
-
[99]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdel rahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Genera- tion, Translation, and Comprehension. In Annual Meeting of t...
2019
-
[100]
Xiangru Lin, Yuyang Chen, Guanbin Li, and Yizhou Yu. 2022. A causal inference look at unsupervised video anomaly detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 1620–1629
2022
-
[101]
KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355 (2023)
2023 arXiv
-
[102]
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. 2024. Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22195–22206
2024
-
[103]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruction Tuning. arXiv:2304.08485 [cs.CV] https://arxiv.org/abs/2304.08485
2023 arXiv
-
[104]
Jiajun Liu, Yibing Wang, Hanghang Ma, Xiaoping Wu, Xiaoqi Ma, Xiaoming Wei, Jianbin Jiao, Enhua Wu, and Jie Hu. 2024. Kangaroo: A powerful video-language model supporting long-context video input. arXiv preprint arXiv:2408.15542 (2024)
2024 arXiv
-
[105]
Ruyang Liu, Chen Li, Haoran Tang, Yixiao Ge, Ying Shan, and Ge Li. 2025. St-llm: Large language models are effective temporal learners. In European Conference on Computer Vision. Springer, 1–18
2025
-
[107]
Ye Liu, Siyuan Li, Yang Wu, Chang-Wen Chen, Ying Shan, and Xiaohu Qie
-
[108]
Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. 2024. Vila: On pre-training for visual language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 26689–26699
2024
-
[109]
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. 2022. Video swin transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 3202–3211
2022
-
[110]
Ding Liu, Zhaowen Wang, Yuchen Fan, Xianming Liu, Zhangyang Wang, Shiyu Chang, Xinchao Wang, and Thomas S. Huang. 2018. Learning Temporal Dynam- ics for Video Super-Resolution: A Deep Learning Approach. IEEE Transactions on Image Processing 27, 7 (2018), 3432–3445. https://doi...
2018 doi
-
[111]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved Baselines with Visual Instruction Tuning. arXiv:2310.03744 [cs.CV] https: //arxiv.org/abs/2310.03744
2024 arXiv
-
[112]
Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Da Li, Pengcheng Lu, Tao Wang, Linmei Hu, Minghui Qiu, and Zhongyu Wei. 2023. Valley: Video assistant with large language model enhanced ability. arXiv preprint arXiv:2306.07207 (2023)
2023 arXiv
-
[113]
Chenyang Lyu, Minghao Wu, Longyue Wang, Xinting Huang, Bingshuai Liu, Zefeng Du, Shuming Shi, and Zhaopeng Tu. 2023. Macaw-llm: Multi-modal language modeling with image, audio, video, and text integration.arXiv preprint arXiv:2306.09093 (2023)
2023 arXiv
-
[114]
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Khan. 2024. VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Under- standing. arXiv preprint arXiv:2406.09418 (2024)
2024 arXiv
-
[115]
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan
-
[116]
Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. 2023. Egoschema: A diagnostic benchmark for very long-form video language un- derstanding. Advances in Neural Information Processing Systems 36 (2023), 46212–46244
2023
-
[117]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3042–3051
-
[118]
Zuyan Liu, Yuhao Dong, Ziwei Liu, Winston Hu, Jiwen Lu, and Yongming Rao. 2024. Oryx mllm: On-demand spatial-temporal understanding at arbitrary resolution. arXiv preprint arXiv:2409.12961 (2024)
2024 arXiv
-
[119]
Ming Nie, Dan Ding, Chunwei Wang, Yuanfan Guo, Jianhua Han, Hang Xu, and Li Zhang. 2024. SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=FOkKndty5B
2024
-
[120]
Fuchen Long, Zhaofan Qiu, Ting Yao, and Tao Mei. 2024. Videodrafter: Content- consistent multi-scene video generation with llm.arXiv preprint arXiv:2401.01256 (2024)
2024 arXiv
-
[121]
Cewu Lu, Jianping Shi, and Jiaya Jia. 2013. Abnormal Event Detection at 150 FPS in MATLAB. In 2013 IEEE International Conference on Computer Vision . 2720–2727. https://doi.org/10.1109/ICCV.2013.338
2013 doi
-
[122]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al
-
[123]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9
2019
-
[124]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140 (2020), 1–67
2020
-
[125]
Arjun Raj, Lei Wang, and Tom Gedeon. 2024. TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention Maps. arXiv preprint arXiv:2409.14543 (2024)
2024 arXiv
-
[126]
arXiv preprint arXiv:2306.05424 (2023)
Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424 (2023)
2023 arXiv
-
[127]
Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor. 2021. ImageNet-21K Pretraining for the Masses. arXiv:2104.10972 [cs.CV]
2021 arXiv
-
[128]
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF international conference on computer vision ....
2019
-
[129]
Thong Nguyen, Yi Bin, Junbin Xiao, Leigang Qu, Yicong Li, Jay Zhangjie Wu, Cong-Duy Nguyen, See-Kiong Ng, and Luu Anh Tuan. 2024. Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives. arXiv preprint arXiv:2406.05615 (2024)
2024 arXiv
-
[130]
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural I...
2022
-
[131]
Mandela Patrick, Dylan Campbell, Yuki Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Joao F Henriques. 2021. Keeping your eye on the ball: Trajectory attention in video transformers. Advances in neural information processing systems 34 (2021), ...
2021
-
[132]
Zhenyue Qin, Yang Liu, Pan Ji, Dongwoo Kim, Lei Wang, Saeed Anwar, and Tom Gedeon. 2022. Fusing higher-order features in graph neural networks for skeleton-based action recognition. IEEE Transactions on Neural Networks and Learning Systems 35, 4 (2022), 4783–4797
2022
-
[133]
Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, De- An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, et al. 2024. Eagle: Exploring the design space for multimodal llms with mixture of encoders. arXiv preprint arXiv:2408.15998 (2024)
2024 arXiv
-
[134]
In International conference on machine learning
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
-
[135]
Yan Shu, Peitian Zhang, Zheng Liu, Minghao Qin, Junjie Zhou, Tiejun Huang, and Bo Zhao. 2024. Video-xl: Extra-long vision language model for hour-scale video understanding. arXiv preprint arXiv:2409.14485 (2024)
2024 arXiv
-
[136]
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In Computer Vision–ECCV 2016: 14th European Confer- ence, Amsterdam, The Netherlands, October 11–14, ...
2016
-
[137]
K Soomro. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)
2012 arXiv
-
[138]
Bharathkumar Ramachandra and Michael Jones. 2020. Street Scene: A new dataset and evaluation protocol for video anomaly detection. arXiv:1902.05872 [cs.CV] https://arxiv.org/abs/1902.05872
2020 arXiv
-
[139]
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid
-
[140]
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele. 2017. Movie descrip- tion. International Journal of Computer Vision 123 (2017), 94–120
2017
-
[141]
Arka Sadhu, Tanmay Gupta, Mark Yatskar, Ram Nevatia, and Aniruddha Kemb- havi. 2021. Visual semantic role labeling for video understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5589– 5600
2021
-
[142]
Yunlong Tang, Jing Bi, Siting Xu, Luchuan Song, Susan Liang, Teng Wang, Daoan Zhang, Jie An, Jingyang Lin, Rongyi Zhu, et al. 2023. Video understanding with large language models: A survey. arXiv preprint arXiv:2312.17432 (2023)
2023
-
[143]
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Con- ceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...
2018
-
[144]
Alex Sherstinsky. 2020. Fundamentals of recurrent neural network (RNN) and long short-term memory (LSTM) network. Physica D: Nonlinear Phenomena 404 (2020), 132306
2020
-
[145]
Together.xyz. 2023. Releasing 3b and 7b redpajama incite family of models including base, instruction-tuned and chat models. https://www.together.xyz/ blog/redpajama-models-v1
2023
-
[146]
Fangxun Shu, Lei Zhang, Hao Jiang, and Cihang Xie. 2023. Audio-visual llm for video understanding. arXiv preprint arXiv:2312.06720 (2023)
2023 arXiv
-
[147]
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri
-
[148]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. ...
2017
-
[149]
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. 2016. Generating videos with scene dynamics. Advances in neural information processing systems 29 (2016)
2016
-
[150]
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork. 2021. WIT: Wikipedia-Based Image Text Dataset for Multimodal Multi- lingual Machine Learning. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Inform...
2021
-
[151]
Junke Wang, Dongdong Chen, Chong Luo, Xiyang Dai, Lu Yuan, Zuxuan Wu, and Yu-Gang Jiang. 2023. Chatvideo: A tracklet-centric multimodal and versatile video understanding system. arXiv preprint arXiv:2304.14407 (2023)
2023 arXiv
-
[152]
Junke Wang, Dongdong Chen, Chong Luo, Bo He, Lu Yuan, Zuxuan Wu, and Yu-Gang Jiang. 2024. Omnivid: A generative framework for universal video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18209–18220
2024
-
[153]
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. 2023. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389 (2023)
2023 arXiv
-
[154]
Mingtian Tan, Mike A Merrill, Vinayak Gupta, Tim Althoff, and Thomas Hartvigsen. 2024. Are Language Models Actually Useful for Time Series Forecast- ing?. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=DV15UbHCY1
2024
-
[155]
Jiaqi Wang, Hanqi Jiang, Yiheng Liu, Chong Ma, Xu Zhang, Yi Pan, Mengyuan Liu, Peiran Gu, Sichen Xia, Wenjun Li, et al. 2024. A comprehensive review of multimodal large language models: Performance and challenges across different tasks. arXiv preprint arXiv:2408.01319 (2024)
2024 arXiv
-
[156]
Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. 2019. Coin: A large-scale dataset for comprehen- sive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1207–1216
2019
-
[157]
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016. Movieqa: Understanding stories in movies through question-answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4631–4640
2016
-
[158]
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao. 2023. Videomae v2: Scaling video masked autoencoders with dual masking. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14549–14560
2023
-
[159]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)
2023 arXiv
-
[160]
Lei Wang, Du Q Huynh, and Moussa Reda Mansour. 2019. Loss switching fusion with similarity search for video classification. In 2019 IEEE international conference on image processing (ICIP) . IEEE, 974–978
2019
-
[161]
Lei Wang and Piotr Koniusz. 2021. Self-supervising action recognition by statistical moment and subspace descriptors. In Proceedings of the 29th ACM international conference on multimedia . 4324–4333. Do Language Models Understand Time? WWW Companion ’25, April 28-May 2, 2025,...
2021
-
[162]
Lei Wang and Piotr Koniusz. 2022. Temporal-viewpoint transportation plan for skeletal few-shot action recognition. In Proceedings of the Asian Conference on Computer Vision. 4176–4193
2022
-
[163]
Lei Wang and Piotr Koniusz. 2022. Uncertainty-dtw for time series and se- quences. In European Conference on Computer Vision . Springer, 176–195
2022
-
[164]
Alex Jinpeng Wang, Linjie Li, Kevin Qinghong Lin, Jianfeng Wang, Kevin Lin, Zhengyuan Yang, Lijuan Wang, and Mike Zheng Shou. 2024. COSMO: COn- trastive Streamlined MultimOdal Model with Interleaved Pre-Training. arXiv preprint arXiv:2401.00849 (2024)
2024 arXiv
-
[165]
Lei Wang and Piotr Koniusz. 2024. Flow dynamics correction for action recog- nition. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3795–3799
2024
-
[166]
Lei Wang, Piotr Koniusz, and Du Q Huynh. 2019. Hallucinating idt descriptors and i3d optical flow features for action recognition with cnns. In Proceedings of the IEEE/CVF international conference on computer vision . 8698–8708
2019
-
[167]
Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, and Lu Yuan. 2022. Omnivl: One foundation model for image-language and video-language tasks. Advances in neural information processing systems 35 (2022), 5696–5710
2022
-
[168]
Jiexin Wang, Adam Jatowt, and Yi Cai. 2024. Towards Effective Time-Aware Language Representation: Exploring Enhanced Temporal Understanding in Language Models. arXiv preprint arXiv:2406.01863 (2024)
2024 arXiv
-
[169]
Lei Wang, Ke Sun, and Piotr Koniusz. 2024. High-order tensor pooling with at- tention for action recognition. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3885–3889
2024
-
[170]
Lei Wang. 2021. Analysis and evaluation of Kinect-based action recognition algorithms. arXiv preprint arXiv:2112.08626 (2021)
2021 arXiv
-
[171]
Lei Wang. 2023. Robust human action modelling . Ph. D. Dissertation. The Australian National University (Australia)
2023
-
[172]
Shaojie Wang, Wentian Zhao, Ziyi Kou, Jing Shi, and Chenliang Xu. 2021. How to make a blt sandwich? learning vqa towards understanding web instructional videos. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 1130–1139
2021
-
[173]
Lei Wang, Du Q Huynh, and Piotr Koniusz. 2019. A comparative review of recent kinect-based action recognition algorithms. IEEE Transactions on Image Processing 29 (2019), 15–28
2019
-
[174]
Weiyun Wang, Zhe Chen, Wenhai Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Jinguo Zhu, Xizhou Zhu, Lewei Lu, Yu Qiao, and Jifeng Dai. 2024. Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization. arXiv preprint arXiv:2411.10442 (2024)
2024 arXiv
-
[175]
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. 2019. Vatex: A large-scale, high-quality multilingual dataset for video- and-language research. In Proceedings of the IEEE/CVF international conference on computer vision. 4581–4591
2019
-
[176]
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Conghui He, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Yu Qiao. 2024. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation. ar...
2024 arXiv
-
[177]
Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Jilan Xu, Zun Wang, et al. 2024. Internvideo2: Scaling video foundation models for multimodal video understanding. arXiv e-prints (2024), arXiv–2403
2024
-
[178]
Lei Wang and Piotr Koniusz. 2023. 3mformer: Multi-order multi-mode trans- former for skeletal action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5620–5631
2023
-
[179]
Yuqing Wang, Tianwei Xiong, Daquan Zhou, Zhijie Lin, Yang Zhao, Bingyi Kang, Jiashi Feng, and Xihui Liu. 2024. Loong: Generating Minute-level Long Videos with Autoregressive Language Models. arXiv preprint arXiv:2410.02757 (2024)
2024 arXiv
-
[180]
Zhanyu Wang, Longyue Wang, Zhen Zhao, Minghao Wu, Chenyang Lyu, Huayang Li, Deng Cai, Luping Zhou, Shuming Shi, and Zhaopeng Tu. 2024. Gpt4video: A unified multimodal large language model for lnstruction-followed understanding and safety-aware generation. In Proceedings of the...
2024
-
[181]
Lei Wang, Jun Liu, and Piotr Koniusz. 2021. 3D Skeleton-based Few-shot Action Recognition with JEANIE is not so Naïve.arXiv preprint arXiv:2112.12668 (2021)
2021 arXiv
-
[182]
Lei Wang, Jun Liu, Liang Zheng, Tom Gedeon, and Piotr Koniusz. 2024. Meet JEANIE: a Similarity Measure for 3D Skeleton Sequences via Temporal- Viewpoint Alignment. International Journal of Computer Vision (2024), 1–32
2024
-
[183]
Peng Wu, Jing Liu, Yujia Shi, Yujia Sun, Fangtao Shao, Zhaoyang Wu, and Zhiwei Yang. 2020. Not only Look, but also Listen: Learning Multimodal Violence Detection under Weak Supervision. arXiv:2007.04687 [cs.CV] https://arxiv.org/ abs/2007.04687
2020 arXiv
-
[184]
Lei Wang, Xiuyuan Yuan, Tom Gedeon, and Liang Zheng. [n. d.]. Taylor Videos for Action Recognition. In Forty-first International Conference on Machine Learn- ing
-
[185]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. 2024. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv preprint arXiv:2409.12191 (2024)
2024 arXiv
-
[186]
Bo Xu and Mu-ming Poo. 2023. Large language models and brain-inspired general intelligence. National Science Review 10, 10 (2023), nwad267
2023
-
[187]
Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al. 2024. Visionllm: Large language model is also an open-ended decoder for vision-centric tasks.Advances in Neural Information Processing Systems 36 (2024)
2024
-
[188]
Haiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li, Bin Bi, Qi Qian, Wei Wang, et al. 2023. mplug-2: A modularized multi-modal foundation model across text, image and video. In International Conference on Machine Learning. PMLR, 38728–38748
2023
-
[189]
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5288–5296
2016
-
[190]
Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng
-
[191]
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont- Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. 2023. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. InProceedings of the IEEE/CVF Conference on Computer V...
2023
-
[192]
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al. 2022. Internvideo: General video foundation models via generative and discriminative learning. arXiv preprint arXiv:2212.03191 (2022)
2022 arXiv
-
[193]
Dongjie Yang, Suyuan Huang, Chengqiang Lu, Xiaodong Han, Haoxin Zhang, Yan Gao, Yao Hu, and Hai Zhao. 2024. Vript: A Video Is Worth Thousands of Words. arXiv preprint arXiv:2406.06040 (2024)
2024 arXiv
-
[194]
Qu Yang, Mang Ye, and Bo Du. 2024. Emollm: Multimodal emotional under- standing meets large language models. arXiv preprint arXiv:2406.16442 (2024)
2024 arXiv
-
[195]
Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan
-
[196]
arXiv preprint arXiv:2405.09711 (2024)
Star: A benchmark for situated reasoning in real-world videos. arXiv preprint arXiv:2405.09711 (2024)
2024 arXiv
-
[197]
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. 2023. Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models. arXiv:2303.04671 [cs.CV] https://arxiv.org/abs/2303.04671
2023 arXiv
-
[198]
Luca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang, and Elisa Ricci. 2024. Harnessing Large Language Models for Training-free Video Anom- aly Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18527–18536
2024
-
[199]
Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2023. Next-gpt: Any-to-any multimodal llm. arXiv preprint arXiv:2309.05519 (2023)
2023 arXiv
-
[200]
Weijia Wu, Yuzhong Zhao, Zhuang Li, Jiahong Li, Hong Zhou, Mike Zheng Shou, and Xiang Bai. 2023. A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension. arXiv:2305.03347 [cs.CV] https://arxiv.org/abs/2305.03347
2023 arXiv
-
[201]
Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction- tuned Audio-Visual Language Model for Video Understanding. In Conference on Empirical Methods in Natural Language Processing . https://api.semanticscholar. org/CorpusID:259075356
2023
-
[202]
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. [n. d.]. Video Question Answering via Gradually Refined Attention over Appearance and Motion. In ACM Multimedia
-
[203]
Jieyu Zhang, Weikai Huang, Zixian Ma, Oscar Michel, Dong He, Tanmay Gupta, Wei-Chiu Ma, Ali Farhadi, Aniruddha Kembhavi, and Ranjay Krishna. 2024. Task Me Anything. arXiv preprint arXiv:2406.11775 (2024)
2024 arXiv
-
[204]
Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen. 2023. T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations. arXiv:2301.06052 [cs.CV] https://arxiv.org/abs/2301.06052
2023 arXiv
-
[205]
Zhihan Zhang, Yixin Cao, Chenchen Ye, Yunshan Ma, Lizi Liao, and Tat-Seng Chua. 2024. Analyzing Temporal Complex Events with Large Language Models? A Benchmark towards Temporal, Long Context Understanding. arXiv preprint arXiv:2406.02472 (2024)
2024 arXiv
-
[206]
arXiv preprint arXiv:2404.16994 (2024)
Pllava: Parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994 (2024)
2024 arXiv
-
[207]
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar. 2023. Learning video representations from large language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6586–6597
2023
-
[208]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al . 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671 (2024)
2024 arXiv
-
[209]
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. 2018. Temporal relational reasoning in videos. In Proceedings of the European conference on computer vision (ECCV). 803–818
2018
-
[210]
Pengyuan Zhou, Lin Wang, Zhi Liu, Yanbin Hao, Pan Hui, Sasu Tarkoma, and Jussi Kangasharju. 2024. A survey on generative ai and llm for video generation, understanding, and streaming. arXiv preprint arXiv:2404.16038 (2024)
2024 arXiv
-
[211]
Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G
Lijun Yu, José Lezama, Nitesh B. Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G. Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A. Ross, and Lu Jiang. 2024. Language Model Beats Diffusion – To...
2024 arXiv
-
[212]
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. 2019. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 9127–9134
2019
-
[213]
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. 2021. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432 (2021)
2021 arXiv
-
[215]
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 11975–11986
2023
-
[216]
Chen-Lin Zhang, Jianxin Wu, and Yin Li. 2022. Actionformer: Localizing mo- ments of actions with transformers. In European Conference on Computer Vision . Springer, 492–510
2022
-
[218]
Huaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo, Chuchu Han, Xiaonan Huang, Changxin Gao, Yuehuan Wang, and Nong Sang. 2024. Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM. arXiv preprint arXiv:2406.12235 (2024)
2024 arXiv
-
[222]
Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. 2024. Expel: Llm agents are experiential learners. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 19632–19642. WWW Companion ’25, April 28-May 2, 2025, Sydney, NSW, A...
2024
-
[224]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT- Bench and Chatbot Arena. arXiv:2306.05685 [cs.CL] https://a...
2023 arXiv
-
[227]
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al. 2023. Languagebind: Ex- tending video-language pretraining to n-modality by language-based semantic alignment. arXiv preprint arXiv:2310.01852 (2023)
2023 arXiv
-
[228]
Liyun Zhu, Lei Wang, Arjun Raj, Tom Gedeon, and Chen Chen. 2024. Advanc- ing Video Anomaly Detection: A Concise Review and a New Dataset. In The Thirty-eighth Conference on Neural Information Processing Systems Datasets and Benchmarks Track
2024
-
[229]
Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Serena Yeung-Levy, and Xide Xia. 2024. Apollo: An Exploration of Video Understanding in Large Multimodal Models. arXiv:2412.10360 [cs...
2024 arXiv
-
[2015]
In Proceedings of the IEEE international conference on computer vision
Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision . 4489–4497
-
[2017]
arXiv preprint arXiv:1705.06950 (2017)
The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017)
2017 arXiv
-
[2019]
In Proceedings of the IEEE/CVF international conference on computer vision
Videobert: A joint model for video and language representation learning. In Proceedings of the IEEE/CVF international conference on computer vision . 7464– 7473
-
[2020]
arXiv preprint arXiv:2005.00200 (2020)
Hero: Hierarchical encoder for video+ language omni-representation pre-training. arXiv preprint arXiv:2005.00200 (2020)
2020 arXiv
-
[2021]
Transactions of the Association for Compu- tational Linguistics 9 (2021), 978–994
Multimodal pretraining unmasked: A meta-analysis and a unified frame- work of vision-and-language BERTs. Transactions of the Association for Compu- tational Linguistics 9 (2021), 978–994
2021
-
[2022]
arXiv preprint arXiv:2212.12017 (2022)
Opt-iml: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint arXiv:2212.12017 (2022)
2022 arXiv
-
[2023]
See https://vicuna
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023) 2, 3 (2023), 6
2023
-
[2024]
arXiv preprint arXiv:2408.14023 (2024)
Video-ccam: Enhancing video-language understanding with causal cross- attention masks for short and long videos. arXiv preprint arXiv:2408.14023 (2024)
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.