Pith. sign in

REVIEW 3 major objections 6 minor 2 cited by

Do Language Models Understand Time?

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper argues that today's video-LLMs achieve benchmark success without genuinely understanding time, because temporal structure is supplied by pretrained encoders and datasets, not learned by the LLM.

desk verdict A useful survey of video-LLMs and datasets whose central claim about temporal awareness is asserted rather than tested. read the letter →

arxiv 2412.13845 v3 pith:G3V2S6PG submitted 2024-12-18 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords temporalreasoningvideounderstandinglargelanguagemodelsvideo-LLMspretrainedvisualencodersannotationsbenchmarksmultimodalfusion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to answer whether large language models used for video understanding actually understand time. Its answer is no: LLMs paired with pretrained visual encoders can label actions, answer questions, and caption clips, but they lack direct temporal awareness, in the sense that they never see a sequence they have to reason about as a sequence; the encoder supplies tokens and motion features, and the LLM attends over them without an inherent model of event order, causality, or duration. The paper supports this by reviewing recent video-LLMs, their encoder-LLM fusion mechanisms, and video datasets, arguing that both the encoders and the datasets are biased toward short-term motion and static appearance, and that evaluations mix incompatible protocols. If this is right, current benchmark success is not evidence of temporal reasoning, and progress depends on joint encoder-LLM training, explicit temporal annotations, and long-context evaluation.

What carries the argument

The load-bearing object is the interaction/fusion mechanism between a pretrained visual encoder and an LLM: projection layers, cross-attention modules, Q-Formers, and temporal-specific connectors such as temporal samplers and scene-level sequential alignment. The paper's argument is that this coupling is where temporal information is filtered or lost. Encoders such as CLIP, I3D, TimeSformer, and Video Swin supply frame-level tokens and short-range motion features; the LLM then aligns those with language without any inherent model of sequence order, so any true temporal structure must come from outside. The distribution of fusion mechanisms across recent models, and the benchmark plots built on this coupling, are the evidence the review leans on.

What would settle it

Run one controlled experiment: take a current video-LLM and ask the same set of order-sensitive questions on the same videos in correct and in temporally shuffled frame order, with subtitles off and a fixed prompt. If accuracy on the shuffled version stays at the same level, the model is not using order and the claim of lacking temporal awareness is supported; if accuracy collapses on the shuffled version, the claim that LLMs lack direct temporal awareness would be falsified.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a negative result about current architectures: state-of-the-art video-LLMs achieve competitive task performance through the interaction of pretrained visual encoders and LLMs, but this performance does not amount to temporal comprehension. The paper claims that LLMs lack direct temporal awareness, that encoders focus on short-term patterns and fragmented cues, and that datasets rarely carry the temporal annotations—event order, duration, causality—that would let either component learn long-term dependencies. The conclusion is that no current video-LLM excels across video tasks and that the apparent progress is partly an artifact of inconsistent evaluation.

Load-bearing premise

The argument depends on the benchmark numbers plotted across Figures 4-6 being comparable; if those models were evaluated under different prompts, subtitles settings, sampling, or metric versions, the performance gaps and the conclusion that no video-LLM excels across tasks do not follow.

Editorial extensions

If this is right

  • Benchmark numbers on Video-MME, MSVD-QA, MSRVTT-QA, ActivityNet-QA, and retrieval or captioning sets should not be read as measures of temporal reasoning, because they mix encoder quality, language priors, and short-term cues.
  • Improving temporal understanding requires training encoders and LLMs jointly on temporally annotated data, not just scaling model size or context length.
  • New benchmarks must include explicit temporal labels such as event order, duration, and causality, and must use long-form videos; otherwise progress in temporal reasoning will be mismeasured.
  • Most current models use projection layers and cross-attention rather than temporal-specific fusion mechanisms, so temporal structure is imposed only weakly and inconsistently.
  • Comparisons of video-LLMs against traditional video models such as I3D, SlowFast, and Video Swin are unfair; within-paradigm benchmarking against other video-LLMs is needed to reveal real progress.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper is right, a cheap diagnostic follows: shuffling frame order or reversing event order in existing test videos should produce a large accuracy drop on order-sensitive questions, and this could be measured today without new training.
  • The same short-term bias would likely carry over to video generation and long-context models, because a longer context window does not by itself supply training signal about causality or event progression.
  • One way to test the dataset bottleneck directly is to take a fixed video-LLM and fine-tune it on a small set of temporally annotated long videos; if temporal reasoning improves substantially, the bottleneck is data rather than architecture.
  • The paper's critique generalizes to time-series and event-forecasting uses of LLMs, where surface task success may again come from textual priors rather than from an internal model of time.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper is a survey/critical review of video-based large language models (video-LLMs) and their temporal reasoning capabilities. It reviews recent video-LLM architectures and their interaction mechanisms with pretrained visual encoders (Tables 1–2), catalogs video datasets across tasks such as action recognition, video QA, captioning, retrieval, and anomaly detection (Table 3), and presents benchmark comparisons of video-LLMs (Figures 4–6). The central claim is that LLMs, despite achieving strong task performance when paired with pretrained video encoders, lack direct temporal awareness and fall short in understanding long-term temporal dependencies. The paper identifies dataset limitations (lack of temporal annotations, short-term bias, low diversity) and proposes future directions including joint encoder-LLM training, richer temporal annotations, and new fusion architectures.

Significance. The paper is a broad and timely survey that usefully organizes a fast-moving literature on video-LLMs, fusion mechanisms, and video datasets. Its strength is the critical framing: it explicitly argues that benchmark comparisons are often inconsistent and that dataset design shapes temporal reasoning, which is a valuable message for the community. The GitHub repository and the extensive tables (Tables 1–3) provide a convenient reference. However, the paper's central claim about the absence of direct temporal awareness is asserted rather than demonstrated, and the benchmark figures are presented without a documented extraction protocol. If the central claim were supported by controlled evidence or a systematic benchmark analysis, the paper would be a significant position piece; as it stands, its conclusions outrun its evidence. The paper contains no new experiments, code, or machine-checked derivations, so its value depends on the accuracy and representativeness of the plotted numbers and the force of its conceptual argument.

major comments (3)
  1. [Section 3, 'Analysis and Discussion'] The central thesis that LLMs 'lack direct temporal awareness' is load-bearing for the paper, but it is neither operationalized nor tested. The same section states that LLMs 'can infer temporal relationships through contextual cues such as "first", "then", and "after"', which is itself a form of temporal inference, and the paper cites Ref. [64] (Gurnee and Tegmark) and Ref. [106] (TempCompass) without reconciling their evidence of temporal representation in LLMs with the negative claim. Please define what would count as 'direct' temporal awareness (e.g., whether positional embeddings, frame-order tokens, or learned temporal projections count), state the claim as a falsifiable hypothesis, and support it with either a controlled experiment or a systematic analysis of existing temporal-reasoning benchmarks such as TempCompass, TGIF-QA, or MVBench.
  2. [Figures 4–6] The benchmark figures are presented without the extraction protocol needed to support the claim that no video-LLM excels across tasks. For each plotted number, the paper should report the exact model checkpoint, prompt template, subtitles-on/off condition, frame sampling, metric version, and original source. As written, the numbers mix models, tasks, and possibly evaluation settings, so the performance gaps in Figures 4–6 do not logically follow. This is a particular problem because the paper itself (Section 'Fair evaluation is needed') argues that inconsistent evaluations produce misleading conclusions; the charts should not reproduce the very inconsistency they criticize.
  3. [Section 5, Conclusion] The conclusion states as a finding that LLMs 'fall short in understanding long-term temporal dependencies', but the evidence consists of architectural observations and aggregate benchmark charts rather than a comparison that isolates temporal reasoning from other factors. A model could fail on long videos because of context-window limits, token subsampling, or dataset bias rather than because it lacks temporal awareness. The causal attribution to 'the encoders' focus on short-term patterns' is underdetermined. Please either add an analysis that isolates the temporal component (e.g., comparing shuffled vs. ordered frames, short vs. long clips, or temporal-order questions vs. content questions) or explicitly soften the conclusion to a research hypothesis.
minor comments (6)
  1. [Keywords] The keyword 'Language language models' contains a typo; it should be 'Large language models'.
  2. [Table 3] Some dataset facts appear questionable: MSVD-QA is listed as having 'Start and end timestamps provided', but MSVD-QA questions are not typically timestamped; please verify and clarify the annotation type. Also, the average length listed for EPIC-KITCHENS (~458 s) refers to full raw videos rather than the annotated segments used in most benchmarks; please clarify.
  3. [References] References [26], [170], and [187] lack complete publication data (venue/year); please complete these entries.
  4. [Figure 4] The caption says 'Performance (accuracy) comparison', but the exact metric (overall accuracy vs. per-category accuracy, with or without subtitles) should be specified in the caption or text, and the definition should match the source benchmark.
  5. [Section 3, 'State-of-the-Art video LLMs'] The description of ActionFormer as part of 'video LLMs' is imprecise; ActionFormer is a temporal action localization transformer, not an LLM-based video model. Please tighten terminology to avoid conflating video transformers with video LLMs.
  6. [Section 2 and Section 3] The discussion of dataset limitations appears twice with overlapping content (Section 2 'Datasets for video understanding' and Section 3 'Video datasets: an enabler or bottleneck?'); consider consolidating to reduce redundancy.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a critical survey whose central claim is asserted, not derived, so there is no input-output equivalence to collapse.

full rationale

This is a position/survey paper, not a derivation or fitting exercise. It contains no equations, fitted parameters, or benchmark experiments through which a prediction could reduce by construction to an input. The load-bearing negative claim that LLMs "lack direct temporal awareness" and "fall short in understanding long-term temporal dependencies" is asserted from architectural observation and literature citations, not derived from any quantity defined in terms of itself. Unsupported or under-evidenced claims are an epistemic weakness, but unsupportedness is not circularity. The paper does cite the authors' own prior work (e.g., refs. [26], [125], [157], [212]), but those citations support auxiliary points about motion prompts, tracking, skeleton action recognition, and anomaly-detection datasets; they are not used to force the central conclusion about LLM temporal awareness, and no uniqueness theorem or fitted value is imported from them. The Figures 4-6 benchmark comparisons are drawn from external model reports, and the paper itself warns that evaluations are inconsistent; that weakens the empirical grounding of the survey but is a correctness/comparability concern, not a circular one. Accordingly, no circular step meets the quoted-reduction standard, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper's central argument rests on broad domain assumptions about LLM temporal awareness, encoder priorities, and dataset annotation quality. These are plausible but asserted rather than established here, and they do the work of making the survey's conclusions coherent. There are no fitted parameters or invented entities.

assumptions (3)
  • domain assumption Standard LLMs do not inherently model the flow of time unless explicitly trained on sequential video data.
    Stated in Section 3; unsupported by direct experiments and conflicts with cited evidence [64] that LLMs encode spatial and temporal representations.
  • domain assumption Pretrained visual encoders generally emphasize short-term motion and spatial content over long-term temporal structure.
    Used throughout Sections 3 and 4; supported only by selected citations, not by a systematic comparison of encoders on temporal reasoning.
  • domain assumption Existing video datasets largely lack temporal annotations and are biased toward short clips.
    Stated in Section 3 and Table 3; many datasets in Table 3 do provide start and end timestamps, so the claim is overgeneralized.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do Language Models Understand Time?." pith.science (2026). https://pith.science/paper/G3V2S6PG

@misc{pith2026241213845,
  author       = {Pith},
  title        = {Pith review of: Do Language Models Understand Time?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/G3V2S6PG}},
  note         = {Machine review of arXiv:2412.13845}
}
read the original abstract

Large language models (LLMs) have revolutionized video-based computer vision applications, including action recognition, anomaly detection, and video summarization. Videos inherently pose unique challenges, combining spatial complexity with temporal dynamics that are absent in static images or textual data. Current approaches to video understanding with LLMs often rely on pretrained video encoders to extract spatiotemporal features and text encoders to capture semantic meaning. These representations are integrated within LLM frameworks, enabling multimodal reasoning across diverse video tasks. However, the critical question persists: Can LLMs truly understand the concept of time, and how effectively can they reason about temporal relationships in videos? This work critically examines the role of LLMs in video processing, with a specific focus on their temporal reasoning capabilities. We identify key limitations in the interaction between LLMs and pretrained encoders, revealing gaps in their ability to model long-term dependencies and abstract temporal concepts such as causality and event progression. Furthermore, we analyze challenges posed by existing video datasets, including biases, lack of temporal annotations, and domain-specific limitations that constrain the temporal understanding of LLMs. To address these gaps, we explore promising future directions, including the co-evolution of LLMs and encoders, the development of enriched datasets with explicit temporal labels, and innovative architectures for integrating spatial, temporal, and semantic reasoning. By addressing these challenges, we aim to advance the temporal comprehension of LLMs, unlocking their full potential in video analysis and beyond. Our paper's GitHub repository can be found at https://github.com/Darcyddx/Video-LLM.

Figures

Figures reproduced from arXiv: 2412.13845 by the authors.

Figure 1
Figure 1. Do language models understand time? In the kitchen arena, where burritos are rolled, rice waits patiently, and sauce [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Performance comparison of visual encoders. ( [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. The distributions of interaction/fusion mechanisms [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Performance (accuracy) comparison of recent video [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 6
Figure 6. Figure 6: Performance comparison of recent video-LLMs on [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evolving Skeletons: Motion Dynamics in Action Recognition

    cs.CV 2025-01 conditional novelty 5.0 of 10

    Taylor-transformed skeletons improve ST-GCN accuracy but reduce Hyperformer accuracy on NTU-60/120, indicating that motion-injected inputs do not universally benefit skeleton-based action recognition models.

  2. Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight

    cs.CV 2024-12 conditional novelty 2.0 of 10

    A survey of 13 recent LLM/VLM-based video anomaly detection methods, organized by interpretability, temporal modeling, few-shot learning, and open-world detection.

Reference graph

Works this paper leans on

229 extracted references · 6 canonical work pages · cited by 2 Pith papers

  1. [64]

    Wes Gurnee and Max Tegmark. 2024. Language Models Represent Space and Time. In The Twelfth International Conference on Learning Representations. https: //openreview.net/forum?id=jE8xbmvFin

  2. [106]

    Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou. 2024. Tempcompass: Do video llms really understand videos? arXiv preprint arXiv:2403.00476 (2024)

  3. [1]

    Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Weizhu Chen, Yen- Chun Chen, Yi-Ling Chen, Hao Cheng, Parul Chopra, Xiyang Dai, Mat...

  4. [2]

    Amit Adam, Ehud Rivlin, Ilan Shimshoni, and Daviv Reinitz. 2008. Robust Real- Time Unusual Event Detection using Multiple Fixed-Location Monitors. IEEE Transactions on Pattern Analysis and Machine Intelligence 30, 3 (2008), 555–560. https://doi.org/10.1109/TPAMI.2007.70825

  5. [3]

    01. AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and...

  6. [4]

    Meta AI. 2024. Introducing Meta Llama 3: The most capable openly available LLM to date. https://ai.meta.com/blog/meta-llama-3/

  7. [5]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35 (2022), 23716–23736

  8. [6]

    Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Deyao Zhu, Jian Ding, and Mohamed Elhoseiny. 2024. Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens. arXiv preprint arXiv:2404.03413 (2024)

Show all 229 references
  1. [7]

    Piyush Bagad, Makarand Tapaswi, and Cees GM Snoek. 2023. Test of time: Instilling video-language models with a sense of time. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2503–2516

  2. [8]

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609 (2023)

  3. [9]

    Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time atten- tion all you need for video understanding?. In ICML, Vol. 2. 4

  4. [10]

    Andy Brock, Soham De, Samuel L Smith, and Karen Simonyan. 2021. High- performance large-scale image recognition without normalization. In Interna- tional conference on machine learning . PMLR, 1059–1071

  5. [11]

    Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, and Desmond Elliott

  6. [12]

    Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. 2022. COYO-700M: Image-Text Pair Dataset. https://github. com/kakaobrain/coyo-dataset

  7. [13]

    Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Car- los Niebles. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pat- tern recognition. 961–970

  8. [14]

    Yuxuan Cai, Yizhuang Zhou, Qi Han, Jianjian Sun, Xiangwen Kong, Jun Li, and Xiangyu Zhang. 2022. Reversible column networks. arXiv preprint arXiv:2212.11696 (2022)

  9. [15]

    Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiaoyi Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting Huang, Tao Jiang,...

  10. [16]

    Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. 2018. A short note about kinetics-600.arXiv preprint arXiv:1808.01340 (2018)

  11. [17]

    Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. 2019. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987 (2019)

  12. [18]

    Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6299–6308

  13. [19]

    Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021. Con- ceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3558–3568

  14. [20]

    Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, et al. 2023. Videollm: Modeling video sequence with large language models. arXiv preprint arXiv:2305.13292 (2023)

  15. [21]

    Huilin Chen, Lei Wang, Yifan Chen, Tom Gedeon, and Piotr Koniusz. 2024. When Spatial meets Temporal in Action Recognition. arXiv preprint arXiv:2411.15284 (2024)

  16. [22]

    Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, and Mike Zheng Shou. 2024. VideoLLM-online: Online Video Large Language Model for Streaming Video. In Proceedings of the IEEE/CVF Conference on Computer Vision...

  17. [23]

    Jin Chen, Xinxiao Wu, Yao Hu, and Jiebo Luo. 2021. Spatial-temporal causal inference for partial image-to-video adaptation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 1027–1035

  18. [24]

    Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, et al . 2024. Sharegpt4video: Improving video understanding and generation with better captions. arXiv preprint arXiv:2406.04325 (2024)

  19. [25]

    Ling-Hao Chen, Shunlin Lu, Ailing Zeng, Hao Zhang, Benyou Wang, Ruimao Zhang, and Lei Zhang. 2024. MotionLLM: Understanding Human Behaviors from Human Motions and Videos. arXiv preprint arXiv:2405.20340 (2024)

  20. [26]

    Qixiang Chen, Lei Wang, Piotr Koniusz, and Tom Gedeon. [n. d.]. Motion meets attention: Video motion prompts. In The 16th Asian Conference on Machine Learning (Conference Track)

  21. [27]

    Sihan Chen, Handong Li, Qunbo Wang, Zijia Zhao, Mingzhen Sun, Xinxin Zhu, and Jing Liu. 2023. Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset. Advances in Neural Information Processing Systems 36 (2023), 72842–72866

  22. [28]

    Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, and Furu Wei. 2022. Beats: Audio pre-training with acoustic tokenizers. arXiv preprint arXiv:2212.09058 (2022)

  23. [29]

    Wenshuo Chen, Hongru Xiao, Erhang Zhang, Lijie Hu, Lei Wang, Mengyuan Liu, and Chen Chen. 2024. SATO: Stable Text-to-Motion Framework. In Proceedings of the 32nd ACM International Conference on Multimedia . 6989–6997

  24. [30]

    Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. 2022. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794 (2022)

  25. [31]

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al . 2024. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling. arXiv preprint arXiv:2412.05271 (2024)

  26. [32]

    Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al . 2024. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites. arXiv preprint arXiv:2404.16821 (2024)

  27. [33]

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Pro- ceedings of the IEEE/CVF Conference on Comp...

  28. [34]

    Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. 2024. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs. arXiv preprint arXiv:2406.07476 (2024)

  29. [35]

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al

  30. [36]

    Clément Christophe, Praveen K Kanithi, Prateek Munjal, Tathagata Raha, Nasir Hayat, Ronnie Rajan, Ahmed Al-Mahrooqi, Avani Gupta, Muhammad Umar Salman, Gurpreet Gosal, et al. 2024. Med42–Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient...

  31. [37]

    StableLM contributors. 2023. StableLM: Stability AI language models. https: //github.com/stability-AI/stableLM

  32. [38]

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2018. Scaling Egocentric Vision: The EPIC- KITCHENS Dataset. ArXiv abs/1804.02748 (2018). http...

  33. [39]

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al . 2022. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International J...

  34. [40]

    Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso. 2013. A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2634–2641

  35. [41]

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition . Ieee, 248–255

  36. [42]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Hum...

  37. [43]

    Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W

    Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2022. Time-Aware Language Models as Temporal Knowledge Bases. Transactions of the Association for Computational Linguistics 10 (03 2022), 257–273. https://doi.org...

  38. [44]

    Dexuan Ding, Lei Wang, Liyun Zhu, Tom Gedeon, and Piotr Koniusz. 2024. Lego: Learnable expansion of graph operators for multi-modal feature fusion. arXiv preprint arXiv:2410.01506 (2024)

  39. [45]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognit...

  40. [46]

    Hang Du, Sicheng Zhang, Binzhu Xie, Guoshun Nan, Jiayang Zhang, Junrui Xu, Hangyu Liu, Sicong Leng, Jiangming Liu, Hehe Fan, Dajiu Huang, Jing Feng, Linli Chen, Can Zhang, Xuhuan Li, Hao Zhang, Jianhang Chen, Qimei Cui, and Xiaofeng Tao. 2024. Uncovering What, Why and How: A C...

  41. [47]

    Jiajun Fei, Dian Li, Zhidong Deng, Zekun Wang, Gang Liu, and Hui Wang

  42. [48]

    Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019. Slow- fast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision . 6202–6211

  43. [49]

    Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. 2024. Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075 (2024)

  44. [50]

    Chaoyou Fu, Haojia Lin, Zuwei Long, Yunhang Shen, Meng Zhao, Yifan Zhang, Shaoqi Dong, Xiong Wang, Di Yin, Long Ma, et al. 2024. Vita: Towards open- source interactive omni multimodal llm. arXiv preprint arXiv:2408.05211 (2024)

  45. [51]

    Zhangwei Gao, Zhe Chen, Erfei Cui, Yiming Ren, Weiyun Wang, Jinguo Zhu, Hao Tian, Shenglong Ye, Junjun He, Xizhou Zhu, et al. 2024. Mini-internvl: A flexible- transfer pocket multimodal model with 5% parameters and 90% performance. arXiv preprint arXiv:2410.16261 (2024)

  46. [52]

    Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Au- dio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal proc...

  47. [53]

    Akash Ghosh, Arkadeep Acharya, Sriparna Saha, Vinija Jain, and Aman Chadha

  48. [54]

    Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15180–15190

  49. [55]

    Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh, Kalyan Vasudev Alwala, Ar- mand Joulin, and Ishan Misra. 2023. Omnimae: Single model masked pretraining on images and videos. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10406–10417

  50. [56]

    Rohit Girdhar and Deva Ramanan. 2019. CATER: A diagnostic dataset for Com- positional Actions and TEmporal Reasoning. arXiv preprint arXiv:1910.04744 (2019)

  51. [57]

    arXiv preprint arXiv:2404.07214 (2024)

    Exploring the frontier of vision-language models: A survey of current methodologies and future directions. arXiv preprint arXiv:2404.07214 (2024)

  52. [58]

    something something

    Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. 2017. The" something something" video database for learning and evaluating visual common sense....

  53. [59]

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. 2022. Ego4d: Around the world in 3,000 hours of egocentric video. In Pro- ceedings of the IEEE/CVF Conference on Computer...

  54. [60]

    Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. 2018. Ava: A video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the IEEE c...

  55. [61]

    Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens Van Der Maaten, Armand Joulin, and Ishan Misra. 2022. Omnivore: A single model for many visual modalities. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 16102–16112

  56. [62]

    Yongxin Guo, Jingyu Liu, Mingda Li, Xiaoying Tang, Xi Chen, and Bo Zhao. 2024. VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding. arXiv preprint arXiv:2405.13382 (2024)

  57. [63]

    Yongxin Guo, Jingyu Liu, Mingda Li, Xiaoying Tang, Qingbin Liu, and Xi Chen

  58. [65]

    Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou, Niu Minzhe, Xiaodan Liang, Lewei Yao, Runhui Huang, Wei Zhang, Xin Jiang, et al. 2022. Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark. Advances in Neural Information Processing Systems 35 (2022), 26418–26431

  59. [66]

    Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, and Andrew Zisserman. 2024. AutoAD III: The Prequel-Back to the Pixels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18164– 18174

  60. [67]

    Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim. 2024. Ma-lmm: Memory-augmented large multimodal model for long-term video understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...

  61. [68]

    arXiv preprint arXiv:2410.05643 (2024)

    Trace: Temporal grounding video llm via causal event modeling. arXiv preprint arXiv:2410.05643 (2024)

  62. [69]

    Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017. Localizing Moments in Video with Natural Language. arXiv:1708.01641 [cs.CV] https://arxiv.org/abs/1708.01641

  63. [70]

    Tengda Han, Max Bain, Arsha Nagrani, Gul Varol, Weidi Xie, and Andrew Zisserman. 2023. Autoad ii: The sequel-who, when, and what in movie audio description. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 13645–13655

  64. [71]

    Hang Hua, Yunlong Tang, Chenliang Xu, and Jiebo Luo. 2024. V2xum-llm: Cross-modal video summarization with temporal prompt instruction tuning. arXiv preprint arXiv:2404.12353 (2024)

  65. [72]

    Bin Huang, Xin Wang, Hong Chen, Zihan Song, and Wenwu Zhu. 2024. Vtimellm: Empower llm to grasp video moments. InProceedings of the IEEE/CVF Conference Do Language Models Understand Time? WWW Companion ’25, April 28-May 2, 2025, Sydney, NSW, Australia on Computer Vision and Pa...

  66. [73]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition . 770–778

  67. [74]

    Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, and Gedas Bertasius. 2024. Video ReCap: Recursive Captioning of Hour-Long Videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18198–18208

  68. [75]

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556 (2022)

  69. [76]

    Raghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha, Adam Jatowt, and Sandipan Dandapat. 2023. Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language Models. In The 2023 Conference on Empirical Met...

  70. [77]

    Yunseok Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2019. Video question answering with spatio-temporal reasoning. International Journal of Computer Vision 127 (2019), 1385–1412

  71. [78]

    Thomas Hummel, Shyamgopal Karthik, Mariana-Iuliana Georgescu, and Zeynep Akata. 2024. EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval. arXiv preprint arXiv:2407.16658 (2024)

  72. [79]

    Cheonsu Jeong. 2024. Fine-tuning and utilization methods of domain-specific llms. arXiv preprint arXiv:2401.02981 (2024)

  73. [80]

    Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Daniel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, et al

  74. [81]

    Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al . 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088 (2024)

  75. [82]

    Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheen- dra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al

  76. [83]

    Byoungjip Kim, Dasol Hwang, Sungjun Cho, Youngsoo Jang, Honglak Lee, and Moontae Lee. 2024. Show Think and Tell: Thought-Augmented Fine-Tuning of Large Language Models for Video Captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . ...

  77. [84]

    Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2758–2766

  78. [85]

    Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, and Ioannis Kompatsiaris. 2019. FIVR: Fine-grained incident video retrieval. IEEE Transac- tions on Multimedia 21, 10 (2019), 2638–2652

  79. [86]

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al . 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)

  80. [87]

    Kuehne, H

    H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. 2011. HMDB: a large video database for human motion recognition. InProceedings of the International Conference on Computer Vision (ICCV)

  81. [88]

    Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre. 2011. HMDB: a large video database for human motion recogni- tion. In 2011 International conference on computer vision . IEEE, 2556–2563

  82. [89]

    Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. 2018. Tvqa: Localized, compositional video question answering. arXiv preprint arXiv:1809.01696 (2018)

  83. [90]

    Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. 2020. Tvr: A large-scale dataset for video-subtitle moment retrieval. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI 16. Springer, 447–463

  84. [91]

    Piotr Koniusz, Lei Wang, and Anoop Cherian. 2021. Tensor representations for action recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 2 (2021), 648–665

  85. [92]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning . PMLR, 19730– 19742

  86. [93]

    Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles

  87. [94]

    In Proceedings of the IEEE international conference on computer vision

    Dense-captioning events in videos. In Proceedings of the IEEE international conference on computer vision . 706–715

  88. [95]

    Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu

  89. [96]

    Yanwei Li, Chengyao Wang, and Jiaya Jia. 2025. Llama-vid: An image is worth 2 tokens in large language models. In European Conference on Computer Vision . Springer, 323–340

  90. [97]

    Ruotong Liao, Max Erler, Huiyu Wang, Guangyao Zhai, Gengyuan Zhang, Yunpu Ma, and Volker Tresp. 2024. VideoINSTA: Zero-shot Long Video Understand- ing via Informative Spatial-Temporal Reasoning with LLMs. arXiv preprint arXiv:2409.20365 (2024)

  91. [98]

    Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. 2023. Video-llava: Learning united visual representation by alignment before projec- tion. arXiv preprint arXiv:2311.10122 (2023)

  92. [99]

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdel rahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Genera- tion, Translation, and Comprehension. In Annual Meeting of t...

  93. [100]

    Xiangru Lin, Yuyang Chen, Guanbin Li, and Yizhou Yu. 2022. A causal inference look at unsupervised video anomaly detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 1620–1629

  94. [101]

    KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355 (2023)

  95. [102]

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. 2024. Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22195–22206

  96. [103]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruction Tuning. arXiv:2304.08485 [cs.CV] https://arxiv.org/abs/2304.08485

  97. [104]

    Jiajun Liu, Yibing Wang, Hanghang Ma, Xiaoping Wu, Xiaoqi Ma, Xiaoming Wei, Jianbin Jiao, Enhua Wu, and Jie Hu. 2024. Kangaroo: A powerful video-language model supporting long-context video input. arXiv preprint arXiv:2408.15542 (2024)

  98. [105]

    Ruyang Liu, Chen Li, Haoran Tang, Yixiao Ge, Ying Shan, and Ge Li. 2025. St-llm: Large language models are effective temporal learners. In European Conference on Computer Vision. Springer, 1–18

  99. [107]

    Ye Liu, Siyuan Li, Yang Wu, Chang-Wen Chen, Ying Shan, and Xiaohu Qie

  100. [108]

    Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. 2024. Vila: On pre-training for visual language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 26689–26699

  101. [109]

    Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. 2022. Video swin transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 3202–3211

  102. [110]

    Ding Liu, Zhaowen Wang, Yuchen Fan, Xianming Liu, Zhangyang Wang, Shiyu Chang, Xinchao Wang, and Thomas S. Huang. 2018. Learning Temporal Dynam- ics for Video Super-Resolution: A Deep Learning Approach. IEEE Transactions on Image Processing 27, 7 (2018), 3432–3445. https://doi...

  103. [111]

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved Baselines with Visual Instruction Tuning. arXiv:2310.03744 [cs.CV] https: //arxiv.org/abs/2310.03744

  104. [112]

    Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Da Li, Pengcheng Lu, Tao Wang, Linmei Hu, Minghui Qiu, and Zhongyu Wei. 2023. Valley: Video assistant with large language model enhanced ability. arXiv preprint arXiv:2306.07207 (2023)

  105. [113]

    Chenyang Lyu, Minghao Wu, Longyue Wang, Xinting Huang, Bingshuai Liu, Zefeng Du, Shuming Shi, and Zhaopeng Tu. 2023. Macaw-llm: Multi-modal language modeling with image, audio, video, and text integration.arXiv preprint arXiv:2306.09093 (2023)

  106. [114]

    Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Khan. 2024. VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Under- standing. arXiv preprint arXiv:2406.09418 (2024)

  107. [115]

    Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan

  108. [116]

    Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. 2023. Egoschema: A diagnostic benchmark for very long-form video language un- derstanding. Advances in Neural Information Processing Systems 36 (2023), 46212–46244

  109. [117]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3042–3051

  110. [118]

    Zuyan Liu, Yuhao Dong, Ziwei Liu, Winston Hu, Jiwen Lu, and Yongming Rao. 2024. Oryx mllm: On-demand spatial-temporal understanding at arbitrary resolution. arXiv preprint arXiv:2409.12961 (2024)

  111. [119]

    Ming Nie, Dan Ding, Chunwei Wang, Yuanfan Guo, Jianhua Han, Hang Xu, and Li Zhang. 2024. SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=FOkKndty5B

  112. [120]

    Fuchen Long, Zhaofan Qiu, Ting Yao, and Tao Mei. 2024. Videodrafter: Content- consistent multi-scene video generation with llm.arXiv preprint arXiv:2401.01256 (2024)

  113. [121]

    Cewu Lu, Jianping Shi, and Jiaya Jia. 2013. Abnormal Event Detection at 150 FPS in MATLAB. In 2013 IEEE International Conference on Computer Vision . 2720–2727. https://doi.org/10.1109/ICCV.2013.338

  114. [122]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al

  115. [123]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9

  116. [124]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140 (2020), 1–67

  117. [125]

    Arjun Raj, Lei Wang, and Tom Gedeon. 2024. TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention Maps. arXiv preprint arXiv:2409.14543 (2024)

  118. [126]

    arXiv preprint arXiv:2306.05424 (2023)

    Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424 (2023)

  119. [127]

    Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor. 2021. ImageNet-21K Pretraining for the Masses. arXiv:2104.10972 [cs.CV]

  120. [128]

    Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF international conference on computer vision ....

  121. [129]

    Thong Nguyen, Yi Bin, Junbin Xiao, Leigang Qu, Yicong Li, Jay Zhangjie Wu, Cong-Duy Nguyen, See-Kiong Ng, and Luu Anh Tuan. 2024. Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives. arXiv preprint arXiv:2406.05615 (2024)

  122. [130]

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural I...

  123. [131]

    Mandela Patrick, Dylan Campbell, Yuki Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Joao F Henriques. 2021. Keeping your eye on the ball: Trajectory attention in video transformers. Advances in neural information processing systems 34 (2021), ...

  124. [132]

    Zhenyue Qin, Yang Liu, Pan Ji, Dongwoo Kim, Lei Wang, Saeed Anwar, and Tom Gedeon. 2022. Fusing higher-order features in graph neural networks for skeleton-based action recognition. IEEE Transactions on Neural Networks and Learning Systems 35, 4 (2022), 4783–4797

  125. [133]

    Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, De- An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, et al. 2024. Eagle: Exploring the design space for multimodal llms with mixture of encoders. arXiv preprint arXiv:2408.15998 (2024)

  126. [134]

    In International conference on machine learning

    Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763

  127. [135]

    Yan Shu, Peitian Zhang, Zheng Liu, Minghao Qin, Junjie Zhou, Tiejun Huang, and Bo Zhao. 2024. Video-xl: Extra-long vision language model for hour-scale video understanding. arXiv preprint arXiv:2409.14485 (2024)

  128. [136]

    Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In Computer Vision–ECCV 2016: 14th European Confer- ence, Amsterdam, The Netherlands, October 11–14, ...

  129. [137]

    K Soomro. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)

  130. [138]

    Bharathkumar Ramachandra and Michael Jones. 2020. Street Scene: A new dataset and evaluation protocol for video anomaly detection. arXiv:1902.05872 [cs.CV] https://arxiv.org/abs/1902.05872

  131. [139]

    Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid

  132. [140]

    Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele. 2017. Movie descrip- tion. International Journal of Computer Vision 123 (2017), 94–120

  133. [141]

    Arka Sadhu, Tanmay Gupta, Mark Yatskar, Ram Nevatia, and Aniruddha Kemb- havi. 2021. Visual semantic role labeling for video understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5589– 5600

  134. [142]

    Yunlong Tang, Jing Bi, Siting Xu, Luchuan Song, Susan Liang, Teng Wang, Daoan Zhang, Jie An, Jingyang Lin, Rongyi Zhu, et al. 2023. Video understanding with large language models: A survey. arXiv preprint arXiv:2312.17432 (2023)

  135. [143]

    Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Con- ceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...

  136. [144]

    Alex Sherstinsky. 2020. Fundamentals of recurrent neural network (RNN) and long short-term memory (LSTM) network. Physica D: Nonlinear Phenomena 404 (2020), 132306

  137. [145]

    Together.xyz. 2023. Releasing 3b and 7b redpajama incite family of models including base, instruction-tuned and chat models. https://www.together.xyz/ blog/redpajama-models-v1

  138. [146]

    Fangxun Shu, Lei Zhang, Hao Jiang, and Cihang Xie. 2023. Audio-visual llm for video understanding. arXiv preprint arXiv:2312.06720 (2023)

  139. [147]

    Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri

  140. [148]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. ...

  141. [149]

    Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. 2016. Generating videos with scene dynamics. Advances in neural information processing systems 29 (2016)

  142. [150]

    Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork. 2021. WIT: Wikipedia-Based Image Text Dataset for Multimodal Multi- lingual Machine Learning. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Inform...

  143. [151]

    Junke Wang, Dongdong Chen, Chong Luo, Xiyang Dai, Lu Yuan, Zuxuan Wu, and Yu-Gang Jiang. 2023. Chatvideo: A tracklet-centric multimodal and versatile video understanding system. arXiv preprint arXiv:2304.14407 (2023)

  144. [152]

    Junke Wang, Dongdong Chen, Chong Luo, Bo He, Lu Yuan, Zuxuan Wu, and Yu-Gang Jiang. 2024. Omnivid: A generative framework for universal video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18209–18220

  145. [153]

    Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. 2023. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389 (2023)

  146. [154]

    Mingtian Tan, Mike A Merrill, Vinayak Gupta, Tim Althoff, and Thomas Hartvigsen. 2024. Are Language Models Actually Useful for Time Series Forecast- ing?. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=DV15UbHCY1

  147. [155]

    Jiaqi Wang, Hanqi Jiang, Yiheng Liu, Chong Ma, Xu Zhang, Yi Pan, Mengyuan Liu, Peiran Gu, Sichen Xia, Wenjun Li, et al. 2024. A comprehensive review of multimodal large language models: Performance and challenges across different tasks. arXiv preprint arXiv:2408.01319 (2024)

  148. [156]

    Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. 2019. Coin: A large-scale dataset for comprehen- sive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1207–1216

  149. [157]

    Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016. Movieqa: Understanding stories in movies through question-answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4631–4640

  150. [158]

    Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao. 2023. Videomae v2: Scaling video masked autoencoders with dual masking. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14549–14560

  151. [159]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)

  152. [160]

    Lei Wang, Du Q Huynh, and Moussa Reda Mansour. 2019. Loss switching fusion with similarity search for video classification. In 2019 IEEE international conference on image processing (ICIP) . IEEE, 974–978

  153. [161]

    Lei Wang and Piotr Koniusz. 2021. Self-supervising action recognition by statistical moment and subspace descriptors. In Proceedings of the 29th ACM international conference on multimedia . 4324–4333. Do Language Models Understand Time? WWW Companion ’25, April 28-May 2, 2025,...

  154. [162]

    Lei Wang and Piotr Koniusz. 2022. Temporal-viewpoint transportation plan for skeletal few-shot action recognition. In Proceedings of the Asian Conference on Computer Vision. 4176–4193

  155. [163]

    Lei Wang and Piotr Koniusz. 2022. Uncertainty-dtw for time series and se- quences. In European Conference on Computer Vision . Springer, 176–195

  156. [164]

    Alex Jinpeng Wang, Linjie Li, Kevin Qinghong Lin, Jianfeng Wang, Kevin Lin, Zhengyuan Yang, Lijuan Wang, and Mike Zheng Shou. 2024. COSMO: COn- trastive Streamlined MultimOdal Model with Interleaved Pre-Training. arXiv preprint arXiv:2401.00849 (2024)

  157. [165]

    Lei Wang and Piotr Koniusz. 2024. Flow dynamics correction for action recog- nition. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3795–3799

  158. [166]

    Lei Wang, Piotr Koniusz, and Du Q Huynh. 2019. Hallucinating idt descriptors and i3d optical flow features for action recognition with cnns. In Proceedings of the IEEE/CVF international conference on computer vision . 8698–8708

  159. [167]

    Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, and Lu Yuan. 2022. Omnivl: One foundation model for image-language and video-language tasks. Advances in neural information processing systems 35 (2022), 5696–5710

  160. [168]

    Jiexin Wang, Adam Jatowt, and Yi Cai. 2024. Towards Effective Time-Aware Language Representation: Exploring Enhanced Temporal Understanding in Language Models. arXiv preprint arXiv:2406.01863 (2024)

  161. [169]

    Lei Wang, Ke Sun, and Piotr Koniusz. 2024. High-order tensor pooling with at- tention for action recognition. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3885–3889

  162. [170]

    Lei Wang. 2021. Analysis and evaluation of Kinect-based action recognition algorithms. arXiv preprint arXiv:2112.08626 (2021)

  163. [171]

    Lei Wang. 2023. Robust human action modelling . Ph. D. Dissertation. The Australian National University (Australia)

  164. [172]

    Shaojie Wang, Wentian Zhao, Ziyi Kou, Jing Shi, and Chenliang Xu. 2021. How to make a blt sandwich? learning vqa towards understanding web instructional videos. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 1130–1139

  165. [173]

    Lei Wang, Du Q Huynh, and Piotr Koniusz. 2019. A comparative review of recent kinect-based action recognition algorithms. IEEE Transactions on Image Processing 29 (2019), 15–28

  166. [174]

    Weiyun Wang, Zhe Chen, Wenhai Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Jinguo Zhu, Xizhou Zhu, Lewei Lu, Yu Qiao, and Jifeng Dai. 2024. Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization. arXiv preprint arXiv:2411.10442 (2024)

  167. [175]

    Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. 2019. Vatex: A large-scale, high-quality multilingual dataset for video- and-language research. In Proceedings of the IEEE/CVF international conference on computer vision. 4581–4591

  168. [176]

    Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Conghui He, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Yu Qiao. 2024. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation. ar...

  169. [177]

    Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Jilan Xu, Zun Wang, et al. 2024. Internvideo2: Scaling video foundation models for multimodal video understanding. arXiv e-prints (2024), arXiv–2403

  170. [178]

    Lei Wang and Piotr Koniusz. 2023. 3mformer: Multi-order multi-mode trans- former for skeletal action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5620–5631

  171. [179]

    Yuqing Wang, Tianwei Xiong, Daquan Zhou, Zhijie Lin, Yang Zhao, Bingyi Kang, Jiashi Feng, and Xihui Liu. 2024. Loong: Generating Minute-level Long Videos with Autoregressive Language Models. arXiv preprint arXiv:2410.02757 (2024)

  172. [180]

    Zhanyu Wang, Longyue Wang, Zhen Zhao, Minghao Wu, Chenyang Lyu, Huayang Li, Deng Cai, Luping Zhou, Shuming Shi, and Zhaopeng Tu. 2024. Gpt4video: A unified multimodal large language model for lnstruction-followed understanding and safety-aware generation. In Proceedings of the...

  173. [181]

    Lei Wang, Jun Liu, and Piotr Koniusz. 2021. 3D Skeleton-based Few-shot Action Recognition with JEANIE is not so Naïve.arXiv preprint arXiv:2112.12668 (2021)

  174. [182]

    Lei Wang, Jun Liu, Liang Zheng, Tom Gedeon, and Piotr Koniusz. 2024. Meet JEANIE: a Similarity Measure for 3D Skeleton Sequences via Temporal- Viewpoint Alignment. International Journal of Computer Vision (2024), 1–32

  175. [183]

    Peng Wu, Jing Liu, Yujia Shi, Yujia Sun, Fangtao Shao, Zhaoyang Wu, and Zhiwei Yang. 2020. Not only Look, but also Listen: Learning Multimodal Violence Detection under Weak Supervision. arXiv:2007.04687 [cs.CV] https://arxiv.org/ abs/2007.04687

  176. [184]

    Lei Wang, Xiuyuan Yuan, Tom Gedeon, and Liang Zheng. [n. d.]. Taylor Videos for Action Recognition. In Forty-first International Conference on Machine Learn- ing

  177. [185]

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. 2024. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv preprint arXiv:2409.12191 (2024)

  178. [186]

    Bo Xu and Mu-ming Poo. 2023. Large language models and brain-inspired general intelligence. National Science Review 10, 10 (2023), nwad267

  179. [187]

    Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al. 2024. Visionllm: Large language model is also an open-ended decoder for vision-centric tasks.Advances in Neural Information Processing Systems 36 (2024)

  180. [188]

    Haiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li, Bin Bi, Qi Qian, Wei Wang, et al. 2023. mplug-2: A modularized multi-modal foundation model across text, image and video. In International Conference on Machine Learning. PMLR, 38728–38748

  181. [189]

    Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5288–5296

  182. [190]

    Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng

  183. [191]

    Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont- Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. 2023. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. InProceedings of the IEEE/CVF Conference on Computer V...

  184. [192]

    Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al. 2022. Internvideo: General video foundation models via generative and discriminative learning. arXiv preprint arXiv:2212.03191 (2022)

  185. [193]

    Dongjie Yang, Suyuan Huang, Chengqiang Lu, Xiaodong Han, Haoxin Zhang, Yan Gao, Yao Hu, and Hai Zhao. 2024. Vript: A Video Is Worth Thousands of Words. arXiv preprint arXiv:2406.06040 (2024)

  186. [194]

    Qu Yang, Mang Ye, and Bo Du. 2024. Emollm: Multimodal emotional under- standing meets large language models. arXiv preprint arXiv:2406.16442 (2024)

  187. [195]

    Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan

  188. [196]

    arXiv preprint arXiv:2405.09711 (2024)

    Star: A benchmark for situated reasoning in real-world videos. arXiv preprint arXiv:2405.09711 (2024)

  189. [197]

    Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. 2023. Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models. arXiv:2303.04671 [cs.CV] https://arxiv.org/abs/2303.04671

  190. [198]

    Luca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang, and Elisa Ricci. 2024. Harnessing Large Language Models for Training-free Video Anom- aly Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18527–18536

  191. [199]

    Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2023. Next-gpt: Any-to-any multimodal llm. arXiv preprint arXiv:2309.05519 (2023)

  192. [200]

    Weijia Wu, Yuzhong Zhao, Zhuang Li, Jiahong Li, Hong Zhou, Mike Zheng Shou, and Xiang Bai. 2023. A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension. arXiv:2305.03347 [cs.CV] https://arxiv.org/abs/2305.03347

  193. [201]

    Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction- tuned Audio-Visual Language Model for Video Understanding. In Conference on Empirical Methods in Natural Language Processing . https://api.semanticscholar. org/CorpusID:259075356

  194. [202]

    Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. [n. d.]. Video Question Answering via Gradually Refined Attention over Appearance and Motion. In ACM Multimedia

  195. [203]

    Jieyu Zhang, Weikai Huang, Zixian Ma, Oscar Michel, Dong He, Tanmay Gupta, Wei-Chiu Ma, Ali Farhadi, Aniruddha Kembhavi, and Ranjay Krishna. 2024. Task Me Anything. arXiv preprint arXiv:2406.11775 (2024)

  196. [204]

    Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen. 2023. T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations. arXiv:2301.06052 [cs.CV] https://arxiv.org/abs/2301.06052

  197. [205]

    Zhihan Zhang, Yixin Cao, Chenchen Ye, Yunshan Ma, Lizi Liao, and Tat-Seng Chua. 2024. Analyzing Temporal Complex Events with Large Language Models? A Benchmark towards Temporal, Long Context Understanding. arXiv preprint arXiv:2406.02472 (2024)

  198. [206]

    arXiv preprint arXiv:2404.16994 (2024)

    Pllava: Parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994 (2024)

  199. [207]

    Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar. 2023. Learning video representations from large language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6586–6597

  200. [208]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al . 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671 (2024)

  201. [209]

    Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. 2018. Temporal relational reasoning in videos. In Proceedings of the European conference on computer vision (ECCV). 803–818

  202. [210]

    Pengyuan Zhou, Lin Wang, Zhi Liu, Yanbin Hao, Pan Hui, Sasu Tarkoma, and Jussi Kangasharju. 2024. A survey on generative ai and llm for video generation, understanding, and streaming. arXiv preprint arXiv:2404.16038 (2024)

  203. [211]

    Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G

    Lijun Yu, José Lezama, Nitesh B. Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G. Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A. Ross, and Lu Jiang. 2024. Language Model Beats Diffusion – To...

  204. [212]

    Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. 2019. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 9127–9134

  205. [213]

    Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. 2021. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432 (2021)

  206. [215]

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 11975–11986

  207. [216]

    Chen-Lin Zhang, Jianxin Wu, and Yin Li. 2022. Actionformer: Localizing mo- ments of actions with transformers. In European Conference on Computer Vision . Springer, 492–510

  208. [218]

    Huaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo, Chuchu Han, Xiaonan Huang, Changxin Gao, Yuehuan Wang, and Nong Sang. 2024. Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM. arXiv preprint arXiv:2406.12235 (2024)

  209. [222]

    Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. 2024. Expel: Llm agents are experiential learners. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 19632–19642. WWW Companion ’25, April 28-May 2, 2025, Sydney, NSW, A...

  210. [224]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT- Bench and Chatbot Arena. arXiv:2306.05685 [cs.CL] https://a...

  211. [227]

    Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al. 2023. Languagebind: Ex- tending video-language pretraining to n-modality by language-based semantic alignment. arXiv preprint arXiv:2310.01852 (2023)

  212. [228]

    Liyun Zhu, Lei Wang, Arjun Raj, Tom Gedeon, and Chen Chen. 2024. Advanc- ing Video Anomaly Detection: A Concise Review and a New Dataset. In The Thirty-eighth Conference on Neural Information Processing Systems Datasets and Benchmarks Track

  213. [229]

    Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Serena Yeung-Levy, and Xide Xia. 2024. Apollo: An Exploration of Video Understanding in Large Multimodal Models. arXiv:2412.10360 [cs...

  214. [2015]

    In Proceedings of the IEEE international conference on computer vision

    Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision . 4489–4497

  215. [2017]

    arXiv preprint arXiv:1705.06950 (2017)

    The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017)

  216. [2019]

    In Proceedings of the IEEE/CVF international conference on computer vision

    Videobert: A joint model for video and language representation learning. In Proceedings of the IEEE/CVF international conference on computer vision . 7464– 7473

  217. [2020]

    arXiv preprint arXiv:2005.00200 (2020)

    Hero: Hierarchical encoder for video+ language omni-representation pre-training. arXiv preprint arXiv:2005.00200 (2020)

  218. [2021]

    Transactions of the Association for Compu- tational Linguistics 9 (2021), 978–994

    Multimodal pretraining unmasked: A meta-analysis and a unified frame- work of vision-and-language BERTs. Transactions of the Association for Compu- tational Linguistics 9 (2021), 978–994

  219. [2022]

    arXiv preprint arXiv:2212.12017 (2022)

    Opt-iml: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint arXiv:2212.12017 (2022)

  220. [2023]

    See https://vicuna

    Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023) 2, 3 (2023), 6

  221. [2024]

    arXiv preprint arXiv:2408.14023 (2024)

    Video-ccam: Enhancing video-language understanding with causal cross- attention masks for short and long videos. arXiv preprint arXiv:2408.14023 (2024)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.