Pith. sign in

REVIEW 3 minor 2 cited by

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

T0 review · 0 major / 3 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read Video MLLMs acquire evidence, preserve context, and produce outputs through watching, remembering, and reasoning.

desk verdict This is a survey that organizes video MLLM work around watching, remembering, and reasoning but introduces no new results or derivations. read the letter →

arxiv 2606.07433 v1 pith:W7NYHLPE submitted 2026-06-05 cs.CV cs.AIcs.MM

classification cs.CVcs.AIcs.MM
keywords videounderstandingmultimodallargelanguagemodelswatchingrememberingreasoninglongprocessingmemorymodelingegocentricstreamingfaithful
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces a unified human-view framework for LLM-based video understanding built around three functional abilities: watching for evidence acquisition, remembering for context preservation, and reasoning for grounded outputs. Instead of isolated task benchmarks, this structure analyzes how models handle sparse evidence, long-range dependencies, and multimodal alignment under computational limits. It formulates video systems through perceptual representations, memory states, reasoning traces, and final predictions, then maps representative methods, challenges, application domains, datasets, and benchmarks onto these elements. The work reviews current approaches across perception, memory, and reasoning while highlighting open problems for scalable video intelligence.

What carries the argument

The three functional abilities—watching, remembering, and reasoning—that partition video MLLM behavior and link the four system components (perceptual representations, memory states, reasoning traces, final predictions) into a unified analysis structure.

What would settle it

A deployed video MLLM whose accuracy, efficiency, or failure modes on long videos cannot be improved or explained by separately measuring or modifying its watching, remembering, or reasoning components.

Watch

Extended reading notes

Core claim

Video understanding with MLLMs is best characterized by a formulation that decomposes systems into perceptual representations, memory states, reasoning traces, and final predictions, which in turn map onto the three abilities of watching, remembering, and reasoning; this decomposition supplies a single lens for organizing methods, identifying challenges in spatio-temporal perception and memory modeling, and covering domains from egocentric to narrative videos.

Load-bearing premise

Every video understanding system can be usefully described by four fixed components and partitioned without overlap or remainder into the three abilities of watching, remembering, and reasoning.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 3 minor

Summary. This survey paper proposes a human-view organizational framework for video understanding with multimodal large language models (MLLMs), structured around three functional abilities (watching, remembering, reasoning) and four system components (perceptual representations, memory states, reasoning traces, final predictions). It uses this framing to categorize methods for fine-grained perception, memory modeling (offline/streaming), and reasoning (text-only or video-grounded), while surveying challenges in spatio-temporal perception and long-video processing, application domains (egocentric, sports, medical, etc.), datasets, benchmarks, and open problems.

Significance. If the proposed partition proves useful for synthesis, the work could help researchers map existing methods onto a common structure for identifying gaps in memory-aware and evidence-grounded video intelligence, particularly as the field shifts toward long, multimodal scenarios. Its value is primarily in organization and coverage rather than new derivations or measurements.

minor comments (3)
  1. The formulation characterizing systems by perceptual representations, memory states, reasoning traces, and final predictions is introduced in the abstract and presumably detailed early in the manuscript; if this is presented only descriptively without a diagram or explicit mapping table to the three abilities, it risks remaining informal for readers attempting to apply the framework to new papers.
  2. The abstract states that representative methods are 'organized by their roles in video MLLM systems' under the three abilities, but without an explicit cross-reference table or section that lists which cited works map to which component, the organizational claim is harder to verify.
  3. The GitHub link for continuously traced related works is mentioned but not cited as a reference; adding a formal citation or footnote would improve traceability.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their positive summary of our survey and the recommendation of minor revision. No specific major comments were provided in the report, so we have no individual points requiring point-by-point rebuttal. We will incorporate minor improvements for clarity and completeness in the revised version.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; organizational survey framing is self-contained

full rationale

The paper is a literature survey that proposes an organizational perspective on video MLLMs structured around three functional abilities (watching, remembering, reasoning) and four system components (perceptual representations, memory states, reasoning traces, final predictions). The central claim is that this supplies a unified structure for analyzing existing methods, challenges, datasets, and benchmarks rather than introducing new empirical results or a falsifiable model. The partition is presented as a useful framing for the survey; no stronger claim of exhaustiveness, disjointness, or predictive power is required for the work to fulfill its stated purpose. No equations, fitted parameters, predictions, or self-citation chains appear in the derivation chain. All cited works are external and the framework does not reduce to its inputs by construction.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The paper is a survey and introduces no new free parameters, mathematical axioms, or invented physical entities; it synthesizes prior literature under an organizational framing.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Watch, Remember, Reason: Human-View Video Understanding with MLLMs." pith.science (2026). https://pith.science/paper/W7NYHLPE

@misc{pith2026260607433,
  author       = {Pith},
  title        = {Pith review of: Watch, Remember, Reason: Human-View Video Understanding with MLLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/W7NYHLPE}},
  note         = {Machine review of arXiv:2606.07433}
}
read the original abstract

Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios. These scenarios require models to handle sparse evidence, long-range dependencies, multimodal alignment, and reliable inference under limited computational budgets. This work presents a human-view perspective on LLM-based video understanding, organized around three functional abilities: watching, remembering, and reasoning. Rather than treating video tasks as isolated benchmarks, this view provides a unified structure for analyzing how video MLLMs acquire evidence, preserve context, and produce grounded outputs. We introduce a formulation that characterizes video understanding systems by their perceptual representations, memory states, reasoning traces, and final predictions. Based on this formulation, we identify challenges in spatio-temporal perception, efficient long-video processing, memory modeling, streaming understanding, and faithful reasoning. Representative methods are organized by their roles in video MLLM systems. Watching covers fine-grained, comprehensive, audio-visual, and efficient perception. Remembering includes offline and streaming memory, while reasoning covers text-only reasoning and thinking with videos. We further examine application domains such as egocentric, sports, instructional, medical, and narrative videos, and cover training datasets and evaluation benchmarks across task types, supervision formats, modalities, and capability dimensions. Finally, we outline open problems and future directions for scalable, memory-aware, and evidence-grounded video intelligence. Related works will be continuously traced at https://github.com/marinero4972/Awesome-HumanView-VideoUnderstanding.

Figures

Figures reproduced from arXiv: 2606.07433 by the authors.

Figure 1
Figure 1. Overview of our survey. Left: the survey pipeline. Right: our Watch–Remember–Reason taxonomy for MLLM-based video understanding. Watch (Sec. 3.1) covers fine-grained grounding, captioning, audio-visual perception, and efficient processing. Remember (Sec. 3.2) includes offline and streaming memory. Reason (Sec. 3.3) covers text-only reasoning and thinking with videos, with both agentic and non-agent approaches. Repre… view at source ↗
Figure 2
Figure 2. Overview of methods related to ”How to Watch?”. Fine-grained watching localizes task-relevant evidence in time and space. Comprehensive watching abstracts videos into summaries, and segment-level or region-level descriptions. Audio-visual watching aligns visual and acoustic streams for omni-modal perception. Efficient watching reduces redundancy through frame selection, token compression, and efficient model process… view at source ↗
Figure 3
Figure 3. Overview of methods related to ”How to Remember?”. Agentic offline memory constructs and updates external memory through LLM/VLM agents. Non-agentic offline memory builds structured short-term and long-term memory via event extraction, frame selection, token compression, and event clustering. Streaming memory maintains and retrieves memory online through sliding windows, recent memory, and long-term memory banks. me… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Overview of methods related to ”How to Reason?”. Agentic text-only reasoning methods decompose reasoning into modular steps such as clip summarization, adaptive search, memory retrieval, reflection, and answer verification. Non￾agent text-only reasoning methods perform…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A streaming-video memory system that filters redundant frames before encoding, stores older video as latent tokens, and re-injects them when visual attention drops reaches 67.73% on StreamingBench.

  2. LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents

    cs.CV 2026-08 conditional novelty 5.0 of 10

    LAVE lets video agents reuse pre-verbal visual hidden states from past tool calls via timestamp-aligned residual injection, improving Video-MME by 3.76 points over the strongest baseline without extra training or frames.

Reference graph

Works this paper leans on

298 extracted references · 161 canonical work pages · cited by 2 Pith papers

  1. [1]

    Qwen3.5: Towards native multimodal agents,

    Qwen Team, “Qwen3.5: Towards native multimodal agents,” February 2026. [Online]. Available: https://qwen.ai/blog?id= qwen3.5

  2. [2]

    Qwen3-Omni Technical Report

    J. Xu, Z. Guo, H. Hu, Y. Chu, X. Wang, J. He, Y. Wang, X. Shi, T. He, X. Zhu, Y. Lv, Y. Wang, D. Guo, H. Wang, L. Ma, P . Zhang, X. Zhang, H. Hao, Z. Guo, B. Yang, B. Zhang, Z. Ma, X. Wei, S. Bai, K. Chen, X. Liu, P . Wang, M. Yang, D. Liu, X. Ren, B. Zheng, R. Men, F. Zhou, B. Yu, J. Yang, L. Yu, J. Zhou, and J. Lin, “Qwen3- omni technical report,” arXiv...

  3. [3]

    Qwen3-vl technical report,

    S. Bai, Y. Cai, R. Chen, K. Chen, X. Chenet al., “Qwen3-vl technical report,” Nov. 2025. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 24

  4. [4]

    Qwen2.5-Omni Technical Report

    J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Danget al., “Qwen2. 5-omni technical report,” arXiv preprint arXiv:2503.20215, 2025

  5. [5]

    Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.arXiv preprint arXiv:2506.19225, 2025

    M. Qin, X. Liu, Z. Liang, Y. Shu, H. Yuan, J. Zhou, S. Xiao, B. Zhao, and Z. Liu, “Video-xl-2: Towards very long-video under- standing through task-aware kv sparsification,” arXiv preprint arXiv:2506.19225, 2025

  6. [6]

    Msr-vtt: A large video descrip- tion dataset for bridging video and language,

    J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video descrip- tion dataset for bridging video and language,” inProceedings of the IEEE conference on computer vision and pattern recognition. Los Alamitos, CA, USA: IEEE Computer Society, 2016, pp. 5288–5296

  7. [7]

    Video question answering via gradually refined attention over appearance and motion,

    D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, “Video question answering via gradually refined attention over appearance and motion,” inProceedings of the 25th ACM interna- tional conference on Multimedia. New York, NY, USA: Association for Computing Machinery, 2017, pp. 1645–1653

  8. [8]

    Tgif-qa: Toward spatio-temporal reasoning in visual question answering,

    Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “Tgif-qa: Toward spatio-temporal reasoning in visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition. Los Alamitos, CA, USA: IEEE Computer Society, 2017, pp. 2758–2766

Show all 298 references
  1. [9]

    LongVU: Spatiotemporal adaptive compression for long video- language understanding,

    X. Shen, Y. Xiong, C. Zhao, L. Wu, J. Chen, C. Zhu, Z. Liu, F. Xiao, B. Varadarajan, F. Bordes, Z. Liu, H. Xu, H. J. Kim, B. Soran, R. Krishnamoorthi, M. Elhoseiny, and V . Chandra, “LongVU: Spatiotemporal adaptive compression for long video- language understanding,” inForty-s...

  2. [10]

    Longvideo-r1: Smart navigation for low-cost long video understanding,

    J. Qiu, L. Xie, X. Huo, Q. Tian, and Q. Ye, “Longvideo-r1: Smart navigation for low-cost long video understanding,” arXiv preprint arXiv:2602.20913, 2026

  3. [11]

    Video-o3: Native inter- leaved clue seeking for long video multi-hop reasoning,

    X. Zeng, Z. Zhang, Y. Zhu, X. Li, Z. Wang, C. Ma, Q. Zhang, Z. Huang, K. Ouyang, T. Jianget al., “Video-o3: Native inter- leaved clue seeking for long video multi-hop reasoning,” arXiv preprint arXiv:2601.23224, 2026

  4. [12]

    Ego4d: Around the world in 3,000 hours of egocentric video,

    K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liuet al., “Ego4d: Around the world in 3,000 hours of egocentric video,” inCVPR. Los Alamitos, CA, USA: IEEE Computer Society, 2022

  5. [14]

    A visually grounded language model for fetal ultrasound understanding,

    X. Guo, M. Alsharid, H. Zhao, Y. Wang, J. Lander, A. T. Pa- pageorghiou, and J. A. Noble, “A visually grounded language model for fetal ultrasound understanding,” Nature Biomedical Engineering, advance online publication, 2026

  6. [15]

    Streamingvlm: Real-time understanding for infinite video streams,

    R. Xu, G. Xiao, Y. Chen, L. He, K. Peng, Y. Lu, and S. Han, “Streamingvlm: Real-time understanding for infinite video streams,” 2025. [Online]. Available: https://arxiv.org/abs/2510. 09608

  7. [16]

    Timechat: A time- sensitive multimodal large language model for long video un- derstanding,

    S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, “Timechat: A time- sensitive multimodal large language model for long video un- derstanding,” arXiv preprint arXiv:2312.02051, 2023

  8. [17]

    Adaptive keyframe sampling for long video understanding,

    X. Tang, J. Qiu, L. Xie, Y. Tian, J. Jiao, and Q. Ye, “Adaptive keyframe sampling for long video understanding,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition. Los Alamitos, CA, USA: IEEE Computer Society, 2025, pp. 29 118–29 128

  9. [18]

    FrameFusion: Combining similarity and importance for video token reduction on large vision language models,

    T. Fu, T. Liu, Q. Han, G. Dai, S. Yan, H. Yang, X. Ning, and Y. Wang, “FrameFusion: Combining similarity and importance for video token reduction on large vision language models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision. Los Alamitos, CA, USA...

  10. [19]

    Moviechat: From dense token to sparse memory for long video understanding,

    E. Song, W. Chai, G. Wang, Y. Zhang, H. Zhou, F. Wu, H. Chi, X. Guo, T. Ye, Y. Zhanget al., “Moviechat: From dense token to sparse memory for long video understanding,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2024, pp. 18 221–18 232

  11. [20]

    Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,

    B. He, H. Li, Y. K. Jang, M. Jia, X. Cao, A. Shah, A. Shrivastava, and S.-N. Lim, “Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13 504–13 514

  12. [21]

    Memory-enhanced retrieval augmentation for long video understanding,

    H. Yuan, Z. Liu, M. Qin, H. Qian, Y. Shu, Z. Dou, J.-R. Wen, and N. Sebe, “Memory-enhanced retrieval augmentation for long video understanding,” 2025. [Online]. Available: https://arxiv.org/abs/2503.09149

  13. [22]

    Flash-vstream: Memory-based real-time understanding for long video streams,

    H. Zhang, Y. Wang, Y. Tang, Y. Liu, J. Feng, J. Dai, and X. Jin, “Flash-vstream: Memory-based real-time understanding for long video streams,” 2024. [Online]. Available: https: //arxiv.org/abs/2406.08085

  14. [23]

    Streammem: Query-agnostic kv cache memory for streaming video understanding,

    Y. Yang, Z. Zhao, S. N. Shukla, A. Singh, S. K. Mishra, L. Zhang, and M. Ren, “Streammem: Query-agnostic kv cache memory for streaming video understanding,” 2025. [Online]. Available: https://arxiv.org/abs/2508.15717

  15. [24]

    Video-r1: Reinforcing video reasoning in mllms,

    K. Feng, K. Gong, B. Li, Z. Guo, Y. Wang, T. Peng, J. Wu, X. Zhang, B. Wang, and X. Yue, “Video-r1: Reinforcing video reasoning in mllms,” arXiv preprint arXiv:2503.21776, 2025

  16. [25]

    Videorft: In- centivizing video reasoning capability in mllms via reinforced fine-tuning,

    Q. Wang, Y. Yu, Y. Yuan, R. Mao, and T. Zhou, “Videorft: In- centivizing video reasoning capability in mllms via reinforced fine-tuning,” arXiv preprint arXiv:2505.12434, 2025

  17. [26]

    Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence,

    J. Meng, X. Li, H. Wang, Y. Tan, T. Zhang, L. Kong, Y. Tong, A. Wang, Z. Teng, Y. Wanget al., “Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence,” arXiv preprint arXiv:2510.20579, 2025

  18. [27]

    Thinking with videos: Multimodal tool- augmented reinforcement learning for long video reasoning,

    H. Zhang, X. Gu, J. Li, C. Ma, S. Bai, C. Zhang, B. Zhang, Z. Zhou, D. He, and Y. Tang, “Thinking with videos: Multimodal tool- augmented reinforcement learning for long video reasoning,” arXiv preprint arXiv:2508.04416, 2025

  19. [28]

    Video-language understanding: A survey from model architecture, model training, and data per- spectives,

    T. Nguyen, Y. Bin, J. Xiao, L. Qu, Y. Li, J. Z. Wu, C.-D. Nguyen, S. K. Ng, and L. A. Tuan, “Video-language understanding: A survey from model architecture, model training, and data per- spectives,” inFindings of the Association for Computational Lin- guistics: ACL 2024. Strou...

  20. [29]

    Video understanding with large language models: A survey,

    Y. Tang, J. Bi, S. Xu, L. Song, S. Liang, T. Wang, D. Zhang, J. An, J. Lin, R. Zhu, A. Vosoughi, C. Huang, Z. Zhang, P . Liu, M. Feng, F. Zheng, J.-L. Gaudiot, P . Luo, J. Luo, and C. Xu, “Video understanding with large language models: A survey,” IEEE Transactions on Circuits...

  21. [30]

    A survey on video temporal grounding with multimodal large lan- guage model,

    J. Wu, W. Liu, Y. Liu, M. Liu, L. Nie, Z. Lin, and C. W. Chen, “A survey on video temporal grounding with multimodal large lan- guage model,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 2, pp. 1521–1541, 2026

  22. [31]

    Video-LMM post-training: A deep dive into video reasoning with large multimodal models,

    Y. Tang, J. Bi, P . Liu, Z. Pan, Z. Tan, Q. Shen, J. Liu, H. Hua, J. Guo, Y. Xiaoet al., “Video-LMM post-training: A deep dive into video reasoning with large multimodal models,” arXiv preprint arXiv:2510.05034, 2025

  23. [32]

    A survey of reinforcement learning for large reasoning models,

    K. Zhang, Y. Zuo, B. He, Y. Sun, R. Liu, C. Jiang, Y. Fan, K. Tian, G. Jia, P . Liet al., “A survey of reinforcement learning for large reasoning models,” arXiv preprint arXiv:2509.08827, 2025

  24. [33]

    Memory in the age of ai agents,

    Y. Hu, S. Liu, Y. Yue, G. Zhang, B. Liu, F. Zhu, J. Lin, H. Guo, S. Dou, Z. Xiet al., “Memory in the age of ai agents,” arXiv preprint arXiv:2512.13564, 2025

  25. [34]

    Token reduction should go beyond efficiency in generative models–from vision, language to multimodality,

    Z. Kong, Y. Li, F. Zeng, L. Xin, S. Messica, X. Lin, P . Zhao, M. Kellis, H. Tang, and M. Zitnik, “Token reduction should go beyond efficiency in generative models–from vision, language to multimodality,” arXiv preprint arXiv:2505.18227, 2025

  26. [35]

    Perception, reason, think, and plan: A survey on large multimodal reasoning models,

    Y. Li, Z. Liu, Z. Li, X. Zhang, Z. Xu, X. Chen, H. Shi, S. Jiang, X. Wang, J. Wanget al., “Perception, reason, think, and plan: A survey on large multimodal reasoning models,” arXiv preprint arXiv:2505.04921, 2025

  27. [36]

    Lita: Language instructed temporal- localization assistant,

    D.-A. Huang, S. Liao, S. Radhakrishnan, H. Yin, P . Molchanov, Z. Yu, and J. Kautz, “Lita: Language instructed temporal- localization assistant,” inEuropean Conference on Computer Vision (ECCV). Cham, Switzerland: Springer, 2024

  28. [37]

    Universal video temporal grounding with generative multi-modal large language models,

    Z. Li, S. Di, Z. Zhai, W. Huang, Y. Wang, and W. Xie, “Universal video temporal grounding with generative multi-modal large language models,” inAdvances in Neural Information Processing Systems (NeurIPS). Red Hook, NY, USA: Curran Associates, Inc., 2025, affiliations: Shanghai...

  29. [38]

    Timelens: Rethinking video temporal grounding with multi- modal llms,

    J. Zhang, T. Wang, Y. Ge, Y. Ge, X. Li, Y. Shan, and L. Wang, “Timelens: Rethinking video temporal grounding with multi- modal llms,” arXiv preprint arXiv:2512.14698, 2025, affiliations: Nanjing University; ARC Lab, Tencent PCG; Shanghai AI Lab

  30. [39]

    Towards one-to-many temporal grounding,

    Q. Xu, T. Yue, S. Chen, J. Meng, A. Wang, S. Ji, H. Fei, and X. Li, “Towards one-to-many temporal grounding,” inProceedings of the 43rd International Conference on Machine Learning (ICML). Brookline, MA, USA: PMLR, 2026

  31. [40]

    Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos,

    H. Yuan, X. Li, T. Zhang, Y. Sun, Z. Huang, S. Xu, S. Ji, Y. Tong, L. Qi, J. Fenget al., “Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos,” arXiv preprint arXiv:2501.04001, 2025. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 25

  32. [41]

    Sama: Towards multi-turn referential grounded video chat with large language models,

    Y. Sun, H. Zhang, H. Ding, T. Zhang, X. Ma, and Y.-G. Jiang, “Sama: Towards multi-turn referential grounded video chat with large language models,” inAdvances in Neural Information Pro- cessing Systems. Red Hook, NY, USA: Curran Associates, Inc., 2025

  33. [42]

    Streaming dense video captioning,

    G. Zhou, X. Xiong, A. Bhattacharyya, and J. J. Corso, “Streaming dense video captioning,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Los Alamitos, CA, USA: IEEE Computer Society, 2024, pp. 18 486–18 496

  34. [43]

    Do you remember? dense video captioning with cross-modal memory retrieval,

    M. Kim, H. B. Kim, J. Moon, J. Choi, and S. T. Kim, “Do you remember? dense video captioning with cross-modal memory retrieval,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Los Alamitos, CA, USA: IEEE Computer Society, 2024, pp. 13 894–13 904

  35. [44]

    Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary en- richment and online refinement,

    H. Wu, H. Liu, Y. Qiao, and X. Sun, “Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary en- richment and online refinement,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition. Los Alamitos, CA, USA: IEEE Computer ...

  36. [45]

    PLLaVA: Parameter-free LLaVA extension from images to videos for video dense captioning,

    L. Xu, Y. Huang, S. Xie, W. Wei, T. Li, B. Pan, Y. Zhao, and J. Yuan, “PLLaVA: Parameter-free LLaVA extension from images to videos for video dense captioning,” arXiv preprint arXiv:2404.16994, 2024

  37. [46]

    AuroraCap: Efficient, performant video detailed captioning and a new benchmark,

    W. Chai, E. Song, Y. Du, C. Meng, V . Madhavan, O. Bar-Tal, J.- N. Hwang, S. Xie, and C. D. Manning, “AuroraCap: Efficient, performant video detailed captioning and a new benchmark,” arXiv preprint arXiv:2410.03051, 2024

  38. [47]

    Tarsier2: Advancing large vision-language models from detailed video de- scription to comprehensive video understanding,

    L. Yuan, J. Wang, H. Sun, Y. Zhang, and Y. Lin, “Tarsier2: Advancing large vision-language models from detailed video de- scription to comprehensive video understanding,” arXiv preprint arXiv:2501.07888, 2025

  39. [48]

    Baichuan-omni technical report,

    Y. Li, H. Sun, M. Lin, T. Li, G. Dong, T. Zhang, B. Ding, W. Song, Z. Cheng, Y. Huoet al., “Baichuan-omni technical report,” arXiv preprint arXiv:2410.08565, 2024

  40. [49]

    Ming-omni: A unified mul- timodal model for perception and generation,

    I. AI, B. Gong, C. Zou, C. Zheng, C. Zhou, C. Yan, C. Jin, C. Shen, D. Zheng, F. Wanget al., “Ming-omni: A unified mul- timodal model for perception and generation,” arXiv preprint arXiv:2506.09344, 2025

  41. [50]

    Llama- omni: Seamless speech interaction with large language models,

    Q. Fang, S. Guo, Y. Zhou, Z. Ma, S. Zhang, and Y. Feng, “Llama- omni: Seamless speech interaction with large language models,” arXiv preprint arXiv:2409.06666, 2024

  42. [51]

    Stream-omni: Simultaneous multimodal interactions with large language- vision-speech model,

    S. Zhang, S. Guo, Q. Fang, Y. Zhou, and Y. Feng, “Stream-omni: Simultaneous multimodal interactions with large language- vision-speech model,” arXiv preprint arXiv:2506.13642, 2025

  43. [52]

    Omnicaptioner: One captioner to rule them all,

    Y. Lu, J. Yuan, Z. Li, S. Zhao, Q. Qin, X. Li, L. Zhuo, L. Wen, D. Liu, Y. Caoet al., “Omnicaptioner: One captioner to rule them all,” arXiv preprint arXiv:2504.07089, 2025

  44. [53]

    Omnivinci: Enhancing ar- chitecture and data for omni-modal understanding llm,

    H. Ye, C.-H. H. Yang, A. Goel, W. Huang, L. Zhu, Y. Su, S. Lin, A.-C. Cheng, Z. Wan, J. Tianet al., “Omnivinci: Enhancing ar- chitecture and data for omni-modal understanding llm,” arXiv preprint arXiv:2510.15870, 2025

  45. [54]

    Q-Frame: Query- aware frame selection and multi-resolution adaptation for video- LLMs,

    S. Zhang, J. Yang, J. Yin, Z. Luo, and J. Luan, “Q-Frame: Query- aware frame selection and multi-resolution adaptation for video- LLMs,” arXiv preprint arXiv:2506.22139, 2025

  46. [55]

    DyCoke: Dynamic compression of tokens for fast video large language models,

    K. Tao, C. Qin, H. You, Y. Sui, and H. Wang, “DyCoke: Dynamic compression of tokens for fast video large language models,” in Proceedings of the Computer Vision and Pattern Recognition Confer- ence. Los Alamitos, CA, USA: IEEE Computer Society, 2025, pp. 18 992–19 001

  47. [56]

    Videonsa: Native sparse attention scales video un- derstanding,

    E. Song, W. Chai, S. Yang, E. Armand, X. Shan, H. Xu, J. Xie, and Z. Tu, “Videonsa: Native sparse attention scales video un- derstanding,” arXiv preprint arXiv:2510.02295, 2025

  48. [57]

    Vtimellm: Empower llm to grasp video moments,

    B. Huang, X. Wang, H. Chen, Z. Song, and W. Zhu, “Vtimellm: Empower llm to grasp video moments,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Los Alamitos, CA, USA: IEEE Computer Society, 2024

  49. [58]

    Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,

    Y. Guo, J. Liu, M. Li, D. Cheng, X. Tang, D. Sui, Q. Liu, X. Chen, and K. Zhao, “Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,” inProceed- ings of the AAAI Conference on Artificial Intelligence. Palo Alto, CA, USA: AAAI Press, 2025

  50. [59]

    Distime: Distribution-based time representation for video large language models,

    Y. Zeng, Z. Huang, Y. Zhong, C. Feng, J. Hu, L. Ma, and Y. Liu, “Distime: Distribution-based time representation for video large language models,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Los Alamitos, CA, USA: IEEE Computer Society, 2025

  51. [60]

    Self-chained image- language model for video localization and question answering,

    S. Yu, J. Cho, P . Yadav, and M. Bansal, “Self-chained image- language model for video localization and question answering,” inAdvances in Neural Information Processing Systems (NeurIPS). Red Hook, NY, USA: Curran Associates, Inc., 2023, affiliation: UNC Chapel Hill

  52. [61]

    Llava-mr: Large language-and-vision assistant for video moment retrieval,

    W. Lu, J. Li, A. Yu, M.-C. Chang, S. Ji, and M. Xia, “Llava-mr: Large language-and-vision assistant for video moment retrieval,” arXiv preprint arXiv:2411.14505, 2024, affiliations: Peking Univer- sity; Tencent Youtu; University at Albany; Zhejiang University. (* indicates cor...

  53. [62]

    Timesuite: Improving mllms for long video understanding via grounded tuning,

    X. Zeng, K. Li, C. Wang, X. Li, T. Jiang, Z. Yan, S. Li, Y. Shi, Z. Yue, Y. Wang, Y. Wang, Y. Qiao, and L. Wang, “Timesuite: Improving mllms for long video understanding via grounded tuning,” inInternational Conference on Learning Representations (ICLR). Online: OpenReview.net, 2025

  54. [63]

    Scanning only once: An end-to-end framework for fast temporal grounding in long videos,

    Y. Pan, X. He, B. Gong, Y. Lv, Y. Shen, Y. Peng, and D. Zhao, “Scanning only once: An end-to-end framework for fast temporal grounding in long videos,” inProceedings of the IEEE/CVF Interna- tional Conference on Computer Vision (ICCV). Los Alamitos, CA, USA: IEEE Computer Soci...

  55. [64]

    Trace: Temporal grounding video llm via causal event modeling,

    Y. Guo, J. Liu, M. Li, Q. Liu, X. Chen, and X. Tang, “Trace: Temporal grounding video llm via causal event modeling,” in International Conference on Learning Representations (ICLR). On- line: OpenReview.net, 2025, affiliations: School of Science and Engineering, The Chinese Un...

  56. [65]

    Tar-tvg: Enhancing vlms with timestamp anchor-constrained reasoning for temporal video grounding,

    C. Guo, X. Mo, Y. Nie, X. Xu, C. Xu, F. Yu, and C. Long, “Tar-tvg: Enhancing vlms with timestamp anchor-constrained reasoning for temporal video grounding,” arXiv preprint arXiv:2508.07683, 2025

  57. [66]

    Grounded-videollm: Sharpening fine- grained temporal grounding in video large language models,

    H. Wang, Z. Xu, Y. Cheng, S. Diao, Y. Zhou, Y. Cao, Q. Wang, W. Ge, and L. Huang, “Grounded-videollm: Sharpening fine- grained temporal grounding in video large language models,” arXiv preprint arXiv:2410.03290, 2024

  58. [67]

    Momentor: Advancing video large language model with fine-grained temporal reasoning,

    L. Qian, J. Li, Y. Wu, Y. Ye, H. Fei, T.-S. Chua, Y. Zhuang, and S. Tang, “Momentor: Advancing video large language model with fine-grained temporal reasoning,” inProceedings of the 41st International Conference on Machine Learning (ICML). Brookline, MA, USA: PMLR, 2024

  59. [68]

    Videoperceiver: Enhancing fine-grained temporal perception in video multimodal large language models,

    F. Zhao, L. Zhang, D. Shi, Y. Gao, C. Ye, Y. Cai, J. Gao, and D. Yan, “Videoperceiver: Enhancing fine-grained temporal perception in video multimodal large language models,” arXiv preprint arXiv:2511.18823, 2025

  60. [69]

    Time-r1: Post-training large vision language model for temporal video grounding,

    Y. Wang, Z. Wang, B. Xu, Y. Du, K. Lin, Z. Xiao, Z. Yue, J. Ju, L. Zhang, D. Yanget al., “Time-r1: Post-training large vision language model for temporal video grounding,” arXiv preprint arXiv:2503.13377, 2025

  61. [70]

    Video-opd: Efficient post-training of multimodal large language models for temporal video grounding via on-policy distillation,

    J. Li, H. Yin, H. Xu, B. Xu, W. Tan, Z. He, J. Ju, Z. Luo, and J. Luan, “Video-opd: Efficient post-training of multimodal large language models for temporal video grounding via on-policy distillation,” arXiv preprint arXiv:2602.02994, 2026

  62. [71]

    Videozoomer: Reinforcement-learned temporal focusing for long video reason- ing,

    Y. Ding, Y. Zhang, X. Lai, R. Chu, and Y. Yang, “Videozoomer: Reinforcement-learned temporal focusing for long video reason- ing,” arXiv preprint arXiv:2512.22315, 2025

  63. [72]

    Datasets and recipes for video temporal grounding via reinforcement learning,

    R. Chen, T. Luo, Z. Fan, H. Zou, Z. Feng, G. Xie, H. Zhang, Z. Wang, Z. Liu, and Z. Huaijian, “Datasets and recipes for video temporal grounding via reinforcement learning,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Trac...

  64. [73]

    Museg: Reinforc- ing video temporal understanding via timestamp-aware multi- segment grounding,

    F. Luo, S. Lou, C. Chen, Z. Wang, C. Li, W. Shen, J. Guo, P . Li, M. Yan, J. Zhang, F. Huang, and Y. Liu, “Museg: Reinforc- ing video temporal understanding via timestamp-aware multi- segment grounding,” arXiv preprint arXiv:2505.20715, 2025

  65. [74]

    Detect anything via next point prediction,

    Q. Jiang, J. Huo, X. Chen, Y. Xiong, Z. Zeng, Y. Chen, T. Ren, J. Yu, and L. Zhang, “Detect anything via next point prediction,” arXiv preprint arXiv:2510.12798, 2025

  66. [75]

    Thinking with bounding boxes: Enhancing spatio-temporal video grounding via reinforcement fine-tuning,

    X. Gu, H. Zhang, Q. Fan, J. Niu, Z. Zhang, L. Zhang, G. Chen, F. Chen, L. Wen, and S. Zhu, “Thinking with bounding boxes: Enhancing spatio-temporal video grounding via reinforcement fine-tuning,” arXiv preprint arXiv:2511.21375, 2025

  67. [76]

    Universal instance perception as object discovery and retrieval,

    B. Yan, Y. Jiang, J. Wu, D. Wang, Z. Yuan, P . Luo, and H. Lu, IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 26 “Universal instance perception as object discovery and retrieval,” inCVPR. Los Alamitos, CA, USA: IEEE Computer Society, 2023

  68. [77]

    Multimodal referring segmentation: A survey,

    H. Ding, S. Tang, S. He, C. Liu, Z. Wu, and Y.-G. Jiang, “Multimodal referring segmentation: A survey,” arXiv preprint arXiv:2508.00265, 2025

  69. [78]

    Deformable detr: Deformable transformers for end-to-end object detection,

    X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” arXiv preprint arXiv:2010.04159, 2020

  70. [79]

    Lisa: Reasoning segmentation via large language model,

    X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia, “Lisa: Reasoning segmentation via large language model,” inCVPR. Los Alamitos, CA, USA: IEEE Computer Society, 2024

  71. [80]

    Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding,

    T. Zhang, X. Li, H. Fei, H. Yuan, S. Wu, S. Ji, C. L. Chen, and S. Yan, “Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding,” inNeurIPS. Red Hook, NY, USA: Curran Associates, Inc., 2024

  72. [81]

    Generalizable entity grounding via assistance of large language model,

    L. Qi, Y.-W. Chen, L. Yang, T. Shen, X. Li, W. Guo, Y. Xu, and M.- H. Yang, “Generalizable entity grounding via assistance of large language model,” arXiv preprint arXiv:2402.02555, 2024

  73. [82]

    Sam 2: Segment anything in images and videos,

    N. Ravi, V . Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R ¨adle, C. Rolland, L. Gustafsonet al., “Sam 2: Segment anything in images and videos,” arXiv preprint arXiv:2408.00714, 2024

  74. [83]

    Unipixel: Unified object referring and segmentation for pixel- level visual reasoning,

    Y. Liu, Z. Ma, J. Pu, Z. Qi, Y. Wu, Y. Shan, and C. W. Chen, “Unipixel: Unified object referring and segmentation for pixel- level visual reasoning,” inAdvances in Neural Information Process- ing Systems. Red Hook, NY, USA: Curran Associates, Inc., 2025

  75. [84]

    Samtok: Representing any mask with two words,

    Y. Zhou, T. Zhang, D. Gong, Y. Wu, Y. Tian, H. Wang, H. Yuan, J. Wang, L. Qi, H. Feiet al., “Samtok: Representing any mask with two words,” arXiv preprint arXiv:2601.16093, 2026

  76. [85]

    Collecting highly parallel data for paraphrase evaluation,

    D. Chen and W. B. Dolan, “Collecting highly parallel data for paraphrase evaluation,” inProceedings of the 49th annual meeting of the association for computational linguistics: human language tech- nologies. Stroudsburg, PA, USA: Association for Computational Linguistics, 2011...

  77. [87]

    Video-llava: Learning united visual representation by alignment before projection,

    B. Lin, Y. Ye, B. Zhu, J. Cui, M. Ning, P . Jin, and L. Yuan, “Video-llava: Learning united visual representation by alignment before projection,” inProceedings of the 2024 conference on empirical methods in natural language processing. Stroudsburg, PA, USA: Association for Co...

  78. [88]

    Llava- video: Video instruction tuning with synthetic data,

    Y. Zhang, J. Wu, W. Li, B. Li, Z. Ma, Z. Liu, and C. Li, “Llava- video: Video instruction tuning with synthetic data,” arXiv preprint arXiv:2410.02713, 2024

  79. [89]

    Video recap: Recursive captioning of hour-long videos,

    M. M. Islam, N. Ho, X. Yang, T. Nagarajan, L. Torresani, and G. Bertasius, “Video recap: Recursive captioning of hour-long videos,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Los Alamitos, CA, USA: IEEE Computer Society, 2024, pp. 18 1...

  80. [90]

    Longcaptioning: Unlocking the power of long video caption generation in large multimodal models,

    H. Wei, Z. Tan, Y. Hu, C. W. Chen, and Z. Chen, “Longcaptioning: Unlocking the power of long video caption generation in large multimodal models,” arXiv preprint arXiv:2502.15393, 2025

  81. [91]

    Fine-grained captioning of long videos through scene graph consolidation,

    S. Chu, S. Seo, and B. Han, “Fine-grained captioning of long videos through scene graph consolidation,” arXiv preprint arXiv:2502.16427, 2025

  82. [92]

    Tarsier: Recipes for training and evaluating large video description models,

    J. Wang, L. Yuan, Y. Zhang, and H. Sun, “Tarsier: Recipes for training and evaluating large video description models,” arXiv preprint arXiv:2407.00634, 2024

  83. [93]

    video-salmonn 2: Caption-enhanced audio-visual large language models,

    C. Tang, Y. Li, Y. Yang, J. Zhuang, G. Sun, W. Li, Z. Ma, and C. Zhang, “video-salmonn 2: Caption-enhanced audio-visual large language models,” arXiv preprint arXiv:2506.15220, 2025

  84. [94]

    Videocap-r1: Enhancing mllms for video captioning via structured thinking,

    D. Meng, R. Huang, Z. Dai, X. Li, Y. Xu, J. Zhang, Z. Huang, M. Zhang, L. Zhang, Y. Liuet al., “Videocap-r1: Enhancing mllms for video captioning via structured thinking,” arXiv preprint arXiv:2506.01725, 2025

  85. [95]

    OwlCap: Harmonizing motion-detail for video caption- ing via HMD-270K and caption set equivalence reward,

    C. Zhong, Q. Hou, Z. Zhou, S. Hao, H. Lu, Y. Zhang, H. Tang, and X. Bai, “OwlCap: Harmonizing motion-detail for video caption- ing via HMD-270K and caption set equivalence reward,” arXiv preprint arXiv:2508.18634, 2025

  86. [96]

    Towards fine-grained human motion video captioning,

    G. Song, G. Wang, Z. Huang, J. Lin, X. Zhe, J. Li, and H. Wang, “Towards fine-grained human motion video captioning,” inACM International Conference on Multimedia. New York, NY, USA: Association for Computing Machinery, 2025, pp. 846–855

  87. [97]

    Sharegpt4video: Improving video understanding and generation with better captions,

    L. Chen, X. Wei, J. Li, X. Dong, P . Zhang, Y. Zang, Z. Chen, H. Duan, B. Lin, Z. Tanget al., “Sharegpt4video: Improving video understanding and generation with better captions,”Advances in Neural Information Processing Systems, vol. 37, pp. 19 472–19 495, 2024

  88. [98]

    Panda- 70m: Captioning 70m videos with multiple cross-modality teach- ers,

    T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H.-w. Chao, B. E. Jeon, Y. Fang, H.-Y. Lee, J. Ren, M.-H. Yanget al., “Panda- 70m: Captioning 70m videos with multiple cross-modality teach- ers,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  89. [99]

    Vript: A video is worth thousands of words,

    D. Yang, S. Huang, C. Lu, X. Han, H. Zhang, Y. Gao, Y. Hu, and H. Zhao, “Vript: A video is worth thousands of words,”Advances in Neural Information Processing Systems, vol. 37, pp. 57 240–57 261, 2024

  90. [100]

    IF-VidCap: Can video caption models follow instructions?

    S. Li, Y. Zhang, J. Wu, Z. Lei, Y. He, R. Wen, C. Liao, C. Jiang, A. Ping, S. Gaoet al., “IF-VidCap: Can video caption models follow instructions?” arXiv preprint arXiv:2510.18726, 2025

  91. [101]

    AnyCap project: A unified framework, dataset, and benchmark for controllable omni-modal captioning,

    Y. Ren, Z. Lin, Y. Li, G. Meng, W. Wang, J. Wang, Z. Lin, J. Dai, Y. Yang, W. Wanget al., “AnyCap project: A unified framework, dataset, and benchmark for controllable omni-modal captioning,” arXiv preprint arXiv:2507.12841, 2025

  92. [102]

    Intentvcnet: Bridging spatio-temporal gaps for intention-oriented controllable video captioning,

    T. Qiu, J. Gao, J. Li, H. Leong, X. Huang, X. Wang, X. Zhang, K. Xu, and L. Zhang, “Intentvcnet: Bridging spatio-temporal gaps for intention-oriented controllable video captioning,” inProceed- ings of the 33rd ACM International Conference on Multimedia. New York, NY, USA: Asso...

  93. [103]

    Dense-captioning events in videos,

    R. Krishnaet al., “Dense-captioning events in videos,” arXiv preprint arXiv:1705.00754, 2017

  94. [104]

    End-to-end dense video captioning with masked transformer,

    L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, “End-to-end dense video captioning with masked transformer,” inProceedings of the IEEE conference on computer vision and pattern recognition. Los Alamitos, CA, USA: IEEE Computer Society, 2018, pp. 8739– 8748

  95. [105]

    End-to-end dense video captioning with parallel decoding,

    T. Wang, R. Zhang, Z. Lu, F. Zheng, R. Cheng, and P . Luo, “End-to-end dense video captioning with parallel decoding,” in Proceedings of the IEEE/CVF international conference on computer vision. Los Alamitos, CA, USA: IEEE Computer Society, 2021, pp. 6847–6857

  96. [106]

    Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,

    A. Yang, A. Nagrani, P . H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid, “Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,” inProceed- ings of the IEEE/CVF conference on computer vision and pattern recognition. Los Al...

  97. [107]

    Videollm knows when to speak: Enhancing time- sensitive video comprehension with video-text duet interaction format,

    Y. Wang, X. Meng, Y. Wang, J. Liang, J. Wei, H. Zhang, and D. Zhao, “Videollm knows when to speak: Enhancing time- sensitive video comprehension with video-text duet interaction format,”arXiv preprint arXiv:2411.17991, vol. 1, no. 3, p. 5, 2024

  98. [108]

    Hicm 2: Hierarchical compact memory modeling for dense video caption- ing,

    M. Kim, H. B. Kim, J. Moon, J. Choi, and S. T. Kim, “Hicm 2: Hierarchical compact memory modeling for dense video caption- ing,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39. Palo Alto, CA, USA: AAAI Press, 2025, pp. 4293–4301

  99. [109]

    VideoRefer suite: Advancing spatial-temporal object understanding with video LLM,

    Y. Yuan, H. Zhang, W. Li, Z. Cheng, B. Zhang, L. Li, X. Li, D. Zhao, W. Zhang, Y. Zhuang, J. Zhu, and L. Bing, “VideoRefer suite: Advancing spatial-temporal object understanding with video LLM,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...

  100. [110]

    PixelRefer: A unified framework for spatio-temporal object referring with arbitrary granularity,

    Alibaba DAMO Academy, “PixelRefer: A unified framework for spatio-temporal object referring with arbitrary granularity,” arXiv preprint arXiv:2510.23603, 2025

  101. [111]

    Describe anything: Detailed localized image and video captioning,

    L. Lian, Y. Ding, Y. Ge, S. Liu, H. Mao, B. Li, M. Pavone, M.-Y. Liu, T. Darrell, A. Yalaet al., “Describe anything: Detailed localized image and video captioning,” inProceedings of the IEEE/CVF International Conference on Computer Vision. Los Alamitos, CA, USA: IEEE Computer ...

  102. [112]

    Omni-rgpt: Unifying image and video region-level understanding via token marks,

    M. Heo, M.-H. Chen, D.-A. Huang, S. Liu, S. Radhakrishnan, S. J. Kim, Y.-C. F. Wang, and R. Hachiuma, “Omni-rgpt: Unifying image and video region-level understanding via token marks,” inProceedings of the Computer Vision and Pattern Recognition Con- ference. Los Alamitos, CA, ...

  103. [113]

    Caption anything in IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 27 video: Fine-grained object-centric captioning via spatiotemporal multimodal prompting,

    Y. Y. Tang, J. Bi, C. Huang, S. Liang, D. Shimada, H. Hua, Y. Xiao, Y. Song, P . Liu, M. Feng, J. Guo, Z. Liu, L. Song, A. Vosoughi, J. He, L. He, Z. Zhang, J. Luo, and C. Xu, “Caption anything in IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 27 video: Fine-g...

  104. [114]

    Perceive anything: Recognize, explain, caption, and segment anything in images and videos,

    W. Lin, X. Wei, R. An, T. Ren, T. Chen, R. Zhang, Z. Guo, W. Zhang, L. Zhang, and H. Li, “Perceive anything: Recognize, explain, caption, and segment anything in images and videos,” arXiv preprint arXiv:2506.05302, 2025

  105. [115]

    Artemis: Towards referential understanding in complex videos,

    J. Qiu, Y. Zhang, X. Tang, L. Xie, T. Ma, P . Yan, D. Doermann, Q. Ye, and Y. Tian, “Artemis: Towards referential understanding in complex videos,” inAdvances in Neural Information Processing Systems, vol. 37. Red Hook, NY, USA: Curran Associates, Inc., 2024, pp. 114 321–114 347

  106. [116]

    Strefer: Empowering video LLMs with space- time referring and reasoning via synthetic instruction data,

    H. Zhou, X. Peng, S. Kendre, M. S. Ryoo, S. Savarese, C. Xiong, and J. C. Niebles, “Strefer: Empowering video LLMs with space- time referring and reasoning via synthetic instruction data,” in Proceedings of the IEEE/CVF International Conference on Computer Vision. Los Alamitos...

  107. [117]

    Elysium: Ex- ploring object-level perception in videos via MLLM,

    H. Wang, Y. Ye, Y. Wang, Y. Nie, and C. Huang, “Elysium: Ex- ploring object-level perception in videos via MLLM,” inEuropean Conference on Computer Vision, Springer. Cham, Switzerland: Springer, 2024, pp. 166–185

  108. [118]

    Dense video object captioning from disjoint supervision,

    X. Zhou, A. Arnab, C. Sun, and C. Schmid, “Dense video object captioning from disjoint supervision,” arXiv preprint arXiv:2306.11729, 2023

  109. [119]

    MaskCaptioner: Learning to jointly seg- ment and caption object trajectories in videos,

    G. Fiastreet al., “MaskCaptioner: Learning to jointly seg- ment and caption object trajectories in videos,” arXiv preprint arXiv:2510.14904, 2025

  110. [120]

    Videoglamm: A large multimodal model for pixel-level visual grounding in videos,

    S. Munasinghe, H. Gani, W. Zhu, J. Cao, E. Xing, F. S. Khan, and S. Khan, “Videoglamm: A large multimodal model for pixel-level visual grounding in videos,” inProceedings of the Computer Vision and Pattern Recognition Conference. Los Alamitos, CA, USA: IEEE Computer Society, 2...

  111. [121]

    VoCap: Video object captioning and segmen- tation from any prompt,

    Google DeepMind, “VoCap: Video object captioning and segmen- tation from any prompt,” arXiv preprint arXiv:2508.21809, 2025

  112. [122]

    ViCaS: A dataset for combin- ing holistic and pixel-level video understanding using captions with grounded segmentation,

    A. Athar, X. Deng, and L.-C. Chen, “ViCaS: A dataset for combin- ing holistic and pixel-level video understanding using captions with grounded segmentation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Los Alamitos, CA, USA: IEEE...

  113. [123]

    Gpt-4o system card,

    A. Hurst, A. Lerer, A. P . Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radfordet al., “Gpt-4o system card,” arXiv preprint arXiv:2410.21276, 2024

  114. [124]

    Omni-r1: Towards the unified generative paradigm for multimodal reasoning,

    D. Cheng, Y. Li, Z. Ma, H. Cai, Y. Hu, W. Wang, L. Nie, and W. Li, “Omni-r1: Towards the unified generative paradigm for multimodal reasoning,” arXiv preprint arXiv:2601.09536, 2026

  115. [125]

    Megrez-omni technical report,

    B. Li, Y. Li, Z. Li, C. Liu, W. Liu, G. Niu, Z. Tan, H. Xu, Z. Yao, T. Yuanet al., “Megrez-omni technical report,” arXiv preprint arXiv:2502.15803, 2025

  116. [126]

    Interactiveomni: A unified omni-modal model for audio-visual multi-turn dialogue,

    W. Tong, H. Guo, D. Ran, J. Chen, J. Lu, K. Wang, K. Li, X. Zhu, J. Li, K. Liet al., “Interactiveomni: A unified omni-modal model for audio-visual multi-turn dialogue,” arXiv preprint arXiv:2510.13747, 2025

  117. [127]

    Llama-omni2: Llm-based real-time spoken chatbot with autoregressive stream- ing speech synthesis,

    Q. Fang, Y. Zhou, S. Guo, S. Zhang, and Y. Feng, “Llama-omni2: Llm-based real-time spoken chatbot with autoregressive stream- ing speech synthesis,” arXiv preprint arXiv:2505.02625, 2025

  118. [128]

    Adaretake: Adaptive redundancy reduction to perceive longer for video- language understanding,

    X. Wang, Q. Si, S. Zhu, J. Wu, L. Cao, and L. Nie, “Adaretake: Adaptive redundancy reduction to perceive longer for video- language understanding,” inFindings of the Association for Compu- tational Linguistics: ACL 2025. Stroudsburg, PA, USA: Association for Computational Ling...

  119. [129]

    Logic-in-frames: Dynamic keyframe search via visual semantic- logical verification for long video understanding,

    W. Guo, Z. Chen, S. Wang, J. He, Y. Xu, J. Ye, Y. Sun, and H. Xiong, “Logic-in-frames: Dynamic keyframe search via visual semantic- logical verification for long video understanding,” arXiv preprint arXiv:2503.13139, 2025

  120. [130]

    Divide, then ground: Adapting frame selection to query types for long-form video understanding,

    J. Li, B. Li, J. Li, and Y. Lu, “Divide, then ground: Adapting frame selection to query types for long-form video understanding,” arXiv preprint arXiv:2512.04000, 2025

  121. [131]

    Refocus: Reinforcement- guided frame optimization for contextual understanding,

    H. Lee, J. Kim, H. Kim, and Y. M. Ro, “Refocus: Reinforcement- guided frame optimization for contextual understanding,” arXiv preprint arXiv:2506.01274, 2025

  122. [132]

    FrameOracle: Learning what to see and how much to see in videos,

    C. Li, T. Li, F. Tao, Z. Zhao, Z. Wu, M. Zhao, J. Song, C. Niu, and P . Fazli, “FrameOracle: Learning what to see and how much to see in videos,” arXiv preprint arXiv:2510.03584, 2025

  123. [133]

    From frames to clips: Training-free adaptive key clip se- lection for long-form video understanding,

    G. Sun, A. Singhal, B. Uzkent, M. Shah, C. Chen, and G. Kessler, “From frames to clips: Training-free adaptive key clip se- lection for long-form video understanding,” arXiv preprint arXiv:2510.02262, 2025

  124. [134]

    K-frames: Scene-driven any-k keyframe selection for long video understanding,

    Y. Yao, Y. Yun, J. Wang, H. Zhang, D. Zhao, K. Tian, Z. Wang, M. Qiu, and T. Wang, “K-frames: Scene-driven any-k keyframe selection for long video understanding,” arXiv preprint arXiv:2510.13891, 2025

  125. [135]

    HoliTom: Holistic token merging for fast video large language models,

    K. Shao, K. Tao, C. Qin, H. You, Y. Sui, and H. Wang, “HoliTom: Holistic token merging for fast video large language models,” arXiv preprint arXiv:2505.21334, 2025

  126. [136]

    Video compression commander: Plug-and-play inference acceleration for video large language models,

    X. Liu, Y. Wang, J. Ma, and L. Zhang, “Video compression commander: Plug-and-play inference acceleration for video large language models,” arXiv preprint arXiv:2505.14454, 2025

  127. [137]

    Seeing more, saying more: Lightweight language experts are dynamic video token compressors,

    X. Wang, J. Zhang, T. Wang, H. Zhang, and F. Zheng, “Seeing more, saying more: Lightweight language experts are dynamic video token compressors,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for Com...

  128. [138]

    Less is more, but where? dynamic token compression via LLM-guided keyframe prior,

    Y. Li, H. Gui, Z. Fan, J. Wang, B. Kang, B. Chen, and Z. Tian, “Less is more, but where? dynamic token compression via LLM-guided keyframe prior,” arXiv preprint arXiv:2512.06866, 2025

  129. [139]

    Videollm-mod: Efficient video- language streaming with mixture-of-depths vision computation,

    S. Wu, J. Chen, K. Q. Lin, Q. Wang, Y. Gao, Q. Xu, T. Xu, Y. Hu, E. Chen, and M. Z. Shou, “Videollm-mod: Efficient video- language streaming with mixture-of-depths vision computation,” Advances in Neural Information Processing Systems, vol. 37, pp. 109 922–109 947, 2024

  130. [140]

    Adaptive video understanding agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning,

    S. Jeoung, G. Huybrechts, B. Ganesh, A. Galstyan, and S. Bodap- ati, “Adaptive video understanding agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning,” arXiv preprint arXiv:2410.20252, 2024

  131. [141]

    Videoagent: A memory-augmented multimodal agent for video understand- ing,

    Y. Fan, X. Ma, R. Wu, Y. Du, J. Li, Z. Gao, and Q. Li, “Videoagent: A memory-augmented multimodal agent for video understand- ing,” inEuropean Conference on Computer Vision, pp. 75–92

  132. [142]

    Lvagent: Long video understanding by multi-round dynamical collaboration of mllm agents,

    B. Chen, Z. Yue, S. Chen, Z. Wang, Y. Liu, P . Li, and Y. Wang, “Lvagent: Long video understanding by multi-round dynamical collaboration of mllm agents,” 2025. [Online]. Available: https://arxiv.org/abs/2503.10200

  133. [143]

    Adavideorag: Omni-contextual adaptive retrieval-augmented efficient long video understanding,

    Z. Xue, J. Zhang, X. Xie, Y. Cai, Y. Liu, X. Li, and D. Tao, “Adavideorag: Omni-contextual adaptive retrieval-augmented efficient long video understanding,” inThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025

  134. [144]

    Videolucy: Deep memory backtracking for long video understanding,

    J. Zuo, Y. Deng, L. Kong, J. Yang, R. Jin, Y. Zhang, N. Sang, L. Pan, Z. Liu, and C. Gao, “Videolucy: Deep memory backtracking for long video understanding,” inThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025

  135. [145]

    Gcagent: Long-video understanding via schematic and narrative episodic memory,

    J. H. Yeo, S. Chung, S. Park, D. H. Kim, J. Moon, and Y. M. Ro, “Gcagent: Long-video understanding via schematic and narrative episodic memory,” 2025. [Online]. Available: https://arxiv.org/abs/2511.12027

  136. [146]

    Agentic very long video understanding,

    A. Rege, A. Sadhu, Y. Li, K. Li, R. K. Vinayak, Y. Chai, Y. J. Lee, and H. J. Kim, “Agentic very long video understanding,” 2026. [Online]. Available: https://arxiv.org/abs/2601.18157

  137. [147]

    Memgen: Weaving genera- tive latent memory for self-evolving agents,

    G. Zhang, M. Fu, and S. Yan, “Memgen: Weaving genera- tive latent memory for self-evolving agents,”arXiv preprint arXiv:2509.24704, 2025

  138. [148]

    Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory,

    L. Long, Y. He, W. Ye, Y. Pan, Y. Lin, H. Li, J. Zhao, and W. Li, “Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory,” inThe Fourteenth International Conference on Learning Representations, 2026

  139. [149]

    Vidcompress: Memory- enhanced temporal compression for video understanding in large language models,

    X. Lan, Y. Yuan, Z. Jie, and L. Ma, “Vidcompress: Memory- enhanced temporal compression for video understanding in large language models,” 2024. [Online]. Available: https: //arxiv.org/abs/2410.11417

  140. [150]

    Rewind: Understanding long videos with instructed learnable memory,

    A. Diko, T. Wang, W. Swaileh, S. Sun, and I. Patras, “Rewind: Understanding long videos with instructed learnable memory,” inProceedings of the Computer Vision and Pattern Recognition Con- ference, 2025, pp. 13 734–13 743

  141. [151]

    Enhancing long video understanding via hierarchical event-based memory,

    D. Cheng, M. Li, J. Liu, Y. Guo, B. Jiang, Q. Liu, X. Chen, and B. Zhao, “Enhancing long video understanding via hierarchical event-based memory,” in2025 IEEE International Conference on Multimedia and Expo (ICME), 2025, pp. 1–6

  142. [152]

    ∞- video: A training-free approach to long video understanding via continuous-time memory consolidation,

    S. Santos, A. Farinhas, D. C. McNamee, and A. Martins, “∞- video: A training-free approach to long video understanding via continuous-time memory consolidation,” inInternational Confer- ence on Machine Learning, 2025, pp. 52 877–52 893. IEEE TRANSACTIONS ON PATTERN ANAL YSIS A...

  143. [153]

    Videollamb: Long streaming video understanding with recurrent memory bridges,

    Y. Wang, Y. Song, C. Xie, Y. Liu, and Z. Zheng, “Videollamb: Long streaming video understanding with recurrent memory bridges,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 24 170–24 181

  144. [154]

    Hermes: temporal-coherent long-form understanding with episodes and semantics,

    G. J. Faure, J.-F. Yeh, M.-H. Chen, H.-T. Su, S.-H. Lai, and W. H. Hsu, “Hermes: temporal-coherent long-form understanding with episodes and semantics,” inProceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, 2025, pp. 22 911–22 921

  145. [155]

    Hierarq: Task-aware hi- erarchical q-former for enhanced video understanding,

    S. Azad, V . Vineet, and Y. S. Rawat, “Hierarq: Task-aware hi- erarchical q-former for enhanced video understanding,” inPro- ceedings of the IEEE/CVF Computer Vision and Pattern Recognition Conference, 2025, pp. 8545–8556

  146. [156]

    MARC: Memory-augmented RL token compression for efficient video understanding,

    P . Wu, Z. Yu, Y. Liu, C.-H. Wu, E. Zhou, and J. Shen, “MARC: Memory-augmented RL token compression for efficient video understanding,” inThe Fourteenth International Conference on Learning Representations, 2026

  147. [157]

    Streaming long video understanding with large language mod- els,

    R. Qian, X. Dong, P . Zhang, Y. Zang, S. Ding, D. Lin, and J. Wang, “Streaming long video understanding with large language mod- els,” inAdvances in Neural Information Processing Systems, 2024, pp. 119 336–119 360

  148. [158]

    Streaming video understanding and multi-round interaction with memory-enhanced knowledge,

    H. Xiong, Z. Yang, J. Yu, Y. Zhuge, L. Zhang, J. Zhu, and H. Lu, “Streaming video understanding and multi-round interaction with memory-enhanced knowledge,” inThe Thirteenth Interna- tional Conference on Learning Representations, 2025

  149. [159]

    Recurrent attention-based token selection for efficient streaming video- llms,

    E. Dorovatas, S. Seifi, G. Gupta, and R. Aljundi, “Recurrent attention-based token selection for efficient streaming video- llms,” inAdvances in Neural Information Processing Systems, 2026, pp. 144 088–144 114

  150. [160]

    Memory-efficient streaming videollms for real-time procedural video understanding,

    D. Chatterjee, E. Remelli, Y. Song, B. Tekin, A. Mittal, B. Bhatnagar, N. C. Camg ¨oz, S. Hampali, E. Sauser, S. Ma, A. Yao, and F. Sener, “Memory-efficient streaming videollms for real-time procedural video understanding,” 2025. [Online]. Available: https://arxiv.org/abs/2504.13915

  151. [161]

    Infinipot-v: Memory- constrained KV cache compression for streaming video under- standing,

    M. Kim, K. Shim, J. Choi, and S. Chang, “Infinipot-v: Memory- constrained KV cache compression for streaming video under- standing,” inThe Thirty-ninth Annual Conference on Neural Infor- mation Processing Systems, 2025

  152. [162]

    Unleashing hour-scale video training for long video-language understanding,

    J. Lin, J. Wu, X. Sun, Z. Wang, J. Liu, Y. Su, X. Yu, H. Chen, J. Luo, Z. Liu, and E. Barsoum, “Unleashing hour-scale video training for long video-language understanding,” 2025. [Online]. Available: https://arxiv.org/abs/2506.05332

  153. [163]

    Question- guided visual compression with memory feedback for long- term video understanding,

    S. Yamao, N. Miyahara, Y. Qi, and S. Takeuchi, “Question- guided visual compression with memory feedback for long- term video understanding,” 2026. [Online]. Available: https: //arxiv.org/abs/2603.15167

  154. [164]

    Learning compact video representations for efficient long-form video understanding in large multimodal models,

    Y. Chen, J. Wang, Z. Zhang, J. Yi, X. Zhang, Y. Zou, Z. Cai, J. Yuan, X. Li, H. Yanget al., “Learning compact video representations for efficient long-form video understanding in large multimodal models,” inProceedings of the IEEE/CVF Winter Conference on Applications of Compu...

  155. [165]

    See more, store less: Memory-efficient resolution for video moment retrieval,

    M. Jeon, S. Han, J. Hwang, M. Kwon, J. Kim, and J. Kim, “See more, store less: Memory-efficient resolution for video moment retrieval,” arXiv preprint arXiv:2601.09350, 2026

  156. [166]

    Streamingtom: Streaming token compression for efficient video understanding,

    X. Chen, K. Tao, K. Shao, and H. Wang, “Streamingtom: Streaming token compression for efficient video understanding,”

  157. [167]

    Available: https://arxiv.org/abs/2510.18269

    [Online]. Available: https://arxiv.org/abs/2510.18269

  158. [168]

    video- salmonn s: Memory-enhanced streaming audio-visual llm,

    G. Sun, Y. Li, X. Wu, Y. Yang, W. Li, Z. Ma, and C. Zhang, “video- salmonn s: Memory-enhanced streaming audio-visual llm,”

  159. [169]

    Available: https://arxiv.org/abs/2510.11129

    [Online]. Available: https://arxiv.org/abs/2510.11129

  160. [170]

    Flash-vstream: Efficient real-time understanding for long video streams,

    H. Zhang, Y. Wang, Y. Tang, Y. Liu, J. Feng, and X. Jin, “Flash-vstream: Efficient real-time understanding for long video streams,” 2025. [Online]. Available: https://arxiv.org/abs/2506. 23825

  161. [171]

    Streamforest: Efficient online video understanding with persistent event memory,

    X. Zeng, K. Qiu, Q. Zhang, X. Li, J. Wang, J. Li, Z. Yan, K. Tian, M. Tian, X. Zhao, Y. Wang, and L. Wang, “Streamforest: Efficient online video understanding with persistent event memory,”

  162. [172]

    Available: https://arxiv.org/abs/2509.24871

    [Online]. Available: https://arxiv.org/abs/2509.24871

  163. [173]

    Quickvideo: Real-time long video understanding with system algorithm co-design,

    B. Schneider, D. Jiang, C. Du, T. Pang, and W. Chen, “Quickvideo: Real-time long video understanding with system algorithm co-design,” 2025. [Online]. Available: https://arxiv. org/abs/2505.16175

  164. [174]

    Livevlm: Efficient online video understanding via streaming-oriented kv cache and retrieval,

    Z. Ning, G. Liu, Q. Jin, W. Ding, M. Guo, and J. Zhao, “Livevlm: Efficient online video understanding via streaming-oriented kv cache and retrieval,” arXiv preprint arXiv:2505.15269, 2025

  165. [175]

    Streambridge: Turning your offline video large language model into a proactive streaming assistant,

    H. Wang, B. Feng, Z. Lai, M. Xu, S. Li, W. Ge, A. Dehghan, M. Cao, and P . Huang, “Streambridge: Turning your offline video large language model into a proactive streaming assistant,” arXiv preprint arXiv:2505.05467, 2025

  166. [176]

    Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction,

    R. Qian, S. Ding, X. Dong, P . Zhang, Y. Zang, Y. Cao, D. Lin, and J. Wang, “Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction,” inProceedings of the Computer Vision and Pattern Recognition Con- ference. Los Ala...

  167. [177]

    Dorae- mongpt: Toward understanding dynamic scenes with large lan- guage models (exemplified as a video agent),

    Z. Yang, G. Chen, X. Li, W. Wang, and Y. Yang, “Dorae- mongpt: Toward understanding dynamic scenes with large lan- guage models (exemplified as a video agent),” arXiv preprint arXiv:2401.08392, 2024

  168. [178]

    Video-of-thought: Step-by-step video reasoning from perception to cognition,

    H. Fei, S. Wu, W. Ji, H. Zhang, M. Zhang, M.-L. Lee, and W. Hsu, “Video-of-thought: Step-by-step video reasoning from perception to cognition,” arXiv preprint arXiv:2501.03230, 2024

  169. [179]

    Vca: Video curious agent for long video understanding,

    Z. Yang, D. Chen, X. Yu, M. Shen, and C. Gan, “Vca: Video curious agent for long video understanding,” inProceedings of the IEEE/CVF International Conference on Computer Vision. Los Alamitos, CA, USA: IEEE Computer Society, 2025, pp. 20 168– 20 179

  170. [180]

    Flow4agent: Long- form video understanding via motion prior from optical flow,

    R. Liu, S. Sun, H. Tang, W. Gao, and G. Li, “Flow4agent: Long- form video understanding via motion prior from optical flow,” in Proceedings of the IEEE/CVF International Conference on Computer Vision. Los Alamitos, CA, USA: IEEE Computer Society, 2025, pp. 23 817–23 827

  171. [181]

    Deep video discovery: Agentic search with tool use for long-form video understanding,

    X. Zhang, Z. Jia, Z. Guo, J. Li, B. Li, H. Li, and Y. Lu, “Deep video discovery: Agentic search with tool use for long-form video understanding,” arXiv preprint arXiv:2505.18079, 2025

  172. [182]

    Videoa- gent2: Enhancing the llm-based agent system for long-form video understanding by uncertainty-aware cot,

    Z. Zhi, Q. Wu, W. Li, Y. Li, K. Shao, K. Zhouet al., “Videoa- gent2: Enhancing the llm-based agent system for long-form video understanding by uncertainty-aware cot,” arXiv preprint arXiv:2504.04471, 2025

  173. [183]

    Cot-vid: Dynamic chain-of-thought routing with self verification for training-free video reasoning,

    H. Jin, R. Liu, W. Zhang, G. Luo, and G. Li, “Cot-vid: Dynamic chain-of-thought routing with self verification for training-free video reasoning,” arXiv preprint arXiv:2505.11830, 2025

  174. [184]

    Reinforcing video reasoning with focused thinking,

    J. Dang, J. Wu, T. Wang, X. Lin, N. Zhu, H. Chen, W.-S. Zheng, M. Wang, and T.-S. Chua, “Reinforcing video reasoning with focused thinking,” arXiv preprint arXiv:2505.24718, 2025

  175. [185]

    Vistadpo: Video hierarchical spatial-temporal direct preference optimization for large video models,

    H. Huang, H. Chen, S. Wu, M. Luo, J. Fu, X. Du, H. Zhang, and H. Fei, “Vistadpo: Video hierarchical spatial-temporal direct preference optimization for large video models,” arXiv preprint arXiv:2504.13122, 2025

  176. [186]

    Veripo: Cultivating long reasoning in video-llms via verifier-gudied iterative policy optimization,

    Y. Li, X. Chen, Z. Li, Z. Liu, L. Wang, W. Luo, B. Hu, and M. Zhang, “Veripo: Cultivating long reasoning in video-llms via verifier-gudied iterative policy optimization,” arXiv preprint arXiv:2505.19000, 2025

  177. [187]

    Deepvideo-r1: Video reinforcement fine-tuning via difficulty-aware regressive grpo,

    J. Park, J. Na, J. Kim, and H. J. Kim, “Deepvideo-r1: Video reinforcement fine-tuning via difficulty-aware regressive grpo,” arXiv preprint arXiv:2506.07464, 2025

  178. [188]

    Video-cot: A comprehensive dataset for spa- tiotemporal understanding of videos based on chain-of-thought,

    S. Zhang, X. Hao, Y. Tang, L. Zhang, P . Wang, Z. Wang, H. Ma, and S. Zhang, “Video-cot: A comprehensive dataset for spa- tiotemporal understanding of videos based on chain-of-thought,” inProceedings of the 33rd ACM International Conference on Multime- dia. New York, NY, USA: ...

  179. [189]

    Spacer: Reinforcing mllms in video spatial reasoning,

    K. Ouyang, Y. Liu, H. Wu, Y. Liu, H. Zhou, J. Zhou, F. Meng, and X. Sun, “Spacer: Reinforcing mllms in video spatial reasoning,” arXiv preprint arXiv:2504.01805, 2025

  180. [190]

    Videochat-r1. 5: Visual test-time scaling to re- inforce multimodal reasoning by iterative perception,

    Z. Yan, X. Li, Y. He, Z. Yue, X. Zeng, Y. Wang, Y. Qiao, L. Wang, and Y. Wang, “Videochat-r1. 5: Visual test-time scaling to re- inforce multimodal reasoning by iterative perception,” arXiv preprint arXiv:2509.21100, 2025

  181. [191]

    Pixel reasoner: Incentivizing pixel-space reasoning with curiosity-driven rein- forcement learning,

    H. Wang, A. Su, W. Ren, F. Lin, and W. Chen, “Pixel reasoner: Incentivizing pixel-space reasoning with curiosity-driven rein- forcement learning,” arXiv preprint arXiv:2505.15966, 2025

  182. [192]

    Framemind: Frame-interleaved video reasoning via reinforcement learning,

    H. Ge, Y. Wang, K.-W. Chang, H. Wu, and Y. Cai, “Framemind: Frame-interleaved video reasoning via reinforcement learning,” arXiv preprint arXiv:2509.24008, 2025

  183. [193]

    Love-r1: Advancing long video understanding with an adaptive zoom-in mechanism via multi-step reasoning,

    S. Fu, Q. Yang, Y.-M. Li, X. Wei, X. Xie, and W.-S. Zheng, “Love-r1: Advancing long video understanding with an adaptive zoom-in mechanism via multi-step reasoning,” arXiv preprint arXiv:2509.24786, 2025

  184. [194]

    Conan: Progressive learning to reason like a detective over multi-scale visual evidence,

    K. Ouyang, Y. Liu, L. Yao, Y. Cai, H. Zhou, J. Zhou, F. Meng, and X. Sun, “Conan: Progressive learning to reason like a detective over multi-scale visual evidence,” arXiv preprint arXiv:2510.20470, 2025. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 29

  185. [195]

    Videotemp-o3: Harmonizing temporal grounding and video understanding in agentic thinking-with- videos,

    W. Liu, Y. Wang, S. Ma, M. Liu, Q. Su, T. Zhang, H. Fan, C. Liu, K. Jiang, J. Chenet al., “Videotemp-o3: Harmonizing temporal grounding and video understanding in agentic thinking-with- videos,” arXiv preprint arXiv:2602.07801, 2026

  186. [196]

    Videoseek: Long-horizon video agent with tool- guided seeking,

    J. Lin, J. Wu, J. Liu, X. Sun, Z. Wang, X. Yu, J. Luo, Z. Liu, and E. Barsoum, “Videoseek: Long-horizon video agent with tool- guided seeking,” arXiv preprint arXiv:2603.20185, 2026

  187. [197]

    Video-thinker: Sparking

    S. Wang, J. Jin, X. Wang, L. Song, R. Fu, H. Wang, Z. Ge, Y. Lu, and X. Cheng, “Video-thinker: Sparking” thinking with videos” via reinforcement learning,” arXiv preprint arXiv:2510.23473, 2025

  188. [198]

    Rewatch-r1: Boosting complex video reasoning in large vision-language models through agentic data synthesis,

    C. Zhang, Z. Wang, Y. Ma, J. Peng, Y. Wang, Q. Zhou, J. Song, and B. Zheng, “Rewatch-r1: Boosting complex video reasoning in large vision-language models through agentic data synthesis,” arXiv preprint arXiv:2509.23652, 2025

  189. [199]

    OpenAI-o3,

    OpenAI, “OpenAI-o3,” https://openai.com/index/ introducing-o3-and-o4-mini/, 2025

  190. [200]

    DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning,

    D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P . Wang, X. Biet al., “DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning,” arXiv preprint arXiv:2501.12948, 2025

  191. [201]

    Video-star: Reinforcing open-vocabulary action recognition with tools,

    Z. Yuan, X. Qu, C. Qian, R. Chen, J. Tang, L. Sun, X. Chu, D. Zhang, Y. Wang, Y. Caiet al., “Video-star: Reinforcing open-vocabulary action recognition with tools,” arXiv preprint arXiv:2510.08480, 2025

  192. [202]

    Omni-r1: Reinforcement learning for om- nimodal reasoning via two-system collaboration,

    H. Zhong, M. Zhu, Z. Du, Z. Huang, C. Zhao, M. Liu, W. Wang, H. Chen, and C. Shen, “Omni-r1: Reinforcement learning for om- nimodal reasoning via two-system collaboration,” arXiv preprint arXiv:2505.20256, 2025

  193. [203]

    Video-in-the-loop: Span-grounded long video qa with interleaved reasoning,

    C. Wang, D. Bai, Y. Yang, X. Jin, A. Zhang, R. Wang, S. Jiang, Y. Yang, H. Wu, Q. Daiet al., “Video-in-the-loop: Span-grounded long video qa with interleaved reasoning,” arXiv preprint arXiv:2510.04022, 2025

  194. [204]

    Vit- cot: Video-text interleaved chain-of-thought for boosting video understanding in large language models,

    Y. Zhang, X. Liu, R. Tao, Q. Chen, H. Fei, W. Che, and L. Qin, “Vit- cot: Video-text interleaved chain-of-thought for boosting video understanding in large language models,” inProceedings of the 33rd ACM International Conference on Multimedia. New York, NY, USA: Association fo...

  195. [205]

    Videomind: A chain-of-lora agent for long video reasoning,

    Y. Liu, K. Q. Lin, C. W. Chen, and M. Z. Shou, “Videomind: A chain-of-lora agent for long video reasoning,” arXiv preprint arXiv:2503.13444, 2025

  196. [206]

    Per- ceive, reflect and understand long video: Progressive multi- granular clue exploration with interactive agents,

    J. Li, K. Wei, Z. Xu, Z. Su, X. Yang, and C. Deng, “Per- ceive, reflect and understand long video: Progressive multi- granular clue exploration with interactive agents,” arXiv preprint arXiv:2509.24943, 2025

  197. [207]

    Videoagent: Long-form video understanding with large language model as agent,

    X. Wang, Y. Zhang, O. Zohar, and S. Yeung-Levy, “Videoagent: Long-form video understanding with large language model as agent,” inEuropean Conference on Computer Vision, 2024, pp. 58– 76

  198. [208]

    Vide- orag: Retrieval-augmented generation with extreme long-context videos,

    X. Ren, L. Xu, L. Xia, S. Wang, D. Yin, and C. Huang, “Vide- orag: Retrieval-augmented generation with extreme long-context videos,” arXiv preprint arXiv:2502.01549, 2025

  199. [209]

    Videoforest: Person-anchored hierarchical reasoning for cross- video question answering,

    Y. Meng, J. Ye, W. Zhou, G. Yue, X. Mao, R. Wang, and B. Zhao, “Videoforest: Person-anchored hierarchical reasoning for cross- video question answering,” inProceedings of the 33rd ACM In- ternational Conference on Multimedia. New York, NY, USA: Association for Computing Machin...

  200. [210]

    Drvideo: Document retrieval based long video understanding,

    Z. Ma, C. Gou, H. Shi, B. Sun, S. Li, H. Rezatofighi, and J. Cai, “Drvideo: Document retrieval based long video understanding,” inProceedings of the Computer Vision and Pattern Recognition Con- ference. Los Alamitos, CA, USA: IEEE Computer Society, 2025, pp. 18 936–18 946

  201. [211]

    Vgent: Graph- based retrieval-reasoning-augmented generation for long video understanding,

    X. Shen, W. Zhang, J. Chen, and M. Elhoseiny, “Vgent: Graph- based retrieval-reasoning-augmented generation for long video understanding,” arXiv preprint arXiv:2510.14032, 2025

  202. [212]

    Videotree: Adaptive tree-based video represen- tation for llm reasoning on long videos,

    Z. Wang, S. Yu, E. Stengel-Eskin, J. Yoon, F. Cheng, G. Bertasius, and M. Bansal, “Videotree: Adaptive tree-based video represen- tation for llm reasoning on long videos,” inProceedings of the Computer Vision and Pattern Recognition Conference. Los Alamitos, CA, USA: IEEE Comp...

  203. [213]

    Streamagent: Towards anticipa- tory agents for streaming video understanding,

    H. Yang, F. Tang, L. Zhao, X. An, M. Hu, H. Li, X. Zhuang, Y. Lu, X. Zhang, A. Swikiret al., “Streamagent: Towards anticipa- tory agents for streaming video understanding,” arXiv preprint arXiv:2508.01875, 2025

  204. [214]

    Viqagent: Zero-shot video question an- swering via agent with open-vocabulary grounding validation,

    T. Montes and F. Lozano, “Viqagent: Zero-shot video question an- swering via agent with open-vocabulary grounding validation,” arXiv preprint arXiv:2505.15928, 2025

  205. [215]

    Chain-of-frames: Advancing video understanding in multimodal llms via frame-aware reasoning,

    S. Ghazanfari, F. Croce, N. Flammarion, P . Krishnamurthy, F. Khorrami, and S. Garg, “Chain-of-frames: Advancing video understanding in multimodal llms via frame-aware reasoning,” arXiv preprint arXiv:2506.00318, 2025

  206. [216]

    Kwai keye-vl 1.5 technical report,

    B. Yang, B. Wen, B. Ding, C. Liu, C. Chu, C. Song, C. Rao, C. Yi, D. Li, D. Zanget al., “Kwai keye-vl 1.5 technical report,” arXiv preprint arXiv:2509.01563, 2025

  207. [217]

    video-salmonn-o1: Reasoning-enhanced audio-visual large language model,

    G. Sun, Y. Yang, J. Zhuang, C. Tang, Y. Li, W. Li, Z. Ma, and C. Zhang, “video-salmonn-o1: Reasoning-enhanced audio-visual large language model,” arXiv preprint arXiv:2502.11775, 2025

  208. [218]

    Vidbridge-r1: Bridging qa and captioning for rl-based video understanding models with intermediate proxy tasks,

    X. Chen, Y. Zhang, Y. Guan, W. Lin, Z. Wang, B. Zeng, Y. Shi, S. Yang, Q. Liu, P . Wan, L. Wang, and T. Tan, “Vidbridge-r1: Bridging qa and captioning for rl-based video understanding models with intermediate proxy tasks,” 2025

  209. [219]

    Improved visual-spatial reasoning via r1-zero-like training,

    Z. Liao, Q. Xie, Y. Zhang, Z. Kong, H. Lu, Z. Yang, and Z. Deng, “Improved visual-spatial reasoning via r1-zero-like training,” arXiv preprint arXiv:2504.00883, 2025

  210. [220]

    Video-str: Reinforcing mllms in video spatio-temporal reasoning with relation graph,

    W. Wang, H. Zou, T. Luo, R. Huang, Y. Zhao, Z. Wang, H. Zhang, C. Qin, Y. Wang, L. Zhaoet al., “Video-str: Reinforcing mllms in video spatio-temporal reasoning with relation graph,” arXiv preprint arXiv:2510.10976, 2025

  211. [221]

    Spatialladder: Progressive training for spatial reasoning in vision-language models,

    H. Li, D. Li, Z. Wang, Y. Yan, H. Wu, W. Zhang, Y. Shen, W. Lu, J. Xiao, and Y. Zhuang, “Spatialladder: Progressive training for spatial reasoning in vision-language models,” arXiv preprint arXiv:2510.08531, 2025

  212. [222]

    Cambrian-s: Towards spatial supersensing in video,

    S. Yang, J. Yang, P . Huang, E. Brown, Z. Yang, Y. Yu, S. Tong, Z. Zheng, Y. Xu, M. Wanget al., “Cambrian-s: Towards spatial supersensing in video,” arXiv preprint arXiv:2511.04670, 2025

  213. [223]

    Busterx: Mllm-powered ai- generated video forgery detection and explanation,

    H. Wen, Y. He, Z. Huang, T. Li, Z. Yu, X. Huang, L. Qi, B. Wu, X. Li, and G. Cheng, “Busterx: Mllm-powered ai- generated video forgery detection and explanation,” arXiv preprint arXiv:2505.12620, 2025

  214. [224]

    Omni-fake: Bench- marking unified multimodal social media deepfake detection,

    T. Li, Z. Huang, H. Wen, Y. He, X. Li, B. Zhu, W. Duan, C. Chen, Z. Fu, Y. Dong, B. Wu, J. Li, and G. Cheng, “Omni-fake: Bench- marking unified multimodal social media deepfake detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Los...

  215. [225]

    Longvt: Incentivizing

    Z. Yang, S. Wang, K. Zhang, K. Wu, S. Leng, Y. Zhang, C. Qin, S. Lu, X. Li, and L. Bing, “Longvt: Incentivizing” think- ing with long videos” via native tool calling,” arXiv preprint arXiv:2511.20785, 2025

  216. [226]

    Frame- thinker: Learning to think with long videos via multi-turn frame spotlighting,

    Z. He, X. Qu, Y. Li, S. Huang, D. Liu, and Y. Cheng, “Frame- thinker: Learning to think with long videos via multi-turn frame spotlighting,” arXiv preprint arXiv:2509.24304, 2025

  217. [227]

    Videoexplorer: Think with videos for agentic long-video understanding,

    H. Yuan, Z. Liu, J. Zhou, H. Qian, Y. Shu, N. Sebe, J.-R. Wen, and Z. Dou, “Videoexplorer: Think with videos for agentic long-video understanding,” 2025

  218. [228]

    Reinforcing spatial reasoning in vision-language models with interwoven thinking and visual drawing,

    J. Wu, J. Guan, K. Feng, Q. Liu, S. Wu, L. Wang, W. Wu, and T. Tan, “Reinforcing spatial reasoning in vision-language models with interwoven thinking and visual drawing,” arXiv preprint arXiv:2506.09965, 2025

  219. [229]

    Cyberv: Cybernetics for test-time scaling in video understanding,

    J. Meng, S. Sun, Y. Tan, L. Qi, Y. Tong, X. Li, and L. Wen, “Cyberv: Cybernetics for test-time scaling in video understanding,” arXiv preprint arXiv:2506.07971, 2025

  220. [230]

    Active video perception: Iterative evidence seeking for agentic long video understanding,

    Z. Wang, H. Zhou, S. Wang, J. Li, C. Xiong, S. Savarese, M. Bansal, M. S. Ryoo, and J. C. Niebles, “Active video perception: Iterative evidence seeking for agentic long video understanding,” arXiv preprint arXiv:2512.05774, 2025

  221. [231]

    Video-r2: Reinforcing consistent and grounded reasoning in multimodal language models,

    M. Maaz, H. Rasheed, F. S. Khan, and S. Khan, “Video-r2: Reinforcing consistent and grounded reasoning in multimodal language models,” arXiv preprint arXiv:2511.23478, 2025

  222. [232]

    Anticipative video transformer,

    R. Girdhar and K. Grauman, “Anticipative video transformer,” inProceedings of the IEEE/CVF international conference on computer vision. Los Alamitos, CA, USA: IEEE Computer Society, 2021, pp. 13 505–13 515

  223. [233]

    Is space-time attention all you need for video understanding?

    G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” inIcml, vol. 2. Brookline, MA, USA: PMLR, 2021, p. 4

  224. [234]

    Egocentric video-language pretraining,

    K. Q. Lin, J. Wang, M. Soldan, M. Wray, R. Yan, E. Z. Xu, D. Gao, R.-C. Tu, W. Zhao, W. Konget al., “Egocentric video-language pretraining,”Advances in Neural Information Processing Systems, vol. 35, pp. 7575–7586, 2022

  225. [235]

    Egovlpv2: Egocentric video- language pre-training with fusion in the backbone,

    S. Pramanick, Y. Song, S. Nag, K. Q. Lin, H. Shah, M. Z. Shou, R. Chellappa, and P . Zhang, “Egovlpv2: Egocentric video- language pre-training with fusion in the backbone,” inProceed- ings of the IEEE/CVF International Conference on Computer Vision. IEEE TRANSACTIONS ON PATTER...

  226. [236]

    Fine-grained spatiotemporal grounding on egocentric videos,

    S. Liang, Y. Zhong, Z.-Y. Hu, Y. Tao, and L. Wang, “Fine-grained spatiotemporal grounding on egocentric videos,” inProceedings of the IEEE/CVF International Conference on Computer Vision. Los Alamitos, CA, USA: IEEE Computer Society, 2025, pp. 9385–9395

  227. [237]

    Dmc3: Dual-modal coun- terfactual contrastive construction for egocentric video question answering,

    J. Zou, C. Chen, B.-K. Bao, and C. Xu, “Dmc3: Dual-modal coun- terfactual contrastive construction for egocentric video question answering,” inProceedings of the 33rd ACM International Conference on Multimedia. New York, NY, USA: Association for Computing Machinery, 2025, pp. ...

  228. [238]

    St-think: How multimodal large language models reason about 4d worlds from ego-centric videos,

    P . Wu, Y. Liu, M. Liu, and J. Shen, “St-think: How multimodal large language models reason about 4d worlds from ego-centric videos,” arXiv preprint arXiv:2503.12542, 2025

  229. [239]

    Vln-r1: Vision-language navigation via reinforcement fine-tuning,

    Z. Qi, Z. Zhang, Y. Yu, J. Wang, and H. Zhao, “Vln-r1: Vision-language navigation via reinforcement fine-tuning,” arXiv preprint arXiv:2506.17221, 2025

  230. [240]

    Ego-r1: Chain-of-tool-thought for ultra-long egocentric video reasoning,

    S. Tian, R. Wang, H. Guo, P . Wu, Y. Dong, X. Wang, J. Yang, H. Zhang, H. Zhu, and Z. Liu, “Ego-r1: Chain-of-tool-thought for ultra-long egocentric video reasoning,” arXiv preprint arXiv:2506.13654, 2025

  231. [241]

    Eyes wide open: Ego proactive video-llm for streaming video,

    Y. Zhang, C. Shi, Y. Wang, and S. Yang, “Eyes wide open: Ego proactive video-llm for streaming video,” arXiv preprint arXiv:2510.14560, 2025

  232. [242]

    Egosocial: Benchmarking proactive intervention ability of omnimodal llms via egocentric social interaction per- ception,

    X. Wang, T. Sharma, A. Kulshrestha, A. Meka, A. Purohit, and D. Manocha, “Egosocial: Benchmarking proactive intervention ability of omnimodal llms via egocentric social interaction per- ception,” 2025

  233. [243]

    Are vision llms road-ready? a comprehensive benchmark for safety-critical driving video understanding,

    T. Zeng, L. Wu, L. Shi, D. Zhou, and F. Guo, “Are vision llms road-ready? a comprehensive benchmark for safety-critical driving video understanding,” 2025

  234. [244]

    Sportu: A comprehensive sports under- standing benchmark for multimodal large language models,

    H. Xia, Z. Yang, J. Zou, R. Tracy, Y. Wang, C. Lu, C. Lai, Y. He, X. Shao, Z. Xieet al., “Sportu: A comprehensive sports under- standing benchmark for multimodal large language models,” arXiv preprint arXiv:2410.08474, 2024

  235. [245]

    Towards universal soccer video understanding,

    J. Rao, H. Wu, H. Jiang, Y. Zhang, Y. Wang, and W. Xie, “Towards universal soccer video understanding,” inProceedings of the Com- puter Vision and Pattern Recognition Conference. Los Alamitos, CA, USA: IEEE Computer Society, 2025, pp. 8384–8394

  236. [246]

    Domain adaptation of vlm for soccer video understanding,

    T. Jiang, H. Wang, M. S. Salekin, P . Atighehchian, and S. Zhang, “Domain adaptation of vlm for soccer video understanding,” in Proceedings of the Computer Vision and Pattern Recognition Confer- ence. Los Alamitos, CA, USA: IEEE Computer Society, 2025, pp. 6111–6121

  237. [247]

    Deepsport: A multimodal large language model for comprehensive sports video reasoning via agentic reinforcement learning,

    J. Zou, H. Xia, Z. Ye, S. Zhang, C. Lai, V . Ordonez, W. Shen, and H. Chen, “Deepsport: A multimodal large language model for comprehensive sports video reasoning via agentic reinforcement learning,” arXiv preprint arXiv:2511.12908, 2025

  238. [248]

    Finequest: Adaptive knowledge-assisted sports video understanding via agent-of- thoughts reasoning,

    H. Chen, H. Huang, X. Yin, and D. Shao, “Finequest: Adaptive knowledge-assisted sports video understanding via agent-of- thoughts reasoning,” inProceedings of the 33rd ACM International Conference on Multimedia. New York, NY, USA: Association for Computing Machinery, 2025, pp....

  239. [249]

    Tennistv: Do multimodal large lan- guage models understand tennis rallies?

    Z. Bao and L. Zhang, “Tennistv: Do multimodal large lan- guage models understand tennis rallies?” arXiv preprint arXiv:2509.15602, 2025

  240. [250]

    Learning consistent temporal ground- ing between related tasks in sports coaching,

    A. Rai and A. Kovashka, “Learning consistent temporal ground- ing between related tasks in sports coaching,” arXiv preprint arXiv:2603.18453, 2026

  241. [251]

    Video-mmmu: Evaluating knowledge acquisition from multi- discipline professional videos,

    K. Hu, P . Wu, F. Pu, W. Xiao, Y. Zhang, X. Yue, B. Li, and Z. Liu, “Video-mmmu: Evaluating knowledge acquisition from multi- discipline professional videos,” arXiv preprint arXiv:2501.13826, 2025

  242. [252]

    Video- mmlu: A massive multi-discipline lecture understanding bench- mark,

    E. Song, W. Chai, W. Xu, J. Xie, Y. Liu, and G. Wang, “Video- mmlu: A massive multi-discipline lecture understanding bench- mark,” arXiv preprint arXiv:2504.14693, 2025

  243. [253]

    Noteit: A system converting instructional videos to interactable notes through multimodal video understanding,

    R. Zhao, Z. Jiang, X. Zhang, C. Chang, H. Chen, W. Deng, L. Jin, X. Qi, X. Qian, and E. C. Ngai, “Noteit: A system converting instructional videos to interactable notes through multimodal video understanding,” inProceedings of the 38th Annual ACM Symposium on User Interface So...

  244. [254]

    Instructionbench: An instructional video understanding benchmark,

    H. Wei, Y. Yuan, X. Lan, W. Ke, and L. Ma, “Instructionbench: An instructional video understanding benchmark,” arXiv preprint arXiv:2504.05040, 2025

  245. [255]

    Docvideoqa: Towards comprehen- sive understanding of document-centric videos through question answering,

    H. Wang, K. Hu, and L. Gao, “Docvideoqa: Towards comprehen- sive understanding of document-centric videos through question answering,” inICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE. Piscat- away, NJ, USA: IEEE, 2025, pp. 1–5

  246. [256]

    Install: Context-aware instructional task assistance with multi-modal large language models,

    P . Nguyen, S. Sengupta, G. Malik, A. Gupta, and B. Min, “Install: Context-aware instructional task assistance with multi-modal large language models,” arXiv preprint arXiv:2501.12231, 2025

  247. [257]

    Tecno: Surgical phase recognition with multi-stage temporal convolutional networks,

    T. Czempiel, M. Paschali, M. Keicher, W. Simson, H. Feussner, S. T. Kim, and N. Navab, “Tecno: Surgical phase recognition with multi-stage temporal convolutional networks,” inInterna- tional conference on medical image computing and computer-assisted intervention, Springer. Ch...

  248. [258]

    Multi-task temporal convolutional networks for joint recognition of surgical phases and steps in gastric bypass procedures,

    S. Ramesh, D. Dall’Alba, C. Gonzalez, T. Yu, P . Mascagni, D. Mut- ter, J. Marescaux, P . Fiorini, and N. Padoy, “Multi-task temporal convolutional networks for joint recognition of surgical phases and steps in gastric bypass procedures,”International journal of computer assis...

  249. [259]

    Rendezvous: Attention mecha- nisms for the recognition of surgical action triplets in endoscopic videos,

    C. I. Nwoye, T. Yu, C. Gonzalez, B. Seeliger, P . Mascagni, D. Mut- ter, J. Marescaux, and N. Padoy, “Rendezvous: Attention mecha- nisms for the recognition of surgical action triplets in endoscopic videos,”Medical Image Analysis, vol. 78, p. 102433, 2022

  250. [260]

    Cholec- triplet2022: Show me a tool and tell me the triplet—an endoscopic vision challenge for surgical action triplet detection,

    C. I. Nwoye, T. Yu, S. Sharma, A. Murali, D. Alapatt, A. Var- dazaryan, K. Yuan, J. Hajek, W. Reiter, A. Yamlahiet al., “Cholec- triplet2022: Show me a tool and tell me the triplet—an endoscopic vision challenge for surgical action triplet detection,”Medical Image Analysis, vo...

  251. [261]

    Dis- secting self-supervised learning methods for surgical computer vision,

    S. Ramesh, V . Srivastav, D. Alapatt, T. Yu, A. Murali, L. Sestini, C. I. Nwoye, I. Hamoud, S. Sharma, A. Fleurentinet al., “Dis- secting self-supervised learning methods for surgical computer vision,”Medical Image Analysis, vol. 88, p. 102844, 2023

  252. [262]

    Learning multi-modal representations by watching hundreds of surgical video lec- tures,

    K. Yuan, V . Srivastav, T. Yu, J. L. Lavanchy, J. Marescaux, P . Mascagni, N. Navab, and N. Padoy, “Learning multi-modal representations by watching hundreds of surgical video lec- tures,”Medical Image Analysis, vol. 105, p. 103644, 2025

  253. [263]

    Large-scale self-supervised video foundation model for intelligent surgery,

    S. Yang, F. Zhou, L. Mayer, F. Huang, Y. Chen, Y. Wang, S. He, Y. Nie, X. Wang, Y. Jinet al., “Large-scale self-supervised video foundation model for intelligent surgery,”npj Digital Medicine, 2026

  254. [264]

    Mm-or: A large multimodal operating room dataset for semantic under- standing of high-intensity surgical environments,

    E. ¨Ozsoy, C. Pellegrini, T. Czempiel, F. Tristram, K. Yuan, D. Bani- Harouni, U. Eck, B. Busam, M. Keicher, and N. Navab, “Mm-or: A large multimodal operating room dataset for semantic under- standing of high-intensity surgical environments,” inProceedings of the Computer Vis...

  255. [265]

    Llava-surg: towards multimodal surgical assistant via structured surgical video learning,

    J. Li, G. Skinner, G. Yang, B. R. Quaranto, S. D. Schwaitzberg, P . C. Kim, and J. Xiong, “Llava-surg: towards multimodal surgical assistant via structured surgical video learning,” arXiv preprint arXiv:2408.07981, 2024

  256. [266]

    Surgical-llava: Toward surgical scenario understanding via large language and vision models,

    J. Jin and C. W. Jeong, “Surgical-llava: Toward surgical scenario understanding via large language and vision models,” arXiv preprint arXiv:2410.09750, 2024

  257. [267]

    Endochat: Grounded multimodal large language model for endoscopic surgery,

    G. Wang, L. Bai, J. Wang, K. Yuan, Z. Li, T. Jiang, X. He, J. Wu, Z. Chen, Z. Leiet al., “Endochat: Grounded multimodal large language model for endoscopic surgery,”Medical Image Analysis, p. 103789, 2025

  258. [268]

    Surgvlm: A large vision-language model and systematic evaluation benchmark for surgical intelli- gence,

    Z. Zeng, Z. Zhuo, X. Jia, E. Zhang, J. Wu, J. Zhang, Y. Wang, C. H. Low, J. Jiang, Z. Zhenget al., “Surgvlm: A large vision-language model and systematic evaluation benchmark for surgical intelli- gence,” arXiv preprint arXiv:2506.02555, 2025

  259. [269]

    Surgvidlm: Towards multi-grained surgical video understanding with large language model,

    G. Wang, J. Wang, W. Mo, L. Bai, K. Yuan, M. Hu, J. Wu, J. He, Y. Huang, N. Padoyet al., “Surgvidlm: Towards multi-grained surgical video understanding with large language model,” arXiv preprint arXiv:2506.17873, 2025

  260. [270]

    Surgvivqa: Temporally-grounded video question answering for surgical scene understanding,

    M. O. Drago, L. Carlini, P . C. Balyemez, D. Pierantozzi, C. Lena, C. Hassan, D. Stoyanov, E. De Momi, S. Bano, and M. I. Hoque, “Surgvivqa: Temporally-grounded video question answering for surgical scene understanding,” arXiv preprint arXiv:2511.03325, 2025

  261. [271]

    Vision–language foundation model for echocardiogram inter- pretation,

    M. Christensen, M. Vukadinovic, N. Yuan, and D. Ouyang, “Vision–language foundation model for echocardiogram inter- pretation,”Nature Medicine, vol. 30, no. 5, pp. 1481–1488, 2024

  262. [272]

    Mmsummary: multimodal summary generation for fetal ultrasound video,

    X. Guo, Q. Men, and J. A. Noble, “Mmsummary: multimodal summary generation for fetal ultrasound video,” inInternational Conference on Medical Image Computing and Computer-Assisted In- IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 31 tervention, Springer. Cham...

  263. [273]

    Movieqa: Understanding stories in movies through question-answering,

    M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler, “Movieqa: Understanding stories in movies through question-answering,” inProceedings of the IEEE conference on computer vision and pattern recognition. Los Alamitos, CA, USA: IEEE Computer Society, 20...

  264. [274]

    Movienet: A holistic dataset for movie understanding,

    Q. Huang, Y. Xiong, A. Rao, J. Wang, and D. Lin, “Movienet: A holistic dataset for movie understanding,” inEuropean conference on computer vision, Springer. Cham, Switzerland: Springer, 2020, pp. 709–727

  265. [275]

    Mad: A scalable dataset for language ground- ing in videos from movie audio descriptions,

    M. Soldan, A. Pardo, J. L. Alc ´azar, F. Caba, C. Zhao, S. Giancola, and B. Ghanem, “Mad: A scalable dataset for language ground- ing in videos from movie audio descriptions,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Los Alamitos, CA...

  266. [276]

    Movqa: A benchmark of versatile question-answering for long-form movie understanding,

    H. Zhang, Y. Liu, L. Dong, Y. Huang, Z.-H. Ling, Y. Wang, L. Wang, and Y. Qiao, “Movqa: A benchmark of versatile question-answering for long-form movie understanding,” arXiv e-prints, pp. arXiv–2312, 2023

  267. [277]

    Short film dataset (sfd): A benchmark for story-level video understanding,

    R. Ghermi, X. Wang, V . Kalogeiton, and I. Laptev, “Short film dataset (sfd): A benchmark for story-level video understanding,” arXiv preprint arXiv:2406.10221, vol. 2, no. 3, p. 6, 2024

  268. [278]

    Scvbench: A benchmark with multi-turn dialogues for story-centric video understanding,

    S. You, B. Yuan, and B.-K. Bao, “Scvbench: A benchmark with multi-turn dialogues for story-centric video understanding,” in Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. California, USA: International Joint Confer- ences on Artific...

  269. [279]

    Vrbench: A benchmark for multi-step reasoning in long narrative videos,

    J. Yu, Y. Wu, M. Chu, Z. Ren, Z. Huang, P . Chu, R. Zhang, Y. He, Q. Li, S. Liet al., “Vrbench: A benchmark for multi-step reasoning in long narrative videos,” arXiv preprint arXiv:2506.10857, 2025

  270. [280]

    Seriesbench: A benchmark for narrative-driven drama series understanding,

    C. Zhang, Y. Lei, Z. Liu, H. Leng, S. Liu, T. Gao, Q. Liu, and Y. Wang, “Seriesbench: A benchmark for narrative-driven drama series understanding,” inProceedings of the Computer Vision and Pattern Recognition Conference. Los Alamitos, CA, USA: IEEE Computer Society, 2025, pp. ...

  271. [281]

    Cin\’{e} aste: A fine-grained contextual movie question answering bench- mark,

    N. A. Shah, A. Ziai, C. Ekanadham, and V . M. Patel, “Cin\’{e} aste: A fine-grained contextual movie question answering bench- mark,” arXiv preprint arXiv:2509.14227, 2025

  272. [282]

    Moviecore: Cognitive reasoning in movies,

    G. J. Faure, M.-H. Chen, J.-F. Yeh, Y. Cheng, H.-T. Su, Y.-H. Tang, S.-H. Lai, and W. H. Hsu, “Moviecore: Cognitive reasoning in movies,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for Computation...

  273. [283]

    Arc-chapter: Structuring hour-long videos into navigable chapters and hierar- chical summaries,

    J. Pu, T. Wang, Y. Ge, Y. Ge, C. Li, and Y. Shan, “Arc-chapter: Structuring hour-long videos into navigable chapters and hierar- chical summaries,” arXiv preprint arXiv:2511.14349, 2025

  274. [284]

    Mvbench: A comprehensive multi-modal video understanding benchmark,

    K. Li, Y. Wang, Y. He, Y. Li, Y. Wanget al., “Mvbench: A comprehensive multi-modal video understanding benchmark,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Los Alamitos, CA, USA: IEEE Computer Society, 2024

  275. [285]

    Videocot: A video chain-of-thought dataset with active annotation tool,

    Y. Wang, Y. Zeng, J. Zheng, X. Xing, J. Xu, and X. Xu, “Videocot: A video chain-of-thought dataset with active annotation tool,” arXiv preprint arXiv:2407.05355, 2024

  276. [286]

    Videoespresso: A large-scale chain- of-thought dataset for fine-grained video reasoning via core frame selection,

    S. Han, W. Huang, H. Shi, L. Zhuo, X. Su, S. Zhang, X. Zhou, X. Qi, Y. Liao, and S. Liu, “Videoespresso: A large-scale chain- of-thought dataset for fine-grained video reasoning via core frame selection,” inProceedings of the Computer Vision and Pattern Recognition Conference....

  277. [287]

    Scaling rl to long videos,

    Y. Chen, W. Huang, B. Shi, Q. Hu, H. Ye, L. Zhu, Z. Liu, P . Molchanov, J. Kautz, X. Qiet al., “Scaling rl to long videos,” arXiv preprint arXiv:2507.07966, 2025

  278. [288]

    MiraData: A large-scale video dataset with long durations and structured captions,

    X. Ju, Y. Gao, Z. Zhang, Z. Yuan, X. Wang, A. Zeng, Y. Xiong, Q. Xu, and Y. Shan, “MiraData: A large-scale video dataset with long durations and structured captions,” inAdvances in Neural Information Processing Systems, vol. 37. Red Hook, NY, USA: Curran Associates, Inc., 2024...

  279. [289]

    FineVideo,

    M. Farr ´e, A. Marafioti, L. Tunstall, L. Von Werra, and T. Wolf, “FineVideo,” https://huggingface.co/datasets/ HuggingFaceFV/finevideo, 2024

  280. [290]

    UltraVideo: High-quality UHD video dataset with comprehensive captions,

    Z. Xue, J. Zhang, T. Hu, H. He, Y. Chen, Y. Cai, Y. Wang, C. Wang, Y. Liu, X. Li, and D. Tao, “UltraVideo: High-quality UHD video dataset with comprehensive captions,” inThirty-ninth Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Red Hook, N...

  281. [291]

    Timechat-captioner: Scripting multi-scene videos with time-aware and structural audio-visual captions,

    L. Yao, Y. Wei, Y. Zhang, L. Li, X. Chen, F. Song, Z. Wang, K. Ouyang, Y. Liu, L. Kong, Q. Liu, P . Wan, K. Gai, Y. Zhang, and X. Sun, “Timechat-captioner: Scripting multi-scene videos with time-aware and structural audio-visual captions,” 2026

  282. [292]

    E.t. bench: Towards open-ended event-level video-language under- standing,

    Y. Liu, Z. Ma, Z. Qi, Y. Wu, Y. Shan, and C. W. Chen, “E.t. bench: Towards open-ended event-level video-language under- standing,” inAdvances in Neural Information Processing Systems (NeurIPS). Red Hook, NY, USA: Curran Associates, Inc., 2024

  283. [293]

    Vid-morp: Video moment retrieval pretraining from unlabeled videos in the wild,

    P . Bao, C. Xia, Z. Xu, W. Yang, S.-K. Ng, M. Kankanhalli, A. C. Kot, and B. Wen, “Vid-morp: Video moment retrieval pretraining from unlabeled videos in the wild,” arXiv preprint arXiv:2412.00811, 2024

  284. [294]

    Videoitg: Multimodal video understanding with instructed tem- poral grounding,

    S. Wang, G. Zhao, H. Yin, P . Molchanov, J. Kautz, Y. Lu, and Z. Yu, “Videoitg: Multimodal video understanding with instructed tem- poral grounding,” arXiv preprint arXiv:2507.13353, 2025

  285. [295]

    Activitynet-qa: A dataset for understanding complex web videos via question answering,

    Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao, “Activitynet-qa: A dataset for understanding complex web videos via question answering,” inProceedings of the AAAI confer- ence on artificial intelligence, vol. 33. Palo Alto, CA, USA: AAAI Press, 2019, pp. 9127–9134

  286. [296]

    Tvqa: Localized, composi- tional video question answering,

    J. Lei, L. Yu, M. Bansal, and T. Berg, “Tvqa: Localized, composi- tional video question answering,” inProceedings of the 2018 confer- ence on empirical methods in natural language processing. Strouds- burg, PA, USA: Association for Computational Linguistics, 2018, pp. 1369–1379

  287. [297]

    Next-qa: Next phase of question-answering to explaining temporal actions,

    J. Xiao, X. Shang, A. Yao, and T.-S. Chua, “Next-qa: Next phase of question-answering to explaining temporal actions,” inPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition. Los Alamitos, CA, USA: IEEE Computer Society, 2021, pp. 9777–9786

  288. [298]

    Clevrer: Collision events for video representation and reasoning,

    K. Yi, C. Gan, Y. Li, P . Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum, “Clevrer: Collision events for video representation and reasoning,” arXiv preprint arXiv:1910.01442, 2019

  289. [299]

    Just ask: Learning to answer questions from millions of narrated videos,

    A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Just ask: Learning to answer questions from millions of narrated videos,” inProceedings of the IEEE/CVF international conference on computer vision. Los Alamitos, CA, USA: IEEE Computer Society, 2021, pp. 1686–1697

  290. [300]

    Video-chatgpt: Towards detailed video understanding via large vision and language models,

    M. Maaz, H. Rasheed, S. Khan, and F. Khan, “Video-chatgpt: Towards detailed video understanding via large vision and language models,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg, PA, USA: Assoc...

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.