Pith. sign in

REVIEW 5 minor 176 references

Video translation works better as three coordinated MLLM roles than as a cascade of ASR, translation, speech, and lip-sync.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 22:03 UTC pith:3EQ76VC7

load-bearing objection Solid role-oriented survey that cleanly maps MLLM video translation without overclaiming; useful map, not a new result.

arxiv 2604.11283 v2 pith:3EQ76VC7 submitted 2026-04-13 cs.CV

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

classification cs.CV
keywords multimodal large language modelsvideo translationvideo understandingtext-to-speechlip synchronizationtalking-head generationspeech-to-speech translationrole-oriented taxonomy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This survey argues that multimodal large language models are turning video translation from a brittle pipeline of separate modules into one integrated multimodal reasoning-and-generation problem. High-quality results need more than correct words: they need timing, speaker identity, and emotional tone held steady across sight, sound, and language. The authors reorganize the scattered literature into three functional roles—Semantic Reasoner for video understanding and translation planning, Expressive Performer for controllable target speech, and Visual Synthesizer for lip sync and coherent speaker rendering—and map methods, datasets, and metrics onto those roles. They show that today’s evaluations are mostly role-level proxies and do not measure complete translated videos. The practical upshot is a roadmap for systems that keep meaning, voice, face, and timing aligned for natural cross-lingual video.

Core claim

MLLM-enabled video translation should be framed as an integrated multimodal reasoning and generation task organized by three functional roles—Semantic Reasoner, Expressive Performer, and Visual Synthesizer—rather than as an isolated cascade of ASR, machine translation, TTS, and lip synchronization. That taxonomy links semantic grounding, expressive speech, and visual rendering and exposes why current proxy benchmarks fall short of end-to-end translated-video quality.

What carries the argument

The three-role taxonomy (Semantic Reasoner for multimodal grounding, temporal/event reasoning, cross-lingual planning, and efficient adaptation; Expressive Performer for speech-token LMs, prompt/emotion control, and diffusion/flow rendering; Visual Synthesizer for lip sync, talking-head animation, and enabling video backbones). It re-cuts the literature by function in the translation pipeline instead of by isolated subtask names.

Load-bearing premise

That sorting studies by these three roles and the survey’s inclusion rules gives a complete enough map that role-level proxy scores can stand in for what end-to-end video translation still lacks.

What would settle it

Build or publish a public end-to-end multilingual video-translation benchmark with aligned source video, target speech, visual outputs, and multi-axis human ratings; if systems strong only on the paper’s role proxies still rank poorly there, or if strong end-to-end systems cannot be usefully described by the three roles, the taxonomy’s organizing claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. This survey argues that MLLM progress is reshaping video translation from a cascaded ASR–MT–TTS–lip-sync pipeline into an integrated multimodal reasoning-and-generation problem. It organizes MLLM-enabled and MLLM-relevant work into three functional roles—Semantic Reasoner (multimodal grounding, temporal/event reasoning, cross-lingual planning, efficient adaptation), Expressive Performer (speech-token LMs, prompt/emotion control, diffusion/flow rendering), and Visual Synthesizer (lip sync, talking-head animation, enabling visual backbones)—and reviews role-level datasets, benchmarks, and metrics while arguing that current proxy evaluations fall short of end-to-end translated-video quality. Open challenges include long-form understanding, temporal/cross-modal alignment, multilingual robustness, real-time deployment, and responsible use.

Significance. If accepted as a framing contribution, the paper offers a useful map for a fragmented area that spans video-MLLMs, speech translation/TTS, and talking-head/video generation. Strengths include an explicit literature protocol (§II-A), a clear comparison with prior surveys (Table I), a role taxonomy that separates core MLLM reasoning from supporting generation modules (Fig. 1 and caption), and a consistent gap analysis that does not overclaim end-to-end systems from proxy VideoQA/TTS/lip-sync scores (§V–VI). For a survey venue, this is a coherent organizational contribution rather than a new empirical result.

minor comments (5)
  1. In Table III, PromptTTS is cited as [161] while the taxonomy and earlier text use [90]; unify the PromptTTS citation and ensure Table III entries match the bibliography consistently.
  2. Fig. 2–4 captions and body text are clear, but several figure placeholders in the source appear as garbled character blocks; ensure final production figures are clean and that architecture diagrams remain readable at print scale.
  3. §V-A Table II reports published zero-shot VideoQA numbers as proxy evidence; a short caveat in the table caption that scores are not re-evaluated under a single protocol would further reduce cross-paper comparability risk.
  4. Minor typography/consistency: mixed spellings such as CosyVoice/CosyV oice, AV2AV/A V2A V, and occasional spacing artifacts (e.g., F .) should be normalized in copyediting.
  5. §II-A states coverage up to May 2026; confirm that the final arXiv/journal version’s cutoff date matches the bibliography and that any post-cutoff highly relevant systems are either included or explicitly scoped out.

Circularity Check

0 steps flagged

No significant circularity: role taxonomy is organizational mapping of external work, not a fitted or self-referential derivation.

full rationale

This is a survey paper whose central contribution is a role-oriented taxonomy (Semantic Reasoner, Expressive Performer, Visual Synthesizer) for organizing existing MLLM-enabled and MLLM-relevant video-translation literature. The taxonomy is definitional organization of external methods, datasets, and metrics; it does not claim a first-principles derivation, uniqueness theorem, or quantitative prediction that could reduce to its own inputs by construction. Inclusion criteria (§II-A) and the explicit separation of core MLLM reasoning from supporting TTS/visual backbones (Fig. 1 caption; §I) are stated as scoping choices, not as empirical results forced by self-citation. Table I honestly marks partial coverage of prior surveys; §V–VI repeatedly treat VideoQA, MOS/WER, and SyncNet-style scores as proxies that do not measure end-to-end translated-video quality. Evaluation tables report published scores from other papers rather than redefining success via fitted parameters. There is no self-definitional loop, no fitted input called a prediction, no load-bearing uniqueness imported from the authors’ prior work, and no renaming of a known empirical law as a new derivation. Self-citations, if any, are ordinary survey practice and not load-bearing for a forced result. Score 0 is therefore the correct, proportionate finding.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 3 invented entities

As a survey, the load-bearing content is definitional framing plus domain premises about what video translation requires and how MLLMs relate to cascaded pipelines. No fitted free parameters. Invented entities are the three role labels used as the taxonomy. Axioms are standard domain assumptions drawn from the literature rather than ad-hoc physics-style postulates.

axioms (4)
  • domain assumption High-quality video translation requires joint semantic fidelity, temporal alignment, speaker consistency, and emotional expressiveness across visual, acoustic, and linguistic streams.
    Stated in Abstract and §I as the problem definition that motivates the three-role taxonomy.
  • domain assumption Cascaded ASR–MT–TTS–lip-sync pipelines suffer error propagation and weak cross-modal coordination.
    §I and §III-A; used to justify reframing as unified multimodal reasoning/generation.
  • ad hoc to paper MLLMs and related generation modules can be productively analyzed by functional role (reason, speak, render) rather than only by classical task boundaries.
    Core organizing premise of §IV; the taxonomy is the paper’s contribution, not an external theorem.
  • domain assumption Role-level proxy benchmarks (VideoQA, TTS metrics, lip-sync scores) are informative diagnostics but insufficient for end-to-end translated-video quality.
    §V–VI; underpins the evaluation-gap claims.
invented entities (3)
  • Semantic Reasoner (role) no independent evidence
    purpose: Bucket for multimodal grounding, temporal/event reasoning, cross-lingual planning, and efficient adaptation in video translation.
    Taxonomy label introduced in Abstract/§IV-A; organizational, not a physical entity.
  • Expressive Performer (role) no independent evidence
    purpose: Bucket for speech-token LMs, prompt/emotion control, and diffusion/flow speech rendering under video constraints.
    Taxonomy label in §IV-B.
  • Visual Synthesizer (role) no independent evidence
    purpose: Bucket for lip sync, talking-head animation, and enabling video backbones for identity-consistent rendering.
    Taxonomy label in §IV-C.

pith-pipeline@v1.1.0-grok45 · 32177 in / 2844 out tokens · 32120 ms · 2026-07-12T22:03:58.376557+00:00 · methodology

0 comments
read the original abstract

Recent progress in multimodal large language models (MLLMs) is reshaping video translation from a cascaded pipeline of automatic speech recognition, machine translation, text-to-speech, and lip synchronization into a unified multimodal reasoning and generation problem. High-quality video translation requires not only semantic fidelity, but also temporal alignment, speaker consistency, and emotional expressiveness across visual, acoustic, and linguistic streams. This survey provides a focused review of MLLM-enabled video translation through a role-oriented taxonomy. We organize MLLM-enabled and MLLM-relevant studies into three functional roles: Semantic Reasoner, which grounds translation in video understanding, temporal reasoning, and multimodal fusion; Expressive Performer, which supports controllable and context-aware speech generation; and Visual Synthesizer, which enables lip synchronization and visually coherent speaker rendering. We further summarize representative datasets, benchmarks, and metrics for each role, and discuss how current evaluation protocols fall short of end-to-end video translation requirements. Finally, we identify open challenges in long-form video understanding, temporal modeling, multimodal alignment, multilingual robustness, and responsible deployment, outlining future directions for natural and trustworthy cross-lingual video communication.

Figures

Figures reproduced from arXiv: 2604.11283 by Bingzheng QU, Kehai Chen, Min Zhang, Xuefeng Bai.

Figure 1
Figure 1. Figure 1: Taxonomy of MLLMs-based video translation, encompassing three primary dimensions: The Semantic Reasoner, Expressive Performer, and Visual [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Typical architecture of an MLLMs-based video understanding model. The text, audio, and video encoders can be either learnable or frozen. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

176 extracted references · 62 linked inside Pith

  1. [1]

    Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,

    G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. rahman Mohamed, N. Jaitly, A. Senior, V . Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”IEEE Signal Process. Mag., vol. 29, no. 6, pp. 82–97, 2012

  2. [2]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS, 2017, pp. 5998–6008

  3. [3]

    Tacotron: Towards end-to-end speech synthesis,

    Y . Wang, R. J. Skerry-Ryan, D. Stanton, Y . Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y . Xiao, Z. Chen, S. Bengio, Q. Le, Y . Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” inInterspeech, 2017, pp. 4006–4010

  4. [4]

    A lip sync expert is all you need for speech to lip generation in the wild,

    K. R. Prajwal, R. Mukhopadhyay, V . P. Namboodiri, and C. V . Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” inACM MM, 2020, pp. 484–492

  5. [5]

    Direct speech-to-speech translation with a sequence-to- sequence model,

    Y . Jia, R. J. Weiss, F. Biadsy, W. Macherey, M. Johnson, Z. Chen, and Y . Wu, “Direct speech-to-speech translation with a sequence-to- sequence model,” inInterspeech, 2019, pp. 1123–1127

  6. [6]

    Learning transferable visual models from natural language supervi- sion,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervi- sion,” inICML, 2021, pp. 8748–8763

  7. [7]

    Flamingo: A visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Bi ´nkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan, “Flamingo: A visual languag...

  8. [8]

    Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

    J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,” inICML, 2023, pp. 19 730–19 742

  9. [9]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” in NeurIPS, vol. 36, 2023, pp. 34 892–34 916

  10. [10]

    mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,

    Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang, “mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,” inCVPR, 2024, pp. 13 040–13 051

  11. [11]

    Video-chatgpt: Towards detailed video understanding via large vision and language models,

    M. Maaz, H. Rasheed, S. Khan, and F. Khan, “Video-chatgpt: Towards detailed video understanding via large vision and language models,” in ACL, 2024, pp. 12 585–12 602

  12. [12]

    Video-llama: An instruction-tuned audio-visual language model for video understanding,

    H. Zhang, X. Li, and L. Bing, “Video-llama: An instruction-tuned audio-visual language model for video understanding,” inEMNLP, 2023, pp. 543–553

  13. [13]

    Llama-vid: An image is worth 2 tokens in large language models,

    Y . Li, C. Wang, and J. Jia, “Llama-vid: An image is worth 2 tokens in large language models,” inECCV, 2024, pp. 323–340

  14. [14]

    Internvideo2: Scaling foundation models for multimodal video understanding,

    Y . Wang, K. Li, X. Li, J. Yu, Y . He, G. Chen, B. Pei, R. Zheng, Z. Wang, Y . Shi, T. Jiang, S. Li, J. Xu, H. Zhang, Y . Huang, Y . Qiao, Y . Wang, and L. Wang, “Internvideo2: Scaling foundation models for multimodal video understanding,” inECCV, 2024, pp. 396–416

  15. [15]

    Cosyvoice 2: Scalable streaming speech synthesis with large language models,

    Z. Du, Y . Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y . Yang, C. Gao, H. Wang, F. Yu, H. Liu, Z. Sheng, Y . Gu, C. Deng, W. Wang, S. Zhang, Z. Yan, and J. Zhou, “Cosyvoice 2: Scalable streaming speech synthesis with large language models,”arXiv:2412.10117, 2024

  16. [16]

    F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,

    Y . Chen, Z. Niu, Z. Ma, K. Deng, C. Wang, J. Zhao, K. Yu, and X. Chen, “F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,” inACL, 2025, pp. 6255–6271

  17. [17]

    A survey on video diffusion models,

    Z. Xing, Q. Feng, H. Chen, Q. Dai, H. Hu, H. Xu, Z. Wu, and Y .- G. Jiang, “A survey on video diffusion models,”ACM Comput. Surv., vol. 57, no. 2, pp. 1–42, 2024

  18. [18]

    Omnisync: Towards universal lip synchronization via diffusion transformers,

    Z. Peng, J. Liu, H. Zhang, X. Liu, S. Tang, P. Wan, D. Zhang, H. Liu, and J. He, “Omnisync: Towards universal lip synchronization via diffusion transformers,”arXiv:2505.21448, 2025

  19. [19]

    A survey on multi-modal machine translation: Tasks, methods and challenges,

    H. Shen, L. Shao, W. Li, Z. Lan, Z. Liu, and J. Su, “A survey on multi-modal machine translation: Tasks, methods and challenges,” arXiv:2405.12669, 2024

  20. [20]

    Video-guided machine translation: A survey of models, datasets, and challenges,

    P. Das, V . Singh, P. Bhattacharyya, and G. Haffari, “Video-guided machine translation: A survey of models, datasets, and challenges,” inIJCNLP-AACL, 2025, pp. 3346–3356

  21. [21]

    A survey on multimodal large language models,

    S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen, “A survey on multimodal large language models,”Natl. Sci. Rev., vol. 11, no. 12, p. nwae403, 2024

  22. [22]

    Video understanding with large language models: A survey,

    Y . Tang, J. Bi, S. Xu, L. Song, S. Liang, T. Wang, D. Zhang, J. An, J. Lin, R. Zhu, A. V osoughi, C. Huang, Z. Zhang, P. Liu, M. Feng, F. Zheng, J. Zhang, P. Luo, J. Luo, and C. Xu, “Video understanding with large language models: A survey,”TCSVT, 2025

  23. [23]

    A survey on video temporal grounding with multimodal large language model,

    J. Wu, W. Liu, Y . Liu, M. Liu, L. Nie, Z. Lin, and C. W. Chen, “A survey on video temporal grounding with multimodal large language model,”IEEE Trans. Pattern Anal. Mach. Intell., 2025

  24. [24]

    Direct speech-to-speech neural machine translation: A survey,

    M. Gupta, M. Dutta, and C. K. Maurya, “Direct speech-to-speech neural machine translation: A survey,”arXiv:2411.14453, 2024

  25. [25]

    Towards controllable speech synthesis in the era of large language models: A systematic survey,

    T. Xie, Y . Rong, P. Zhang, W. Wang, and L. Liu, “Towards controllable speech synthesis in the era of large language models: A systematic survey,” inEMNLP, 2025, pp. 764–791

  26. [26]

    Advancing talking head generation: A comprehensive survey of multi-modal methodologies, datasets, evaluation metrics, and loss functions,

    V . K. Rakesh, S. Mazumdar, R. P. Maity, S. Pal, A. Das, and T. Samanta, “Advancing talking head generation: A comprehensive survey of multi-modal methodologies, datasets, evaluation metrics, and loss functions,”arXiv:2507.02900, 2025

  27. [27]

    Rabiner and B.-H

    L. Rabiner and B.-H. Juang,Fundamentals of Speech Recognition. Prentice-Hall, 1993

  28. [28]

    Connection- ist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,

    A. Graves, S. Fern ´andez, F. Gomez, and J. Schmidhuber, “Connection- ist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” inICML, 2006, pp. 369–376

  29. [29]

    Deep speech: Scaling up end-to-end speech recognition,

    A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, and A. Y . Ng, “Deep speech: Scaling up end-to-end speech recognition,”arXiv:1412.5567, 2014

  30. [30]

    Listen, attend and spell,

    W. Chan, N. Jaitly, Q. V . Le, and O. Vinyals, “Listen, attend and spell,” inICASSP, 2016, pp. 4960–4964

  31. [31]

    The mathematics of statistical machine translation: Parameter estimation,

    P. F. Brown, S. A. D. Pietra, V . J. D. Pietra, and R. L. Mercer, “The mathematics of statistical machine translation: Parameter estimation,” Comput. Linguist., vol. 19, no. 2, pp. 263–311, 1993

  32. [32]

    Statistical phrase-based transla- tion,

    P. Koehn, F. J. Och, and D. Marcu, “Statistical phrase-based transla- tion,” inNAACL, 2003, pp. 127–133

  33. [33]

    Sequence to sequence learning with neural networks,

    I. Sutskever, O. Vinyals, and Q. V . Le, “Sequence to sequence learning with neural networks,” inNeurIPS, 2014, pp. 3104–3112

  34. [34]

    Speech parameter generation algorithms for hmm-based speech syn- thesis,

    K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for hmm-based speech syn- thesis,” inICASSP, vol. 3, 2000, pp. 1315–1318

  35. [35]

    Statistical parametric speech synthesis,

    H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,”Speech Commun., vol. 51, no. 11, pp. 1039–1064, 2009

  36. [36]

    Wavenet: A generative model for raw audio,

    A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,”arXiv:1609.03499, 2016

  37. [37]

    Photo-realistic talking-heads from image samples,

    E. Cosatto and H. P. Graf, “Photo-realistic talking-heads from image samples,”IEEE Trans. Multimedia, vol. 2, no. 3, pp. 152–163, 2002

  38. [38]

    Lip reading sentences in the wild,

    J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” inCVPR, 2017, pp. 6447–6456

  39. [39]

    Few-shot adversarial learning of realistic neural talking head models,

    E. Zakharov, A. Shysheya, E. Burkov, and V . Lempitsky, “Few-shot adversarial learning of realistic neural talking head models,” inICCV, 2019, pp. 9459–9468

  40. [40]

    Onellm: One framework to align all modalities with language,

    J. Han, K. Gong, Y . Zhang, J. Wang, K. Zhang, D. Lin, Y . Qiao, P. Gao, and X. Yue, “Onellm: One framework to align all modalities with language,” inCVPR, 2024, pp. 26 584–26 595

  41. [41]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” inICLR, 2021

  42. [42]

    Scaling up visual and vision-language representation learning with noisy text supervision,

    C. Jia, Y . Yang, Y . Xia, Y .-T. Chen, Z. Parekh, H. Pham, Q. Le, Y .-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” inICML, 2021, pp. 4904–4916

  43. [43]

    Videobert: A joint model for video and language representation learning,

    C. Sun, A. Myers, C. V ondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” inICCV, 2019, pp. 7464–7473

  44. [44]

    Actbert: Learning global-local video-text repre- sentations,

    L. Zhu and Y . Yang, “Actbert: Learning global-local video-text repre- sentations,” inCVPR, 2020, pp. 8746–8755

  45. [45]

    Merlot: Multimodal neural script knowledge models,

    R. Zellers, X. Lu, J. Hessel, Y . Yu, J. S. Park, J. Cao, A. Farhadi, and Y . Choi, “Merlot: Multimodal neural script knowledge models,” in NeurIPS, vol. 34, 2021, pp. 23 634–23 651

  46. [46]

    Frozen in time: A joint video and image encoder for end-to-end retrieval,

    M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” inICCV, 2021, pp. 1728–1738

  47. [47]

    Violet: End-to-end video-language transformers with masked visual- token modeling,

    T.-J. Fu, L. Li, Z. Gan, K. Lin, W. Y . Wang, L. Wang, and Z. Liu, “Violet: End-to-end video-language transformers with masked visual- token modeling,” inNeurIPS, 2021

  48. [48]

    Omnivl: One foundation model for image- language and video-language tasks,

    J. Wang, D. Chen, Z. Wu, C. Luo, L. Zhou, Y . Zhao, Y . Xie, C. Liu, Y .-G. Jiang, and L. Yuan, “Omnivl: One foundation model for image- language and video-language tasks,” inNeurIPS, vol. 35, 2022, pp. 5696–5710

  49. [49]

    Internvideo: General video foundation models via generative and discriminative learning,

    Y . Wang, K. Li, Y . Li, Y . He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y . Liu, Z. Wang, Y . Shi, Z. Zhang, Z. Li, Y . Wang, and L. Wang, “Internvideo: General video foundation models via generative and discriminative learning,”arXiv:2212.03191, 2022

  50. [50]

    Unival: Unified model for image, video, audio and language tasks,

    M. Shukor, C. Dancette, A. Rame, and M. Cord, “Unival: Unified model for image, video, audio and language tasks,”TMLR, 2023

  51. [51]

    Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,

    J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi, “Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,” inCVPR, 2024, pp. 26 439–26 455

  52. [52]

    Pali-x: On scaling up a multilingual vision and language model,

    X. Chen, J. Djolonga, P. Padlewski, B. Mustafa, S. Changpinyo, J. Wu, C. R. Ruiz, S. Goodman, X. Wang, and Y . Tay, “Pali-x: On scaling up a multilingual vision and language model,”arXiv:2305.18565, 2023

  53. [53]

    Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond,

    J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou, “Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond,”arXiv:2308.12966, 2023

  54. [54]

    Videollama 3: Frontier multimodal foundation models for image and video understanding,

    B. Zhang, K. Li, Z. Cheng, Z. Hu, Y . Yuan, G. Chen, S. Leng, Y . Jiang, H. Zhang, X. Li, P. Jin, W. Zhang, F. Wang, L. Bing, and D. Zhao, “Videollama 3: Frontier multimodal foundation models for image and video understanding,”arXiv:2501.13106, 2025

  55. [55]

    Videochat: Chat-centric video understanding,

    K. Li, Y . He, Y . Wang, Y . Li, W. Wang, P. Luo, Y . Wang, L. Wang, and Y . Qiao, “Videochat: Chat-centric video understanding,” arXiv:2305.06355, 2023

  56. [56]

    Valley: Video assistant with large language model enhanced ability,

    R. Luo, Z. Zhao, M. Yang, J. Dong, D. Li, P. Lu, T. Wang, L. Hu, M. Qiu, and Z. Wei, “Valley: Video assistant with large language model enhanced ability,”arXiv:2306.07207, 2023

  57. [57]

    Mvbench: A comprehensive multi- modal video understanding benchmark,

    K. Li, Y . Wang, Y . He, Y . Li, Y . Wang, Y . Liu, Z. Wang, J. Xu, G. Chen, P. Luo, L. Wang, and Y . Qiao, “Mvbench: A comprehensive multi- modal video understanding benchmark,” inCVPR, 2024, pp. 22 195– 22 206

  58. [58]

    Groundinggpt: Language enhanced multi-modal grounding model,

    Z. Li, Q. Xu, D. Zhang, H. Song, Y . Cai, Q. Qi, R. Zhou, J. Pan, Z. Li, V . T. Vu, Z. Huang, and T. Wang, “Groundinggpt: Language enhanced multi-modal grounding model,” inACL, 2024, pp. 6657–6678. 13

  59. [59]

    Moviechat: From dense token to sparse memory for long video understanding,

    E. Song, W. Chai, G. Wang, Y . Zhang, H. Zhou, F. Wu, H. Chi, X. Guo, T. Ye, Y . Lu, J.-N. Hwang, and G. Wang, “Moviechat: From dense token to sparse memory for long video understanding,” inCVPR, 2024, pp. 18 221–18 232

  60. [60]

    Moviellm: Enhancing long video understanding with ai-generated movies,

    Z. Song, C. Wang, J. Sheng, C. Zhang, G. Yu, J. Fan, and T. Chen, “Moviellm: Enhancing long video understanding with ai-generated movies,”arXiv:2403.01422, 2024

  61. [61]

    Longvlm: Efficient long video understanding via large language models,

    Y . Weng, M. Han, H. He, X. Chang, and B. Zhuang, “Longvlm: Efficient long video understanding via large language models,” in ECCV, 2024, pp. 453–470

  62. [62]

    Timechat: A time-sensitive multimodal large language model for long video understanding,

    S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, “Timechat: A time-sensitive multimodal large language model for long video understanding,” in CVPR, 2024, pp. 14 313–14 323

  63. [63]

    Vtimellm: Empower llm to grasp video moments,

    B. Huang, X. Wang, H. Chen, Z. Song, and W. Zhu, “Vtimellm: Empower llm to grasp video moments,” inCVPR, 2024, pp. 14 271– 14 280

  64. [64]

    Momentor: Advancing video large language model with fine-grained temporal reasoning,

    L. Qian, J. Li, Y . Wu, Y . Ye, H. Fei, T.-S. Chua, Y . Zhuang, and S. Tang, “Momentor: Advancing video large language model with fine-grained temporal reasoning,”arXiv:2402.11435, 2024

  65. [65]

    Videollm-online: Online video large language model for streaming video,

    J. Chen, Z. Lv, S. Wu, K. Q. Lin, C. Song, D. Gao, J.-W. Liu, Z. Gao, D. Mao, and M. Z. Shou, “Videollm-online: Online video large language model for streaming video,” inCVPR, 2024, pp. 18 407– 18 418

  66. [66]

    Lita: Language instructed temporal-localization assistant,

    D.-A. Huang, S. Liao, S. Radhakrishnan, H. Yin, P. Molchanov, Z. Yu, and J. Kautz, “Lita: Language instructed temporal-localization assistant,” inECCV, 2024, pp. 202–218

  67. [67]

    Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,

    Y . Guo, J. Liu, M. Li, D. Cheng, X. Tang, D. Sui, Q. Liu, X. Chen, and K. Zhao, “Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,” inAAAI, vol. 39, no. 3, 2025, pp. 3302–3310

  68. [68]

    Video-xl: Extra-long vision language model for hour-scale video understanding,

    Y . Shu, Z. Liu, P. Zhang, M. Qin, J. Zhou, Z. Liang, T. Huang, and B. Zhao, “Video-xl: Extra-long vision language model for hour-scale video understanding,” inCVPR, 2025, pp. 26 160–26 169

  69. [69]

    Time-r1: Post-training large vision language model for temporal video grounding,

    Y . Wang, Z. Wang, B. Xu, Y . Du, K. Lin, Z. Xiao, Z. Yue, J. Ju, L. Zhang, D. Yang, X. Fang, Z. He, Z. Luo, W. Wang, J. Lin, J. Luan, and Q. Jin, “Time-r1: Post-training large vision language model for temporal video grounding,”arXiv:2503.13377, 2025

  70. [70]

    Seam- lessm4t: Massively multilingual & multimodal machine translation,

    Seamless Communication, L. Barrault, Y .-A. Chung, M. C. Meglioli, D. Dale, N. Dong, P.-A. Duquenne, H. Elsahar, H. Gong, K. Heffernan, J. Hoffman, C. Klaiber, P. Li, D. Licht, J. Maillard, A. Rakotoarison, K. R. Sadagopan, G. Wenzek, E. Ye, B. Akula, P.-J. Chen, N. E. Hachem, B. Ellis, G. M. Gonzalez, J. Haaheim, P. Hansanti, R. Howes, B. Huang, M.-J. Hw...

  71. [71]

    Av2av: Direct audio- visual speech to audio-visual speech translation with unified audio- visual speech representation,

    J. Choi, S. J. Park, M. Kim, and Y . M. Ro, “Av2av: Direct audio- visual speech to audio-visual speech translation with unified audio- visual speech representation,” inCVPR, 2024, pp. 27 325–27 337

  72. [72]

    Adaptive inner speech-text alignment for llm-based speech trans- lation,

    H. Liu, A. Chen, K. Chen, X. Bai, M. Zhong, Y . Qiu, and M. Zhang, “Adaptive inner speech-text alignment for llm-based speech trans- lation,” inNatural Language Processing and Chinese Computing (NLPCC 2025), 2025

  73. [73]

    Inimagetrans: Multimodal llm-based text image machine translation,

    F. Zuo, K. Chen, Y . Zhang, Z. Xue, and M. Zhang, “Inimagetrans: Multimodal llm-based text image machine translation,” inFindings of the Association for Computational Linguistics: ACL 2025, 2025

  74. [74]

    Llama-adapter: Efficient fine-tuning of language models with zero-init attention,

    R. Zhang, J. Han, C. Liu, P. Gao, A. Zhou, X. Hu, S. Yan, P. Lu, H. Li, and Y . Qiao, “Llama-adapter: Efficient fine-tuning of language models with zero-init attention,” inICLR, 2024

  75. [75]

    Bt-adapter: Video conversation is feasible without video instruction tuning,

    R. Liu, C. Li, Y . Ge, T. H. Li, Y . Shan, and G. Li, “Bt-adapter: Video conversation is feasible without video instruction tuning,” inCVPR, 2024, pp. 13 658–13 667

  76. [76]

    Otter: A multi-modal model with in-context instruction tuning,

    B. Li, Y . Zhang, L. Chen, J. Wang, F. Pu, J. A. Cahyono, J. Yang, C. Li, and Z. Liu, “Otter: A multi-modal model with in-context instruction tuning,”IEEE Trans. Pattern Anal. Mach. Intell., 2025

  77. [77]

    Pllava: Parameter-free llava extension from images to videos for video dense captioning,

    L. Xu, Y . Zhao, D. Zhou, Z. Lin, S. K. Ng, and J. Feng, “Pllava: Parameter-free llava extension from images to videos for video dense captioning,”arXiv:2404.16994, 2024

  78. [78]

    From image to video, what do we need in multimodal llms?

    S. Huang, H. Zhang, L. Zhong, H. Chen, Y . Gao, Y . Hu, and Z. Qin, “From image to video, what do we need in multimodal llms?” arXiv:2404.11865, 2024

  79. [79]

    Reef: Relevance-aware and efficient llm adapter for video understanding,

    S. Reza, X. Song, H. Yu, Z. Lin, M. Moghaddam, and O. Camps, “Reef: Relevance-aware and efficient llm adapter for video understanding,” in CVPR, 2025, pp. 2592–2603

  80. [80]

    Vlog: Video-language models by generative retrieval of narration vocabulary,

    K. Q. Lin and M. Z. Shou, “Vlog: Video-language models by generative retrieval of narration vocabulary,” inCVPR, 2025, pp. 3218–3228

Showing first 80 references.