REVIEW 5 minor 176 references
Video translation works better as three coordinated MLLM roles than as a cascade of ASR, translation, speech, and lip-sync.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 22:03 UTC pith:3EQ76VC7
load-bearing objection Solid role-oriented survey that cleanly maps MLLM video translation without overclaiming; useful map, not a new result.
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
MLLM-enabled video translation should be framed as an integrated multimodal reasoning and generation task organized by three functional roles—Semantic Reasoner, Expressive Performer, and Visual Synthesizer—rather than as an isolated cascade of ASR, machine translation, TTS, and lip synchronization. That taxonomy links semantic grounding, expressive speech, and visual rendering and exposes why current proxy benchmarks fall short of end-to-end translated-video quality.
What carries the argument
The three-role taxonomy (Semantic Reasoner for multimodal grounding, temporal/event reasoning, cross-lingual planning, and efficient adaptation; Expressive Performer for speech-token LMs, prompt/emotion control, and diffusion/flow rendering; Visual Synthesizer for lip sync, talking-head animation, and enabling video backbones). It re-cuts the literature by function in the translation pipeline instead of by isolated subtask names.
Load-bearing premise
That sorting studies by these three roles and the survey’s inclusion rules gives a complete enough map that role-level proxy scores can stand in for what end-to-end video translation still lacks.
What would settle it
Build or publish a public end-to-end multilingual video-translation benchmark with aligned source video, target speech, visual outputs, and multi-axis human ratings; if systems strong only on the paper’s role proxies still rank poorly there, or if strong end-to-end systems cannot be usefully described by the three roles, the taxonomy’s organizing claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey argues that MLLM progress is reshaping video translation from a cascaded ASR–MT–TTS–lip-sync pipeline into an integrated multimodal reasoning-and-generation problem. It organizes MLLM-enabled and MLLM-relevant work into three functional roles—Semantic Reasoner (multimodal grounding, temporal/event reasoning, cross-lingual planning, efficient adaptation), Expressive Performer (speech-token LMs, prompt/emotion control, diffusion/flow rendering), and Visual Synthesizer (lip sync, talking-head animation, enabling visual backbones)—and reviews role-level datasets, benchmarks, and metrics while arguing that current proxy evaluations fall short of end-to-end translated-video quality. Open challenges include long-form understanding, temporal/cross-modal alignment, multilingual robustness, real-time deployment, and responsible use.
Significance. If accepted as a framing contribution, the paper offers a useful map for a fragmented area that spans video-MLLMs, speech translation/TTS, and talking-head/video generation. Strengths include an explicit literature protocol (§II-A), a clear comparison with prior surveys (Table I), a role taxonomy that separates core MLLM reasoning from supporting generation modules (Fig. 1 and caption), and a consistent gap analysis that does not overclaim end-to-end systems from proxy VideoQA/TTS/lip-sync scores (§V–VI). For a survey venue, this is a coherent organizational contribution rather than a new empirical result.
minor comments (5)
- In Table III, PromptTTS is cited as [161] while the taxonomy and earlier text use [90]; unify the PromptTTS citation and ensure Table III entries match the bibliography consistently.
- Fig. 2–4 captions and body text are clear, but several figure placeholders in the source appear as garbled character blocks; ensure final production figures are clean and that architecture diagrams remain readable at print scale.
- §V-A Table II reports published zero-shot VideoQA numbers as proxy evidence; a short caveat in the table caption that scores are not re-evaluated under a single protocol would further reduce cross-paper comparability risk.
- Minor typography/consistency: mixed spellings such as CosyVoice/CosyV oice, AV2AV/A V2A V, and occasional spacing artifacts (e.g., F .) should be normalized in copyediting.
- §II-A states coverage up to May 2026; confirm that the final arXiv/journal version’s cutoff date matches the bibliography and that any post-cutoff highly relevant systems are either included or explicitly scoped out.
Circularity Check
No significant circularity: role taxonomy is organizational mapping of external work, not a fitted or self-referential derivation.
full rationale
This is a survey paper whose central contribution is a role-oriented taxonomy (Semantic Reasoner, Expressive Performer, Visual Synthesizer) for organizing existing MLLM-enabled and MLLM-relevant video-translation literature. The taxonomy is definitional organization of external methods, datasets, and metrics; it does not claim a first-principles derivation, uniqueness theorem, or quantitative prediction that could reduce to its own inputs by construction. Inclusion criteria (§II-A) and the explicit separation of core MLLM reasoning from supporting TTS/visual backbones (Fig. 1 caption; §I) are stated as scoping choices, not as empirical results forced by self-citation. Table I honestly marks partial coverage of prior surveys; §V–VI repeatedly treat VideoQA, MOS/WER, and SyncNet-style scores as proxies that do not measure end-to-end translated-video quality. Evaluation tables report published scores from other papers rather than redefining success via fitted parameters. There is no self-definitional loop, no fitted input called a prediction, no load-bearing uniqueness imported from the authors’ prior work, and no renaming of a known empirical law as a new derivation. Self-citations, if any, are ordinary survey practice and not load-bearing for a forced result. Score 0 is therefore the correct, proportionate finding.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption High-quality video translation requires joint semantic fidelity, temporal alignment, speaker consistency, and emotional expressiveness across visual, acoustic, and linguistic streams.
- domain assumption Cascaded ASR–MT–TTS–lip-sync pipelines suffer error propagation and weak cross-modal coordination.
- ad hoc to paper MLLMs and related generation modules can be productively analyzed by functional role (reason, speak, render) rather than only by classical task boundaries.
- domain assumption Role-level proxy benchmarks (VideoQA, TTS metrics, lip-sync scores) are informative diagnostics but insufficient for end-to-end translated-video quality.
invented entities (3)
-
Semantic Reasoner (role)
no independent evidence
-
Expressive Performer (role)
no independent evidence
-
Visual Synthesizer (role)
no independent evidence
read the original abstract
Recent progress in multimodal large language models (MLLMs) is reshaping video translation from a cascaded pipeline of automatic speech recognition, machine translation, text-to-speech, and lip synchronization into a unified multimodal reasoning and generation problem. High-quality video translation requires not only semantic fidelity, but also temporal alignment, speaker consistency, and emotional expressiveness across visual, acoustic, and linguistic streams. This survey provides a focused review of MLLM-enabled video translation through a role-oriented taxonomy. We organize MLLM-enabled and MLLM-relevant studies into three functional roles: Semantic Reasoner, which grounds translation in video understanding, temporal reasoning, and multimodal fusion; Expressive Performer, which supports controllable and context-aware speech generation; and Visual Synthesizer, which enables lip synchronization and visually coherent speaker rendering. We further summarize representative datasets, benchmarks, and metrics for each role, and discuss how current evaluation protocols fall short of end-to-end video translation requirements. Finally, we identify open challenges in long-form video understanding, temporal modeling, multimodal alignment, multilingual robustness, and responsible deployment, outlining future directions for natural and trustworthy cross-lingual video communication.
Figures
Reference graph
Works this paper leans on
-
[1]
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. rahman Mohamed, N. Jaitly, A. Senior, V . Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”IEEE Signal Process. Mag., vol. 29, no. 6, pp. 82–97, 2012
2012
-
[2]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS, 2017, pp. 5998–6008
2017
-
[3]
Tacotron: Towards end-to-end speech synthesis,
Y . Wang, R. J. Skerry-Ryan, D. Stanton, Y . Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y . Xiao, Z. Chen, S. Bengio, Q. Le, Y . Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” inInterspeech, 2017, pp. 4006–4010
2017
-
[4]
A lip sync expert is all you need for speech to lip generation in the wild,
K. R. Prajwal, R. Mukhopadhyay, V . P. Namboodiri, and C. V . Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” inACM MM, 2020, pp. 484–492
2020
-
[5]
Direct speech-to-speech translation with a sequence-to- sequence model,
Y . Jia, R. J. Weiss, F. Biadsy, W. Macherey, M. Johnson, Z. Chen, and Y . Wu, “Direct speech-to-speech translation with a sequence-to- sequence model,” inInterspeech, 2019, pp. 1123–1127
2019
-
[6]
Learning transferable visual models from natural language supervi- sion,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervi- sion,” inICML, 2021, pp. 8748–8763
2021
-
[7]
Flamingo: A visual language model for few-shot learning,
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Bi ´nkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan, “Flamingo: A visual languag...
2022
-
[8]
Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,” inICML, 2023, pp. 19 730–19 742
2023
-
[9]
Visual instruction tuning,
H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” in NeurIPS, vol. 36, 2023, pp. 34 892–34 916
2023
-
[10]
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,
Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang, “mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,” inCVPR, 2024, pp. 13 040–13 051
2024
-
[11]
Video-chatgpt: Towards detailed video understanding via large vision and language models,
M. Maaz, H. Rasheed, S. Khan, and F. Khan, “Video-chatgpt: Towards detailed video understanding via large vision and language models,” in ACL, 2024, pp. 12 585–12 602
2024
-
[12]
Video-llama: An instruction-tuned audio-visual language model for video understanding,
H. Zhang, X. Li, and L. Bing, “Video-llama: An instruction-tuned audio-visual language model for video understanding,” inEMNLP, 2023, pp. 543–553
2023
-
[13]
Llama-vid: An image is worth 2 tokens in large language models,
Y . Li, C. Wang, and J. Jia, “Llama-vid: An image is worth 2 tokens in large language models,” inECCV, 2024, pp. 323–340
2024
-
[14]
Internvideo2: Scaling foundation models for multimodal video understanding,
Y . Wang, K. Li, X. Li, J. Yu, Y . He, G. Chen, B. Pei, R. Zheng, Z. Wang, Y . Shi, T. Jiang, S. Li, J. Xu, H. Zhang, Y . Huang, Y . Qiao, Y . Wang, and L. Wang, “Internvideo2: Scaling foundation models for multimodal video understanding,” inECCV, 2024, pp. 396–416
2024
-
[15]
Cosyvoice 2: Scalable streaming speech synthesis with large language models,
Z. Du, Y . Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y . Yang, C. Gao, H. Wang, F. Yu, H. Liu, Z. Sheng, Y . Gu, C. Deng, W. Wang, S. Zhang, Z. Yan, and J. Zhou, “Cosyvoice 2: Scalable streaming speech synthesis with large language models,”arXiv:2412.10117, 2024
Pith/arXiv arXiv 2024
-
[16]
F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,
Y . Chen, Z. Niu, Z. Ma, K. Deng, C. Wang, J. Zhao, K. Yu, and X. Chen, “F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,” inACL, 2025, pp. 6255–6271
2025
-
[17]
A survey on video diffusion models,
Z. Xing, Q. Feng, H. Chen, Q. Dai, H. Hu, H. Xu, Z. Wu, and Y .- G. Jiang, “A survey on video diffusion models,”ACM Comput. Surv., vol. 57, no. 2, pp. 1–42, 2024
2024
-
[18]
Omnisync: Towards universal lip synchronization via diffusion transformers,
Z. Peng, J. Liu, H. Zhang, X. Liu, S. Tang, P. Wan, D. Zhang, H. Liu, and J. He, “Omnisync: Towards universal lip synchronization via diffusion transformers,”arXiv:2505.21448, 2025
arXiv 2025
-
[19]
A survey on multi-modal machine translation: Tasks, methods and challenges,
H. Shen, L. Shao, W. Li, Z. Lan, Z. Liu, and J. Su, “A survey on multi-modal machine translation: Tasks, methods and challenges,” arXiv:2405.12669, 2024
Pith/arXiv arXiv 2024
-
[20]
Video-guided machine translation: A survey of models, datasets, and challenges,
P. Das, V . Singh, P. Bhattacharyya, and G. Haffari, “Video-guided machine translation: A survey of models, datasets, and challenges,” inIJCNLP-AACL, 2025, pp. 3346–3356
2025
-
[21]
A survey on multimodal large language models,
S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen, “A survey on multimodal large language models,”Natl. Sci. Rev., vol. 11, no. 12, p. nwae403, 2024
2024
-
[22]
Video understanding with large language models: A survey,
Y . Tang, J. Bi, S. Xu, L. Song, S. Liang, T. Wang, D. Zhang, J. An, J. Lin, R. Zhu, A. V osoughi, C. Huang, Z. Zhang, P. Liu, M. Feng, F. Zheng, J. Zhang, P. Luo, J. Luo, and C. Xu, “Video understanding with large language models: A survey,”TCSVT, 2025
2025
-
[23]
A survey on video temporal grounding with multimodal large language model,
J. Wu, W. Liu, Y . Liu, M. Liu, L. Nie, Z. Lin, and C. W. Chen, “A survey on video temporal grounding with multimodal large language model,”IEEE Trans. Pattern Anal. Mach. Intell., 2025
2025
-
[24]
Direct speech-to-speech neural machine translation: A survey,
M. Gupta, M. Dutta, and C. K. Maurya, “Direct speech-to-speech neural machine translation: A survey,”arXiv:2411.14453, 2024
Pith/arXiv arXiv 2024
-
[25]
Towards controllable speech synthesis in the era of large language models: A systematic survey,
T. Xie, Y . Rong, P. Zhang, W. Wang, and L. Liu, “Towards controllable speech synthesis in the era of large language models: A systematic survey,” inEMNLP, 2025, pp. 764–791
2025
-
[26]
V . K. Rakesh, S. Mazumdar, R. P. Maity, S. Pal, A. Das, and T. Samanta, “Advancing talking head generation: A comprehensive survey of multi-modal methodologies, datasets, evaluation metrics, and loss functions,”arXiv:2507.02900, 2025
arXiv 2025
-
[27]
Rabiner and B.-H
L. Rabiner and B.-H. Juang,Fundamentals of Speech Recognition. Prentice-Hall, 1993
1993
-
[28]
Connection- ist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,
A. Graves, S. Fern ´andez, F. Gomez, and J. Schmidhuber, “Connection- ist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” inICML, 2006, pp. 369–376
2006
-
[29]
Deep speech: Scaling up end-to-end speech recognition,
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, and A. Y . Ng, “Deep speech: Scaling up end-to-end speech recognition,”arXiv:1412.5567, 2014
Pith/arXiv arXiv 2014
-
[30]
Listen, attend and spell,
W. Chan, N. Jaitly, Q. V . Le, and O. Vinyals, “Listen, attend and spell,” inICASSP, 2016, pp. 4960–4964
2016
-
[31]
The mathematics of statistical machine translation: Parameter estimation,
P. F. Brown, S. A. D. Pietra, V . J. D. Pietra, and R. L. Mercer, “The mathematics of statistical machine translation: Parameter estimation,” Comput. Linguist., vol. 19, no. 2, pp. 263–311, 1993
1993
-
[32]
Statistical phrase-based transla- tion,
P. Koehn, F. J. Och, and D. Marcu, “Statistical phrase-based transla- tion,” inNAACL, 2003, pp. 127–133
2003
-
[33]
Sequence to sequence learning with neural networks,
I. Sutskever, O. Vinyals, and Q. V . Le, “Sequence to sequence learning with neural networks,” inNeurIPS, 2014, pp. 3104–3112
2014
-
[34]
Speech parameter generation algorithms for hmm-based speech syn- thesis,
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for hmm-based speech syn- thesis,” inICASSP, vol. 3, 2000, pp. 1315–1318
2000
-
[35]
Statistical parametric speech synthesis,
H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,”Speech Commun., vol. 51, no. 11, pp. 1039–1064, 2009
2009
-
[36]
Wavenet: A generative model for raw audio,
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,”arXiv:1609.03499, 2016
Pith/arXiv arXiv 2016
-
[37]
Photo-realistic talking-heads from image samples,
E. Cosatto and H. P. Graf, “Photo-realistic talking-heads from image samples,”IEEE Trans. Multimedia, vol. 2, no. 3, pp. 152–163, 2002
2002
-
[38]
Lip reading sentences in the wild,
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” inCVPR, 2017, pp. 6447–6456
2017
-
[39]
Few-shot adversarial learning of realistic neural talking head models,
E. Zakharov, A. Shysheya, E. Burkov, and V . Lempitsky, “Few-shot adversarial learning of realistic neural talking head models,” inICCV, 2019, pp. 9459–9468
2019
-
[40]
Onellm: One framework to align all modalities with language,
J. Han, K. Gong, Y . Zhang, J. Wang, K. Zhang, D. Lin, Y . Qiao, P. Gao, and X. Yue, “Onellm: One framework to align all modalities with language,” inCVPR, 2024, pp. 26 584–26 595
2024
-
[41]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” inICLR, 2021
2021
-
[42]
Scaling up visual and vision-language representation learning with noisy text supervision,
C. Jia, Y . Yang, Y . Xia, Y .-T. Chen, Z. Parekh, H. Pham, Q. Le, Y .-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” inICML, 2021, pp. 4904–4916
2021
-
[43]
Videobert: A joint model for video and language representation learning,
C. Sun, A. Myers, C. V ondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” inICCV, 2019, pp. 7464–7473
2019
-
[44]
Actbert: Learning global-local video-text repre- sentations,
L. Zhu and Y . Yang, “Actbert: Learning global-local video-text repre- sentations,” inCVPR, 2020, pp. 8746–8755
2020
-
[45]
Merlot: Multimodal neural script knowledge models,
R. Zellers, X. Lu, J. Hessel, Y . Yu, J. S. Park, J. Cao, A. Farhadi, and Y . Choi, “Merlot: Multimodal neural script knowledge models,” in NeurIPS, vol. 34, 2021, pp. 23 634–23 651
2021
-
[46]
Frozen in time: A joint video and image encoder for end-to-end retrieval,
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” inICCV, 2021, pp. 1728–1738
2021
-
[47]
Violet: End-to-end video-language transformers with masked visual- token modeling,
T.-J. Fu, L. Li, Z. Gan, K. Lin, W. Y . Wang, L. Wang, and Z. Liu, “Violet: End-to-end video-language transformers with masked visual- token modeling,” inNeurIPS, 2021
2021
-
[48]
Omnivl: One foundation model for image- language and video-language tasks,
J. Wang, D. Chen, Z. Wu, C. Luo, L. Zhou, Y . Zhao, Y . Xie, C. Liu, Y .-G. Jiang, and L. Yuan, “Omnivl: One foundation model for image- language and video-language tasks,” inNeurIPS, vol. 35, 2022, pp. 5696–5710
2022
-
[49]
Internvideo: General video foundation models via generative and discriminative learning,
Y . Wang, K. Li, Y . Li, Y . He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y . Liu, Z. Wang, Y . Shi, Z. Zhang, Z. Li, Y . Wang, and L. Wang, “Internvideo: General video foundation models via generative and discriminative learning,”arXiv:2212.03191, 2022
Pith/arXiv arXiv 2022
-
[50]
Unival: Unified model for image, video, audio and language tasks,
M. Shukor, C. Dancette, A. Rame, and M. Cord, “Unival: Unified model for image, video, audio and language tasks,”TMLR, 2023
2023
-
[51]
Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,
J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi, “Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,” inCVPR, 2024, pp. 26 439–26 455
2024
-
[52]
Pali-x: On scaling up a multilingual vision and language model,
X. Chen, J. Djolonga, P. Padlewski, B. Mustafa, S. Changpinyo, J. Wu, C. R. Ruiz, S. Goodman, X. Wang, and Y . Tay, “Pali-x: On scaling up a multilingual vision and language model,”arXiv:2305.18565, 2023
Pith/arXiv arXiv 2023
-
[53]
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou, “Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond,”arXiv:2308.12966, 2023
Pith/arXiv arXiv 2023
-
[54]
Videollama 3: Frontier multimodal foundation models for image and video understanding,
B. Zhang, K. Li, Z. Cheng, Z. Hu, Y . Yuan, G. Chen, S. Leng, Y . Jiang, H. Zhang, X. Li, P. Jin, W. Zhang, F. Wang, L. Bing, and D. Zhao, “Videollama 3: Frontier multimodal foundation models for image and video understanding,”arXiv:2501.13106, 2025
Pith/arXiv arXiv 2025
-
[55]
Videochat: Chat-centric video understanding,
K. Li, Y . He, Y . Wang, Y . Li, W. Wang, P. Luo, Y . Wang, L. Wang, and Y . Qiao, “Videochat: Chat-centric video understanding,” arXiv:2305.06355, 2023
Pith/arXiv arXiv 2023
-
[56]
Valley: Video assistant with large language model enhanced ability,
R. Luo, Z. Zhao, M. Yang, J. Dong, D. Li, P. Lu, T. Wang, L. Hu, M. Qiu, and Z. Wei, “Valley: Video assistant with large language model enhanced ability,”arXiv:2306.07207, 2023
Pith/arXiv arXiv 2023
-
[57]
Mvbench: A comprehensive multi- modal video understanding benchmark,
K. Li, Y . Wang, Y . He, Y . Li, Y . Wang, Y . Liu, Z. Wang, J. Xu, G. Chen, P. Luo, L. Wang, and Y . Qiao, “Mvbench: A comprehensive multi- modal video understanding benchmark,” inCVPR, 2024, pp. 22 195– 22 206
2024
-
[58]
Groundinggpt: Language enhanced multi-modal grounding model,
Z. Li, Q. Xu, D. Zhang, H. Song, Y . Cai, Q. Qi, R. Zhou, J. Pan, Z. Li, V . T. Vu, Z. Huang, and T. Wang, “Groundinggpt: Language enhanced multi-modal grounding model,” inACL, 2024, pp. 6657–6678. 13
2024
-
[59]
Moviechat: From dense token to sparse memory for long video understanding,
E. Song, W. Chai, G. Wang, Y . Zhang, H. Zhou, F. Wu, H. Chi, X. Guo, T. Ye, Y . Lu, J.-N. Hwang, and G. Wang, “Moviechat: From dense token to sparse memory for long video understanding,” inCVPR, 2024, pp. 18 221–18 232
2024
-
[60]
Moviellm: Enhancing long video understanding with ai-generated movies,
Z. Song, C. Wang, J. Sheng, C. Zhang, G. Yu, J. Fan, and T. Chen, “Moviellm: Enhancing long video understanding with ai-generated movies,”arXiv:2403.01422, 2024
Pith/arXiv arXiv 2024
-
[61]
Longvlm: Efficient long video understanding via large language models,
Y . Weng, M. Han, H. He, X. Chang, and B. Zhuang, “Longvlm: Efficient long video understanding via large language models,” in ECCV, 2024, pp. 453–470
2024
-
[62]
Timechat: A time-sensitive multimodal large language model for long video understanding,
S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, “Timechat: A time-sensitive multimodal large language model for long video understanding,” in CVPR, 2024, pp. 14 313–14 323
2024
-
[63]
Vtimellm: Empower llm to grasp video moments,
B. Huang, X. Wang, H. Chen, Z. Song, and W. Zhu, “Vtimellm: Empower llm to grasp video moments,” inCVPR, 2024, pp. 14 271– 14 280
2024
-
[64]
Momentor: Advancing video large language model with fine-grained temporal reasoning,
L. Qian, J. Li, Y . Wu, Y . Ye, H. Fei, T.-S. Chua, Y . Zhuang, and S. Tang, “Momentor: Advancing video large language model with fine-grained temporal reasoning,”arXiv:2402.11435, 2024
Pith/arXiv arXiv 2024
-
[65]
Videollm-online: Online video large language model for streaming video,
J. Chen, Z. Lv, S. Wu, K. Q. Lin, C. Song, D. Gao, J.-W. Liu, Z. Gao, D. Mao, and M. Z. Shou, “Videollm-online: Online video large language model for streaming video,” inCVPR, 2024, pp. 18 407– 18 418
2024
-
[66]
Lita: Language instructed temporal-localization assistant,
D.-A. Huang, S. Liao, S. Radhakrishnan, H. Yin, P. Molchanov, Z. Yu, and J. Kautz, “Lita: Language instructed temporal-localization assistant,” inECCV, 2024, pp. 202–218
2024
-
[67]
Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,
Y . Guo, J. Liu, M. Li, D. Cheng, X. Tang, D. Sui, Q. Liu, X. Chen, and K. Zhao, “Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,” inAAAI, vol. 39, no. 3, 2025, pp. 3302–3310
2025
-
[68]
Video-xl: Extra-long vision language model for hour-scale video understanding,
Y . Shu, Z. Liu, P. Zhang, M. Qin, J. Zhou, Z. Liang, T. Huang, and B. Zhao, “Video-xl: Extra-long vision language model for hour-scale video understanding,” inCVPR, 2025, pp. 26 160–26 169
2025
-
[69]
Time-r1: Post-training large vision language model for temporal video grounding,
Y . Wang, Z. Wang, B. Xu, Y . Du, K. Lin, Z. Xiao, Z. Yue, J. Ju, L. Zhang, D. Yang, X. Fang, Z. He, Z. Luo, W. Wang, J. Lin, J. Luan, and Q. Jin, “Time-r1: Post-training large vision language model for temporal video grounding,”arXiv:2503.13377, 2025
Pith/arXiv arXiv 2025
-
[70]
Seam- lessm4t: Massively multilingual & multimodal machine translation,
Seamless Communication, L. Barrault, Y .-A. Chung, M. C. Meglioli, D. Dale, N. Dong, P.-A. Duquenne, H. Elsahar, H. Gong, K. Heffernan, J. Hoffman, C. Klaiber, P. Li, D. Licht, J. Maillard, A. Rakotoarison, K. R. Sadagopan, G. Wenzek, E. Ye, B. Akula, P.-J. Chen, N. E. Hachem, B. Ellis, G. M. Gonzalez, J. Haaheim, P. Hansanti, R. Howes, B. Huang, M.-J. Hw...
Pith/arXiv arXiv 2023
-
[71]
Av2av: Direct audio- visual speech to audio-visual speech translation with unified audio- visual speech representation,
J. Choi, S. J. Park, M. Kim, and Y . M. Ro, “Av2av: Direct audio- visual speech to audio-visual speech translation with unified audio- visual speech representation,” inCVPR, 2024, pp. 27 325–27 337
2024
-
[72]
Adaptive inner speech-text alignment for llm-based speech trans- lation,
H. Liu, A. Chen, K. Chen, X. Bai, M. Zhong, Y . Qiu, and M. Zhang, “Adaptive inner speech-text alignment for llm-based speech trans- lation,” inNatural Language Processing and Chinese Computing (NLPCC 2025), 2025
2025
-
[73]
Inimagetrans: Multimodal llm-based text image machine translation,
F. Zuo, K. Chen, Y . Zhang, Z. Xue, and M. Zhang, “Inimagetrans: Multimodal llm-based text image machine translation,” inFindings of the Association for Computational Linguistics: ACL 2025, 2025
2025
-
[74]
Llama-adapter: Efficient fine-tuning of language models with zero-init attention,
R. Zhang, J. Han, C. Liu, P. Gao, A. Zhou, X. Hu, S. Yan, P. Lu, H. Li, and Y . Qiao, “Llama-adapter: Efficient fine-tuning of language models with zero-init attention,” inICLR, 2024
2024
-
[75]
Bt-adapter: Video conversation is feasible without video instruction tuning,
R. Liu, C. Li, Y . Ge, T. H. Li, Y . Shan, and G. Li, “Bt-adapter: Video conversation is feasible without video instruction tuning,” inCVPR, 2024, pp. 13 658–13 667
2024
-
[76]
Otter: A multi-modal model with in-context instruction tuning,
B. Li, Y . Zhang, L. Chen, J. Wang, F. Pu, J. A. Cahyono, J. Yang, C. Li, and Z. Liu, “Otter: A multi-modal model with in-context instruction tuning,”IEEE Trans. Pattern Anal. Mach. Intell., 2025
2025
-
[77]
Pllava: Parameter-free llava extension from images to videos for video dense captioning,
L. Xu, Y . Zhao, D. Zhou, Z. Lin, S. K. Ng, and J. Feng, “Pllava: Parameter-free llava extension from images to videos for video dense captioning,”arXiv:2404.16994, 2024
Pith/arXiv arXiv 2024
-
[78]
From image to video, what do we need in multimodal llms?
S. Huang, H. Zhang, L. Zhong, H. Chen, Y . Gao, Y . Hu, and Z. Qin, “From image to video, what do we need in multimodal llms?” arXiv:2404.11865, 2024
Pith/arXiv arXiv 2024
-
[79]
Reef: Relevance-aware and efficient llm adapter for video understanding,
S. Reza, X. Song, H. Yu, Z. Lin, M. Moghaddam, and O. Camps, “Reef: Relevance-aware and efficient llm adapter for video understanding,” in CVPR, 2025, pp. 2592–2603
2025
-
[80]
Vlog: Video-language models by generative retrieval of narration vocabulary,
K. Q. Lin and M. Z. Shou, “Vlog: Video-language models by generative retrieval of narration vocabulary,” inCVPR, 2025, pp. 3218–3228
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.