Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T22:04:31.302192Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 100 of 176 outbound references and 0 inbound Pith citation observations for arXiv:2604.11283.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T22:04:31.302192Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 176 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 10768c91-4c8f-48fa-a272-7d1ab7561600 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 160cea28-2304-4f49-b1da-d2adfc3268ab · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Attention is all you need,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc18f9d-2a66-4c58-8660-89aecf196f2d · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Tacotron: Towards end-to-end speech synthesis,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dcdd481-d9f2-4978-bef9-da71125d0718 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey A lip sync expert is all you need for speech to lip generation in the wild,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5303075-5d96-42b2-98ae-d02318d4c7a1 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Direct speech-to-speech translation with a sequence-to- sequence model,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55e557c7-a83c-461d-982a-04df1c1f32ee · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Learning transferable visual models from natural language supervi- sion,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9a5b663-85ae-46c3-9c1d-b639bcbaf8a1 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Flamingo: A visual language model for few-shot learning,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8270dea-65ae-4719-bcf5-da0ce077fb42 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a10440fb-602f-488a-8a5f-a99d21359d6f · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Visual instruction tuning,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4dd3ce0-a03e-43fc-92bc-5c1ff2c46ad4 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c52d8c46-8963-4dcc-a36d-3f38695b628a · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Video-chatgpt: Towards detailed video understanding via large vision and language models,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90cb84c3-4bd3-4607-a61a-152f361024b2 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Video-llama: An instruction-tuned audio-visual language model for video understanding,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64c1ce9a-3c23-4695-9a24-c3e21dd9b449 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Llama-vid: An image is worth 2 tokens in large language models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c54f1da5-4857-4b49-a020-fc06cc90dc37 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Internvideo2: Scaling foundation models for multimodal video understanding,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4899045-5f27-4698-9c8e-33b41a8884b0 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea2a24ff-9863-478d-81c4-6342b93c7be4 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 275fcc5a-7699-43c7-962b-7237ba203c8c · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey A survey on video diffusion models,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfc92ba4-c667-4b87-a752-7c922f749208 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Omnisync: Towards universal lip synchronization via diffusion transformers,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e706ad03-e91b-4dd9-85ae-4b5d210013d1 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey A Survey on Multi-modal Machine Translation: Tasks, Methods and Challenges
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66d72910-9195-4787-bb78-39e8c1b4dd81 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Video-guided machine translation: A survey of models, datasets, and challenges,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8f38568-bab9-4527-9df5-4febdffb687f · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey A survey on multimodal large language models,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db826b02-7e8f-420c-a01c-f7a3abafbed3 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Video understanding with large language models: A survey,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26eccb96-1dd0-44e9-b82f-316a3a735192 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey A survey on video temporal grounding with multimodal large language model,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc0aed22-173e-4283-a805-bdd0a25edd93 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Direct Speech-to-Speech Neural Machine Translation: A Survey
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce5f19c4-0f83-4a90-87fa-0e9f9e34ac11 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Towards controllable speech synthesis in the era of large language models: A systematic survey,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fa792ce-96a7-448d-ad97-39496589d1bd · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Advancing talking head generation: A comprehensive survey of multi-modal methodologies, datasets, evaluation metrics, and loss functions,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa4cfb08-33cb-4d22-bc6b-a9b6e4760899 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Rabiner and B.-H
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e0517a-cf0a-4063-b2f1-c93093946845 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Connection- ist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b95312-4df9-4ba7-af79-1b6243102680 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Deep Speech: Scaling up end-to-end speech recognition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe51874d-d962-4993-9687-15d22b3cea4a · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Listen, attend and spell,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82758b61-a320-4cef-9ec0-d8e837e138c9 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey The mathematics of statistical machine translation: Parameter estimation,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad93a695-7de4-415e-b647-666d5f3d0846 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Statistical phrase-based transla- tion,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfd52df-9c8c-4911-b3e0-31fadbe1e67f · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Sequence to sequence learning with neural networks,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c4693f-be0c-4201-87c0-7809eeb7524e · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Speech parameter generation algorithms for hmm-based speech syn- thesis,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b1a67ab-e196-4150-a9e6-3b888ee3815b · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Statistical parametric speech synthesis,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d11733ce-2223-4cbb-9f7d-d36caa5c0f0b · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey WaveNet: A Generative Model for Raw Audio
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2be9e07-127d-46ac-a2b5-6052be1cb955 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Photo-realistic talking-heads from image samples,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9379d2f9-a24f-47dd-af73-53137e5254d2 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Lip reading sentences in the wild,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3f0000a-9e5b-480a-8ecc-bb5ad6a9b90a · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Few-shot adversarial learning of realistic neural talking head models,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ce64f5f-d3fa-49fc-92df-656f8c3a125e · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Onellm: One framework to align all modalities with language,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d73ef55-7391-4aea-bce7-041f6685bbf2 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey An image is worth 16x16 words: Transformers for image recognition at scale,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daf14975-2cad-485a-93ad-6bf2230210c0 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Scaling up visual and vision-language representation learning with noisy text supervision,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ff4e49-2cc0-47ea-9c56-24f88d298d42 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Videobert: A joint model for video and language representation learning,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b35132a1-d16d-4fe1-91de-d9e115b87974 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Actbert: Learning global-local video-text repre- sentations,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c09d45f8-535c-453f-b463-64b97345b01b · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Merlot: Multimodal neural script knowledge models,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 483d256a-4ef9-4304-af45-e3e508f8d609 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Frozen in time: A joint video and image encoder for end-to-end retrieval,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dfdf966-6e4c-43ee-aaf5-0669857d0008 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Violet: End-to-end video-language transformers with masked visual- token modeling,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa371314-5ff1-4d14-aaf2-f7878025e093 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Omnivl: One foundation model for image- language and video-language tasks,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c999f0-f809-4892-80c9-05670df423ee · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4507d018-e051-4a22-b52d-49fbc6b6d26b · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Unival: Unified model for image, video, audio and language tasks,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6004dae6-7026-4b8e-bafc-53a5143f010a · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa3e50f-30de-4fa4-980a-fbbf94b42c88 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b91d41-a5c1-4b18-9ac9-70901ebbfe73 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b60ed54-80be-45aa-82d0-3f1d280504ee · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f1f047-b84d-4a46-95c5-dbe1724e49d1 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey VideoChat: Chat-Centric Video Understanding
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d2e3842-bbb3-4e92-84ad-800e932d1c89 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d22707e-a9a9-45ad-be56-2c7f62545229 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Mvbench: A comprehensive multi- modal video understanding benchmark,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be48bc47-fefc-4306-8f73-aa1e24715988 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Groundinggpt: Language enhanced multi-modal grounding model,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ca39dc-8743-4b51-afb6-6072a4cf7452 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Moviechat: From dense token to sparse memory for long video understanding,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dde9bc3-9171-4c61-bc3a-f7ffe8c86a11 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54af9688-e32a-4b0d-83c8-0854bf176e86 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Longvlm: Efficient long video understanding via large language models,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f16a7d-ef0c-4731-9d87-6bfb2f9429cf · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Timechat: A time-sensitive multimodal large language model for long video understanding,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b7076a6-220a-4457-b723-32c353d2f7c0 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vtimellm: Empower llm to grasp video moments,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b14dd2fa-b7d8-4e7a-afc2-dc5e2c459a8d · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43c77e3f-1d5a-4a5a-b119-3c912e7a53c2 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Videollm-online: Online video large language model for streaming video,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b1bcd79-5687-4d88-a1fc-899c71622054 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Lita: Language instructed temporal-localization assistant,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5398d75c-2f60-4d21-8592-3986f2665dbd · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ffa213a-6bef-424f-935e-45aa0041ff81 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Video-xl: Extra-long vision language model for hour-scale video understanding,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e8ebaf9-3410-4ae5-bd62-0a945189ece1 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be03c237-266f-43f1-9f75-ee4376385482 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ef2916-36f7-4037-9372-f0039b9b1239 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Av2av: Direct audio- visual speech to audio-visual speech translation with unified audio- visual speech representation,
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d4d6e7a-9647-421a-9d92-bd5f93a675bf · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Adaptive inner speech-text alignment for llm-based speech trans- lation,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8f12530-b7c1-41d5-92f2-7e5223a631f3 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Inimagetrans: Multimodal llm-based text image machine translation,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60853e80-729e-4c6a-8be1-e44a06e385f5 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Llama-adapter: Efficient fine-tuning of language models with zero-init attention,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f6a9045-e26e-4717-b525-43abcce74042 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Bt-adapter: Video conversation is feasible without video instruction tuning,
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa46b280-4fb8-47fe-b285-5ffe9559d136 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Otter: A multi-modal model with in-context instruction tuning,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3ad4f5-f0cb-481c-a1e4-4ea5f04d58bc · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dd0a2ad-7068-4bb3-a3dc-a6079849cc68 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey From Image to Video, what do we need in multimodal LLMs?
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82dee9e6-900f-4f14-84ba-277bccc87674 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Reef: Relevance-aware and efficient llm adapter for video understanding,
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0072e5b-df64-4ba2-be21-c4407ea5164a · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vlog: Video-language models by generative retrieval of narration vocabulary,
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d8b360-d86c-4a6e-b5a3-ad3d28189a0b · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vall-e r: Robust and efficient zero-shot text-to-speech synthesis via monotonic alignment,
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f06165-fdfa-496c-8376-a8fe887675b3 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f348cb99-400d-4a35-b18c-dc165d39c950 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 148aecbb-bd2a-4eb1-b797-1cb8b87f8337 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa24220d-3b3d-4b45-8d3f-e9dc9e36fabd · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb3292c-5288-40fd-be9f-aa3f4257990a · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03fc689e-c71d-42b2-a81e-7ef85ef8b1af · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4c39120-64ca-43b2-befd-ce6d88be8442 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a1f961-7ef8-4cc2-ac74-94e3e8085089 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a850a614-1c5e-462a-93bb-3ff0952954d1 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey PromptTTS 2: Describing and Generating Voices with Text Prompt
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7dca130-2869-4c72-9d71-5b43ab92708f · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06085be6-89e8-469c-8f4a-485aece58c1b · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Styletts 2: Towards human-level text-to-speech through style diffu- sion and adversarial training with large speech language models,
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 008ddd17-1a0e-4145-bed2-d3466210c9d4 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Controlspeech: Towards simultaneous and independent zero- shot speaker cloning and zero-shot language style control,
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de70324d-a926-4f93-ab8b-cc9bd948aa7b · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1674fc9b-f029-4588-b108-cfe923adc7ce · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e6503c-fabf-4965-ab72-5abcbb1424b8 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey V oicebox: Text- guided multilingual universal speech generation at scale,
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c4a353-49c8-4094-9e8c-fc49dce02c09 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20d62fa4-d936-4e00-9ab3-b8291178187d · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e3aaad8-48ce-4b1e-a5c5-1a1eb2beeeec · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae34acb-80fa-4cd5-a6f1-5152c13aca10 · outbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.