Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:10.491851Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2505.16279.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:10.491851Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:07.961535Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T15:07:10.886803Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c74c4cec-69f9-42aa-861a-f36edf271657 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Exist- ing dubbing methods can be categorized into two groups, each focusing on learning different styles of key prior information to generate high-quality voices
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 74e62b26-d73e-4307-998b-7010f449b1a2 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 907a7eee-cea2-48df-a5a4-d728b32c66ec · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Datasets Emilia is a comprehensive multilingual speech generation dataset containing a total of 101,654 hours of speech data across six languages [21]
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 375117c2-b77b-41b1-86d7-a80476fd686f · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing To as- sess pronunciation accuracy, we use Word Error Rate (WER) with Whisper-V3[24] as the ASR model
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 498d685b-d814-4216-89ac-e165952d658e · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Additionally, we have de- veloped a movie dubbing dataset with multi-type annotations to enhance movie understanding and improve dubbing quality
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0aea5d8d-5aee-43db-9ce9-3407e0333e0f · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing V2c: Vi- sual voice cloning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fc319574-6ad6-4352-ad12-923d5cb6d196 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing More than words: In-the-wild visually-driven prosody for text-to-speech,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 575950f6-6b37-4f81-b5a9-6886c6669eb4 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Generalized end-to-end loss for speaker verification,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6475ee3-f567-450d-b6ad-8bbabb0b15de · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Learning to dub movies via hierarchical prosody models,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52d242fa-daf6-4206-9957-d93433bf20b2 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Neu- ral dubber: Dubbing for videos according to scripts,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc5635ba-34aa-4c19-942f-ce1a908c9175 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Imaginary voice: Face- styled diffusion model for text-to-speech,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0912692c-d95d-4188-8855-7853c0c477ef · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Mcdubber: Multimodal context-aware expressive video dubbing,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c9175e7-1513-408d-8ac9-b6bb9cec84cd · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Audiopedia: Audio qa with knowledge,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e501b4aa-ab4b-48e9-bb12-9dc7b57355a7 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8be3954-69be-438e-9ad5-4432531ef75d · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f78254c-f256-495b-aee1-8f76d4b0dadb · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Learning to dub movies via hierarchical prosody models,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 37ce9657-f384-4ed4-a705-5791653bab47 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a0d21ad-481e-4ab7-aa36-f256148f54d8 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing From speaker to dubber: Movie dubbing with prosody and duration consistency learning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a2ca5ff3-e452-482c-a4cc-9c0d122a62cd · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Visual instruction tuning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb820020-97f2-4de1-9232-86c7858f292d · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b22e3ad7-cc45-48fe-b3c1-c8e53935a9a1 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 297f42d7-bd54-44b4-aed5-644aafba4cb8 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Flow matching for generative modeling,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e275e094-90a4-4712-9e20-fe0c7af3ff27 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing An audio-visual corpus for speech perception and automatic speech recognition,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d6cc985-0e82-459d-850e-bb737c7a2e05 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Exploring the limits of transfer learning with a unified text-to-text transformer,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0cdbc5a6-b3fb-4b09-aebe-f4d579bae2a7 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Learning transferable visual models from natural language supervision,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 710d353d-232f-42a6-875b-bfde1a713ea6 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing DiVE: Dit-based video generation with enhanced control,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f3fe19ef-6c80-4bbd-9835-209e6f266201 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e30b8f5-a9a6-4481-8b01-4d2924fa4b96 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing V2C: Visual Voice Cloning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 092bb8b5-818c-403a-9762-39b6ad937b9a · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Robust Speech Recognition via Large-Scale Weak Supervision
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3fd5fe-5aa3-4ab2-b700-eb822cd43a8a · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9edd3850-dac4-4e3a-a2ff-8b7a03a45d5b · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Tem- poral modeling matters: A novel temporal emotional modeling approach for speech emotion recognition,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c2c5a5c-a7cc-448a-aa58-3a770be52bba · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Out of time: automated lip sync in the wild,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afeebc99-ff03-4cc8-8e88-c14dd6001225 · outbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Available: https://openreview.net/forum?id= PqvMRDCJT9t
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74e62b26-d73e-4307-998b-7010f449b1a2 · inbound
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.