Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:34:07.920960Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2412.18748.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:34:07.920960Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1764f373-12b8-4dc3-9752-73911e284520 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Neural dubber: Dubbing for videos according to scripts,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ec54db63-1103-45d6-a4df-f48b18939241 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Prosody Modeling with 3D Visual Information for Expressive Video Dubbing,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 532688cd-8e25-4821-9be8-8dccc87e6896 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e07d3b-dd93-47d1-a79c-27271ebc6bc6 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction V2c: Visual voice cloning,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6d8cef30-e21c-4608-acba-195e24dab3aa · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction From speaker to dubber: Movie dubbing with prosody and duration consistency learning,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3d42e29e-c8bc-407d-8f4b-c075f902d825 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 076a3700-5358-47d4-b854-4be31dabfaa0 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a9d75b0b-ad46-4407-adea-1b9b0a18b657 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c9ed64-3d2a-46ff-baba-9e02046012bd · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction More than words: In-the-wild visually-driven prosody for text-to-speech,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec185281-1a87-4f95-95d9-0d25343fd108 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Learning to dub movies via hierarchical prosody models,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eb396945-041d-4a16-9cc0-c33c93495eeb · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction MCDubber: Multimodal Context-Aware Expressive Video Dubbing
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b6c150b-4206-40f6-9818-d1ee99cd7d13 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction To- wards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0d71bcdc-de81-40a7-9b32-2f5b9f125c49 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Msstyletts: Multi-scale style modeling with hierarchical context infor- mation for expressive speech synthesis,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 688ab624-d08f-4aa7-9a80-136595686d56 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Unsupervised multi-scale expressive speaking style modeling with hierarchical context information for audiobook speech synthesis,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation edfbc176-425b-45b1-8274-b302725aed01 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37cda586-b2a8-43b7-a36e-3a46b127c0eb · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5f69df-dd96-4e6b-99c6-b53bed2baedc · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Estimation of continuous valence and arousal levels from faces in naturalistic conditions,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce456576-3b5c-4c7f-83c9-3b088757a057 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Towards Multi-Scale Style Control for Expressive Speech Synthesis
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f6da0022-bbc2-459a-b689-5e853cf2a636 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Fastspeech: Fast, robust and controllable text to speech,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60da494e-3e91-4d06-9a2b-37f6263c0626 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emotion-english-roberta-large,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8bd7cb49-60bb-4b7c-af54-203db6ed751b · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5a25f9-586a-43d5-b5cb-77c6bbd199c2 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6185fe-dba4-47dd-b786-bceae8be4482 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Iemocap: Interactive emotional dyadic motion capture database,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f05115e-8c98-41a0-8900-5ef51a094696 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 98ea45bd-36ee-4d5e-83fa-c11bcb31e692 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Graph Attention Networks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8d6a142-b366-41a0-9501-e5f13c3e9403 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f32eebf-0417-427d-917a-9b4282893bbb · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd2a6eaf-eeac-401a-9bc7-d319b635a0c3 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6543205f-33e0-49c2-a923-6013d7cf04da · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Out of time: automated lip sync in the wild,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05c1dd58-9811-46ec-964f-baf9744ac89a · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction A lip sync expert is all you need for speech to lip generation in the wild,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 94fe868d-8b6d-4530-a256-6b0d5e4926d7 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Robust speech recognition via large-scale weak supervi- sion,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49774c0-f3e4-4156-ad62-43b9ea9708e5 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0deab94c-ff93-4883-8c05-0eb527827fc1 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ece83f0-cdae-48d8-b707-a9faaa591b34 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6be81104-14f3-459f-ab46-f3adbaabcb2e · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ba2496-6766-4eea-8122-51f259e4ce91 · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Generative expressive conversational speech synthesis,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee44814d-3815-4c23-8edb-3d0aa133750c · outbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Fctalker: Fine and coarse grained context modeling for expressive conversational speech synthesis,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.