Pith. sign in

Paper Citation Record · LEDGER

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

As of 16 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2412.18748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18748 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:34:07.920960Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact5
  • verified fuzzy14
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1764f373-12b8-4dc3-9752-73911e284520 · outbound

This paper cites Neural dubber: Dubbing for videos according to scripts,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Neural dubber: Dubbing for videos according to scripts,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.436892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.758539Z digest=sha256:5ba4c7b668c22f741f5d6380e095b178be76adbb1c30469165ccee6e96da15a0

Observation ec54db63-1103-45d6-a4df-f48b18939241 · outbound

This paper cites Prosody Modeling with 3D Visual Information for Expressive Video Dubbing,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Prosody Modeling with 3D Visual Information for Expressive Video Dubbing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.423769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.763464Z digest=sha256:10915f5e8947e96ffa347634b2720026488247e81824750eef2367b6d408a4ca

Observation 532688cd-8e25-4821-9be8-8dccc87e6896 · outbound

This paper cites StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.768994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.768994Z digest=sha256:54a260527f94057a1c8d70380ea748e6b7b9b7fbfb05e2a4aff5b0c5c0a89c80

Observation d8e07d3b-dd93-47d1-a79c-27271ebc6bc6 · outbound

This paper cites V2c: Visual voice cloning,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction V2c: Visual voice cloning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.410226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.774060Z digest=sha256:f664bcd9b21b658ac5d7368dc48f2abb14ba322b8653e00bd588e3d28bdaf776

Observation 6d8cef30-e21c-4608-acba-195e24dab3aa · outbound

This paper cites From speaker to dubber: Movie dubbing with prosody and duration consistency learning,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction From speaker to dubber: Movie dubbing with prosody and duration consistency learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.397469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.778636Z digest=sha256:79c614a301516177f93e13df806b9f299eeeccb8f265a806cb7727cd2baeb0ef

Observation 3d42e29e-c8bc-407d-8f4b-c075f902d825 · outbound

This paper cites EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.783191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.783191Z digest=sha256:a61ce841e51a408872fefbe355692fa8ab0157a0b89d9ca947a9cc32a4e66b29

Observation 076a3700-5358-47d4-b854-4be31dabfaa0 · outbound

This paper cites High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.115409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.788553Z digest=sha256:259845b57b50c2e3da591096f9159531680fb1a6e9c45f656836e824b40b775d

Observation a9d75b0b-ad46-4407-adea-1b9b0a18b657 · outbound

This paper cites DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.793247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.793247Z digest=sha256:df9d244a2bd834bb7b63b06e69a4bc12c7cfed12e3b2a21f7427a90fde5f3a8c

Observation 10c9ed64-3d2a-46ff-baba-9e02046012bd · outbound

This paper cites More than words: In-the-wild visually-driven prosody for text-to-speech,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction More than words: In-the-wild visually-driven prosody for text-to-speech,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.797858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.797858Z digest=sha256:83df1a6b084dfc0c28ddf9391055c6b09c6d6ae981721fe69382090dafcbc233

Observation ec185281-1a87-4f95-95d9-0d25343fd108 · outbound

This paper cites Learning to dub movies via hierarchical prosody models,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Learning to dub movies via hierarchical prosody models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.374498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.803591Z digest=sha256:eff497670dee42e9db91af071478a5bbe1808f125fc27caadd369cef60cfc4d2

Observation eb396945-041d-4a16-9cc0-c33c93495eeb · outbound

This paper cites MCDubber: Multimodal Context-Aware Expressive Video Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction MCDubber: Multimodal Context-Aware Expressive Video Dubbing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.808156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.808156Z digest=sha256:bd5e292adda585ed15070a79639d319aa0d3c4f0be89a33b35f11263273f86f9

Observation 5b6c150b-4206-40f6-9818-d1ee99cd7d13 · outbound

This paper cites To- wards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction To- wards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.359596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.813123Z digest=sha256:a732f8d9995372ae4b58dc576a4187f449195f663ca606c9ed0c22333986e675

Observation 0d71bcdc-de81-40a7-9b32-2f5b9f125c49 · outbound

This paper cites Msstyletts: Multi-scale style modeling with hierarchical context infor- mation for expressive speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Msstyletts: Multi-scale style modeling with hierarchical context infor- mation for expressive speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.345484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.817279Z digest=sha256:361f201eac28ce8e4028da46f6cf33d68820393b5e3a934f0ddf65831fdd5ae7

Observation 688ab624-d08f-4aa7-9a80-136595686d56 · outbound

This paper cites Unsupervised multi-scale expressive speaking style modeling with hierarchical context information for audiobook speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Unsupervised multi-scale expressive speaking style modeling with hierarchical context information for audiobook speech synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.331259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.821403Z digest=sha256:30a2bae97da3fad64afd31a85e3f652f6cc8c50c83b6c7dd966654a2cdcba525

Observation edfbc176-425b-45b1-8274-b302725aed01 · outbound

This paper cites Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.825600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.825600Z digest=sha256:b2208bd6cf239db8f600b4fec01f01d2843c030cfe799604b88e0c50051d8dfe

Observation 37cda586-b2a8-43b7-a36e-3a46b127c0eb · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.830150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.830150Z digest=sha256:4a21a47c45549bed3b05a76ef354676ff761f15b7e03f1bc3b1432e1d5a292d0

Observation 0a5f69df-dd96-4e6b-99c6-b53bed2baedc · outbound

This paper cites Estimation of continuous valence and arousal levels from faces in naturalistic conditions,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Estimation of continuous valence and arousal levels from faces in naturalistic conditions,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.834265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.834265Z digest=sha256:e5e41c8288331088489d5047b22b1860132569c12e16fe96e7115b706ebaef02

Observation ce456576-3b5c-4c7f-83c9-3b088757a057 · outbound

This paper cites Towards Multi-Scale Style Control for Expressive Speech Synthesis.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Towards Multi-Scale Style Control for Expressive Speech Synthesis

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.066385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.838476Z digest=sha256:f56ef44c961a3107b28fd4348845bbe2b9c7e68f575a55a4eeaadb1e36819cd2

Observation f6da0022-bbc2-459a-b689-5e853cf2a636 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Fastspeech: Fast, robust and controllable text to speech,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.843110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.843110Z digest=sha256:80aecabfabac61a24f91a66e2cb5aa37f59342ce5cc194faf5fe3517bed59c12

Observation 60da494e-3e91-4d06-9a2b-37f6263c0626 · outbound

This paper cites Emotion-english-roberta-large,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emotion-english-roberta-large,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.281979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.847568Z digest=sha256:d4ce7da0041a9d6d0acde548abfc3152d89d33c3f8858df0a0c6d2b0b21901ea

Observation 8bd7cb49-60bb-4b7c-af54-203db6ed751b · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.851676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.851676Z digest=sha256:7e7f948aaa7e2495147bf1e23af423cacdbfe94466fa94b55cac7c30b77aa5b0

Observation bf5a25f9-586a-43d5-b5cb-77c6bbd199c2 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.856108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.856108Z digest=sha256:0d530e7b7e73a77b09503099848d9892cbde67c65f660d46e151f80055f0da15

Observation 0f6185fe-dba4-47dd-b786-bceae8be4482 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Iemocap: Interactive emotional dyadic motion capture database,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.860218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.860218Z digest=sha256:8f72d9e5afce977d5de4049fccfb251997713edeb7a78bbc750d980e4cf1e146

Observation 8f05115e-8c98-41a0-8900-5ef51a094696 · outbound

This paper cites Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.031588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.863878Z digest=sha256:55be8c7d8c52f8395efc0f057e3d3919b222e20426119a15ac1d2b5f4371865e

Observation 98ea45bd-36ee-4d5e-83fa-c11bcb31e692 · outbound

This paper cites Graph Attention Networks.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Graph Attention Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.868501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.868501Z digest=sha256:0d894386a2e48942b2dae9e43b33f5c760ec76133d4389eabcf2d295c2942f92

Observation e8d6a142-b366-41a0-9501-e5f13c3e9403 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.872926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.872926Z digest=sha256:4b8a82d62852365b482942a980422d818c79b6f1c802366e1259d8437d867d83

Observation 2f32eebf-0417-427d-917a-9b4282893bbb · outbound

This paper cites A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.241427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.877073Z digest=sha256:70496b4af51d0f444ad9e218263c3fb7603dfe9754d65c04c8a228721fc99b05

Observation bd2a6eaf-eeac-401a-9bc7-d319b635a0c3 · outbound

This paper cites Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.226734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.881464Z digest=sha256:d4db176e3d85db147fe834dc8fe61e3bf80edcd639156bc64a82859642b8f6fa

Observation 6543205f-33e0-49c2-a923-6013d7cf04da · outbound

This paper cites Out of time: automated lip sync in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Out of time: automated lip sync in the wild,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.885819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.885819Z digest=sha256:743c60184e48b627e293ebeab2ea8e81f7ac01d7c9c5a97886d7a15ccd5cec92

Observation 05c1dd58-9811-46ec-964f-baf9744ac89a · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction A lip sync expert is all you need for speech to lip generation in the wild,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.203889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.890030Z digest=sha256:d6126cc6ec4ccb483554230016cd3fdcd83b319ba3e0d7d022107bc12de38ff0

Observation 94fe868d-8b6d-4530-a256-6b0d5e4926d7 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Robust speech recognition via large-scale weak supervi- sion,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.894310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.894310Z digest=sha256:dfab9d777fe1d795d289231647df17c24c8ecafd19fc8778d9b19a2781ee2ebc

Observation d49774c0-f3e4-4156-ad62-43b9ea9708e5 · outbound

This paper cites Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.181646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.898590Z digest=sha256:bd83729e93c74a3db7f9d0b9451e3a0eb001199e718f6b51d058def4e1e3d583

Observation 0deab94c-ff93-4883-8c05-0eb527827fc1 · outbound

This paper cites Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:07.997986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.902789Z digest=sha256:28d69112a1de53c57183eb828c9c276f8f474dda9e041efdc67cbd64c7774561

Observation 6ece83f0-cdae-48d8-b707-a9faaa591b34 · outbound

This paper cites Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:07.977433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.907336Z digest=sha256:6b02fe1dae8a3b37c74080f3eaedcca03cdf73bd1934aba02ddf2b89c8ae0014

Observation 6be81104-14f3-459f-ab46-f3adbaabcb2e · outbound

This paper cites Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.911899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.911899Z digest=sha256:e9cfda61c167df82652865cdd3e4fe74fc32dd601d8bd1779242248c4e0c7beb

Observation e5ba2496-6766-4eea-8122-51f259e4ce91 · outbound

This paper cites Generative expressive conversational speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Generative expressive conversational speech synthesis,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.916601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.916601Z digest=sha256:515311123d692c1e801fc036b7269f88a838a769f2cd381fa304fe80fba73d48

Observation ee44814d-3815-4c23-8edb-3d0aa133750c · outbound

This paper cites Fctalker: Fine and coarse grained context modeling for expressive conversational speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Fctalker: Fine and coarse grained context modeling for expressive conversational speech synthesis,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.159255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:34:07.920960Z digest=sha256:1240b6090c335539af5a07067cbed761af800855e34e0b5925403d9d03fb3e51

Pith citing papers

No inbound Pith citation observations are available.