Pith. sign in

Paper Citation Record · LEDGER

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

As of 11 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2412.18748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18748 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:34:07.920960Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact5
  • verified fuzzy14
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1764f373-12b8-4dc3-9752-73911e284520 · outbound

This paper cites Neural dubber: Dubbing for videos according to scripts,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Neural dubber: Dubbing for videos according to scripts,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.436892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.758539Z digest=sha256:f7d24e961c1478aaf9b50ef9285eb62d7523431ffc766d2db26e2bdf6e6ca84a

Observation ec54db63-1103-45d6-a4df-f48b18939241 · outbound

This paper cites Prosody Modeling with 3D Visual Information for Expressive Video Dubbing,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Prosody Modeling with 3D Visual Information for Expressive Video Dubbing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.423769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.763464Z digest=sha256:6d05fd5fcac4d32aeb5c2586c79611ead15911315319d52afb0bb8443d200343

Observation 532688cd-8e25-4821-9be8-8dccc87e6896 · outbound

This paper cites StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.768994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.768994Z digest=sha256:00ba9ead5da6661e2887856e0554a9c6fc76ab69bd5964e854bc2ce451c2f12b

Observation d8e07d3b-dd93-47d1-a79c-27271ebc6bc6 · outbound

This paper cites V2c: Visual voice cloning,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction V2c: Visual voice cloning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.410226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.774060Z digest=sha256:26e186ccb959d49c2c97a37edfa45ff461c2c94f61bb7595df5d90e3b4460a94

Observation 6d8cef30-e21c-4608-acba-195e24dab3aa · outbound

This paper cites From speaker to dubber: Movie dubbing with prosody and duration consistency learning,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction From speaker to dubber: Movie dubbing with prosody and duration consistency learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.397469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.778636Z digest=sha256:f880fab2c7a73560c03414e393d050ee3a6eb98206eb27c781c112fa081edddb

Observation 3d42e29e-c8bc-407d-8f4b-c075f902d825 · outbound

This paper cites EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.783191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.783191Z digest=sha256:90d2970510faa6b7feaf77e943839e6d3acd5409b087f9afaf2546d74d8bd8f4

Observation 076a3700-5358-47d4-b854-4be31dabfaa0 · outbound

This paper cites High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.115409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.788553Z digest=sha256:4f1bcddbf03e47bf3baaa730e732630b2bb90698908f6b61efe5e6151859bb71

Observation a9d75b0b-ad46-4407-adea-1b9b0a18b657 · outbound

This paper cites DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.793247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.793247Z digest=sha256:30ed31ccade7100554ddd5ae614b8f7c3843c82d6ce6fcc96a4677de1775cb58

Observation 10c9ed64-3d2a-46ff-baba-9e02046012bd · outbound

This paper cites More than words: In-the-wild visually-driven prosody for text-to-speech,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction More than words: In-the-wild visually-driven prosody for text-to-speech,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.797858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.797858Z digest=sha256:c8d7d189df1b143e60f72bcba73bcdc88872371464f04301ba8d8c4ba1a6dbca

Observation ec185281-1a87-4f95-95d9-0d25343fd108 · outbound

This paper cites Learning to dub movies via hierarchical prosody models,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Learning to dub movies via hierarchical prosody models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.374498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.803591Z digest=sha256:ddaf3241a414c5d2b4b2d71376f7da5d56109a8aac7eca26b4cfe557360e13f0

Observation eb396945-041d-4a16-9cc0-c33c93495eeb · outbound

This paper cites MCDubber: Multimodal Context-Aware Expressive Video Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction MCDubber: Multimodal Context-Aware Expressive Video Dubbing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.808156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.808156Z digest=sha256:071c67968bf38089d2050950f5a1c3710b781ec03324568e0fbdaf015d27d624

Observation 5b6c150b-4206-40f6-9818-d1ee99cd7d13 · outbound

This paper cites To- wards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction To- wards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.359596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.813123Z digest=sha256:847224ba4768a6c5613de56c8672d9fac08c8bdf2d818d7a647427d13e4bcbe8

Observation 0d71bcdc-de81-40a7-9b32-2f5b9f125c49 · outbound

This paper cites Msstyletts: Multi-scale style modeling with hierarchical context infor- mation for expressive speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Msstyletts: Multi-scale style modeling with hierarchical context infor- mation for expressive speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.345484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.817279Z digest=sha256:2832a7d01d8f0422ec50d32cdb90e420c556f8bdaf8c439f6f0cbe33bae45592

Observation 688ab624-d08f-4aa7-9a80-136595686d56 · outbound

This paper cites Unsupervised multi-scale expressive speaking style modeling with hierarchical context information for audiobook speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Unsupervised multi-scale expressive speaking style modeling with hierarchical context information for audiobook speech synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.331259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.821403Z digest=sha256:ec4f7dd19c77d919407551de4f2b1faa2d8887c70e8bd9e633b51aca748c2b37

Observation edfbc176-425b-45b1-8274-b302725aed01 · outbound

This paper cites Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.825600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.825600Z digest=sha256:e9fb3be4b6badd0c9e53fce727ac570369eb04300b93984325f1ee38ea8f8e6d

Observation 37cda586-b2a8-43b7-a36e-3a46b127c0eb · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.830150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.830150Z digest=sha256:663e246c0cd79c1c99e019eef49add1450057da2317f0ac409e22244cabf425e

Observation 0a5f69df-dd96-4e6b-99c6-b53bed2baedc · outbound

This paper cites Estimation of continuous valence and arousal levels from faces in naturalistic conditions,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Estimation of continuous valence and arousal levels from faces in naturalistic conditions,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.834265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.834265Z digest=sha256:20598b93970869da76412a910d835320279b98c80fc023728f833331b8278f77

Observation ce456576-3b5c-4c7f-83c9-3b088757a057 · outbound

This paper cites Towards Multi-Scale Style Control for Expressive Speech Synthesis.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Towards Multi-Scale Style Control for Expressive Speech Synthesis

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.066385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.838476Z digest=sha256:214d04c7ab0a838bd0dd3551027b046473335fadbd7aedb70373e4f070abb95c

Observation f6da0022-bbc2-459a-b689-5e853cf2a636 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Fastspeech: Fast, robust and controllable text to speech,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.843110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.843110Z digest=sha256:65af9c46e5a759f016f8c2960bb3b24f496f3f2a153b7dd476ba884f42efff4e

Observation 60da494e-3e91-4d06-9a2b-37f6263c0626 · outbound

This paper cites Emotion-english-roberta-large,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emotion-english-roberta-large,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.281979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.847568Z digest=sha256:79a9b5e9480e7013761fc4b0b7304337fdb937aba2ec106b3535dd4d5bd29986

Observation 8bd7cb49-60bb-4b7c-af54-203db6ed751b · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.851676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.851676Z digest=sha256:95402ee84640866c4636b4207c16dc27e6c7389021f17a2f0e4287829a1b1203

Observation bf5a25f9-586a-43d5-b5cb-77c6bbd199c2 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.856108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.856108Z digest=sha256:bbaba712e3c017c773cc6457f5cd16a8d72699a73c6ecc54e7701f41b45f0375

Observation 0f6185fe-dba4-47dd-b786-bceae8be4482 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Iemocap: Interactive emotional dyadic motion capture database,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.860218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.860218Z digest=sha256:076fd9ae8f0056a4a85b02c68887b0e6550ae2d9dbe843db38097a01387dadc1

Observation 8f05115e-8c98-41a0-8900-5ef51a094696 · outbound

This paper cites Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.031588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.863878Z digest=sha256:f3340db176381c40e444b6c28dc9423932992d59d9c853f59796c8210ea8fd3d

Observation 98ea45bd-36ee-4d5e-83fa-c11bcb31e692 · outbound

This paper cites Graph Attention Networks.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Graph Attention Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.868501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.868501Z digest=sha256:7be409ef3f1708df1ed5b9467ddaf997dc10be471e03bcb34c591d4f226d8f59

Observation e8d6a142-b366-41a0-9501-e5f13c3e9403 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.872926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.872926Z digest=sha256:ea710b5c454887da00c39f11c71ce260abf9577bd060eee45222fffefe9a1339

Observation 2f32eebf-0417-427d-917a-9b4282893bbb · outbound

This paper cites A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.241427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.877073Z digest=sha256:7ab5e3d32fae112a136553456d3b10f32d96d47781f9974a16f00a6e04a99cbb

Observation bd2a6eaf-eeac-401a-9bc7-d319b635a0c3 · outbound

This paper cites Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.226734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.881464Z digest=sha256:f55fc4adcf00c031dad53007a3b0771ec49c7672d0c779530173c3fc349e9fbd

Observation 6543205f-33e0-49c2-a923-6013d7cf04da · outbound

This paper cites Out of time: automated lip sync in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Out of time: automated lip sync in the wild,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.885819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.885819Z digest=sha256:09e0e05bafcce5d5e1182a499ba1be03c2696ebd4cf5c5ffee569e0ab0c26cb3

Observation 05c1dd58-9811-46ec-964f-baf9744ac89a · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction A lip sync expert is all you need for speech to lip generation in the wild,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.203889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.890030Z digest=sha256:57b01444ef0eba69457ce890f50e39ff33b47635b08edd94a99850e59bdbb4ba

Observation 94fe868d-8b6d-4530-a256-6b0d5e4926d7 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Robust speech recognition via large-scale weak supervi- sion,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.894310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.894310Z digest=sha256:b95247e583290b49fcff4d9f411e16d7a80075ef386705fc2e7c759c86bbf9e8

Observation d49774c0-f3e4-4156-ad62-43b9ea9708e5 · outbound

This paper cites Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.181646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.898590Z digest=sha256:0c62301ea1ec10774022fafc8f4bb942bbb28b3f1615ad714a4b669a9e64212e

Observation 0deab94c-ff93-4883-8c05-0eb527827fc1 · outbound

This paper cites Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:07.997986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.902789Z digest=sha256:ca7b5da9af7bc7b0dd126928662875c61fb151c8b2fa384c3f5e42d162cc3b15

Observation 6ece83f0-cdae-48d8-b707-a9faaa591b34 · outbound

This paper cites Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:07.977433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.907336Z digest=sha256:debbc68ea83c3e263cdabfef92d78fd3a4e7e1c55a4f2f4a9ad620c7fb3db025

Observation 6be81104-14f3-459f-ab46-f3adbaabcb2e · outbound

This paper cites Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.911899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.911899Z digest=sha256:14a8bf4ed8e71f49943184ae9eb5848215f5236c79cd6603c025fb19d79c8a08

Observation e5ba2496-6766-4eea-8122-51f259e4ce91 · outbound

This paper cites Generative expressive conversational speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Generative expressive conversational speech synthesis,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.916601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.916601Z digest=sha256:f2c1f95b3d330bce911b2c31550e043d6ade0ea0e78160d1ff2cb6a10cb19b09

Observation ee44814d-3815-4c23-8edb-3d0aa133750c · outbound

This paper cites Fctalker: Fine and coarse grained context modeling for expressive conversational speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Fctalker: Fine and coarse grained context modeling for expressive conversational speech synthesis,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.159255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.920960Z digest=sha256:2e5ed07ba815afdf3a09e2bbc7bf6caa9255e49622ea96652519517889e570f3

Pith citing papers

No inbound Pith citation observations are available.