Pith. sign in

Paper Citation Record · LEDGER

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video

As of 10 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2501.19258.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.19258 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:50:17.390376Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:50:17.213254Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-09T20:50:17.551912Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a1cde93-cc78-44a7-a999-ef99aa50718d · outbound

This paper cites Despite the high-quality output, the prosody of the generated speech sometimes be- comes inappropriate for the context.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Despite the high-quality output, the prosody of the generated speech sometimes be- comes inappropriate for the context

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.994955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.207333Z digest=sha256:834ad6377ab09f1136be2ecc263d5329b1252ce175ae391c4506cff220cd598b

Observation 46264d7c-d2c7-418f-8164-aa4dcda14f9b · outbound

This paper cites VisualSpeech: Enhancing Prosody Modeling in TTS Using Video.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video VisualSpeech: Enhancing Prosody Modeling in TTS Using Video

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T20:50:17.559362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.213254Z digest=sha256:b4382cf0e749bc928a85cc0a9513ef4125a2f6a772ffb4ab91d5dd94ad9bca8b

Observation 614772b2-901a-40fa-901d-0bc9060921b3 · outbound

This paper cites Experimental Setup To assess the impact of visual information on speech generation performance, a dataset containing both diverse prosodic and vi- sual data is essential.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Experimental Setup To assess the impact of visual information on speech generation performance, a dataset containing both diverse prosodic and vi- sual data is essential

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.978937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.218991Z digest=sha256:69f6e492e41f69a8b54d61372ee7662ab24b5879831dcc4430e51e3b234de188

Observation b65af53c-99b6-4490-8153-84449ea9e73f · outbound

This paper cites Using two dis- tinct video feature extractors, we demonstrate that these visual features encapsulate prosodic information.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Using two dis- tinct video feature extractors, we demonstrate that these visual features encapsulate prosodic information

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.944262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.229624Z digest=sha256:f23763627e4bdca707ae532dad5ee7ed1aec945a3326cfb0f3289a9b4e70f5e1

Observation 5116b954-f3ef-435a-bc36-ba1e01d24db1 · outbound

This paper cites FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.254499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.254499Z digest=sha256:a57e697d70e6bee4b15be892069aa4b98f104fd0e3cb345d9efea648cba77df5

Observation 9638055a-4317-40f2-8b01-ae94bdd9e1ea · outbound

This paper cites Deep mixture density networks for acous- tic modeling in statistical parametric speech synthesis,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Deep mixture density networks for acous- tic modeling in statistical parametric speech synthesis,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.925404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.234534Z digest=sha256:02461d1594f92e765a3db9660156cf590e81608afb5c53ce0286225399c858ab

Observation 6ef20369-d217-483d-b630-cd32c8fcf63c · outbound

This paper cites an unresolved cited work.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:50:17.961376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.224287Z digest=sha256:0d4135eb4a6b3154680f4911cbf3e959a0f31af3e002ddbc6732650123085f6e

Observation 88fe2c5f-9cc4-4a1b-b191-bdac925e8e76 · outbound

This paper cites Unidirectional long short-term memory re- current neural network with recurrent output layer for low-latency speech synthesis,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Unidirectional long short-term memory re- current neural network with recurrent output layer for low-latency speech synthesis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.906544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.239311Z digest=sha256:2616213125c4a563844aec63ae1987bdb044873235dd9a3a410d2adfd4b03808

Observation 7bc7dcca-935c-4562-86a2-3d334cd4375a · outbound

This paper cites NaturalSpeech: End-to-end text-to- speech synthesis with human-level quality,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video NaturalSpeech: End-to-end text-to- speech synthesis with human-level quality,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.888796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.244348Z digest=sha256:1c853102784fe77c0b08f899da91d1868dafcbb930276d2db66dd716d6adbdf6

Observation 9d9b9273-93e7-4dd2-bef1-3602fcccba1b · outbound

This paper cites FastSpeech: Fast, robust and controllable text to speech,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastSpeech: Fast, robust and controllable text to speech,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.871652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.249595Z digest=sha256:260899e07004c0eb8357b5b5682ca99048b623b192fa77374fe0089ebe6d6c7b

Observation 6bf8f256-6251-4a71-921f-55f767396b09 · outbound

This paper cites On granularity of prosodic representations in expressive text-to-speech,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video On granularity of prosodic representations in expressive text-to-speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.854616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.260279Z digest=sha256:c7a24340d3670900df88e1c4d5d5ae7834cc77b11ce0602b037b3471eed991a7

Observation b3401b66-54ee-4800-a4b7-f3fed282a958 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.265052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.265052Z digest=sha256:cc0736295bdb868315cf41aa80d7089a1af4d2d7219b1f9cec0ffaa63efe77e5

Observation 0c0bffc1-dbce-4135-85f1-968810ebaf38 · outbound

This paper cites Chive: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Chive: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.838085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.275553Z digest=sha256:1d40fbe35f660deb2f645699ce872aa8bd47f0fbfa7f662a5266beb43deb2d47

Observation d0eea874-c8af-416c-b8dc-55541afae4b7 · outbound

This paper cites Mellotron: Mul- tispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Mellotron: Mul- tispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.280420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.280420Z digest=sha256:d788b1a56be4cc83a14afae3b4e2293397de1dc344331a0151aa540b552a4bb8

Observation 94ffe127-0a7d-4bff-af9a-4a3d08338dee · outbound

This paper cites ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.284979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.284979Z digest=sha256:7ca18b7c5ca6327e6f75360235ac05438a03df25f3c571da76c2f29c86c42aca

Observation 46b094d9-a8b9-4d94-819b-d3989ab742f9 · outbound

This paper cites The LJ Speech Dataset,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video The LJ Speech Dataset,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.289999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.289999Z digest=sha256:3f109e913bafb0aaf4322d7b254db4f9645531a23f30b66a1beba3adac3c7544

Observation ed07cf1c-e7a6-4559-b713-9ea96e7bb626 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.294686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.294686Z digest=sha256:66b5ad4dfe66390b77f89bbc3db44a8270091c25a7cb705a265d9167422d3158

Observation 9a95dddd-a602-4907-a7dd-a1b6023ca39c · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.801073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.300516Z digest=sha256:9ec8bb36026b39e2d78d9e643967e9597eb4b40e5eb35f2b3a3b1ac2a069f2a7

Observation 3fab26cc-05a9-4223-80fa-deed08cd241a · outbound

This paper cites Condensed movies: Story based retrieval with contextual embeddings,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Condensed movies: Story based retrieval with contextual embeddings,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.784934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.305703Z digest=sha256:87c8bb42b099a3043e24e0dbbcd35b838ba0e36893af74ec1df62fb3a7432f9a

Observation c26d7058-5e60-4d59-8f14-ba95f6a3a9f1 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Robust speech recognition via large-scale weak supervision,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.769915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.310629Z digest=sha256:ee7e0a7a0fadf0f2c8af2c7d6e7b62355888b31f79a6aef150ef4eb7b7fa1169

Observation f03ce37d-88da-4aec-a4cd-b3e2861d5cc6 · outbound

This paper cites CMD+: A D.I.Y . Audiovisual Dataset for Multi- Speaker TTS,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video CMD+: A D.I.Y . Audiovisual Dataset for Multi- Speaker TTS,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.754642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.316038Z digest=sha256:7267d4ab3eb80c0f7be175b869784fba718f36b8e59c0234a82920499af8df33

Observation ec659a45-2c03-4b10-bd2e-179f4d146a1d · outbound

This paper cites The Sound Demixing Challenge 2023 $\unicode{x2013}$ Music Demixing Track.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video The Sound Demixing Challenge 2023 $\unicode{x2013}$ Music Demixing Track

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.320971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.320971Z digest=sha256:8d80779249f613460f0fd498c5f6633d4b8c9c04be4703d125c0b4ea5758059e

Observation 77c18112-6f8c-482e-b89b-e2c661ee6d81 · outbound

This paper cites Resemblyzer,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Resemblyzer,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.738232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.326282Z digest=sha256:a5bdb29b50245ee4dddd34919da5e4c9fde680c2ea419f50adc0e0e80f722c39

Observation ba086c40-95b3-425b-8253-8c5412fe4d0d · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.331087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.331087Z digest=sha256:3fcfd63d78b439e33f128566276a2d14a5b295a8fb841cd0984fcc57dcb69b96

Observation d23f2617-43ca-43af-9b6f-d7c6a5031199 · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.721365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.336097Z digest=sha256:099ce3788fbef7b67e83774fe7fdd9b57139ab9364b20cc89ba9cf0a85380373

Observation 1fc38aa8-fb53-4dd2-9a6f-7a97c3cbc26a · outbound

This paper cites Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.705037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.341247Z digest=sha256:3cf44238f20867325b04307c5a554aa9a5ac8b3b8ecefa02f9b152874b6dd96a

Observation 8ac8f403-5036-45c5-b5a3-c09bb07886bd · outbound

This paper cites SDR– half-baked or well done?.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video SDR– half-baked or well done?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.686765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.346185Z digest=sha256:3174595d9215f1c0b006102db9d4763d317a048f66f94e06229dd6614dcef495

Observation 0f45bc67-4145-4ae8-ac5e-f4f90fd1eec3 · outbound

This paper cites Torchaudio-Squim: Reference-less speech quality and intelligibility measures in torchaudio,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Torchaudio-Squim: Reference-less speech quality and intelligibility measures in torchaudio,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.670742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.350971Z digest=sha256:637866d15a19d4c7d6290abdc2bb1f5218f04f3035b09b0aa2249aae1391cda1

Observation d918af71-c765-4722-8453-a3ce1a8793bc · outbound

This paper cites Omnivore: A single model for many visual modalities,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Omnivore: A single model for many visual modalities,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.651564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.356165Z digest=sha256:49a2c80c2537927a7208ea5b6d86bba47080988768840df84af2f96e0215a3d3

Observation a12931ea-8921-4708-afa1-3bc265c6a23e · outbound

This paper cites Deep residual learning for image recognition,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Deep residual learning for image recognition,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.360860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.360860Z digest=sha256:cfdf630d38e3edc0ef650fd01fa36a389d22e5c312264a3062decab1e03cd3cd

Observation de6c5db6-6bc8-4274-87e0-efa7ea411131 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Adam: A Method for Stochastic Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.365484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.365484Z digest=sha256:401548bc11e8ceadad3a3ecef5f71651099698515119bcde3b81b0a75f87511c

Observation dfe24b21-8c07-4d94-97e1-6b4b4d3a5c46 · outbound

This paper cites FastPitch: Parallel text-to-speech with pitch pre- diction,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastPitch: Parallel text-to-speech with pitch pre- diction,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.624813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.370920Z digest=sha256:ef35ac0751372093187acdccbfa1ce3f8438fedeb71f5d560b135535c61469a7

Observation e3e3b893-f7fc-4721-9143-ac3f359b7a27 · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Sonicvisionlm: Playing sound with vision language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.608653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.375699Z digest=sha256:2d86de7090f05596ae07e1d31641adbef0bedc79fa394055917de08adca784e5

Observation 4eeae4f1-ef67-4de6-8687-25c0e93965a3 · outbound

This paper cites End-to-end video-to-speech synthesis using gener- ative adversarial networks,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video End-to-end video-to-speech synthesis using gener- ative adversarial networks,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.593035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.380409Z digest=sha256:b2a182b1b418949e77935b4a80e3e96dbc618a0e9860d412924488c16eeb8634

Observation dce0b177-93b1-4bad-ac6b-1a923b5ebdff · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Intelligible Lip-to-Speech Synthesis with Speech Units

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.385174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.385174Z digest=sha256:d8c6ef697cefc43336017d8cf33b85c30a35b1f2a8103525540a440ab4049ef1

Observation f82f58ef-f399-4ae2-8df7-c33c5e669b64 · outbound

This paper cites Camp: a two-stage approach to modelling prosody in context,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Camp: a two-stage approach to modelling prosody in context,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.577819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.390376Z digest=sha256:1e038b0f21c467020998fee1a5d0e0ac2000123d3330944bbe3305b066e7a02a

Pith citing papers

Observation 46264d7c-d2c7-418f-8164-aa4dcda14f9b · inbound

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video cites this paper.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video VisualSpeech: Enhancing Prosody Modeling in TTS Using Video

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T20:50:17.559362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.213254Z digest=sha256:b4382cf0e749bc928a85cc0a9513ef4125a2f6a772ffb4ab91d5dd94ad9bca8b