Pith. sign in

Paper Citation Record · LEDGER

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2501.19258.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.19258 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:50:17.390376Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:50:17.213254Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-09T20:50:17.551912Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a1cde93-cc78-44a7-a999-ef99aa50718d · outbound

This paper cites Despite the high-quality output, the prosody of the generated speech sometimes be- comes inappropriate for the context.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Despite the high-quality output, the prosody of the generated speech sometimes be- comes inappropriate for the context

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.994955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.207333Z digest=sha256:b4c0ff7699fef438e626096abd37b0ceb26f619c6161f894e5f5e6f33e2c7df5

Observation 46264d7c-d2c7-418f-8164-aa4dcda14f9b · outbound

This paper cites VisualSpeech: Enhancing Prosody Modeling in TTS Using Video.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video VisualSpeech: Enhancing Prosody Modeling in TTS Using Video

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T20:50:17.559362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.213254Z digest=sha256:4361afff9e1d5b11cffb56238bc83819c8239c5b4a383c70188a4f99514ee084

Observation 614772b2-901a-40fa-901d-0bc9060921b3 · outbound

This paper cites Experimental Setup To assess the impact of visual information on speech generation performance, a dataset containing both diverse prosodic and vi- sual data is essential.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Experimental Setup To assess the impact of visual information on speech generation performance, a dataset containing both diverse prosodic and vi- sual data is essential

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.978937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.218991Z digest=sha256:2910eb2a961e995a3c7a83fe21a31faa3a0e40e92e7294a6d77a9a20aec2f7cd

Observation b65af53c-99b6-4490-8153-84449ea9e73f · outbound

This paper cites Using two dis- tinct video feature extractors, we demonstrate that these visual features encapsulate prosodic information.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Using two dis- tinct video feature extractors, we demonstrate that these visual features encapsulate prosodic information

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.944262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.229624Z digest=sha256:1d1a92b84cc1d8bf907fada4ebfa00f87dc60dac7afc330b1eaffc42c8bdc98b

Observation 5116b954-f3ef-435a-bc36-ba1e01d24db1 · outbound

This paper cites FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.254499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.254499Z digest=sha256:8a53b5047954673c3d8ac10559f9aba6dcd0393f3f164e727c199959f49b9af7

Observation 9638055a-4317-40f2-8b01-ae94bdd9e1ea · outbound

This paper cites Deep mixture density networks for acous- tic modeling in statistical parametric speech synthesis,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Deep mixture density networks for acous- tic modeling in statistical parametric speech synthesis,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.925404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.234534Z digest=sha256:9871a5f6aa81631f3d24f562b3c61d324c85b9492e1789237ae2a32199365191

Observation 6ef20369-d217-483d-b630-cd32c8fcf63c · outbound

This paper cites an unresolved cited work.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:50:17.961376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.224287Z digest=sha256:a2dbbc346115b87149c9d3924bfd28f1f5c9d2b88aa9759a1f764064056e617e

Observation 88fe2c5f-9cc4-4a1b-b191-bdac925e8e76 · outbound

This paper cites Unidirectional long short-term memory re- current neural network with recurrent output layer for low-latency speech synthesis,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Unidirectional long short-term memory re- current neural network with recurrent output layer for low-latency speech synthesis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.906544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.239311Z digest=sha256:a9741d74ad5d3d5820df7f147ad9d89b286723ac9ad1262994e506f5154b758f

Observation 7bc7dcca-935c-4562-86a2-3d334cd4375a · outbound

This paper cites NaturalSpeech: End-to-end text-to- speech synthesis with human-level quality,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video NaturalSpeech: End-to-end text-to- speech synthesis with human-level quality,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.888796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.244348Z digest=sha256:a49af9e0233ec1e6142e524dd927aa0563b90320138c9d20ddaaead4b3f4a0e5

Observation 9d9b9273-93e7-4dd2-bef1-3602fcccba1b · outbound

This paper cites FastSpeech: Fast, robust and controllable text to speech,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastSpeech: Fast, robust and controllable text to speech,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.871652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.249595Z digest=sha256:8784f05b07e154205adceff8b736e75a548cfad32a1d604840cf330a2ba45d40

Observation 6bf8f256-6251-4a71-921f-55f767396b09 · outbound

This paper cites On granularity of prosodic representations in expressive text-to-speech,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video On granularity of prosodic representations in expressive text-to-speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.854616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.260279Z digest=sha256:448a81a69a2887383c84be5fb00e36c627f13152c22bae5f78df0b8134ae6c4e

Observation b3401b66-54ee-4800-a4b7-f3fed282a958 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.265052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.265052Z digest=sha256:a72817977d14475cf0034d43c052596ce14e3fc3219fb8e3299a87ebe90f6df7

Observation 0c0bffc1-dbce-4135-85f1-968810ebaf38 · outbound

This paper cites Chive: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Chive: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.838085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.275553Z digest=sha256:c3345f4ce9d4a91e090d98cf076b71604316ed932fe62fc5fb75d09b9c328ea1

Observation d0eea874-c8af-416c-b8dc-55541afae4b7 · outbound

This paper cites Mellotron: Mul- tispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Mellotron: Mul- tispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.280420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.280420Z digest=sha256:a6aa9ebf29d3180d16686e345f2e7e6340f50df48d51947eb18791183a008c1d

Observation 94ffe127-0a7d-4bff-af9a-4a3d08338dee · outbound

This paper cites ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.284979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.284979Z digest=sha256:b106a6a2514038e13c616c478007bc9cc6421ff2e31d75ac0e583711af109e12

Observation 46b094d9-a8b9-4d94-819b-d3989ab742f9 · outbound

This paper cites The LJ Speech Dataset,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video The LJ Speech Dataset,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.289999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.289999Z digest=sha256:cf05fd5a003488952ffe15451d879b83fa7af56bcddae50e2dda20527bc9b007

Observation ed07cf1c-e7a6-4559-b713-9ea96e7bb626 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.294686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.294686Z digest=sha256:01ebd1672db045f0a14305c9f221e766af492550095bfe7aa49329bf77b52d17

Observation 9a95dddd-a602-4907-a7dd-a1b6023ca39c · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.801073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.300516Z digest=sha256:6c5efe4cc86437edc5642fc59d01c2a2e3e4b7f056e1d8a0d3cff257c3941a54

Observation 3fab26cc-05a9-4223-80fa-deed08cd241a · outbound

This paper cites Condensed movies: Story based retrieval with contextual embeddings,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Condensed movies: Story based retrieval with contextual embeddings,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.784934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.305703Z digest=sha256:469508fa36e564506e934832fb639e70fce6391fd3b4fb49154754f64d3925e8

Observation c26d7058-5e60-4d59-8f14-ba95f6a3a9f1 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Robust speech recognition via large-scale weak supervision,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.769915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.310629Z digest=sha256:e6a930e83321dadd8e6fd49c12c12726b932aef7c9074aadec458fa9fd6389e8

Observation f03ce37d-88da-4aec-a4cd-b3e2861d5cc6 · outbound

This paper cites CMD+: A D.I.Y . Audiovisual Dataset for Multi- Speaker TTS,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video CMD+: A D.I.Y . Audiovisual Dataset for Multi- Speaker TTS,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.754642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.316038Z digest=sha256:e97efd167366a730718c407ef246f9d136ab95ff465a4cff5f7e10ea2012deb7

Observation ec659a45-2c03-4b10-bd2e-179f4d146a1d · outbound

This paper cites The Sound Demixing Challenge 2023 $\unicode{x2013}$ Music Demixing Track.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video The Sound Demixing Challenge 2023 $\unicode{x2013}$ Music Demixing Track

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.320971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.320971Z digest=sha256:5fcbabafe9129c1ace20344ef9a8d82bd26b69bd21c48fc310c5bdd888a601e8

Observation 77c18112-6f8c-482e-b89b-e2c661ee6d81 · outbound

This paper cites Resemblyzer,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Resemblyzer,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.738232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.326282Z digest=sha256:a8b4d6d66353db5531a4ff6f2c2a5c3d687975e51824ab023060cce3fa4a0dde

Observation ba086c40-95b3-425b-8253-8c5412fe4d0d · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.331087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.331087Z digest=sha256:f6406402c8761e116b32eb3526b8f4b7331f04c6059ab4040ae58da35fd6dba4

Observation d23f2617-43ca-43af-9b6f-d7c6a5031199 · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.721365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.336097Z digest=sha256:d13ca198e226ec9ab372691a170dd09bc584e3812404a8a4a5afb24f7b72f8e8

Observation 1fc38aa8-fb53-4dd2-9a6f-7a97c3cbc26a · outbound

This paper cites Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.705037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.341247Z digest=sha256:0e03ec33a1f1e81cd218522edcbafc1ad79fff956d3954efffcd4057343b0c75

Observation 8ac8f403-5036-45c5-b5a3-c09bb07886bd · outbound

This paper cites SDR– half-baked or well done?.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video SDR– half-baked or well done?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.686765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.346185Z digest=sha256:d122b02bb55dfbdcaa8ec8ce498bc7b56d36552f449dc85ff7d4dfd9d9af70fa

Observation 0f45bc67-4145-4ae8-ac5e-f4f90fd1eec3 · outbound

This paper cites Torchaudio-Squim: Reference-less speech quality and intelligibility measures in torchaudio,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Torchaudio-Squim: Reference-less speech quality and intelligibility measures in torchaudio,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.670742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.350971Z digest=sha256:e881cfa7c80876576c05468766c88d5577314ce72cd4c1909a2df3518d9493a0

Observation d918af71-c765-4722-8453-a3ce1a8793bc · outbound

This paper cites Omnivore: A single model for many visual modalities,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Omnivore: A single model for many visual modalities,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.651564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.356165Z digest=sha256:cdfa0d70a54efcf3e240e99d2bcdca892cf372573f0c1ff3b1d23fd70affdc38

Observation a12931ea-8921-4708-afa1-3bc265c6a23e · outbound

This paper cites Deep residual learning for image recognition,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Deep residual learning for image recognition,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.360860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.360860Z digest=sha256:c6639425bab3fb04e7af9b96999ad72cae873feabd7f2dace38e148d6668e7c2

Observation de6c5db6-6bc8-4274-87e0-efa7ea411131 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Adam: A Method for Stochastic Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.365484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.365484Z digest=sha256:380b1e045bdb8a454939323436517cf4ae6eab3f9d2f9d06feeecba43b00e814

Observation dfe24b21-8c07-4d94-97e1-6b4b4d3a5c46 · outbound

This paper cites FastPitch: Parallel text-to-speech with pitch pre- diction,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastPitch: Parallel text-to-speech with pitch pre- diction,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.624813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.370920Z digest=sha256:e248b54714d021c4d07e90a485c74c91796eaff9f956f2414992735626dfac1e

Observation e3e3b893-f7fc-4721-9143-ac3f359b7a27 · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Sonicvisionlm: Playing sound with vision language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.608653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.375699Z digest=sha256:6c56696f6ae017bb58af179a46b391dd56fd14c5c4740330400eaae72170d774

Observation 4eeae4f1-ef67-4de6-8687-25c0e93965a3 · outbound

This paper cites End-to-end video-to-speech synthesis using gener- ative adversarial networks,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video End-to-end video-to-speech synthesis using gener- ative adversarial networks,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.593035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.380409Z digest=sha256:669d7bebc5259d0d38c70f4422f319fa07e2c291ee54508123a6e3d05e25b98f

Observation dce0b177-93b1-4bad-ac6b-1a923b5ebdff · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Intelligible Lip-to-Speech Synthesis with Speech Units

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.385174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.385174Z digest=sha256:83a5bc50fde51afe183248316edccd7224c13fa8592759918fca5ad036aafb3a

Observation f82f58ef-f399-4ae2-8df7-c33c5e669b64 · outbound

This paper cites Camp: a two-stage approach to modelling prosody in context,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Camp: a two-stage approach to modelling prosody in context,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.577819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.390376Z digest=sha256:24a5a82187303cf98dd12563862dcb844a8dae61d8abb503271526e997c648d4

Pith citing papers

Observation 46264d7c-d2c7-418f-8164-aa4dcda14f9b · inbound

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video cites this paper.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video VisualSpeech: Enhancing Prosody Modeling in TTS Using Video

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T20:50:17.559362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T20:50:17.213254Z digest=sha256:4361afff9e1d5b11cffb56238bc83819c8239c5b4a383c70188a4f99514ee084