Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:12:33.473190Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2411.19486.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:12:33.473190Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:46:13.885497Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T17:46:19.902908Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3a52dc70-4441-4f50-9d00-029164b8b6cb · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Generating intelligible audio speech from visual speech,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f33766a3-e298-470c-bd41-29d07099b41e · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Vid2speech: speech reconstruction from silent video,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9060d26f-f28f-47ae-9fa3-55058727ac35 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lipper: Synthesizing thy speech using multi-view lipreading,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f88e01dc-5a3c-4c9b-97f4-b59df04802e4 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lip-to-speech synthesis in the wild with multi-task learning,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 85c86e60-8d4b-4e4d-9f4c-0c71fb585642 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Revise: Self- supervised speech resynthesis with visual input for universal and gen- eralized speech regeneration,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fc17b11f-8f59-4d2d-8922-ed03051886b1 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Intelligible lip-to-speech synthesis with speech units,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5c4a94b1-356f-4a84-bc08-a0ff09d6ed75 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Let there be sound: Reconstructing high quality speech from silent videos,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b9b2ff64-435b-45ae-b718-1781e3ebbf07 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Uni-dubbing: Zero-shot speech synthesis from visual articulation,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6f3e9730-ee8e-4399-991a-278c771f47b3 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Towards accurate lip-to-speech synthesis in-the-wild,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cf03be55-be98-4811-bd82-de329b49f0e5 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lipvoicer: Generating speech from silent videos guided by lip reading,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 17f0792a-cfc8-401d-b40f-584739977d75 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Svts: Scalable video-to-speech synthesis,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation df7f3d8a-97ab-4371-a337-4af3d0fd2639 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Diffv2s: Diffusion-based video-to- speech synthesis with vision-guided speaker embedding,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 50a6a718-123a-4bfd-92da-95919813e28f · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Flow-based unconstrained lip to speech generation,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f2e1d024-3a0d-484b-a20f-17a96bcda6f1 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lip to speech synthesis with visual context attentional GAN,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 20878490-e349-4577-911d-2a88fb062de2 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow End-to-end video-to-speech synthesis using generative adversarial networks,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 79a59ced-fd91-4f5b-9129-905ceca0a2a3 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lip-to-speech synthesis for arbitrary speakers in the wild,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 70744e8d-c8af-4a01-91fe-52f5a2867e08 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Speech resynthesis from discrete disen- tangled self-supervised representations,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bcc2bcd6-7355-43c5-81e4-b77af5845005 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Neural analysis and synthesis: Reconstructing speech from self-supervised representations,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation af802c87-7747-45dc-8a01-3add204d1411 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Flow straight and fast: Learning to generate and transfer data with rectified flow,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 49449a35-076d-4b15-aa1f-12ccf82a7258 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Scalable diffusion models with transformers,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1192bcb5-5255-4d0a-9e64-5145a254d781 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Distinguishing homophenes using multi-head visual-audio memory for lip reading,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 47afa1c7-b44f-4e01-b338-3f4ebb3879b9 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Hubert: Self-supervised speech representation learning by masked prediction of hidden units,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f24615b4-f340-4133-9466-39aa64e3034d · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow On gener- ative spoken language modeling from raw audio,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 88d69b2b-ebf3-4c37-8831-f75943c4217e · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Textless speech emotion conversion using discrete and decomposed representations,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8aa8e910-2427-41b0-af4d-1624a013542d · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Textless unit-to-unit training for many-to-many multilingual speech-to-speech translation,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3783685d-3e5e-4c10-afb4-243ba6c247f4 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow High fidelity speech regeneration with application to speech enhancement,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a30ab050-b743-4997-9a6f-349a9e707786 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Dddm-vc: Decoupled denoising diffusion models with disentangled representation and prior mixup for verified robust voice conversion,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4fbef8df-9de0-43ac-bbb5-be38500d398b · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Yet another algorithm for pitch tracking,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cdf380bd-2f47-480f-b1df-13399f2db80e · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Neural discrete representation learning,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377651fd-ec5f-4fe5-a240-229d10be024c · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Generalized end-to-end loss for speaker verification,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 32ccd70b-3e79-475a-b6ab-c0a492de142f · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow More than words: In-the-wild visually-driven prosody for text-to-speech,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c11d1386-bca4-4495-ac0f-9fc5ec4b3005 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Learning audio- visual speech representation by masked multimodal cluster prediction,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bbedbd3f-3923-4e8d-b3d4-1ff51109c997 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Hear your face: Face-based voice conversion with f0 estimation,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 039a756d-a5ed-48c2-96b5-4e2e6d973cf2 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Conformer: Convolution-augmented transformer for speech recognition,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 443bb694-5d9a-4410-bc7a-d6d1df1065c0 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Attention is all you need,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 19b04568-a170-432d-8863-3b32cb0ea69b · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Classifier-free diffusion guidance,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 70df1950-c2cd-4403-a39e-693c831e3b1b · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow LRS3-TED: a large-scale dataset for visual speech recognition
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add5692a-6a67-442f-9582-9ec0366d2e3e · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lip reading sentences in the wild,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9005b27f-a0cf-4342-ad62-62d65f32b1cc · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Utmos: Utokyo-sarulab system for voicemos challenge 2022,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c6ea4b-c525-403b-980f-61ba2576ffdd · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow End-to-end audio-visual speech recognition with conformers,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 266fbebb-f5b6-4116-8a0f-bafba30b8061 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow A lip sync expert is all you need for speech to lip generation in the wild,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 23dc3476-4f05-4df1-9bd8-274781cdd4b0 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Out of time: automated lip sync in the wild,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d4a1cfcb-b12d-425c-9b45-e74291bb88f5 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Retinaface: Single-shot multi-level face localisation in the wild,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 85fffd24-6d68-4ffd-8cc9-c0451d2c9f5c · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow How far are we from solving the 2d & 3d face alignment problem?(and a dataset of 230,000 3d facial landmarks),
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5ea883b3-42b9-4f45-93a8-752b5f193dfb · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Scaling rectified flow transformers for high-resolution image synthesis,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation afd33c5b-96a8-4e98-8321-f0f3fd67f996 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e4a61223-6897-4f6e-bdbd-50a92965b291 · outbound
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Denoising diffusion implicit models,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation af46798e-68a2-4bfc-a683-538e14f5b39d · inbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.