Pith. sign in

Paper Citation Record · LEDGER

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow

As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2411.19486.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19486 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:12:33.473190Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:46:13.885497Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:46:19.902908Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3a52dc70-4441-4f50-9d00-029164b8b6cb · outbound

This paper cites Generating intelligible audio speech from visual speech,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Generating intelligible audio speech from visual speech,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.261428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.247857Z digest=sha256:d759eb5f7a33e9cb23526aac36c77652d440fbe2493b4496f9acd32b03d74c71

Observation f33766a3-e298-470c-bd41-29d07099b41e · outbound

This paper cites Vid2speech: speech reconstruction from silent video,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Vid2speech: speech reconstruction from silent video,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.245629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.253545Z digest=sha256:a06cf26beadd25eaadee51357420ea814ab970f7d41d0823d5dc6eb48223ece3

Observation 9060d26f-f28f-47ae-9fa3-55058727ac35 · outbound

This paper cites Lipper: Synthesizing thy speech using multi-view lipreading,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lipper: Synthesizing thy speech using multi-view lipreading,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.229909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.258547Z digest=sha256:8600bb041a9c34b0f246e9d0dfd37c62f84cbf48f0d3dddb18395a823f125000

Observation f88e01dc-5a3c-4c9b-97f4-b59df04802e4 · outbound

This paper cites Lip-to-speech synthesis in the wild with multi-task learning,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lip-to-speech synthesis in the wild with multi-task learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.213839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.263553Z digest=sha256:90d4f67ebd45b9a70ee5e876e02e07f5738a7e848b35059fd9babe00f6a1b003

Observation 85c86e60-8d4b-4e4d-9f4c-0c71fb585642 · outbound

This paper cites Revise: Self- supervised speech resynthesis with visual input for universal and gen- eralized speech regeneration,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Revise: Self- supervised speech resynthesis with visual input for universal and gen- eralized speech regeneration,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.198199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.268842Z digest=sha256:d408427bd6283695d4c3001eac614b743d7e866a0ba99fa7ce5ce7379322b081

Observation fc17b11f-8f59-4d2d-8922-ed03051886b1 · outbound

This paper cites Intelligible lip-to-speech synthesis with speech units,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Intelligible lip-to-speech synthesis with speech units,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.182760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.274673Z digest=sha256:b9db756697cd65d2a600ae58eb12f35d492fff045b80205ab24025ad71db057f

Observation 5c4a94b1-356f-4a84-bc08-a0ff09d6ed75 · outbound

This paper cites Let there be sound: Reconstructing high quality speech from silent videos,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Let there be sound: Reconstructing high quality speech from silent videos,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.166967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.280189Z digest=sha256:9af06e1a6e2996cbd1fb8edee5ad216026dd4f2733c7148201f06045d6472047

Observation b9b2ff64-435b-45ae-b718-1781e3ebbf07 · outbound

This paper cites Uni-dubbing: Zero-shot speech synthesis from visual articulation,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Uni-dubbing: Zero-shot speech synthesis from visual articulation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.152076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.284839Z digest=sha256:1d32fc8c8f26a627b23e67122eec5271296e7e693b1521501dbd9b5d9fc07c42

Observation 6f3e9730-ee8e-4399-991a-278c771f47b3 · outbound

This paper cites Towards accurate lip-to-speech synthesis in-the-wild,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Towards accurate lip-to-speech synthesis in-the-wild,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.135977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.289628Z digest=sha256:d6d67707c0e66d0979caacce247ff446198b1e14ac5620070fa405ddb720fab9

Observation cf03be55-be98-4811-bd82-de329b49f0e5 · outbound

This paper cites Lipvoicer: Generating speech from silent videos guided by lip reading,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lipvoicer: Generating speech from silent videos guided by lip reading,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.120071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.295176Z digest=sha256:a4710e7df2b75dba6be960593ddaf56537bd9b708b597e0e061a2971efdd9118

Observation 17f0792a-cfc8-401d-b40f-584739977d75 · outbound

This paper cites Svts: Scalable video-to-speech synthesis,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Svts: Scalable video-to-speech synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.104592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.299854Z digest=sha256:4f309a35e3d63656c07be32f8b2a1565f70f6a54af0274442c6a8b99b49c3b26

Observation df7f3d8a-97ab-4371-a337-4af3d0fd2639 · outbound

This paper cites Diffv2s: Diffusion-based video-to- speech synthesis with vision-guided speaker embedding,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Diffv2s: Diffusion-based video-to- speech synthesis with vision-guided speaker embedding,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.088978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.304216Z digest=sha256:21e5bd5b2e3e18d48a1eea68b0ee6ef8866ceed471787bfcbfbef55a84f13de3

Observation 50a6a718-123a-4bfd-92da-95919813e28f · outbound

This paper cites Flow-based unconstrained lip to speech generation,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Flow-based unconstrained lip to speech generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.073087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.308873Z digest=sha256:c269751564b8e2c5b6d65a2bac2fc87c185b3dc3f5cdfc9ff3f8b26b4d28dfe3

Observation f2e1d024-3a0d-484b-a20f-17a96bcda6f1 · outbound

This paper cites Lip to speech synthesis with visual context attentional GAN,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lip to speech synthesis with visual context attentional GAN,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.056592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.313801Z digest=sha256:780e55cb25ac27441dc42aa0a79225becd4a35d5dc00b6178bdee7963b115725

Observation 20878490-e349-4577-911d-2a88fb062de2 · outbound

This paper cites End-to-end video-to-speech synthesis using generative adversarial networks,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow End-to-end video-to-speech synthesis using generative adversarial networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.041054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.318369Z digest=sha256:fa999f20b9f41f99fbd9ce62c4c14b907dac7a16826c31ba178b86bbd30e56ab

Observation 79a59ced-fd91-4f5b-9129-905ceca0a2a3 · outbound

This paper cites Lip-to-speech synthesis for arbitrary speakers in the wild,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lip-to-speech synthesis for arbitrary speakers in the wild,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.025105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.322882Z digest=sha256:143b7999e1c806791468939f58d3643fa7dc8db8c7e035015da5db111d1084a0

Observation 70744e8d-c8af-4a01-91fe-52f5a2867e08 · outbound

This paper cites Speech resynthesis from discrete disen- tangled self-supervised representations,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Speech resynthesis from discrete disen- tangled self-supervised representations,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:34.009934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.327445Z digest=sha256:4eab066e0cb2fe34b196b5b7f7a27e2181906f1be10ec41662bfff55ec06d402

Observation bcc2bcd6-7355-43c5-81e4-b77af5845005 · outbound

This paper cites Neural analysis and synthesis: Reconstructing speech from self-supervised representations,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Neural analysis and synthesis: Reconstructing speech from self-supervised representations,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.993345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.331919Z digest=sha256:fd913c68eb58bb78ce6ef475a4f196a396eb5f152d050b27a0513e9f212e0324

Observation af802c87-7747-45dc-8a01-3add204d1411 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Flow straight and fast: Learning to generate and transfer data with rectified flow,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.977527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.336439Z digest=sha256:f7718a5494cf6bd3eca02641119d3394d04fa986257164af7808defd35e872d8

Observation 49449a35-076d-4b15-aa1f-12ccf82a7258 · outbound

This paper cites Scalable diffusion models with transformers,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Scalable diffusion models with transformers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.961351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.340798Z digest=sha256:ceb1c3d4ca170b3fb2a7c9bd64850a7375a42ff78e505e7c4e412f0343d3e288

Observation 1192bcb5-5255-4d0a-9e64-5145a254d781 · outbound

This paper cites Distinguishing homophenes using multi-head visual-audio memory for lip reading,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Distinguishing homophenes using multi-head visual-audio memory for lip reading,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.944494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.345191Z digest=sha256:29df8391961008abe0519555c696612877232627a7aba1b236cb3f9b1659c245

Observation 47afa1c7-b44f-4e01-b338-3f4ebb3879b9 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.928574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.349961Z digest=sha256:51719af516429912a84eddc04fec2314b5e4862b260a4368c21da2e0ba482fa0

Observation f24615b4-f340-4133-9466-39aa64e3034d · outbound

This paper cites On gener- ative spoken language modeling from raw audio,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow On gener- ative spoken language modeling from raw audio,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.911242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.354356Z digest=sha256:ddef1167cab5bda1e755116d8968959aa0e10032cd86cf633e277895e26206d1

Observation 88d69b2b-ebf3-4c37-8831-f75943c4217e · outbound

This paper cites Textless speech emotion conversion using discrete and decomposed representations,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Textless speech emotion conversion using discrete and decomposed representations,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.892658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.358879Z digest=sha256:cc0372fad327cd17b64575880f7729ee2b5996d232a3112e7a48f0fdff35a6ae

Observation 8aa8e910-2427-41b0-af4d-1624a013542d · outbound

This paper cites Textless unit-to-unit training for many-to-many multilingual speech-to-speech translation,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Textless unit-to-unit training for many-to-many multilingual speech-to-speech translation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.876102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.363590Z digest=sha256:2b01e6f0aeafec61ed975dd9c5a03ac101c5b0857af2ade91df6ad7c705574a7

Observation 3783685d-3e5e-4c10-afb4-243ba6c247f4 · outbound

This paper cites High fidelity speech regeneration with application to speech enhancement,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow High fidelity speech regeneration with application to speech enhancement,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.858755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.368165Z digest=sha256:ae2c459575a18bd165827599507304f3723cee60c419d687d9e7933c0c35c338

Observation a30ab050-b743-4997-9a6f-349a9e707786 · outbound

This paper cites Dddm-vc: Decoupled denoising diffusion models with disentangled representation and prior mixup for verified robust voice conversion,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Dddm-vc: Decoupled denoising diffusion models with disentangled representation and prior mixup for verified robust voice conversion,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.841716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.372707Z digest=sha256:dbe95ee16b68474dc412b27e34849ca94d415a9b4ff67614ccd4c9547ba84ebc

Observation 4fbef8df-9de0-43ac-bbb5-be38500d398b · outbound

This paper cites Yet another algorithm for pitch tracking,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Yet another algorithm for pitch tracking,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.824523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.377286Z digest=sha256:c58774407a96580b1c9f42742720c9e5632c3febaa7ef3e4a9c42aecc6f81cf9

Observation cdf380bd-2f47-480f-b1df-13399f2db80e · outbound

This paper cites Neural discrete representation learning,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Neural discrete representation learning,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:12:33.382572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:12:33.382572Z digest=sha256:0d70efd5189d6b4b05de1ffc406d9eed5f42c2c59d27e8a0cc77923082fdafe8

Observation 377651fd-ec5f-4fe5-a240-229d10be024c · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Generalized end-to-end loss for speaker verification,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.796138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.387485Z digest=sha256:ef425d898ff14eb7cbfb5511d15a491e7470ff099b8b1113384651e63e25fd6c

Observation 32ccd70b-3e79-475a-b6ab-c0a492de142f · outbound

This paper cites More than words: In-the-wild visually-driven prosody for text-to-speech,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow More than words: In-the-wild visually-driven prosody for text-to-speech,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.779779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.392335Z digest=sha256:69fdd208561d3d80a1b498805c463d12da6c71d4810569066f438fe56fc05f8a

Observation c11d1386-bca4-4495-ac0f-9fc5ec4b3005 · outbound

This paper cites Learning audio- visual speech representation by masked multimodal cluster prediction,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Learning audio- visual speech representation by masked multimodal cluster prediction,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.759482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.396879Z digest=sha256:33fce4dfe071b8c32c5934998c4203a001122956bdd0e478eb3a463e7dbe5ee4

Observation bbedbd3f-3923-4e8d-b3d4-1ff51109c997 · outbound

This paper cites Hear your face: Face-based voice conversion with f0 estimation,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Hear your face: Face-based voice conversion with f0 estimation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.739865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.401635Z digest=sha256:3dec4acd9ab276aea1ffebd5231f8bccf5031a263ae0c53dabbb08dfd29e6d76

Observation 039a756d-a5ed-48c2-96b5-4e2e6d973cf2 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Conformer: Convolution-augmented transformer for speech recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.723202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.406952Z digest=sha256:5044e5573c6218103855e962fc7c37a01cc9f4584aebdaf7e332345bac3efc15

Observation 443bb694-5d9a-4410-bc7a-d6d1df1065c0 · outbound

This paper cites Attention is all you need,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Attention is all you need,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.706655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.412338Z digest=sha256:404bceef97307d6a1dca97b3d7a93aabcfd913a3f11cf4a91dd2acd87696309f

Observation 19b04568-a170-432d-8863-3b32cb0ea69b · outbound

This paper cites Classifier-free diffusion guidance,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Classifier-free diffusion guidance,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.690301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.417118Z digest=sha256:a2e7ea76dae08e7ec3a5e42944ea597b24c042f0a59c9037180b6050660feeec

Observation 70df1950-c2cd-4403-a39e-693c831e3b1b · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow LRS3-TED: a large-scale dataset for visual speech recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T10:12:33.421735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:12:33.421735Z digest=sha256:c0dfefb365b54765832015d5292e8e93f4398e19978b66d914b4ef1ec2afa623

Observation add5692a-6a67-442f-9582-9ec0366d2e3e · outbound

This paper cites Lip reading sentences in the wild,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Lip reading sentences in the wild,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.674303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.427717Z digest=sha256:a423fa32c293f1cfcd8d5fdeb5a6f2d3e0d253de4922c71c63d5ea5ed3872d11

Observation 9005b27f-a0cf-4342-ad62-62d65f32b1cc · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Utmos: Utokyo-sarulab system for voicemos challenge 2022,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T10:12:33.433471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:12:33.433471Z digest=sha256:8cbfb0b110e53606e893dc05ee63526557f92d0875aea3bef445b94b409ed282

Observation a6c6ea4b-c525-403b-980f-61ba2576ffdd · outbound

This paper cites End-to-end audio-visual speech recognition with conformers,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow End-to-end audio-visual speech recognition with conformers,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.647654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.438302Z digest=sha256:a48118dd7be3fbeb07857718d0c1197b71e00d145fe21fa7f0a2342fd7dbcd38

Observation 266fbebb-f5b6-4116-8a0f-bafba30b8061 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow A lip sync expert is all you need for speech to lip generation in the wild,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.632166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.442996Z digest=sha256:dc84086232accaa3d47962058ce10afb043e786e3719fdeffa05f66221342f2e

Observation 23dc3476-4f05-4df1-9bd8-274781cdd4b0 · outbound

This paper cites Out of time: automated lip sync in the wild,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Out of time: automated lip sync in the wild,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.615548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.447720Z digest=sha256:84efbf4d41124444bd43d3fe99f29ae7db8ab3f270b50d29180a82d2d7c6c8d6

Observation d4a1cfcb-b12d-425c-9b45-e74291bb88f5 · outbound

This paper cites Retinaface: Single-shot multi-level face localisation in the wild,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Retinaface: Single-shot multi-level face localisation in the wild,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.598487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.452607Z digest=sha256:020d9f0180a6a97c627460945712d313c06268416530ee16bd733697f8e2d6fc

Observation 85fffd24-6d68-4ffd-8cc9-c0451d2c9f5c · outbound

This paper cites How far are we from solving the 2d & 3d face alignment problem?(and a dataset of 230,000 3d facial landmarks),.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow How far are we from solving the 2d & 3d face alignment problem?(and a dataset of 230,000 3d facial landmarks),

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.581806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.457725Z digest=sha256:0bf6c8693a56c9d85ea2e71c5dd5854c226a7a41507b91e5f43eb0ca86593d77

Observation 5ea883b3-42b9-4f45-93a8-752b5f193dfb · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Scaling rectified flow transformers for high-resolution image synthesis,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.565533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.462868Z digest=sha256:7dae31112ced2f3e99ef53a1e0b1586e797db0e7e74b2726385ae3ff57f974a5

Observation afd33c5b-96a8-4e98-8321-f0f3fd67f996 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.548946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.468268Z digest=sha256:33c5ce0802a9667ca29a9fe20dae45a94f47f345d96849f3386d93b8936ec34c

Observation e4a61223-6897-4f6e-bdbd-50a92965b291 · outbound

This paper cites Denoising diffusion implicit models,.

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow Denoising diffusion implicit models,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:12:33.532560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T10:12:33.473190Z digest=sha256:3539d37f4bff823306f56515bed8224a4303ff8310a16d0909c743d0f24fd07d

Pith citing papers

Observation af46798e-68a2-4bfc-a683-538e14f5b39d · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:46:19.957973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:46:13.885497Z digest=sha256:46593cbd560e962b67584403d5d42a3fa53fa3c720d376eab18e83c006901d04