Pith. sign in

REVIEW 1 cited by

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.19486 v2 pith:5MUVCBK7 submitted 2024-11-29 cs.CV cs.SDeess.AS

classification cs.CVcs.SDeess.AS
keywords speechv2sflowattributesdatasetsdirectlyflowgeneratemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we introduce V2SFlow, a novel Video-to-Speech (V2S) framework designed to generate natural and intelligible speech directly from silent talking face videos. While recent V2S systems have shown promising results on constrained datasets with limited speakers and vocabularies, their performance often degrades on real-world, unconstrained datasets due to the inherent variability and complexity of speech signals. To address these challenges, we decompose the speech signal into manageable subspaces (content, pitch, and speaker information), each representing distinct speech attributes, and predict them directly from the visual input. To generate coherent and realistic speech from these predicted attributes, we employ a rectified flow matching decoder built on a Transformer architecture, which models efficient probabilistic pathways from random noise to the target speech distribution. Extensive experiments demonstrate that V2SFlow significantly outperforms state-of-the-art methods, even surpassing the naturalness of ground truth utterances. Code and models are available at: https://github.com/kaistmm/V2SFlow

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

    cs.MM 2025-07 conditional novelty 6.0 of 10

    A single multimodal language model can generate intelligible speech and synchronized background audio jointly from a silent video, transcript, and reference voice, outperforming separately concatenated speech and audi...

Pith tools