Pith. sign in

REVIEW 4 cited by

Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03491 v1 pith:SS2BNS3P submitted 2023-12-06 cs.LG cs.SDeess.AS

classification cs.LGcs.SDeess.AS
keywords diffusionsynthesismodelspriorbridge-ttsdeterministicgenerationinformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In text-to-speech (TTS) synthesis, diffusion models have achieved promising generation quality. However, because of the pre-defined data-to-noise diffusion process, their prior distribution is restricted to a noisy representation, which provides little information of the generation target. In this work, we present a novel TTS system, Bridge-TTS, making the first attempt to substitute the noisy Gaussian prior in established diffusion-based TTS methods with a clean and deterministic one, which provides strong structural information of the target. Specifically, we leverage the latent representation obtained from text input as our prior, and build a fully tractable Schrodinger bridge between it and the ground-truth mel-spectrogram, leading to a data-to-data process. Moreover, the tractability and flexibility of our formulation allow us to empirically study the design spaces such as noise schedules, as well as to develop stochastic and deterministic samplers. Experimental results on the LJ-Speech dataset illustrate the effectiveness of our method in terms of both synthesis quality and sampling efficiency, significantly outperforming our diffusion counterpart Grad-TTS in 50-step/1000-step synthesis and strong fast TTS models in few-step scenarios. Project page: https://bridge-tts.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation

    cs.CV 2026-01 conditional novelty 6.0 of 10

    A stochastic bridge model treats video object removal as video-to-video translation, starting from the source video rather than Gaussian noise, with adaptive mask modulation and a new benchmark.

  2. Incorporating Pre-trained Diffusion Models in Solving the Schr\"odinger Bridge Problem

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Schrödinger Bridge models can be trained with diffusion-style mean, terminus, and flow-matching losses and initialized from pretrained diffusion models, improving image generation and unpaired translation.

  3. Schr\"odinger Bridge Mamba for One-Step Speech Enhancement

    cs.SD 2025-10 conditional novelty 5.0 of 10

    A Mamba-based speech enhancer trained with Schrödinger Bridge objectives produces strong denoising and dereverberation in one inference step with a low real-time factor.

  4. Few-step Adversarial Schr\"{o}dinger Bridge for Generative Speech Enhancement

    cs.SD 2025-06 conditional novelty 4.0 of 10

    Adding an adversarial GAN objective to a Schrödinger Bridge speech enhancement model enables high-quality denoising and dereverberation at one to four sampling steps, surpassing slower baselines on full-band benchmarks.

Pith tools