Pith. sign in

REVIEW 3 cited by

Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02243 v1 pith:YYY4FODZ submitted 2024-07-02 cs.CL cs.SDeess.AS

Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

classification cs.CL cs.SDeess.AS
keywords inferencereversespeechoptimizationrobustnesszero-shotgeneratedhuman
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In this paper, we propose reverse inference optimization (RIO), a simple and effective method designed to enhance the robustness of autoregressive-model-based zero-shot text-to-speech (TTS) systems using reinforcement learning from human feedback (RLHF). To assess the quality of speech produced by the TTS system without human annotations, RIO introduces a novel concept termed as reverse inference based on the Bayesian principle, which suggests that a high-quality generated speech should be able to be used as a prompt for subsequent generation using the same TTS model. By leveraging reverse inference as the standard to select exemplars used in RLHF from the speech samples generated by the TTS system itself, RIO steers the subsequent optimization towards a direction of enhancing the TTS robustness. The RIO framework, comprising sampling, automatic annotating, and learning, obviates the need for a reward model or pairwise preference data, and significantly improves the stability of zero-shot TTS performance by reducing the discrepancies between training and inference conditions. Our experimental results verify that RIO can effectively improve both subjective and objective metrics, including mean opinion scores, word error rates, and speaker similarity. Remarkably, RIO can also diminish the incidence of bad outputs to nearly zero percent, rivalling the robustness when using ground-truth speech as the prompt.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection

    cs.SD 2026-03 unverdicted novelty 7.0

    MSpoof-TTS improves zero-shot discrete speech synthesis by integrating multi-resolution token-based spoof detection into a hierarchical decoding process that prunes low-quality candidates.

  2. Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment

    cs.CL 2026-07 conditional novelty 6.0

    BoN TTS verifier rankings reverse across ASR families; same-family pairs recover 2–3× more oracle headroom, and cross-family rank ensembles give the most robust WER gains.

  3. DDPO-VC: Speaker De-Identification via Diffusion Denoising Policy Optimization

    eess.AS 2026-06 unverdicted novelty 6.0

    DDPO-VC applies diffusion denoising policy optimization with dual-teacher rewards to improve speaker de-identification while preserving cognitive utility on dementia speech benchmarks.