Four-step SR-FD fine-tuning cuts VoxCPM2's Seed-TTS English WER from 2.23% to 1.41%, beating the ten-step baseline at 1.74% with no inference-time cost.
ARCHI-TTS: A flow-matching-based text-to-speech model with self-supervised semantic aligner and accelerated inference,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Fr\'echet Distance Loss on Speech Representations for Text-to-Speech Synthesis
Four-step SR-FD fine-tuning cuts VoxCPM2's Seed-TTS English WER from 2.23% to 1.41%, beating the ten-step baseline at 1.74% with no inference-time cost.