Pith. sign in

REVIEW 4 cited by

AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1611.09207 v1 pith:INO4DLMQ submitted 2016-11-28 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords humanautomosratersspeechcorrelationsmodelqualitysynthesized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep recurrent neural network whose inputs consist solely of a raw waveform. Our best models provide utterance-level estimates of MOS only moderately inferior to sampled human ratings, as shown by Pearson and Spearman correlations. When multiple utterances are scored and averaged, a scenario common in synthesizer quality assessment, AutoMOS achieves correlations approaching those of human raters. The AutoMOS model has a number of applications, such as the ability to explore the parameter space of a speech synthesizer without requiring a human-in-the-loop.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

    cs.SD 2025-02 unverdicted novelty 6.0 of 10

    Unified no-reference models assess audio aesthetics across speech, music, and sound via four perceptual axes and achieve performance comparable or superior to human mean opinion scores.

  2. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

  3. SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit

    cs.SD 2025-05 conditional novelty 5.0 of 10

    SHEET provides unified training and evaluation for MOS predictors, and its benchmark shows WavLM large and XLS-R 1b are the best SSL backbones for SSL-MOS on the tested datasets.

  4. SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction

    cs.SD 2025-06 reject novelty 4.0 of 10

    SALF-MOS, a compact U-Net-style model using frozen wav2vec features, claims state-of-the-art MOS prediction on four benchmarks with only 1,574 parameters.

Pith tools