Pith. sign in

REVIEW 1 cited by

SpeechLMScore: Evaluating speech generation using speech language model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.04559 v1 pith:2CEISFTF submitted 2022-12-08 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords speechevaluationhumangenerationmetricspeechlmscoreannotationaverage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While human evaluation is the most reliable metric for evaluating speech generation systems, it is generally costly and time-consuming. Previous studies on automatic speech quality assessment address the problem by predicting human evaluation scores with machine learning models. However, they rely on supervised learning and thus suffer from high annotation costs and domain-shift problems. We propose SpeechLMScore, an unsupervised metric to evaluate generated speech using a speech-language model. SpeechLMScore computes the average log-probability of a speech signal by mapping it into discrete tokens and measures the average probability of generating the sequence of tokens. Therefore, it does not require human annotation and is a highly scalable framework. Evaluation results demonstrate that the proposed metric shows a promising correlation with human evaluation scores on different speech generation tasks including voice conversion, text-to-speech, and speech enhancement.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Model as Loss: A Self-Consistent Training Paradigm

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Using the model's own encoder as a feature loss improves perceptual quality and iterative stability of a speech enhancement model compared with a WavLM-based loss.

Pith tools