Pith. sign in

REVIEW 2 cited by

Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.16321 v1 pith:QNHWB7KG submitted 2024-02-26 cs.SD cs.AIcs.LGeess.AS

classification cs.SDcs.AIcs.LGeess.AS
keywords speechenhancementqualitytrainingcleanestimationmodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expensive for label collection. To solve this problem, we propose VQScore, a self-supervised metric for evaluating speech based on the quantization error of a vector-quantized-variational autoencoder (VQ-VAE). The training of VQ-VAE relies on clean speech; hence, large quantization errors can be expected when the speech is distorted. To further improve correlation with real quality scores, domain knowledge of speech processing is incorporated into the model design. We found that the vector quantization mechanism could also be used for self-supervised speech enhancement (SE) model training. To improve the robustness of the encoder for SE, a novel self-distillation mechanism combined with adversarial training is introduced. In summary, the proposed speech quality estimation method and enhancement models require only clean speech for training without any label requirements. Experimental results show that the proposed VQScore and enhancement model are competitive with supervised baselines. The code will be released after publication.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement

    cs.SD 2026-03 conditional novelty 6.0 of 10

    Using LLM-generated text descriptions of enhanced speech converted to sentiment scores as PPO rewards improves PESQ, STOI, and neural quality scores over supervised and DNSMOS-reward baselines on AVSEC-4.

  2. GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling

    eess.AS 2025-02 conditional novelty 6.0 of 10

    GenSE enhances speech by first denoising semantic tokens with a language model and then generating acoustic tokens from a single-quantizer codec, reporting higher DNSMOS, speaker similarity, and lower WER than prior systems.

Pith tools