Pith. sign in

REVIEW 14 cited by

NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.09494 v1 pith:GVTB2DHF submitted 2021-04-19 eess.AS cs.AIcs.LGcs.SD

classification eess.AScs.AIcs.LGcs.SD
keywords modelspeechqualitydatasetsnisqaoverallpredictiontrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present an update to the NISQA speech quality prediction model that is focused on distortions that occur in communication networks. In contrast to the previous version, the model is trained end-to-end and the time-dependency modelling and time-pooling is achieved through a Self-Attention mechanism. Besides overall speech quality, the model also predicts the four speech quality dimensions Noisiness, Coloration, Discontinuity, and Loudness, and in this way gives more insight into the cause of a quality degradation. Furthermore, new datasets with over 13,000 speech files were created for training and validation of the model. The model was finally tested on a new, live-talking test dataset that contains recordings of real telephone calls. Overall, NISQA was trained and evaluated on 81 datasets from different sources and showed to provide reliable predictions also for unknown speech samples. The code, model weights, and datasets are open-sourced.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CallScreenBench: Benchmarking On-Device Models as Phone Secretaries

    cs.CR 2026-08 conditional novelty 7.0 of 10

    A new phone-secretary benchmark shows quality scaling with capability, no triage scaling once degenerate baselines are subtracted, and more capable models relaying scam callback numbers more often.

  2. Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

    cs.SD 2026-06 unverdicted novelty 7.0 of 10

    Introduces a feature-vocoder adversarial attack on ASR using SSL representations that reports +26.6 WER black-box transfer and +36.2 WER defense resistance over baselines.

  3. VABench: A Comprehensive Benchmark for Audio-Video Generation

    cs.CV 2025-12 unverdicted novelty 7.0 of 10

    VABench is a new multi-dimensional benchmark for evaluating synchronous audio-video generation across text-to-AV, image-to-AV, and stereo tasks.

  4. AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

    cs.CL 2025-07 conditional novelty 7.0 of 10

    With prompt engineering (audio concatenation plus in-context examples), large audio models rank speech synthesis systems in line with human preferences, reaching up to 0.91 Spearman correlation.

  5. Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    Emo-LiPO applies listwise preference optimization to model global emotion intensity ordering in LLM TTS, yielding better accuracy and controllability than supervised or DPO baselines on a new multi-speaker dataset.

  6. JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions

    eess.AS 2026-05 unverdicted novelty 6.0 of 10

    JASTIN is an instruction-driven audio evaluation system that achieves state-of-the-art correlation with human ratings on speech, sound, music, and out-of-domain tasks without task-specific retraining.

  7. Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models

    eess.AS 2026-04 unverdicted novelty 6.0 of 10

    A Conformer-conditioned decoder-only language model generates discrete tokens via a neural audio codec to separate four music stems, reaching near state-of-the-art perceptual quality and top NISQA on vocals in MUSDB18...

  8. LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement

    cs.SD 2026-03 conditional novelty 6.0 of 10

    Using LLM-generated text descriptions of enhanced speech converted to sentiment scores as PPO rewards improves PESQ, STOI, and neural quality scores over supervised and DNSMOS-reward baselines on AVSEC-4.

  9. CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses

    cs.SD 2026-07 reject novelty 5.0 of 10

    CS-ETS applies Lyapunov and detrended-fluctuation-analysis losses inside a Samba encoder, but its headline audio gains are confounded by a DTW alignment step not applied to baselines.

  10. Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    Feature-aligned watermarking embeds a codec-generated pseudo-speech signal into the spectrogram to raise robustness against reconstruction models while keeping imperceptibility comparable to prior methods.

  11. GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

    eess.AS 2025-12 conditional novelty 5.0 of 10

    A two-stage decoder-only language model with continuous embeddings and UTMOS-based preference fine-tuning reports improved target-speaker-extraction scores on Libri2Mix.

  12. Schr\"odinger Bridge Mamba for One-Step Speech Enhancement

    cs.SD 2025-10 conditional novelty 5.0 of 10

    A Mamba-based speech enhancer trained with Schrödinger Bridge objectives produces strong denoising and dereverberation in one inference step with a low real-time factor.

  13. Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment

    eess.AS 2026-04 unverdicted novelty 3.0 of 10

    Voice range indicates TTS model capability with VITS highest, Glow-TTS best at soft phonation, and CPPs of 7-8 dB marking natural quality while values over 10 dB sound robotic.

  14. A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models

    eess.AS 2026-05 unverdicted novelty 2.0 of 10

    A structured survey of audio bandwidth extension that organizes the transition from deterministic discriminative DNNs to generative approaches including GANs, diffusion models, and flow-based methods.

Pith tools