REVIEW 8 cited by
SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In speech generation tasks, human subjective ratings, usually referred to as the opinion score, are considered the "gold standard" for speech quality evaluation, with the mean opinion score (MOS) serving as the primary evaluation metric. Due to the high cost of human annotation, several MOS prediction systems have emerged in the speech domain, demonstrating good performance. These MOS prediction models are trained using annotations from previous speech-related challenges. However, compared to the speech domain, the singing domain faces data scarcity and stricter copyright protections, leading to a lack of high-quality MOS-annotated datasets for singing. To address this, we propose SingMOS, a high-quality and diverse MOS dataset for singing, covering a range of Chinese and Japanese datasets. These synthesized vocals are generated using state-of-the-art models in singing synthesis, conversion, or resynthesis tasks and are rated by professional annotators alongside real vocals. Data analysis demonstrates the diversity and reliability of our dataset. Additionally, we conduct further exploration on SingMOS, providing insights for singing MOS prediction and guidance for the continued expansion of SingMOS.
Forward citations
Cited by 8 Pith papers
-
Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves
SongSQA predicts both an overall singing-quality score and a temporal segment-level score curve for full-length songs, using teacher-generated pseudo labels and a learnable attention aggregator.
-
MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres
Current singing voice synthesis models fail to differentiate musical genres, defaulting to pop-like output regardless of input genre, unless given genre-specific fine-tuning data.
-
An Extensive Analysis of the Singing Voice Conversion Challenge 2025 Evaluation Results
SVCC 2025 shows top systems can match ground truth singer identity but cannot yet match naturalness or singing style, with breathy, glissando, and vibrato as the hardest styles.
-
Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
Speaker-recognition pre-trained models (x-vector, ECAPA) outperform other speech and music models for singing voice MOS prediction, and their fusion via a Bhattacharyya-distance loss sets a new reported state of the a...
-
SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
SingNet is a claimed ~3,000-hour in-the-wild singing voice dataset from internet songs and sample packs, with benchmarks for lyric transcription, vocoders, and singing voice conversion.
-
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
The authors release VERSA, a unified open-source evaluation toolkit covering 65 metrics and 729 variants for speech, audio, and music.
-
Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation
Frozen SSL-Transformer embeddings generalize better than fine-tuned SSL or ViViT for cross-corpus MOS prediction, matching specialized SOTA on URGENT 2024 with MSE 0.36.
-
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
PS-SQA, a fusion of pitch-aware and spectrum-aware SSL MOS predictors with bias correction, achieved the best system-level SRCC on the VoiceMOS 2024 singing track.
Discussion (0). Continue with ORCID to comment.