{"id":"d34f534d-4a38-48bd-8f21-7885eb71c72b","arxiv_id":"2607.10146","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Frozen SSL-Transformer embeddings generalize better than fine-tuned SSL or ViViT for cross-corpus MOS prediction, matching specialized SOTA on URGENT 2024 with MSE 0.36.","lead":"Frozen WavLM features plus a Transformer beat fine-tuned SSL on unseen speech-quality data, hitting 0.36 MSE on URGENT 2024 after training on 130k samples. The result matters because automatic MOS scoring is the bottleneck for evaluating modern TTS and enhancement systems at scale.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"The 5 s / 16 kHz standardization pipeline may discard the very long-range and high-frequency cues that produced the original human MOS labels.","rationale":"The reader correctly identified the preprocessing pipeline as the weakest link for the absolute-performance and universality claims. Under identical preprocessing the controlled FRZ-vs-FT and LODO comparisons remain informative, code and weights are released, and no stronger internal inconsistency appears. The single-hold-out nature of the FT/FRZ OOD comparison and absence of error bars are secondary limitations already cited by the reader as grounds for CONDITIONAL; they do not require a further verdict change.","tokens_in":34046,"tokens_out":502,"duration_ms":30857,"concrete_test":"On a URGENT 2024 or TCD-VoIP subset known to contain long-duration or high-band artifacts, re-extract features at native rate with full-utterance variable-length processing (no 5 s cut/pad) and a length-adaptive head; retrain/evaluate the English SSL-FRZ model. If URGENT MSE rises >0.1 or SRCC falls >0.1 relative to the reported 0.36/0.71, the preprocessing is distorting the claimed generalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim (SSL-FRZ is the most stable universal solution, MSE 0.36 on URGENT 2024 matching domain-optimized SOTA) rests on every waveform having been forcibly resampled to 16 kHz and cut/padded into fixed 5 s segments before either mean-pooling (SSL) or CLS aggregation (ViViT) (§3.2). If packet-loss bursts, high-band codec artifacts (>8 kHz), or utterance-level prosodic quality that listeners originally rated are systematically attenuated or zero-masked, then both the LODO generalization gaps (Tables 4–5) and the absolute URGENT numbers become properties of the standardized feature space rather than of the architectures or of true human perception. Relative FRZ/FT rankings can still be valid under identical preprocessing, but the “universal” and “closely matching SOTA” conclusions do not transfer to unprocessed audio. This assumption is load-bearing because it is irreversible and shared by every reported metric.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript benchmarks three MOS predictors—frozen WavLM-Large + Transformer (SSL-FRZ), fine-tuned WavLM-Large + Transformer (SSL-FT), and a ViViT-MFCC + Transformer—on a consolidated 130 k-utterance corpus (19 datasets) and a purified English-only subset (17 datasets). A six-case Leave-One-Dataset-Out (LODO) protocol quantifies seen-versus-unseen generalization; the best English-only SSL-FRZ model is further compared with 18 ARECHO SOTA metrics. The central claim is that frozen SSL embeddings plus a deep Transformer yield the most stable cross-corpus solution, reaching MSE 0.36 on the held-out URGENT 2024 human-MOS set (versus 0.30 for a domain-optimized SOTA) while outperforming fine-tuned SSL on out-of-distribution data; English-only purification improves precision for all architectures.","tokens_in":34346,"tokens_out":1275,"duration_ms":19470,"significance":"If the reported rankings hold, the work supplies a carefully controlled, large-scale reference point for domain-shift studies in non-intrusive speech quality assessment—an area still dominated by narrow-domain models. The two-part corpus design, systematic LODO, dual utterance/system-level metrics, and public release of the English-only SSL-Transformer weights on Hugging Face are concrete contributions that other groups can build on. The finding that freezing a denoising-pretrained SSL backbone can be preferable to full fine-tuning for OOD robustness is practically useful and falsifiable.","major_comments":[{"comment":"§3.2 (and the pipeline in Fig. 1): every waveform is forcibly resampled to 16 kHz and cut/padded into fixed 5 s segments before either mean-pooling (SSL) or CLS aggregation (ViViT). Several source sets (VoiceMOS 2025 Track 3, URGENT) contain 24–48 kHz material and long-range artifacts (packet-loss bursts, high-band codec distortion) that human listeners originally rated. No ablation or discussion quantifies how much of the original perceptual cue set survives this irreversible standardization. Because the same pipeline is shared by every reported number, both the absolute URGENT MSE of 0.36 and the claim of “universal” assessment risk being properties of the standardized feature space rather than of the architectures. A short ablation (native-rate vs. 16 kHz, or variable-length vs. 5 s) or an explicit limitation paragraph is required for the central claim to transfer beyond the preproces","section":null},{"comment":"§3.4 and Tables 4–5: the LODO protocol that underpins the generalization analysis evaluates only frozen SSL and ViViT; full fine-tuning is performed solely on the single full-corpus split (Table 3). Consequently the headline statement that “SSL-FRZ provides superior robustness on unseen distributions” rests on one held-out test set rather than on the systematic LODO design. Either (a) a reduced-budget LODO with FT (e.g., last-layer or LoRA fine-tuning) or (b) an explicit caveat that the FRZ-vs-FT OOD comparison is limited to the URGENT split is needed before the claim can be considered fully supported.","section":null},{"comment":"Tables 3–6 report only point estimates. No standard deviations across seeds, bootstrap confidence intervals, or statistical significance tests accompany the MSE/LCC/SRCC/KTAU differences that drive the ranking of FRZ over FT and of the proposed model over most ARECHO metrics. Given that several pairwise gaps are small (e.g., 0.36 vs. 0.30 MSE on URGENT, or FRZ vs. FT correlation differences of 0.01–0.03), the absence of uncertainty quantification leaves open the possibility that the reported superiority is within noise. Adding at least multi-seed means ± std or a simple paired test on the primary URGENT and LODO numbers would make the central ranking claims load-bearing.","section":null}],"minor_comments":[{"comment":"Numerous typographical and orthographic inconsistencies remain (e.g., “sescond part”, “V oiceMOS”, “V oIP”, “challange”, “T otal”, “arecho_urgent_mos” capitalization). A careful proof-reading pass is needed.","section":null},{"comment":"Figure 1 is dense; the two parallel training paths and the two LODO subsets are hard to parse at a glance. A simplified schematic or clearer color coding would help.","section":null},{"comment":"§3.3.1 states that mean-pooling was chosen over a CLS token “to ensure that every segment contributes equally,” yet no quantitative comparison of the two aggregation strategies is supplied. A one-sentence ablation or reference would strengthen the design choice.","section":null},{"comment":"Table 1 lists sample counts after deduplication, but the exact number of unique systems or listeners per dataset is not given; this information is useful for interpreting system-level correlations.","section":null},{"comment":"The ARECHO comparison (Table 6) normalizes Audiobox PQ from [1,10] to [1,5]; the linear mapping should be stated explicitly in the table caption or §3.7.","section":null}],"recommendation":"major_revision","confidential_remarks":"The experimental scale and open-source release are genuine strengths; the paper is closer to a solid systems/benchmark contribution than to a purely methodological one. The three major points above are addressable within a revision cycle and do not require new data collection. I would be comfortable recommending acceptance after those points are handled. Fit for a speech-processing or multimedia-quality journal is good; the work is less suited to a pure machine-learning venue that expects theoretical novelty."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing is that they ran a clean head-to-head at 130 k scale and showed frozen SSL embeddings plus a small Transformer beat fine-tuning on held-out human MOS (URGENT 2024 MSE 0.36 vs domain-tuned SOTA at 0.30), while fine-tuning wins only on seen data. English-only purification tightens the numbers further. Code and the English SSL-FRZ weights are public, so the ranking is checkable.\n\nWhat is new is the controlled LODO design across two corpus variants (broad vs purified English), the explicit frozen-vs-FT comparison under identical conditions, and the external ARECHO 18-metric comparison. Most MOS papers stop at one corpus and one training regime. The experimental hygiene is better than average: utterance- and system-level metrics, consistent tables, and a real out-of-distribution test set with human labels. The ViViT-MFCC path is a sensible efficiency baseline and behaves as expected (weaker but stable after purification).\n\nSoft spots are ordinary, not load-bearing. No error bars or significance tests; LODO subsets are hand-chosen for balance rather than random. The 5 s segmentation + 16 kHz resampling + mean/CLS pooling is shared by almost every modern non-intrusive model, so relative FRZ/FT rankings remain valid even if absolute numbers partly reflect the standardized feature space. The “universal” language is a bit strong, but the concrete claim (frozen is more stable OOD) holds under the protocols they actually ran.\n\nThis is for people who build or use automatic MOS predictors for TTS, VC, or enhancement and need a practical, released recipe rather than a new theory. It deserves a serious referee; the design is careful enough and the result is useful enough that desk rejection would be a mistake. I would engage with the numbers and the checkpoint.","headline":"Careful large-scale bake-off: frozen WavLM+Transformer is more robust OOD than fine-tuning, English purification helps, and they release the checkpoint; the 5 s/16 kHz pipeline is the usual field caveat, not a unique flaw.","tokens_in":34996,"tokens_out":505,"would_cite":true,"duration_ms":10686,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Frozen self-supervised speech models with Transformers give the most stable automatic MOS scores across unseen datasets, nearly matching specialized metrics on a hard enhancement benchmark.","keywords":["Mean Opinion Score","MOS prediction","self-supervised learning","WavLM","Video Vision Transformer","Leave-One-Dataset-Out","speech quality assessment","domain shift"],"falsifier":"Retrain the identical English-only frozen SSL-Transformer without 5-second segmentation—using full-utterance embeddings or variable-length attention—and re-measure MSE on URGENT 2024; a clear rise in error would show the reported robustness is an artifact of the fixed-window pipeline.","tokens_in":34903,"feed_emoji":"🎧","tokens_out":681,"duration_ms":16266,"temperature":0.7,"pith_summary":"Automatic MOS predictors routinely fail when the test speech comes from a different source than the training data. This paper trains three architectures—frozen SSL, fine-tuned SSL, and a spectral ViViT—on more than 130,000 clips from 19 datasets, then repeats the work on a cleaned English-only subset. Leave-one-dataset-out trials show that models score far better on seen distributions, yet the frozen SSL-Transformer remains the most reliable on truly unseen material, reaching MSE 0.36 on the human-labeled URGENT 2024 set (close to the 0.30 of domain-tuned SOTA). English-only purification improves precision for all three models. The result matters because large-scale TTS and enhancement systems need a single, public, domain-robust quality score rather than a new model for every corpus.","feed_headline":"Frozen SSL hits 0.36 MSE on unseen speech quality scores","feed_subtitle":"English-only training and static embeddings beat fine-tuning for cross-dataset MOS prediction","key_machinery":"Systematic Leave-One-Dataset-Out (LODO) protocol that contrasts Frozen SSL-Transformer, Fine-Tuned SSL-Transformer, and ViViT-MFCC architectures on a broad 19-dataset corpus versus an English-only 17-dataset subset; the frozen path freezes WavLM Large, feeds its 1024-dim frames into a two-layer Transformer with mean pooling, and uses the resulting score to quantify the seen-versus-unseen generalization gap.","core_discovery":"Frozen WavLM embeddings combined with a deep Transformer encoder, trained on a purified English-only corpus of roughly 123,000 utterances, form the most stable cross-corpus MOS predictor: they achieve MSE 0.36 on the held-out URGENT 2024 human-MOS benchmark, stay competitive with 18 specialized SOTA metrics, and outperform both fine-tuned SSL (which overfits seen data) and the lighter ViViT spectral model on out-of-distribution sets.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Frozen SSL hits 0.36 MSE on unseen MOS scores via LODO","English-only frozen SSL beats fine-tuning on domain-shift MOS","SSL-FRZ with deep Transformer yields robust 0.36 MSE on URGENT","Purified English corpus plus frozen SSL tops cross-corpus MOS","LODO confirms frozen SSL over fine-tuned for out-of-domain audio quality"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The claim assumes that resampling every file to 16 kHz, cutting it into fixed 5-second zero-padded chunks, and pooling those chunks still keeps the perceptual cues human listeners used when they rated the original full-length recordings.","fun_headline_variants_meta":{"raw":{"variants":["Frozen SSL hits 0.36 MSE on unseen MOS scores via LODO","English-only frozen SSL beats fine-tuning on domain-shift MOS","SSL-FRZ with deep Transformer yields robust 0.36 MSE on URGENT","Purified English corpus plus frozen SSL tops cross-corpus MOS","LODO confirms frozen SSL over fine-tuned for out-of-domain audio quality"]},"model":"grok-4.5","effort":"low","cost_usd":0.005984,"raw_usage":{"total_tokens":1662,"prompt_tokens":901,"num_sources_used":0,"completion_tokens":104,"cost_in_usd_ticks":59840000,"prompt_tokens_details":{"text_tokens":901,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":657,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":901,"tokens_out":104,"duration_ms":5332,"temperature":1.0,"reasoning_tokens":657,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T13:55:59.060343+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Retrain the identical English-only frozen SSL-Transformer without 5-second segmentation—using full-utterance embeddings or variable-length attention—and re-measure MSE on URGENT 2024; a clear rise in error would show the reported robustness is an artifact of the fixed-window pipeline.","supporting_citations":[],"review_version":1}