Pith. sign in

REVIEW 1 cited by

A Large-Scale Evaluation of Speech Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09385 v2 pith:WC6NCKSI submitted 2024-04-15 eess.AS cs.CLeess.SP

classification eess.AScs.CLeess.SP
keywords foundationspeechmodelparadigmprocessingsuperbtasksbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The foundation model paradigm leverages a shared foundation model to achieve state-of-the-art (SOTA) performance for various tasks, requiring minimal downstream-specific modeling and data annotation. This approach has proven crucial in the field of Natural Language Processing (NLP). However, the speech processing community lacks a similar setup to explore the paradigm systematically. In this work, we establish the Speech processing Universal PERformance Benchmark (SUPERB) to study the effectiveness of the paradigm for speech. We propose a unified multi-tasking framework to address speech processing tasks in SUPERB using a frozen foundation model followed by task-specialized, lightweight prediction heads. Combining our results with community submissions, we verify that the foundation model paradigm is promising for speech, and our multi-tasking framework is simple yet effective, as the best-performing foundation model shows competitive generalizability across most SUPERB tasks. For reproducibility and extensibility, we have developed a long-term maintained platform that enables deterministic benchmarking, allows for result sharing via an online leaderboard, and promotes collaboration through a community-driven benchmark database to support new development cycles. Finally, we conduct a series of analyses to offer an in-depth understanding of SUPERB and speech foundation models, including information flows across tasks inside the models, the correctness of the weighted-sum benchmarking protocol and the statistical significance and robustness of the benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A 30-second counting task, embedded with speaker-identification models, predicts male sleep apnea (AUC 0.64) and shows gender- and condition-specific model rankings across a new 7,188-recording clinical speech benchmark.

Pith tools