Pith. sign in

REVIEW 3 cited by

USAD: Universal Speech and Audio Representation via Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.18843 v2 pith:LLT7XBLY submitted 2025-06-23 cs.SD cs.CLeess.AS

USAD: Universal Speech and Audio Representation via Distillation

classification cs.SD cs.CLeess.AS
keywords audiospeechusaddistillationbenchmarksdomain-specificlearningmodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distillation (USAD), a unified approach to audio representation learning that integrates diverse audio types - speech, sound, and music - into a single model. USAD employs efficient layer-to-layer distillation from domain-specific SSL models to train a student on a comprehensive audio dataset. USAD offers competitive performance across various benchmarks and datasets, including frame and instance-level speech processing tasks, audio tagging, and sound classification, achieving near state-of-the-art results with a single encoder on SUPERB and HEAR benchmarks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Stage-adaptive audio diffusion modeling

    cs.SD 2026-05 unverdicted novelty 6.0

    A semantic progress signal from SSL discrepancy slope enables three stage-aware mechanisms that improve training efficiency and performance in audio diffusion models over static baselines.

  2. Alethia: A Foundational Encoder for Voice Deepfakes

    cs.SD 2026-04 unverdicted novelty 6.0

    Alethia is a pretrained audio encoder using continuous embedding prediction and generative flow-matching reconstruction that outperforms existing speech foundation models on voice deepfake tasks with better robustness...

  3. SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations

    eess.AS 2025-10 conditional novelty 6.0

    SPEAR unifies speech and audio self-supervised learning by masked prediction of multi-codebook tokens distilled from two domain teachers, beating WavLM Large on 12 of 15 SUPERB tasks and scoring competitively on HEAR.