Pith. sign in

REVIEW 1 cited by

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.13880 v3 pith:YC2R5E6B submitted 2025-05-20 eess.AS cs.SDeess.SP

classification eess.AScs.SDeess.SP
keywords audiou-samlanguagealignmentmodelmodelsmusicspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a comprehensive understanding across diverse audio types, such as speech, general audio events, and music. Furthermore, their exclusive reliance on cross-entropy loss for alignment often falls short, as it treats all tokens equally and fails to account for redundant audio features, leading to weaker cross-modal alignment. To deal with the above challenges, this paper introduces U-SAM, an advanced audio language model that integrates specialized encoders for speech, audio, and music with a pre-trained large language model (LLM). U-SAM employs a Mixture of Experts (MoE) projector for task-aware feature fusion, dynamically routing and integrating the domain-specific encoder outputs. Additionally, U-SAM incorporates a Semantic-Aware Contrastive Loss Module, which explicitly identifies redundant audio features under language supervision and rectifies their semantic and spectral representations to enhance cross-modal alignment. Extensive experiments demonstrate that U-SAM consistently outperforms both specialized models and existing audio language models across multiple benchmarks. Moreover, it exhibits emergent capabilities on unseen tasks, showcasing its generalization potential. Code is available (https://github.com/Honee-W/U-SAM/).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction

    cs.SD 2025-06 conditional novelty 4.0 of 10

    WhisQ uses Whisper and Qwen with co-attention and optimal transport to predict music quality and text-alignment scores, but its reported improvements do not match its own data.

Pith tools