Pith. sign in

REVIEW 1 cited by

Transformation of audio embeddings into interpretable, concept-based representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.14076 v1 pith:BMZ3A64J submitted 2025-04-18 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords audioconcept-basedembeddingsrepresentationsinterpretabilitydownstreamtasksclap
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Advancements in audio neural networks have established state-of-the-art results on downstream audio tasks. However, the black-box structure of these models makes it difficult to interpret the information encoded in their internal audio representations. In this work, we explore the semantic interpretability of audio embeddings extracted from these neural networks by leveraging CLAP, a contrastive learning model that brings audio and text into a shared embedding space. We implement a post-hoc method to transform CLAP embeddings into concept-based, sparse representations with semantic interpretability. Qualitative and quantitative evaluations show that the concept-based representations outperform or match the performance of original audio embeddings on downstream tasks while providing interpretability. Additionally, we demonstrate that fine-tuning the concept-based representations can further improve their performance on downstream tasks. Lastly, we publish three audio-specific vocabularies for concept-based interpretability of audio embeddings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings

    eess.AS 2026-07 accept novelty 6.0 of 10

    RT60, LUFS, and relative pitch are approximately linearly recoverable from frozen CLAP embeddings across noise, speech, and music, while spectral centroid needs non-linear probes; the pattern largely generalizes to ot...

Pith tools