Sparse inversion of text concepts into audio SAE feature space yields audio-faithful concept supports that improve edit-preservation trade-offs in steerable music retrieval.
Audio clips are encoded with the audio towers fA of CLAP [28] and MuQ [33]; from each 10-second frame we extract an embeddingz a ∈R 512
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Steering dense music retrieval with open-vocabulary concept discovery
Sparse inversion of text concepts into audio SAE feature space yields audio-faithful concept supports that improve edit-preservation trade-offs in steerable music retrieval.