REVIEW 2 cited by
A Foundation Model for Music Informatics
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper investigates foundation models tailored for music informatics, a domain currently challenged by the scarcity of labeled data and generalization issues. To this end, we conduct an in-depth comparative study among various foundation model variants, examining key determinants such as model architectures, tokenization methods, temporal resolution, data, and model scalability. This research aims to bridge the existing knowledge gap by elucidating how these individual factors contribute to the success of foundation models in music informatics. Employing a careful evaluation framework, we assess the performance of these models across diverse downstream tasks in music information retrieval, with a particular focus on token-level and sequence-level classification. Our results reveal that our model demonstrates robust performance, surpassing existing models in specific key metrics. These findings contribute to the understanding of self-supervised learning in music informatics and pave the way for developing more effective and versatile foundation models in the field. A pretrained version of our model is publicly available to foster reproducibility and future research.
Forward citations
Cited by 2 Pith papers
-
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
MuQ, trained with masked prediction of Mel-RVQ tokens, beats MERT and MusicFM on the MARBLE average despite a much smaller pre-training set.
-
Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models
Layer-wise probing of MusicFM and MuQ shows acoustic-to-semantic feature progression across layers, and single-layer selection often outperforms all-layer aggregation on MIR tasks.
Discussion (0). Continue with ORCID to comment.