Pith. sign in

REVIEW 3 cited by

Music Foundation Model as Generic Booster for Music Downstream Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.01135 v3 pith:HBY534X7 submitted 2024-11-02 cs.SD cs.IRcs.LGeess.AS

classification cs.SDcs.IRcs.LGeess.AS
keywords musictasksdownstreamfoundationfeaturesmodelsmodelapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We demonstrate the efficacy of using intermediate representations from a single foundation model to enhance various music downstream tasks. We introduce SoniDo, a music foundation model (MFM) designed to extract hierarchical features from target music samples. By leveraging hierarchical intermediate features, SoniDo constrains the information granularity, leading to improved performance across various downstream tasks including both understanding and generative tasks. We specifically evaluated this approach on representative tasks such as music tagging, music transcription, music source separation, and music mixing. Our results reveal that the features extracted from foundation models provide valuable enhancements in training downstream task models. This highlights the capability of using features extracted from music foundation models as a booster for downstream tasks. Our approach not only benefits existing task-specific models but also supports music downstream tasks constrained by data scarcity. This paves the way for more effective and accessible music processing solutions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving BERT for Symbolic Music Understanding Using Token Denoising and Pianoroll Prediction

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A music BERT with bounded token denoising and pianoroll prediction beats standard masked-language pre-training on a new 12-task symbolic music benchmark.

  2. SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A ControlNet branch plus a frequency-aware feature aligner lets a pretrained masked generative TTA model produce video-synchronized foley, beating several from-scratch models on VGGSound.

  3. Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Adding self-supervised objectives to a supervised multi-pitch estimator improves closed-set performance but triggers degeneration to blank predictions on additional, unlabeled data.

Pith tools