Pith. sign in

REVIEW 16 cited by

MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.01108 v2 pith:JB5NOWYF submitted 2025-01-02 cs.SD cs.AIcs.CLcs.LGeess.AS

classification cs.SDcs.AIcs.CLcs.LGeess.AS
keywords musicmodellearningself-supervisedperformancequantizationrepresentationresidual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent years have witnessed the success of foundation models pre-trained with self-supervised learning (SSL) in various music informatics understanding tasks, including music tagging, instrument classification, key detection, and more. In this paper, we propose a self-supervised music representation learning model for music understanding. Distinguished from previous studies adopting random projection or existing neural codec, the proposed model, named MuQ, is trained to predict tokens generated by Mel Residual Vector Quantization (Mel-RVQ). Our Mel-RVQ utilizes residual linear projection structure for Mel spectrum quantization to enhance the stability and efficiency of target extraction and lead to better performance. Experiments in a large variety of downstream tasks demonstrate that MuQ outperforms previous self-supervised music representation models with only 0.9K hours of open-source pre-training data. Scaling up the data to over 160K hours and adopting iterative training consistently improve the model performance. To further validate the strength of our model, we present MuQ-MuLan, a joint music-text embedding model based on contrastive learning, which achieves state-of-the-art performance in the zero-shot music tagging task on the MagnaTagATune dataset. Code and checkpoints are open source in https://github.com/tencent-ailab/MuQ.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

    cs.SD 2026-07 accept novelty 6.0 of 10

    MADB is a 9,999-track music aesthetics benchmark with multi-dimensional professional annotations revealing that current pretrained audio models capture only partial aesthetic information.

  2. FIGMA: Towards FIne-Grained Music retrievAl

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    FIGMA proposes a multi-view contrastive architecture plus the FGMCaps dataset to retrieve music from fine-grained textual descriptions of musical attributes, reporting up to 73.3% relative gains over CLAP baselines.

  3. UniVocal: Unified Speech-Singing Code-Switching Synthesis

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    UniVocal presents a text-context-only framework for speech-singing code-switching synthesis via two-stage curriculum learning and a synthetic data pipeline, claiming SOTA on a new benchmark.

  4. S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation

    eess.AS 2026-05 unverdicted novelty 6.0 of 10

    S2Accompanist is a 402M-parameter semantic-aware diffusion model that achieves SOTA on the ATTM Grand Challenge benchmark for music accompaniment generation via automated data processing and structure-guided VAE fine-tuning.

  5. Leveraging Artist Catalogs for Cold-Start Music Recommendation

    cs.IR 2026-04 unverdicted novelty 6.0 of 10

    ACARec attends over artist catalogs to generate CF embeddings for new tracks, more than doubling recall and NDCG versus content-only baselines in music recommendation.

  6. TADA! Tuning Audio Diffusion Models through Activation Steering

    cs.SD 2026-02 unverdicted novelty 6.0 of 10

    Activation steering at a semantic bottleneck in audio diffusion models achieves state-of-the-art control over musical attributes such as instruments, vocals, and genres.

  7. Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling

    cs.SD 2026-01 conditional novelty 6.0 of 10

    A song-aesthetics model with multi-stem cross-attention and hierarchical interval regression beats two adapted MOS baselines on average, with some dimensions showing ties or losses.

  8. Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation

    eess.AS 2025-11 conditional novelty 6.0 of 10

    A 10.7M-pair audio-caption corpus and systematic comparison show contrastive pretraining is more data-efficient while captioning scales better, and supervised initialization yields diminishing returns.

  9. Qwen3-Omni Technical Report

    cs.CL 2025-09 unverdicted novelty 6.0 of 10

    Qwen3-Omni is a unified multimodal model that achieves open-source SOTA on 32 of 36 audio and audio-visual benchmarks and overall SOTA on 22 without degrading performance on text, image, or video relative to single-mo...

  10. The AudioMOS Challenge 2025

    cs.SD 2025-09 conditional novelty 6.0 of 10

    The first AudioMOS challenge compared automatic predictors of human quality scores for synthetic audio across three tracks, and most of the 24 participating teams outperformed the organizers' baselines.

  11. Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Score-aware training uses alignment scores to route low-quality segments into high-noise regimes as implicit regularizers, enabling a 450M model to rank competitively in a text-to-music challenge with limited data.

  12. Adopting State-of-the-Art Pretrained Audio Representations for Music Recommender Systems

    cs.IR 2026-04 unverdicted novelty 5.0 of 10

    Pretrained audio models show large performance gaps between standard MIR tasks and music recommendation in both hot and cold-start settings.

  13. Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity

    cs.AI 2026-03 unverdicted novelty 5.0 of 10

    Separating acoustic and expectation ANN representations as teacher targets improves EEG music identification beyond baselines and seed ensembles.

  14. Revisiting Content-Based Music Recommendation: Efficient Feature Aggregation from Large-Scale Music Models

    cs.IR 2026-02 unverdicted novelty 5.0 of 10

    TASTE dataset and MuQ-token aggregation enable effective use of audio features from large music models to improve content-based music recommendations over collaborative filtering alone.

  15. SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision

    eess.AS 2025-10 unverdicted novelty 5.0 of 10

    SongFormer achieves state-of-the-art strict boundary detection and functional label accuracy in music structure analysis by fusing SSL representations and using learned source embeddings on a new 14k-song corpus and e...

  16. Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation

    cs.SD 2025-08 conditional novelty 5.0 of 10

    PER-based preference optimization (DPO, PPO, GRPO) reduces lyric-to-song hallucination in an audio language model, with the largest gains from DPO plus reject sampling.

Pith tools