Pith. sign in

REVIEW 9 cited by

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.06660 v1 pith:UD2LZBYG submitted 2024-12-09 cs.SD cs.MMeess.AS

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

classification cs.SD cs.MMeess.AS
keywords musicmulti-modalgenerationimagesmodelsmumu-llamaunderstandingvideos
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Research on large language models has advanced significantly across text, speech, images, and videos. However, multi-modal music understanding and generation remain underexplored due to the lack of well-annotated datasets. To address this, we introduce a dataset with 167.69 hours of multi-modal data, including text, images, videos, and music annotations. Based on this dataset, we propose MuMu-LLaMA, a model that leverages pre-trained encoders for music, images, and videos. For music generation, we integrate AudioLDM 2 and MusicGen. Our evaluation across four tasks--music understanding, text-to-music generation, prompt-based music editing, and multi-modal music generation--demonstrates that MuMu-LLaMA outperforms state-of-the-art models, showing its potential for multi-modal music applications.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

    cs.CL 2026-05 unverdicted novelty 7.0

    MusTBENCH evaluates temporal grounding in large audio-language models via five expert-validated tasks, and MusT improves performance through encoder adaptation, LLM adaptation, supervised fine-tuning, and RL optimization.

  2. TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

    cs.SD 2026-07 conditional novelty 6.0

    A new audio benchmark, TORUS, shows unified audio models are not self-coherent: the best model answers 50.5% of questions about its own generations, below a 63.2% cascaded specialist baseline.

  3. JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

    cs.SD 2026-06 unverdicted novelty 6.0

    JenBridge pretrains a flow-matching Transformer on text-audio data then adapts it with video conditioning and an LLM director to select transitions, claiming better coherence than prior methods on a new LVS benchmark.

  4. Assessing Factual Music Comprehension in Large Audio Language Models

    cs.SD 2025-11 conditional novelty 6.0

    Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.

  5. Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach

    cs.SD 2025-09 unverdicted novelty 6.0

    A zero-training VLM framework generates music from images via ABC notation, multi-modal RAG, and self-refinement while providing text and visual explanations for the outputs.

  6. AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

    cs.SD 2026-06 unverdicted novelty 5.0

    AudioX-Turbo distills a Multimodal Diffusion Transformer into a 4-step student model for efficient multimodal anything-to-audio generation, trained on a new 9.2M-sample dataset IF-caps-Pro.

  7. AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

    cs.SD 2026-06 unverdicted novelty 5.0

    A distilled multimodal diffusion model generates audio from text, video, or audio in four steps with claimed superior quality and ~25× fewer function evaluations.

  8. EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement

    cs.SD 2026-06 unverdicted novelty 4.0

    EntangleCodec unifies semantic and acoustic audio tokenization via caption alignment and flow-matching decoding, reporting competitive reconstruction, +7.4% gains on MMAR understanding, and 0.6B-parameter ALMs surpass...

  9. WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation

    cs.SD 2025-09 reject novelty 4.0

    An open multi-agent system that orchestrates specialized music models for understanding, composition, and synthesis, with local or hosted deployment.