Pith. sign in

REVIEW 2 cited by

MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.08730 v3 pith:T5UZ2QNK submitted 2023-09-15 eess.AS cs.AIcs.CLcs.MMcs.SD

classification eess.AScs.AIcs.CLcs.MMcs.SD
keywords musicdatasetmusilingoaudiobridgingcaptioncaptionsdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have shown immense potential in multimodal applications, yet the convergence of textual and musical domains remains not well-explored. To address this gap, we present MusiLingo, a novel system for music caption generation and music-related query responses. MusiLingo employs a single projection layer to align music representations from the pre-trained frozen music audio model MERT with a frozen LLM, bridging the gap between music audio and textual contexts. We train it on an extensive music caption dataset and fine-tune it with instructional data. Due to the scarcity of high-quality music Q&A datasets, we created the MusicInstruct (MI) dataset from captions in the MusicCaps datasets, tailored for open-ended music inquiries. Empirical evaluations demonstrate its competitive performance in generating music captions and composing music-related Q&A pairs. Our introduced dataset enables notable advancements beyond previous ones.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Piano Transcription by Hierarchical Language Modeling with Pretrained Roll-based Encoders

    cs.SD 2025-01 conditional novelty 6.0 of 10

    Decoding pretrained piano-roll encoder embeddings with three hierarchical language models improves onset-offset-velocity F1 by 0.010 to 0.022 over roll outputs on Maestro.

  2. Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation

    cs.MM 2024-12 conditional novelty 5.0 of 10

    An automatic music-visualizer pipeline that uses genre-guided image generation and audio-energy-controlled frame interpolation beats linear interpolation on a new synchrony metric.

Pith tools