Pith. sign in

Paper Citation Record · LEDGER

M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2311.11255.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.11255 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:14:03.581263Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T12:41:22.627211Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d2c250f6-5c41-4c7e-bc80-741c6dc3529e · inbound

Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding cites this paper.

Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T20:14:03.581263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:14:03.581263Z digest=sha256:bf2cedd1a25b6af77a7767ca6b54198fb957744cdb0510418696732eca44e2cc

Observation 1f26ac52-abd5-4bb1-9adc-3a21d04b7115 · inbound

Controllable Video-to-Music Generation with Multiple Time-Varying Conditions cites this paper.

Controllable Video-to-Music Generation with Multiple Time-Varying Conditions M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:31:12.134837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:31:12.134837Z digest=sha256:d0284a11a6266e6c671772371003d1271a7297f4489dc261afb3b6fc7080305a

Observation 4c5248f2-5d31-4653-b4f5-d77997e48596 · inbound

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach cites this paper.

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.629436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T12:39:56.396231Z digest=sha256:a0ce6730ff3083141144d4ac324edc3d89915191f459a174a8455129f0bdb397

Observation d78c0bdd-9ddd-4138-af4f-b54a2c52f83e · inbound

FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration cites this paper.

FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.478472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.478472Z digest=sha256:57945b9e7f2b5a5a06f2122d7c615a998a00637de2cf7c29c2942197fcc1606e

Observation ed992961-cbf9-4d3e-b2af-7f40eb799d94 · inbound

MusiChat: Vibe Composing for Music Creation cites this paper.

MusiChat: Vibe Composing for Music Creation M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T23:25:22.985308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:25:22.985308Z digest=sha256:d2d2eb2ddbec9b655d43d6c29108ad06b4a338b652ca18e14af0cca16a721d1d