Pith. sign in

Paper Citation Record · LEDGER

MuLan: A Joint Embedding of Music Audio and Natural Language

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2208.12415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2208.12415 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:54:22.907977Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:47:06.209809Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c1290719-249f-4986-bb8b-b4d6ba2eabe9 · inbound

An introduction to pitch strength in contemporary popular music analysis and production cites this paper.

An introduction to pitch strength in contemporary popular music analysis and production MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:43.104009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:36:43.104009Z digest=sha256:07d6edeb31bfd19db30d95d20717e1138c329ad6c531d0e935da44f517eec876

Observation 208e4dfc-25bf-4969-9ca7-b30bbad69288 · inbound

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following cites this paper.

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:53.301995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:59:53.301995Z digest=sha256:aa20d635115e274c14736f9fe3a6ea7a4635cd7607fe300b76b1ab8b7ef81503

Observation 042c3dd8-b624-4796-8de9-02b16caf19cb · inbound

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance cites this paper.

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:16.292351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:16.292351Z digest=sha256:4ee4dc2a626d7c40b7f7661f7ec2a9012a9ac6ddc78458d3497d891e6a4ecda6

Observation 36b63466-0a07-400e-9b86-569e022fa58e · inbound

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization cites this paper.

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:49.324430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:40:49.324430Z digest=sha256:a30f48fb29f0080528e9d664482a61f6e59140112e51b4fbc9083965981bb0bc

Observation a4784200-4947-4cd9-b1c1-29d31da4a85d · inbound

PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music cites this paper.

PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:21:54.216932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:21:54.216932Z digest=sha256:a8453dd9d9b57f1b12fdd22d23bb3ac3713433855b092fca3acafd22f208ef81

Observation 90371326-2b88-4523-a120-6df04801184f · inbound

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity cites this paper.

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:40:03.509778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T11:36:06.967549Z digest=sha256:3fda79678940f75713d8583b0a57d998df4885606fa085765a193688ce5bc8c2

Observation e2b7dfc0-64be-4d15-9608-e93b85d0d60e · inbound

CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval cites this paper.

CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:06.850643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T14:51:16.918157Z digest=sha256:efd982c0a328fd0452c83a87b64000bb314dab2d434559e558cf09ad4879b08e

Observation f47275b6-6b8a-499e-a8c7-cf98ee7c0849 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:16:08.767388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:83ec8d83e5512db5008f7abc720079c6d8467a14e52fc4e33a16b0e1d9298e23

Observation bdb6a208-3b53-4369-92bc-db3c7e9550d7 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:15:07.850673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:3afde09e40f62c2cd3303b23d33b19a219f719ef3324bac8afbc131123f5428d

Observation fb3376ef-1a5c-438b-ae7f-8cd01904bce3 · inbound

Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model cites this paper.

Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T01:44:22.712237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T01:44:20.282896Z digest=sha256:9b4e56d0944b2203fef132c6479101f27f1230806ccdf81b1d300dd4b86dc982

Observation f13f9ee2-56da-4ea2-bada-b5b0c1b782e5 · inbound

MERIT: Learning Disentangled Music Representations for Audio Similarity cites this paper.

MERIT: Learning Disentangled Music Representations for Audio Similarity MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T15:43:32.553133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T15:40:57.719214Z digest=sha256:296baf54b50cd2846a9ba9337988e038a1a4f96c69f7b9f4c88f8ac0a2bd6583

Observation 7a1ea906-89b6-4db6-81a8-c2c9f27944a2 · inbound

SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling cites this paper.

SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:46:38.022078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T08:53:17.009342Z digest=sha256:e9e659a1e0006e763e71697325c99f77b94692842de9cf5a2f97eac598d44c4b

Observation d7d21625-9f67-40e0-adef-03cdb5a713de · inbound

FIGMA: Towards FIne-Grained Music retrievAl cites this paper.

FIGMA: Towards FIne-Grained Music retrievAl MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.211145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T23:32:09.401023Z digest=sha256:ffe9e333ae0b6f12d4b8ff4a6b023a48fe31a5c2ab4f758bfde432645a99c960

Observation 8315dcc1-864c-4feb-82bc-24b654613d1f · inbound

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance cites this paper.

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T06:45:43.330341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:45:43.330341Z digest=sha256:6e4b86820df4d1940e226912431cc98e8e6bb1b92b9f39acd5a2728c996a9676

Observation 5e472183-eca2-4d69-b17c-ae624a4b90de · inbound

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance cites this paper.

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T07:05:05.133233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:05:05.133233Z digest=sha256:c9a022a9a8ac860dd7c8d65dabbd5cf81d0ffa6394f5bca749495f3a80f1e3e7

Observation 4f81140f-afe0-49b6-8345-0031d08cab2b · inbound

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model cites this paper.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.907977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.907977Z digest=sha256:6aa15f87c0250efa1c1cd547a384c9f92d61c31d976fe2dc6b3d0c1b3b024535