Pith. sign in

Paper Citation Record · LEDGER

MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2306.00107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.00107 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:29:46.662633Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 648dad47-c0e2-4a1d-8b0e-2897f07b607f · inbound

SALMONN: Towards Generic Hearing Abilities for Large Language Models cites this paper.

SALMONN: Towards Generic Hearing Abilities for Large Language Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:29:46.289406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T02:29:46.242983Z digest=sha256:15299ce00b758620639d80ad4d8b7b031419284a5222269d3a5e46c21c264e3b

Observation 1aff77a5-2bc7-4da2-9f0f-ec9d7ff429c2 · inbound

Exploring How Audio Effects Alter Emotion with Foundation Models cites this paper.

Exploring How Audio Effects Alter Emotion with Foundation Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:34:52.424527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T12:34:31.239629Z digest=sha256:1b586c09eca180a367a020c164b4ac7e152abc94fafbfdb07b26503f77f9fde4

Observation 466a9129-a629-42a0-99e4-52db9dc7a8e6 · inbound

Assessing Factual Music Comprehension in Large Audio Language Models cites this paper.

Assessing Factual Music Comprehension in Large Audio Language Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T00:29:46.662633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:29:46.662633Z digest=sha256:8de5dc4d273d0b30e6f224e2212cba690c71ae888789b7e886285a0af856c9ba

Observation 7a0a1d41-2097-49f4-9dcc-a14f3e366eb4 · inbound

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity cites this paper.

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T11:40:03.186955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T11:36:06.967549Z digest=sha256:4b584d8cff980ef57198252aa52aa989a2d4736e56a13dfa811f07d07f6c9e47

Observation 5fad3ff9-fb68-4ccc-a8bb-054687c776bd · inbound

Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis cites this paper.

Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T17:17:10.597653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:17:10.597653Z digest=sha256:45660faf29fb49fcd069f240bdf4892c650b2c8236d931dc527d43ba7635547c

Observation cce0c1e3-c4fe-4243-b464-6a55876b087a · inbound

ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics cites this paper.

ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:26:59.772765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T07:22:07.949007Z digest=sha256:370b783df3235bff3f5f26080386da83dcbc83a6fab2fd1e5919d3fb17f9131b

Observation 9c5f152f-f46a-4ba5-857a-118d3f5f3fb7 · inbound

Revisiting Content-Based Music Recommendation: Efficient Feature Aggregation from Large-Scale Music Models cites this paper.

Revisiting Content-Based Music Recommendation: Efficient Feature Aggregation from Large-Scale Music Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:40:30.796961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T02:39:46.434831Z digest=sha256:3f330c2bb80bab111cc99d0ec4f1a5c56f44d75a4ce11d8a78e6a9f2d9db4d74

Observation 82206991-8955-4217-90b9-8f2ba3d275c4 · inbound

Adopting State-of-the-Art Pretrained Audio Representations for Music Recommender Systems cites this paper.

Adopting State-of-the-Art Pretrained Audio Representations for Music Recommender Systems MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:01:13.325918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T07:17:42.922029Z digest=sha256:099e41e2dc5adaa18730de902d8ba1c5e9ed1a4485ee3266ba87fc62e5cb41d8

Observation afb7bdbf-5781-4f19-8bd7-74eddb440120 · inbound

ARIA: A Diagnostic Framework for Music Training Data Attribution cites this paper.

ARIA: A Diagnostic Framework for Music Training Data Attribution MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:27:43.044419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T18:24:29.290904Z digest=sha256:3bc56103a37543b483315abdc2d5f78abfe0c0c115f991dfeb8f21c51f305340

Observation e292391c-f19b-4486-a2fb-a0b7211ec3ad · inbound

MERIT: Learning Disentangled Music Representations for Audio Similarity cites this paper.

MERIT: Learning Disentangled Music Representations for Audio Similarity MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:43:32.564872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T15:40:57.719214Z digest=sha256:8db465d7a3fd7974dda11792b7058e2a497dec4e6cd7a634c7c6599952d75482

Observation e2d95631-def2-45f2-b5ff-21ce2edf6bf4 · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.655976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:e6f8a157ec7199785725e0abe14e63a241b3844bdbbab4cd204adef65a3a0278

Observation 5fbd5ea8-fb63-4e40-b732-416492efbf74 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.890564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:0cbba89efa23524bb5ab00ef47fc52f245960954b0a5181fada14508fc0cc4fd

Observation 6dfb6e8a-eceb-49a8-86e2-e512d2929b61 · inbound

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations cites this paper.

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T22:56:37.693891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T22:50:19.910504Z digest=sha256:4b51f9fee8f29118be42923faa7101d372bfd706528ca8bffae2095981ee8123

Observation d83fa3d8-300e-4e25-b351-e4031f6ae035 · inbound

Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment cites this paper.

Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T10:59:22.914974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:59:22.914974Z digest=sha256:867fef87abdfdbb1c0ccbb3f85bb8acd9d8860868534d491ce14d4393827e8d3

Observation b1f2f3da-f1e6-48bc-88ac-3a7e80071f5a · inbound

StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems cites this paper.

StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T22:48:42.353079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:48:42.353079Z digest=sha256:71c5f755ba84e327fdc373f7fd723de198dbe69b9ede934726764042bd4f7d48

Observation c819db37-9690-4c80-8a36-6ab584be0a12 · inbound

Do Music Foundation Models Embed Pitch in Helical Structure? cites this paper.

Do Music Foundation Models Embed Pitch in Helical Structure? MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T13:57:15.866365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:57:15.866365Z digest=sha256:a4a45047ccf64eb0c7eb197afad6a51a8a6c631cd5ff2b1b293da22e40110ccc