Pith. sign in

Paper Citation Record · LEDGER

Music-JEPA: Learning a World Model of Sound from Action

As of 22 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2607.22000.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22000 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:12:47.611011Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:12:40.962023Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af5362bb-95d7-4d1d-b047-3ad8e452d148 · outbound

This paper cites Music-JEPA: Learning a World Model of Sound from Action.

Music-JEPA: Learning a World Model of Sound from Action Music-JEPA: Learning a World Model of Sound from Action

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:40.962023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:40.962023Z digest=sha256:8e272dfa16ca8618cfd6e0eaa2971d54a6f87a5162e8714d214645d811b9d950

Observation 593cc8b4-7f03-43c8-9805-c02377f9c65b · outbound

This paper cites an unresolved cited work.

Music-JEPA: Learning a World Model of Sound from Action Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:41.072877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:41.072877Z digest=sha256:314784e74202a1ed05bd9fd5ddd42c8225f5bf6853525e122ffa891781f883a7

Observation 9310074f-1716-43ea-aff3-6dd70e1fd63a · outbound

This paper cites an unresolved cited work.

Music-JEPA: Learning a World Model of Sound from Action Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:41.205908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:41.205908Z digest=sha256:6280c880ce2c292d815057e9cd961f50186d2f32c911356b01280f101a281ecf

Observation 8467e091-8626-4951-83e0-390fdb8bb302 · outbound

This paper cites We first describe the dataset, training procedure (Section 4.1), and baselines (Section 4.2).

Music-JEPA: Learning a World Model of Sound from Action We first describe the dataset, training procedure (Section 4.1), and baselines (Section 4.2)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:41.324801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:41.324801Z digest=sha256:6e2b12aff8b7afa0bbd5984a0a7539040de936d882020b6dc2ae8bb0b53fcda5

Observation f5d89169-d190-4730-a375-c8576c0df9aa · outbound

This paper cites We demonstrate three key capabilities.

Music-JEPA: Learning a World Model of Sound from Action We demonstrate three key capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:41.453362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:41.453362Z digest=sha256:9bb5e857fdc818540b69772f928d94f2016085b2d978600a780e6ba21a236734

Observation 462bfbe0-92a8-4f21-8846-349753783729 · outbound

This paper cites an unresolved cited work.

Music-JEPA: Learning a World Model of Sound from Action Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:41.590602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:41.590602Z digest=sha256:b4bbcce2a71b9c32cfef6fb13a162d22368a03f64d440a5e6d71006e4a3fc70f

Observation 23b4aaf0-2ba9-41e0-b687-3f51ddca48c3 · outbound

This paper cites Hohwy,The predictive mind.

Music-JEPA: Learning a World Model of Sound from Action Hohwy,The predictive mind

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:41.687186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:41.687186Z digest=sha256:5bc6ff4cfb2abafd0c2ac44a375ecf97d81a14b54a0bfa27c355a54acb6a9baf

Observation e6f91316-3b83-4dae-886a-c0e6c5b8b914 · outbound

This paper cites A path towards autonomous machine in- telligence version 0.9.2, 2022-06-27,.

Music-JEPA: Learning a World Model of Sound from Action A path towards autonomous machine in- telligence version 0.9.2, 2022-06-27,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:41.809200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:41.809200Z digest=sha256:df84c68cfc8d07a0436e8c4d9a4207bee199538a7613e2befd7de891d10268d9

Observation 32630d7b-c87a-456f-94e3-a1723d7dd23b · outbound

This paper cites Self-supervised learning from im- ages with a joint-embedding predictive architec- ture,.

Music-JEPA: Learning a World Model of Sound from Action Self-supervised learning from im- ages with a joint-embedding predictive architec- ture,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:41.930874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:41.930874Z digest=sha256:7685b00feee2a4272523f6e8205ef1fcd90a5b5017b04eb869c1b3adddc9b2d9

Observation b2cc5bb5-4c14-4045-971a-6203a223747f · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Music-JEPA: Learning a World Model of Sound from Action V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:42.094860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:42.094860Z digest=sha256:e8351e63d8eeca692da6e3542beaed7fcd382bafa4f55a7c8005bc0bcef4e473

Observation 0c49ba69-6036-40ff-885a-caa58fac3cbb · outbound

This paper cites A-JEPA: Joint-Embedding Predictive Architecture Can Listen.

Music-JEPA: Learning a World Model of Sound from Action A-JEPA: Joint-Embedding Predictive Architecture Can Listen

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:42.238243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:42.238243Z digest=sha256:54402045fb183623b0039a75733d025e6f0aa7e35b9bbf3692fb5388c0666b6f

Observation 163f1b05-700e-4916-8794-06d0bdd0842c · outbound

This paper cites Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning.

Music-JEPA: Learning a World Model of Sound from Action Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:42.365990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:42.365990Z digest=sha256:a0cf6af0ed675f060ffed76489cb05f1d81f19254ef491a609c0259d9a24646e

Observation cbeb8660-15de-45da-aa3e-bcea2e426376 · outbound

This paper cites Investigating design choices in joint-embedding predictive architectures for general audio repre- sentation learning,.

Music-JEPA: Learning a World Model of Sound from Action Investigating design choices in joint-embedding predictive architectures for general audio repre- sentation learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:42.454661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:42.454661Z digest=sha256:77f370a4a167b316d0a120121c4dec733bd9aff68edd9d71944ea64f80e12d9e

Observation f0abb9ab-8540-4ab7-83f3-d55642f4f899 · outbound

This paper cites Stem-jepa: A joint-embedding predictive architecture for musical stem compatibility estimation,.

Music-JEPA: Learning a World Model of Sound from Action Stem-jepa: A joint-embedding predictive architecture for musical stem compatibility estimation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:42.597033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:42.597033Z digest=sha256:9981ac40ef5012009af16645232f5fe43006b469b69be7301cdda93a335d35f4

Observation ff2fffe6-2cd4-4508-a5fb-a3a954a1c02a · outbound

This paper cites Emergent musical properties of a transformer under contrastive self- supervised learning,.

Music-JEPA: Learning a World Model of Sound from Action Emergent musical properties of a transformer under contrastive self- supervised learning,

Reference 15

Resolution
verified exact
doi, observed 2026-08-01T06:14:13.236348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-01T06:12:42.753490Z digest=sha256:8df2cdc0d82f8168c7fa7edd1d97d9afa4fbbd2270cc583183ba058f01ad98be

Observation dfe89b82-c828-4a57-9547-4cf97014b229 · outbound

This paper cites Contrastive learning of musical representations,.

Music-JEPA: Learning a World Model of Sound from Action Contrastive learning of musical representations,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:42.864492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:42.864492Z digest=sha256:a68777135bf0e15dfc966e39068f53c0b63f39efcea313cd42ebca1012f99467

Observation c62c7420-ea10-4613-91c8-b1ac2e417055 · outbound

This paper cites Masked autoencoders that listen,.

Music-JEPA: Learning a World Model of Sound from Action Masked autoencoders that listen,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:42.989580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:42.989580Z digest=sha256:ed730a04d687baea7a18215feac3ce94ceefffddc28fe6cefcd10f7f249845e0

Observation 156b46a4-3d01-4dba-b295-8bd853c367dc · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,.

Music-JEPA: Learning a World Model of Sound from Action MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:43.155786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:43.155786Z digest=sha256:e284f3582db8bb623ee8cbad8838376eb97d65bcb0e911bf69da8c40d5a93c1a

Observation 94a03cc4-6a09-411f-ab1a-3ab2fcaa2c62 · outbound

This paper cites Exponential moving av- erage normalization for self-supervised and semi- supervised learning,.

Music-JEPA: Learning a World Model of Sound from Action Exponential moving av- erage normalization for self-supervised and semi- supervised learning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:43.299760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:43.299760Z digest=sha256:1b06917146f4ce4a2563c8a6ad2d3f48eccf09b1df880fc7908ebff56a15ee7b

Observation 246a3af9-8cc6-44cd-ade5-42df6ef4906a · outbound

This paper cites Iterative amortized policy optimization,.

Music-JEPA: Learning a World Model of Sound from Action Iterative amortized policy optimization,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:43.425331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:43.425331Z digest=sha256:0edfd4ed58cf5e2fdd8f303ad71aa30867bfd7615fc0f1219b130a59c3cf59c5

Observation a2cf254a-9ada-4b50-b6b7-880df049d17f · outbound

This paper cites High-resolution sustain pedal depth estimation from piano audio across room acoustics,.

Music-JEPA: Learning a World Model of Sound from Action High-resolution sustain pedal depth estimation from piano audio across room acoustics,

Reference 21

Resolution
verified exact
doi, observed 2026-08-01T06:14:12.970082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-01T06:12:43.570627Z digest=sha256:9d7affc222a08cb06d1ec42cf0678db127fb048f3627ee4733c6ee0a5f509315

Observation f6f71b0d-2307-4902-b745-7082e314e94b · outbound

This paper cites Critique of World Model.

Music-JEPA: Learning a World Model of Sound from Action Critique of World Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:43.725257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:43.725257Z digest=sha256:8c2aaaac87f12cc96fe304069d9488729c661d3acbb954904a9dc53a57eb90b4

Observation 8a428ec8-1cc0-4fc3-8f65-79261938f527 · outbound

This paper cites Variance-Covariance Regularization Improves Representation Learning.

Music-JEPA: Learning a World Model of Sound from Action Variance-Covariance Regularization Improves Representation Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:44.487815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:44.487815Z digest=sha256:f9117fb5218d6628b516642d5232d17226320eb1147882f9c44600d44e4d69e2

Observation 13c6c331-710a-4d62-a806-d0e079823859 · outbound

This paper cites PAN: A world model for general, interactable, and long-horizon world simulation,.

Music-JEPA: Learning a World Model of Sound from Action PAN: A world model for general, interactable, and long-horizon world simulation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:43.975885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:43.975885Z digest=sha256:d39f243c7f19fff0c2653624a2bf2d2d07bf5f0759694e0ebc75106b5db65a57

Observation d89bca9d-453e-47a2-9976-f60f22d1e218 · outbound

This paper cites Pointworld: Scaling 3d world models for in-the-wild robotic manipulation,.

Music-JEPA: Learning a World Model of Sound from Action Pointworld: Scaling 3d world models for in-the-wild robotic manipulation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:44.082214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:44.082214Z digest=sha256:1ea9f1d5b86b785d7916070dc2ca08b188136d0aa685f4e99ac823548f10fa83

Observation dc765bb4-dc8b-4fba-b9c3-8637aa5e78b8 · outbound

This paper cites A theory of cortical responses,.

Music-JEPA: Learning a World Model of Sound from Action A theory of cortical responses,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:44.212978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:44.212978Z digest=sha256:6d03d80114ac2680277fe9a7fd5710e0aa3108451140bc35a548941d1e2e03a6

Observation 6a24290f-0b31-4273-850d-d2c635c0b7bc · outbound

This paper cites an unresolved cited work.

Music-JEPA: Learning a World Model of Sound from Action Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:44.328214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:44.328214Z digest=sha256:9bdf1c5a4c3efcd621cde89c20c9cad6683a7131476db92a736a247ed53cd187

Observation 9b4f6962-1111-4dbe-85bd-83b6011ff5c9 · outbound

This paper cites London,Hearing in Time: Psychological Aspects of Musical Meter.

Music-JEPA: Learning a World Model of Sound from Action London,Hearing in Time: Psychological Aspects of Musical Meter

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:44.391764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:44.391764Z digest=sha256:775b61d22c9a47909c8e37e4b2701ce5ba91f1a3f80d2bd9c6663170ec21b687

Observation 4bc32673-fe35-43c8-b288-f6313b164ceb · outbound

This paper cites Unsupervised disentanglement of content and style via variance- invariance constraints,.

Music-JEPA: Learning a World Model of Sound from Action Unsupervised disentanglement of content and style via variance- invariance constraints,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:45.267246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:45.267246Z digest=sha256:07d583ec88543aeefd100daa4db5db054cdd8290648519ad3e6dc071d6dd0b54

Observation 6514da8e-e5ab-445e-9b71-f24898f1bb08 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Music-JEPA: Learning a World Model of Sound from Action An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:45.368586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:45.368586Z digest=sha256:12b8ce4bea8a5a72df94567b0d175d0ede57736d232f1e6e2b923a89d5c40391

Observation d05699c8-aae3-4f8d-a96c-d4ac62eab95b · outbound

This paper cites Emerging properties in self-supervised vision transformers,.

Music-JEPA: Learning a World Model of Sound from Action Emerging properties in self-supervised vision transformers,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:44.749398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:44.749398Z digest=sha256:a5739250d2d073f58dcd367f78c2f0291d3d101bf89a34ea2526666efac3d334

Observation d966f9b0-0e15-4ce9-bcc6-cdcef251e637 · outbound

This paper cites LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics.

Music-JEPA: Learning a World Model of Sound from Action LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:44.829189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:44.829189Z digest=sha256:6af5da64c85412664cdad5cdacbb419a8980b9ebd94d1db5326ac32ff4984ac4

Observation 9038b9cb-4fe1-40c7-8495-8f074d0d2290 · outbound

This paper cites Balancing information preservation and disentangle- ment in self-supervised music representation learn- ing,.

Music-JEPA: Learning a World Model of Sound from Action Balancing information preservation and disentangle- ment in self-supervised music representation learn- ing,

Reference 33

Resolution
malformed identifier
no resolver link, observed 2026-08-01T06:12:44.941200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:44.941200Z digest=sha256:ff6fd8530a8a42205226367a676f433960acb089e4aa32e517c409eb2eb7164b

Observation 68352f67-509d-4375-9439-a0ade8d4962e · outbound

This paper cites Audio barlow twins: Self-supervised audio represen- tation learning,.

Music-JEPA: Learning a World Model of Sound from Action Audio barlow twins: Self-supervised audio represen- tation learning,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:45.078693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:45.078693Z digest=sha256:aa58076f3cbf85511f7edf688094e749324907a92292475333a65c362e2d210e

Observation 0750932b-266f-4471-9fc8-6cd496787145 · outbound

This paper cites PESTO: pitch estimation with self-supervised transposition-equivariant objective,.

Music-JEPA: Learning a World Model of Sound from Action PESTO: pitch estimation with self-supervised transposition-equivariant objective,

Reference 35

Resolution
verified exact
doi, observed 2026-08-01T06:14:12.590023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-01T06:12:45.176714Z digest=sha256:5a687eb72603158f884984bbc901c197b6af2a202f621be919e9e67fa3aacfcd

Observation 6be93974-6482-47b7-829c-3fced9f5bbe1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Music-JEPA: Learning a World Model of Sound from Action Adam: A Method for Stochastic Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:46.534103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:46.534103Z digest=sha256:e3bdd56b71800ec7ba02c3715e5b9fa925473c5140ef23110101009c35f28d34

Observation 5c7452a7-6724-491c-bea4-87ad6124a6fd · outbound

This paper cites Bootstrap your own latent - a new approach to self-supervised learning,.

Music-JEPA: Learning a World Model of Sound from Action Bootstrap your own latent - a new approach to self-supervised learning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:46.694747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:46.694747Z digest=sha256:3ea4860a57c1f986ada2779a7a7e368baac62e867007c27e5fbaf79af58f027a

Observation 441b6b8d-9ae3-4d50-a9e1-3af44a8d294c · outbound

This paper cites ASAP: a dataset of aligned scores and performances for piano transcription,.

Music-JEPA: Learning a World Model of Sound from Action ASAP: a dataset of aligned scores and performances for piano transcription,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:46.826671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:46.826671Z digest=sha256:834e86147a2377af82636c85f683c933bdd2677e4a70e1a6acc070981bb0b92e

Observation 5271a9ff-aece-468e-b047-8c0cb3beeb9c · outbound

This paper cites Closing the train-test gap in world models for gradient-based planning,.

Music-JEPA: Learning a World Model of Sound from Action Closing the train-test gap in world models for gradient-based planning,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:45.670142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:45.670142Z digest=sha256:0b07ccb5de657e1bb32260b4b39613bc8b69c24fd3744b3ce6657293f3977a6b

Observation 7c58a97f-5e86-4c97-b493-11653bb3657a · outbound

This paper cites Tem- poral straightening for latent planning,.

Music-JEPA: Learning a World Model of Sound from Action Tem- poral straightening for latent planning,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:45.810095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:45.810095Z digest=sha256:bd7a9a141608682bdd052ed425a7fb65bbb61c715d8e23b303eba85aecf04089

Observation 2fbcf5b4-4e5f-43c5-a457-408cb5b44bc8 · outbound

This paper cites Tutorial on amortized optimization for learning to optimize over continuous domains,.

Music-JEPA: Learning a World Model of Sound from Action Tutorial on amortized optimization for learning to optimize over continuous domains,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:46.005798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:46.005798Z digest=sha256:5dedd39875a8b8315efeaf83bd2dbabc9a36c6b074b980f8ea0e1eec576a4483

Observation a9321afb-314f-4ed5-bea3-b7eb62499eda · outbound

This paper cites Latent Geometry Beyond Search: Amortizing Planning in World Models.

Music-JEPA: Learning a World Model of Sound from Action Latent Geometry Beyond Search: Amortizing Planning in World Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:46.195503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:46.195503Z digest=sha256:76720cb3fa2045e6144f6ea53cbb35f1ab60ccbbeacd202f7498bcbaa1f4c392

Observation b1c09ccd-d0d4-401b-9d21-d04dafb52111 · outbound

This paper cites Enabling factorized piano music modeling and generation with the MAE- STRO dataset,.

Music-JEPA: Learning a World Model of Sound from Action Enabling factorized piano music modeling and generation with the MAE- STRO dataset,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:46.366841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:46.366841Z digest=sha256:dad34a9cbb2cd15d95e9fd3d7e6761b197c20dda96ed351c478e3c3513149877

Observation ede7d27e-11cb-4756-958c-b7b6e30ccba9 · outbound

This paper cites Beat this! accurate beat tracking without DBN postprocessing,.

Music-JEPA: Learning a World Model of Sound from Action Beat this! accurate beat tracking without DBN postprocessing,

Reference 48

Resolution
verified exact
doi, observed 2026-08-01T06:14:12.164925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-01T06:12:47.066148Z digest=sha256:7171cbdf8e7a826878e17678dac6da233981243c6b56136bf1fcf616b6796be8

Observation b3cc2071-3d37-4f1b-8568-716014f3b21d · outbound

This paper cites madmom: a new Python Audio and Mu- sic Signal Processing Library,.

Music-JEPA: Learning a World Model of Sound from Action madmom: a new Python Audio and Mu- sic Signal Processing Library,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:47.188376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:47.188376Z digest=sha256:df17b1eb4c8f3fbe13228f98bc126aaebd08ba12c118ceebe5e81371d57f6dd5

Observation eff797c3-b216-4486-bd87-9cf6ac6e3613 · outbound

This paper cites High-resolution piano transcription with pedals by regressing onset and offset times,.

Music-JEPA: Learning a World Model of Sound from Action High-resolution piano transcription with pedals by regressing onset and offset times,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:47.335715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:47.335715Z digest=sha256:449203f3842cfc3e03d0b65344a5844b6a8d8b528b3ac1b1b946e96086cc8941

Observation ed5d03ec-d920-41ad-aec3-e60e99a8f1d3 · outbound

This paper cites Available: https://doi.org/10.1109/ TASLP.2021.3121991.

Music-JEPA: Learning a World Model of Sound from Action Available: https://doi.org/10.1109/ TASLP.2021.3121991

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-01T06:12:47.498370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:47.498370Z digest=sha256:c809e538e01907d2ddae5eadaeeab953899840379ce3fc1862caa881be2f6e8c

Observation 8581351a-fa08-430a-872f-cf75d78ed15a · outbound

This paper cites Scoring time intervals using non-hierarchical transformer for automatic piano transcription,.

Music-JEPA: Learning a World Model of Sound from Action Scoring time intervals using non-hierarchical transformer for automatic piano transcription,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:47.611011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:47.611011Z digest=sha256:957d19487bc15abd4b2882156939fdec7fae6ee323a77a337f5c1fbc62f3d4f5

Observation aa78319c-ff7a-49b9-bd36-49d39fa6d2bc · outbound

This paper cites Available: http://archives.ismir.net/ ismir2020/paper/000127.pdf.

Music-JEPA: Learning a World Model of Sound from Action Available: http://archives.ismir.net/ ismir2020/paper/000127.pdf

Reference 541

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:46.922300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:46.922300Z digest=sha256:4de888c6d7a7c06f255e98b3cbdc2a1f0ace09b586b836cbbabf6a2137295418

Observation 28ab99fc-51d1-4b8f-916b-d5a64efef46c · outbound

This paper cites [Online].

Music-JEPA: Learning a World Model of Sound from Action [Online]

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:45.471023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:45.471023Z digest=sha256:eb5ac1e0b69c73ec2a1aa724903e72c8a0b70da2cbc4ba17eb667454130b8b1a

Observation ea6ddf7b-cf18-4572-b4dc-7a4905e03cbb · outbound

This paper cites Variance-Covariance Regularization Improves Representation Learning.

Music-JEPA: Learning a World Model of Sound from Action Variance-Covariance Regularization Improves Representation Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:44.623623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:44.623623Z digest=sha256:85c6f8218e2ec00196cfc701b4306cc3eb8c850f57d4e6469401077b92cbf0d6

Observation 8c384633-909e-4c97-af7b-3a1a2a9b14f7 · outbound

This paper cites Critique of World Model.

Music-JEPA: Learning a World Model of Sound from Action Critique of World Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:43.862975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:43.862975Z digest=sha256:70650f1eab4d5ccfe039d2f2a28ec1b7dd84200b718a0296fa245521fe389a9a

Pith citing papers

Observation af5362bb-95d7-4d1d-b047-3ad8e452d148 · inbound

Music-JEPA: Learning a World Model of Sound from Action cites this paper.

Music-JEPA: Learning a World Model of Sound from Action Music-JEPA: Learning a World Model of Sound from Action

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:40.962023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:40.962023Z digest=sha256:8e272dfa16ca8618cfd6e0eaa2971d54a6f87a5162e8714d214645d811b9d950