Pith. sign in

Paper Citation Record · LEDGER

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 8 inbound Pith citation observations for arXiv:2507.02915.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02915 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:56:46.342431Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:33.047406Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:50:12.629815Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact12
  • verified fuzzy8
  • unresolved20
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f1a916e-e772-4f3d-8282-a493595b354f · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.224028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.224028Z digest=sha256:bc9bbc1b9b3794c5d73ea9d750e4b445373db7a76a8084e232694a26322f2930

Observation 9ceb76b9-7dbc-4c05-9baa-86471b24b068 · outbound

This paper cites CED: Consistent ensemble distillation for audio tagging.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning CED: Consistent ensemble distillation for audio tagging

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:47.266196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.227882Z digest=sha256:6f1079ce5947586a2981fe3d550c30a98ccee5af934c7643b24b9b124cc9c632

Observation b46a13a5-5276-415c-a05e-6a163b4ef094 · outbound

This paper cites Scaling up masked audio encoder learning for general audio classification.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Scaling up masked audio encoder learning for general audio classification

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.230938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.230938Z digest=sha256:bdb5b75da5f1fe3313cd66f32c1e5b15f3b8f9ec5d6a1ee877de64bf4fe4c472

Observation 79718107-558b-4fc2-916e-de189db3939b · outbound

This paper cites Masked Modeling Duo: Towards a Universal Audio Pre-training Framework.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Masked Modeling Duo: Towards a Universal Audio Pre-training Framework

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:47.251153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.234160Z digest=sha256:6b16912576f9ddef5e2cf49168f88d61a9f874edbea49066872f386f2da1470b

Observation fb09b00d-30d2-47af-9ac9-1eb9bb433875 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.237565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.237565Z digest=sha256:78aba7ae6253a7aba74ca7681a0b25d30636c9c18c9c6015d5bad04978f9183a

Observation 8461dad3-f481-4f9a-8785-b4b94006be8c · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.241459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.241459Z digest=sha256:8a39787c18538ebd13b03d65641811d022bcc0fa2f4c9c262fdff4dc24ae3883

Observation 4615b1ab-4aa0-48a8-b010-a13417829438 · outbound

This paper cites data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.244140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.244140Z digest=sha256:851f2defde1a16629b818474179f00caaeec61d804edc63876246d6434b3b239

Observation 148ab8b8-5667-4689-ba24-ac83137fca14 · outbound

This paper cites Efficient Self-super- vised Learning with Contextualized Target Representations for Vision, Speech and Language.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Efficient Self-super- vised Learning with Contextualized Target Representations for Vision, Speech and Language

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.339338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.247277Z digest=sha256:4bbae0ef666334f6b5baa3ce33221eec749f319baf9fb1cba4ececbe000a8a58

Observation e1896868-56f7-43f6-85f3-e72bd47c4294 · outbound

This paper cites A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.332558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.249998Z digest=sha256:f03874efa8d8e8b6251c5afc729647f9e5ae3e0d41c5972c91eef3622a966a0a

Observation 02ad345b-f45c-48aa-a64c-c36c78f35ebb · outbound

This paper cites Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.252958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.252958Z digest=sha256:dab4e99e2591277de66bc76160cf70bdb56a149baa7cb99d6ca7448d48e5ee9d

Observation ba276dbc-f849-4046-b967-aaad6c0b2f1f · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.256060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.256060Z digest=sha256:3f948e960ad2662d6672e862a3d349cbd1e9391bbedc77dc522d31aca57cce5c

Observation 0220be99-caa1-4cc6-aafe-1062babe05ec · outbound

This paper cites A-JEPA: Joint-Embedding Predictive Architecture Can Listen.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning A-JEPA: Joint-Embedding Predictive Architecture Can Listen

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.325028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.259155Z digest=sha256:4a457a5c38bcfd71a13072d3dd8d57e2449ea795677d19e2f9071a5a76fd3c1c

Observation 4346d857-951e-48c1-a778-886ee9227d0c · outbound

This paper cites Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:47.136557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.261794Z digest=sha256:c19b26cc3745d489c3d4ca00ec5103b0aee8d6ba1a815d969e7113a6dddbf59c

Observation f0292ae4-b77a-4c7d-af57-a5001062a6ac · outbound

This paper cites Masked Autoencoders that Listen.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Masked Autoencoders that Listen

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.264640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.264640Z digest=sha256:acf7fda4d149b6ac0423f84922686e2ae9c6ae5951154b9b37f6e42fa4981248

Observation c7a8b073-5e16-4d78-8d5c-3c06e9d4e34a · outbound

This paper cites A Dataset and Taxonomy for Urban Sound Research,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning A Dataset and Taxonomy for Urban Sound Research,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.267285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.267285Z digest=sha256:4268534fc30a43b182b4828e4d2671d5a1011a45f5af87f84da66608205a90fa

Observation 3f63ad7f-c74a-4ac1-9f15-6748b3153b20 · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning VoxCeleb: a large-scale speaker identification dataset,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.270842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.270842Z digest=sha256:92b0e991c1739a0e6087f516b0ca7fd535e6254f55161ab6ee773cc838d2cb5c

Observation 2c3fbb49-5d1e-4fb5-8a87-4c07d3fb9534 · outbound

This paper cites The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.273343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.273343Z digest=sha256:752cd810c0b8a2cba7212852be8d02abd43bb63440922d8c8cd86035ff57e7e1

Observation 48c91fef-02c3-4ea6-b197-ebc56a608317 · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.317711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.276048Z digest=sha256:262dde605aeebf844aeb3794a8a5b6d449f82d63fbb1f468420140bb441e1f9b

Observation 7e524576-c9a8-49e3-b7e4-fb065bb0e736 · outbound

This paper cites TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:46.989062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.278586Z digest=sha256:2c897fd70ff06ff403492e023463c28dbb84e67c1301c99f39c6ebb4c2e7f5f0

Observation 6463c55d-7c7d-4db6-932f-8ac0213d3ecf · outbound

This paper cites GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:46.976297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.282310Z digest=sha256:47e4d8baa7a2d3f7bffffddf71a7d0903d836618d1ecd212c48d142e8abe496f

Observation 62a306d7-63fc-405c-8557-b8c1a04d2a8a · outbound

This paper cites Bootstrap your own latent: A new approach to self- supervised Learning.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Bootstrap your own latent: A new approach to self- supervised Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.309278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.284987Z digest=sha256:0e8f62812bd6f33d872abdc01f36b9fb0ed917d4e3d8e6155e951c3feb3e998a

Observation a7d97c8e-cda5-4076-80b0-a38f88d9d71d · outbound

This paper cites Decoupled Weight Decay Regularization.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Decoupled Weight Decay Regularization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.287631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.287631Z digest=sha256:74ae074aca009f25d717f9c7f89c58298d9211b7197d1e4092f5d0d870e354c6

Observation 43fa2dff-f79e-4711-a5f3-4176d3548695 · outbound

This paper cites an unresolved cited work.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:56:47.302159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.290010Z digest=sha256:4cb377cced8e9af117cbac24b1460a672ba5e12967a291a9b869bf26e08b2459

Observation 84372658-ed8d-4e29-b5b9-d66d883a3b6b · outbound

This paper cites Clotho: An Audio Captioning Dataset.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Clotho: An Audio Captioning Dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.295352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.295352Z digest=sha256:1f60bdb80b1f2929ec68de44563ae0b13bed1a0ad09961b210d5fc9efc808c98

Observation 0365b6ed-a8e6-48e8-a8ea-880de64dcfe8 · outbound

This paper cites CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.298919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.298919Z digest=sha256:06822be18d0abb94aa7cb81a8504a60cf055081bed7af1f4e707f6b092b590e6

Observation eaef854a-9b7d-40d9-9227-e85ca872837c · outbound

This paper cites Generating an item pool for translational social cognition research: methodology and initial validation,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Generating an item pool for translational social cognition research: methodology and initial validation,

Reference 26

Resolution
verified exact
doi, observed 2026-08-06T22:56:46.365463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.301331Z digest=sha256:fcfef6bf7b984f05d7a965d7289e79527b81269a8872080882153a18eccf12d4

Observation 2c57c113-8cee-43b9-b57c-71e60fb42b58 · outbound

This paper cites Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.287063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.304258Z digest=sha256:90ca37d85a99257df1bf0fe205630dcceac94d8bc25ca0d4ac9a56953a876e42

Observation ad610024-4d77-49cf-a3ca-defdbfe393ad · outbound

This paper cites Sound event detection in synthetic domestic environments,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Sound event detection in synthetic domestic environments,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.279730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.306811Z digest=sha256:bd51512ae88476bfda0d27dd6629b4c71a53165a328ddc3ec76afb1ed58cb9a4

Observation 971823ec-6826-4d3f-b24e-acbbf1fad5a3 · outbound

This paper cites ESC: Dataset for Environmental Sound Classification,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning ESC: Dataset for Environmental Sound Classification,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.309460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.309460Z digest=sha256:2a288fef8438f3da50c0ed6b3df2fcb590f4450594305b42781dc0dc463e96dc

Observation 9226004b-7c4c-4480-ba4e-5aa60fa31cbb · outbound

This paper cites FMA: A Dataset For Music Analysis.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning FMA: A Dataset For Music Analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.311880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.311880Z digest=sha256:2ba73147e9e9478c2d4267947567488b28e15d7d473a781280c41f392745020a

Observation 6e457330-e699-4686-be90-3df5856309b2 · outbound

This paper cites General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.314564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.314564Z digest=sha256:036e889ac9ee4da6d22ca736a95d21689cdc6036e0001280f00c795c39f50ad6

Observation bee0be84-1bd7-4e1e-a8d3-fcb2c2b3ce93 · outbound

This paper cites FSD50K: An Open Dataset of Human-Labeled Sound Events.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning FSD50K: An Open Dataset of Human-Labeled Sound Events

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.317985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.317985Z digest=sha256:650100980cf34573b414b8f628a585cc87d93c391d1d71afec818a1740303fc2

Observation 2d5917d5-1fd9-460d-9a4a-134bdf7d73b6 · outbound

This paper cites LibriCount, a dataset for speaker count estimation.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning LibriCount, a dataset for speaker count estimation

Reference 33

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:56:46.821091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.321292Z digest=sha256:70cc86127d296bb585647074a4fdf15c82918e8b59ec0d7708e88b6d936c8f76

Observation e17815ff-6608-4292-9d68-a67f2e4f04ae · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Librispeech: An ASR corpus based on public domain audio books,

Reference 34

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:56:46.323549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.323549Z digest=sha256:3071b75e473be091786d6832ac5398c3643a4334759b4547e44627ad82e991a7

Observation f75daadd-0aeb-4571-a049-2a7d46df962b · outbound

This paper cites Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.327119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.327119Z digest=sha256:401c5744c89c94a4ccaedeb362d6595d285667cdeafa4d84ce3f703f0a6dc164

Observation 25bd55ea-7fa6-40bd-8e6e-ed354dceb3b8 · outbound

This paper cites The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS).

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS)

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:56:46.703751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.329589Z digest=sha256:2406c451502c33a55aad9de4244e5f4d249fb75f8ee6981cf368319376ed6e3e

Observation 32864c85-9e8c-409f-8116-41152284d54f · outbound

This paper cites Vocal Imitation Set v1.1.3 : Thousands of vocal imitations of hundreds of sounds from the AudioSet ontology.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Vocal Imitation Set v1.1.3 : Thousands of vocal imitations of hundreds of sounds from the AudioSet ontology

Reference 37

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:56:46.623337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.332120Z digest=sha256:9d5b6fedf8fc9453432279249b78a15376671e892a12470801136af115894e86

Observation 614e6a51-cad4-40d4-b9bc-255abbffd4eb · outbound

This paper cites Vocalsound: A Dataset for Im- proving Human Vocal Sounds Recognition,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Vocalsound: A Dataset for Im- proving Human Vocal Sounds Recognition,

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:56:46.334390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.334390Z digest=sha256:03f22a6756088ff70655409703b1e52a346861480e81f46af985dd60c88a869b

Observation 26837dd3-70f9-4a23-89b5-4790b3d9fb14 · outbound

This paper cites voxlingua33 in WebDataset Format.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning voxlingua33 in WebDataset Format

Reference 39

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:56:46.486803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.336713Z digest=sha256:6371639b1852f82a3544b8328a4a98d7e22cd5f846d98e789389de3c75829b13

Observation ab31d341-a210-4f98-a294-fbf0fd3e9782 · outbound

This paper cites ConvFormer: Plug-and-Play CNN-Style Transformers for Improving Medical Image Segmentation.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning ConvFormer: Plug-and-Play CNN-Style Transformers for Improving Medical Image Segmentation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:46.392405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.338990Z digest=sha256:a147a0af3259953a3a83893fcef84cc315bc4bed5a20ce12bc8f03b5ade83b92

Observation dbd8b5b8-6aed-4331-a344-773037ffc781 · outbound

This paper cites MetaFormer Baselines for Vision.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning MetaFormer Baselines for Vision

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:46.381208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.342431Z digest=sha256:d8066518a7f982421228412543c472fcb087893c0f043a149f1848c96c41f43d

Observation b4a2ae34-ed58-4a95-8c26-9d91d823260f · outbound

This paper cites Available: https://datashare.ed.ac.uk/handle/ 10283/853.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Available: https://datashare.ed.ac.uk/handle/ 10283/853

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.294245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:56:46.292291Z digest=sha256:dd3b2b7b8b1ca3ecd520cd39db9823758924de371f6a48692fcdd49383f76d02

Pith citing papers

Observation 72f1a689-e70f-43c9-9a42-817712dc2b54 · inbound

Self-Distillation of Hidden Layers for Self-Supervised Representation Learning cites this paper.

Self-Distillation of Hidden Layers for Self-Supervised Representation Learning Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T18:11:35.194361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:11:35.194361Z digest=sha256:ef6006d8c721ba051a2668506168d87fa46abb5bb372baabc82ddb437623ea00

Observation 27c28636-ee95-4fc2-9ef2-e1271c8ca576 · inbound

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space cites this paper.

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:06:23.815592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T12:54:49.511047Z digest=sha256:006a0658c94c7089d20daf3ac590cfa2be9b0962c0d26310bad58abdd40803bc

Observation d5760ae9-19fb-428a-aad5-0ae7642e0cc4 · inbound

Frequency-Aware Self-Supervised Music Representation Learning cites this paper.

Frequency-Aware Self-Supervised Music Representation Learning Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:12.631693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T19:25:54.322923Z digest=sha256:ae3f901983fad64d90ddb707998235a94a5f74b8be931fa71e05be93f8252844

Observation 1bb2d02d-8af9-44e5-8425-e017f2a60551 · inbound

Frequency-Aware Self-Supervised Music Representation Learning cites this paper.

Frequency-Aware Self-Supervised Music Representation Learning Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:04:35.643472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T10:04:11.233115Z digest=sha256:6290bb39d389ac54fcb98ee78e8fd152a01a28bd101632bee1ad5e3ad74d8b16

Observation 8b7c4702-41a1-4256-b744-ec10d5d77758 · inbound

Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition cites this paper.

Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T22:46:06.299864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:46:06.299864Z digest=sha256:8d2b38d4f11919eb883c220208b945a45cb4ea98a697026ac4cc797592659257

Observation 163f1b05-700e-4916-8794-06d0bdd0842c · inbound

Music-JEPA: Learning a World Model of Sound from Action cites this paper.

Music-JEPA: Learning a World Model of Sound from Action Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:42.365990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:42.365990Z digest=sha256:7d171831a1975d25f894f63762957ca3c0963e21ff3025cab6689c6239d5e685

Observation d67cc45a-d8e6-42ec-9e0e-eab62d207d31 · inbound

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds cites this paper.

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:22.852047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:40:22.852047Z digest=sha256:7ecfaf76e41d839bdd53a2a3c75d32b355b1b28c53cf972627caafe556a11399

Observation 1c325d45-a496-43b1-8277-e3b12ce316d7 · inbound

Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture cites this paper.

Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.047406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:44:33.047406Z digest=sha256:d22e88c6b38c03f4161d02a74ec9366241d7c26a91ae9abfb74e0c1f99b52415