Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:56:46.342431Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 8 inbound Pith citation observations for arXiv:2507.02915.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:56:46.342431Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:33.047406Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:50:12.629815Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2f1a916e-e772-4f3d-8282-a493595b354f · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ceb76b9-7dbc-4c05-9baa-86471b24b068 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning CED: Consistent ensemble distillation for audio tagging
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b46a13a5-5276-415c-a05e-6a163b4ef094 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Scaling up masked audio encoder learning for general audio classification
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79718107-558b-4fc2-916e-de189db3939b · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb09b00d-30d2-47af-9ac9-1eb9bb433875 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8461dad3-f481-4f9a-8785-b4b94006be8c · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4615b1ab-4aa0-48a8-b010-a13417829438 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 148ab8b8-5667-4689-ba24-ac83137fca14 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Efficient Self-super- vised Learning with Contextualized Target Representations for Vision, Speech and Language
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e1896868-56f7-43f6-85f3-e72bd47c4294 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02ad345b-f45c-48aa-a64c-c36c78f35ebb · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba276dbc-f849-4046-b967-aaad6c0b2f1f · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Revisiting Feature Prediction for Learning Visual Representations from Video
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0220be99-caa1-4cc6-aafe-1062babe05ec · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning A-JEPA: Joint-Embedding Predictive Architecture Can Listen
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4346d857-951e-48c1-a778-886ee9227d0c · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0292ae4-b77a-4c7d-af57-a5001062a6ac · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Masked Autoencoders that Listen
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7a8b073-5e16-4d78-8d5c-3c06e9d4e34a · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning A Dataset and Taxonomy for Urban Sound Research,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f63ad7f-c74a-4ac1-9f15-6748b3153b20 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning VoxCeleb: a large-scale speaker identification dataset,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c3fbb49-5d1e-4fb5-8a87-4c07d3fb9534 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c91fef-02c3-4ea6-b197-ebc56a608317 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e524576-c9a8-49e3-b7e4-fb065bb0e736 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6463c55d-7c7d-4db6-932f-8ac0213d3ecf · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 62a306d7-63fc-405c-8557-b8c1a04d2a8a · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Bootstrap your own latent: A new approach to self- supervised Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a7d97c8e-cda5-4076-80b0-a38f88d9d71d · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Decoupled Weight Decay Regularization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43fa2dff-f79e-4711-a5f3-4176d3548695 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 84372658-ed8d-4e29-b5b9-d66d883a3b6b · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Clotho: An Audio Captioning Dataset
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0365b6ed-a8e6-48e8-a8ea-880de64dcfe8 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaef854a-9b7d-40d9-9227-e85ca872837c · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Generating an item pool for translational social cognition research: methodology and initial validation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2c57c113-8cee-43b9-b57c-71e60fb42b58 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad610024-4d77-49cf-a3ca-defdbfe393ad · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Sound event detection in synthetic domestic environments,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 971823ec-6826-4d3f-b24e-acbbf1fad5a3 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning ESC: Dataset for Environmental Sound Classification,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9226004b-7c4c-4480-ba4e-5aa60fa31cbb · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning FMA: A Dataset For Music Analysis
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e457330-e699-4686-be90-3df5856309b2 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee0be84-1bd7-4e1e-a8d3-fcb2c2b3ce93 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning FSD50K: An Open Dataset of Human-Labeled Sound Events
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5917d5-1fd9-460d-9a4a-134bdf7d73b6 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning LibriCount, a dataset for speaker count estimation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e17815ff-6608-4292-9d68-a67f2e4f04ae · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Librispeech: An ASR corpus based on public domain audio books,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75daadd-0aeb-4571-a049-2a7d46df962b · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25bd55ea-7fa6-40bd-8e6e-ed354dceb3b8 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS)
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32864c85-9e8c-409f-8116-41152284d54f · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Vocal Imitation Set v1.1.3 : Thousands of vocal imitations of hundreds of sounds from the AudioSet ontology
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 614e6a51-cad4-40d4-b9bc-255abbffd4eb · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Vocalsound: A Dataset for Im- proving Human Vocal Sounds Recognition,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26837dd3-70f9-4a23-89b5-4790b3d9fb14 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning voxlingua33 in WebDataset Format
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ab31d341-a210-4f98-a294-fbf0fd3e9782 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning ConvFormer: Plug-and-Play CNN-Style Transformers for Improving Medical Image Segmentation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dbd8b5b8-6aed-4331-a344-773037ffc781 · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning MetaFormer Baselines for Vision
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b4a2ae34-ed58-4a95-8c26-9d91d823260f · outbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Available: https://datashare.ed.ac.uk/handle/ 10283/853
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72f1a689-e70f-43c9-9a42-817712dc2b54 · inbound
Self-Distillation of Hidden Layers for Self-Supervised Representation Learning Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27c28636-ee95-4fc2-9ef2-e1271c8ca576 · inbound
Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5760ae9-19fb-428a-aad5-0ae7642e0cc4 · inbound
Frequency-Aware Self-Supervised Music Representation Learning Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1bb2d02d-8af9-44e5-8425-e017f2a60551 · inbound
Frequency-Aware Self-Supervised Music Representation Learning Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8b7c4702-41a1-4256-b744-ec10d5d77758 · inbound
Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 163f1b05-700e-4916-8794-06d0bdd0842c · inbound
Music-JEPA: Learning a World Model of Sound from Action Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d67cc45a-d8e6-42ec-9e0e-eab62d207d31 · inbound
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c325d45-a496-43b1-8277-e3b12ce316d7 · inbound
Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.