Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:26:32.991891Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2505.01237.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:26:32.991891Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ab5ba190-b74b-4850-853f-0225be2800ec · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Self-supervised learning of audio-visual objects from video
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b9163a-48bc-4536-91e9-ba49dd15249f · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Look, listen and learn
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8721d22e-4783-44c8-99e1-4be0d12cd5fd · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Objects that sound
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation abb8d35f-e874-40ce-b211-b62f120a2b8e · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Sound- net: Learning sound representations from unlabeled video
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 59c34e70-af9a-468a-94aa-9aac795b5112 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Emerg- ing properties in self-supervised vision transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9420bd3c-3674-43c4-8681-a06a2d76b89d · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vggsound: A large-scale audio-visual dataset
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f19c7b18-3466-41f9-b614-78e182dda387 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Localiz- ing visual sounds the hard way
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aba398b6-c9ce-4d75-a004-151350fab2e0 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Distilling audio-visual knowledge by com- positional contrastive learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 389a656d-d290-4307-9cbf-a00ce46c72ff · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cc1b77d1-1f77-4738-9521-0773c9fdb621 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vision transformers need registers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 29d9ea60-eb9f-47fc-998d-7f82934b04ed · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio set: An ontology and human- labeled dataset for audio events
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3cd7ed33-9754-4237-a3c0-202be15b27a7 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audiovisual masked autoencoders
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fa9b2a8a-99e5-4572-96a7-4ba46a39fa6e · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Imagebind: One embedding space to bind them all
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6c0f4c33-db18-466c-aa5d-712e26f1a4f8 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7003157a-d225-473d-adb4-ac16868a791a · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Cross- mae: Cross-modality masked autoencoders for region-aware audio-visual pre-training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7485129c-84de-4caa-ae10-4b5e257db1a9 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Separating the” chirp” from the” chat”: Self-supervised visual grounding of sound and language
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1e006760-16ec-420d-bc6b-b7abd2a55357 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Jointly dis- covering visual objects and spoken words from raw sensory input
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 78ea52a8-0f76-417b-8399-76e4059b0a34 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Multi- modal attention for fusion of audio and spatiotemporal fea- tures for video description
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 30f027fb-4e56-4c69-b832-00867cb4bbb8 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a20842a8-ea8c-4cea-b304-763c02564db4 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Mavil: Masked audio-video learners
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 52136e88-5727-4d12-b8b6-e0e58f7aa1f8 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment EquiA V: Leveraging Equivari- ance for Audio-Visual Contrastive Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a6e05e86-f8c8-4a25-98b3-0ccfdbc1768a · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Coopera- tive learning of audio and video models from self-supervised synchronization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4a87a967-a4e1-4ec5-8d90-9d8ce76aca57 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Cross-attentional audio-visual fusion for weakly- supervised action localization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97cfe34f-ac02-4013-af24-f9410fbd6eeb · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Siamese vision transform- ers are scalable audio-visual learners
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 19cb27be-d1c7-44cd-bc3d-7b3d14033aa8 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vision transformers are parameter-efficient audio- visual learners
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4aa27521-bd89-44d0-8dde-73aff2b7ec49 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Active contrastive learning of audio-visual video representa- tions
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c48d9fe1-9393-437e-a92b-2beeafada6ea · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Active contrastive learning of audio-visual video representa- tions
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 261f2fbd-a995-4dd1-8c0e-d4548bedf6c8 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Robust audio-visual instance discrimination
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c4844cf3-ab58-4ec3-88bc-72a32a372b69 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio- visual instance discrimination with cross-modal agreement
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ded765ae-f5e6-4a40-aecb-402998a79a28 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio-visual scene analysis with self-supervised multisensory features
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7e8392aa-f18e-4acf-95a7-682aaddfa66f · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Ambient sound provides supervision for visual learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1c192f66-1479-475c-a088-ee4d811c364f · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment On compositions of transformations in contrastive self-supervised learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 741a9754-e2dc-4f0f-8cf8-db0f6c36e4ed · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Broaden your views for self-supervised video learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6a6e6f64-3bc0-4896-80db-b4cb6659f2fc · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Avlnet: Learning audio-visual language representa- tions from instructional videos
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation edc60b02-4d74-4097-b79c-e28efaf0a2bf · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Self-supervised audio- visual representation learning with relaxed cross-modal syn- chronicity
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6ff6cfa9-1b25-4b6d-bf08-532115f77f59 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Event-specific audio-visual fusion layers: A simple and new perspective on video understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e4892cd6-6be3-40fc-ad05-52d2911f5a50 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment From vision to au- dio and beyond: A unified model for audio-visual representa- tion and generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 047fe13b-cd2f-4c29-82c8-3751fc36d5ed · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Learning audio-visual source localization via false negative aware contrastive learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fc5418b9-2d54-4061-b30e-9488c3298d4a · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Multimodal Self-Supervised Learning of General Audio Representations
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89288667-2be6-4606-8168-8889985cf1d2 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Temporal cue guided video highlight detection with low-rank audio-visual fusion
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8620266b-ec0c-4d5e-9ab1-4257529d94c2 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Con- trastive learning of global and local video representations
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a0311b32-5c5d-4fc8-abad-76a0f5d08865 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment The sound of pixels
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aa27f171-004a-4ceb-90f0-7c343d55ca27 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Scene parsing through ade20k dataset
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ae1a59-bbaf-4163-b650-9a12be9fcba5 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Languagebind: Extending video-language pretraining to n- modality by language-based semantic alignment
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 12498e67-28ca-4db0-b812-1f954bcbcfff · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9bb38671-8523-400e-a3c8-de9b7e29c955 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation da8d6112-834b-4bc6-9fe2-b732aade6310 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment di- agonal mean
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dbef4d5f-5e60-45ce-8820-907cc749079f · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Table 12 shows the performance comparison between register to- kens, patch tokens, and the global token
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e08250c1-bba3-4508-8115-66e6321d4225 · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment writing on blackboard with chalk
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 521aa4e2-2093-47a7-aa0d-aa4e7c0f997c · outbound
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment For this experiment, we manually annotate the occurrence of the classes throughout the video
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.