Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:02:47.086421Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2507.11967.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:02:47.086421Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:02:44.662433Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T17:02:47.349677Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3b82f3c9-ed53-401e-b29b-d85ffdd59aa9 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e914d1db-ca25-4909-abb0-28cfae130dff · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos We also propose to automatically gener- ate audio-visual-text triplets from unlabeled videos, which are subsequently used for training LG-CA V-MAE
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48449906-1dcf-4cea-a384-8751903f7c49 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90ee5293-c267-4109-910d-530c82239868 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90dd38b2-8ab7-4064-bc64-7caa43145675 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Contrastive Audio-Visual Masked Autoencoder
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d8bf79-fc66-4056-a9b1-5e2c4ae02a20 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Mavil: Masked audio-video learners,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0509ec62-d6a4-442b-8677-a82e7db0a7fd · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audiovisual masked autoencoders,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa779b7d-4f7c-489a-a940-a87fd09879ca · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audio set: An ontology and human-labeled dataset for audio events,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d1168e6-07ab-4819-86bd-c915906aaf85 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos The Kinetics Human Action Video Dataset
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 437a2c0e-0f7f-474a-b31a-b038349e3fd5 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Vggsound: A large-scale audio-visual dataset,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a805f781-fdf6-4ca9-833e-05d4e3e771e1 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80c6ccf8-756a-4288-a4d3-17eed4ed3939 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f32d6421-83f9-4433-a70f-2b1f887c27f3 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Clap learning audio concepts from natural language supervision,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 308d9055-154c-49cf-98d4-6832043d1a37 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Dense-captioning events in videos,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67eb9617-c53e-48fe-8b5f-e4b75ad074fa · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51708959-5173-4272-b622-144a7580e5dd · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Vatex: A large-scale, high-quality multilingual dataset for video- and-language research,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ad27a96-7e47-4f48-a2ac-5410af91390a · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Avset-10m: An open large-scale audio-visual dataset with high correspondence,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7a7a21c-6951-4462-ad65-f7a9b733d473 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Sound event envelope estimation in polyphonic mixtures,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e7b2256-3a88-40ee-a052-4f261cc9ebde · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5be9aea-0026-43be-80c2-0534256b5b45 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos LP-MusicCaps: LLM-Based Pseudo Music Captioning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 400238fd-6039-465d-a1a4-3597ab50cd9d · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5b6cb5-3de1-4f72-bfa6-25d959fd7fb7 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos A Short Note on the Kinetics-700 Human Action Dataset
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb31ecb-e7fc-4aca-95e5-5150ce7a52c4 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos A simple framework for contrastive learning of visual representations,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0fef977-f530-4dcc-aef9-3cbe601b25c4 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Representation Learning with Contrastive Predictive Coding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c7a069-6bc6-4f75-95bd-94993b383002 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66b550c5-b6d1-4fed-8b0f-f17f9697ef9c · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Improved baselines with visual instruction tuning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56600a93-ee73-4a74-9978-01870ce512b2 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audiovisual SlowFast Networks for Video Recognition
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 099ed48e-97e1-4aa7-ad83-aa66bface629 · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Masked autoencoders that lis- ten,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24cb7e99-41c5-435f-9db0-83539c10c1bc · outbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Adam: A Method for Stochastic Optimization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b82f3c9-ed53-401e-b29b-d85ffdd59aa9 · inbound
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.