Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:15.094780Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2505.14562.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:15.094780Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:53.551896Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T22:41:14.663492Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 99e7f0db-e75a-41eb-b099-380885e3ee43 · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Learning transferable visual models from natural language supervision,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b93de3d-88e2-43dd-a87b-ba987bbb8709 · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities CLAP learning audio concepts from natural lan- guage supervision,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fabf8d23-e5c5-46de-8acd-fa912c52bd6b · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Multimodal learn- ing with deep boltzmann machines,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e96eb8f0-4007-4a5a-aa42-9441bc759fe2 · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Learning visual features from large weakly supervised data,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db6b673f-9dbc-45ab-9532-31d92af84d9d · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Scaling up visual and vision-language representation learning with noisy text supervision,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32e3baa8-88bc-4130-b9d7-aaa8079621cf · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Florence: A New Foundation Model for Computer Vision
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb57bd20-a237-4aa2-84a2-75e9ed2c7cdd · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Wav2CLIP: Learning robust audio representa- tions from CLIP,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 396192f0-8fa6-4f63-a623-313544346d70 · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Microsoft COCO: Common objects in context,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64a99d06-0983-431a-90d9-1fd6fb60e89d · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Visual genome: Connecting lan- guage and vision using crowdsourced dense image annotations,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07f760c5-d9c6-4ada-86c1-3a2436579f30 · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities YFCC100M: The new data in multimedia research,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f471ed8-3cb0-4790-a9ae-65f3e2d60f93 · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities FSD50K: An open dataset of human-labeled sound events,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68537183-119f-44b4-955e-a282dceffadf · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Clotho: An audio captioning dataset,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67bd587a-e2c3-4fa3-9363-a1501bb6a625 · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities AudioCaps: Generating captions for audios in the wild,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0aeba42d-176e-414e-9584-807e916a9439 · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities What is the ground truth? reliability of multi-annotator data for audio tag- ging,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d814e34b-9f97-4733-98df-56ee70e7696b · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Audio- CLIP: Extending CLIP to image, text and audio,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ece6d4f-4a63-40ea-a543-ceba1cabcdaf · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Audio Set: An ontology and human-labeled dataset for audio events,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3dbc9405-fb1f-4ae3-a221-602119900054 · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Sudarsanam, I
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a0b3938-31fc-49ce-ae4f-2a5ff4c5c74e · outbound
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Representation Learning with Contrastive Predictive Coding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 367aff17-6b04-4fb6-9b8f-f32f317f4eed · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3301051a-947f-4ff9-90c4-2bd38cb9021d · inbound
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.