Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T10:33:04.329045Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2607.24786.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T10:33:04.329045Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 835c4629-c3d2-4f32-84ab-bee119b7687b · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Localizing visual sounds the easy way
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1d469d-193c-4185-8c50-4c281b6261dc · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models A closer look at weakly-supervised audio-visual source localization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5058dc0b-c6cb-4b7e-b9df-e479783d343e · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Learning transferable visual models from natural language supervision
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962945e0-c51d-4de9-a9f3-08b09fba1d04 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Sigmoid loss for language image pre-training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 674ee0e8-fa75-49ea-8d8f-564ae12addc2 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Imagebind: One embedding space to bind them all
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b7909e-9f69-40c9-83e3-e2182c7e6571 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Pushing the frontier of audiovisual perception with large-scale multimodal correspondence learning.arXiv preprint arXiv:2512.19687, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b43dc88-62e8-43f7-a76c-02b8f432bbe1 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Learning audio-visual source localization via false negative aware con- trastive learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1914a910-90cf-4540-9dec-4e7ed1c94bda · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models What’s making that sound right now? video-centric audio-visual localization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c334cf1-f7ff-4899-850b-595d1d4f0471 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Flair: Vlm with fine-grained language-informed image representations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d9c5fc8-d090-4aa0-a83f-433f07159929 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Learning to localize sound source in visual scenes
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc124740-8156-4d87-b69c-7ea9404c3c51 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Exploiting transformation invariance and equivariance for self-supervised sound localisation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dd7d787-5a65-4373-a33a-271be860b14e · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Marginnce: Robust sound localization with a negative margin
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8606333-4089-4702-b633-3409d51f09fe · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Sound source localization is all about cross-modal alignment
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56bc5e8a-24d9-4227-bc06-13b43f5116dd · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Audio–visual segmentation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a061c02-1bc4-455e-9767-df5fd6890f55 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70302193-7606-46af-b827-b1f51e739d43 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Audio visual segmentation through text embeddings
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed72577c-daa9-44b3-9e75-a2373db1cfd4 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Open-vocabulary audio-visual semantic segmentation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e8140a3-35d2-4ed4-ae5b-3128f75467c3 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Taco: Training-free sound prompted segmentation via semantically constrained audio-visual co-factorization.Transactions on Machine Learn- ing Research, 2026
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44173815-9d10-4c8f-ae3a-b9b7237766da · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Perception encoder: The best visual embeddings are not at the output of the network
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc16366d-6174-4812-a251-ea90bd3ecda2 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Vision transformers need registers
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96ae56ee-03eb-4b9d-8084-2d5b29dde6aa · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Gradient-based learning applied to document recognition.Proceedings of the IEEE, 86(11):2278–2324, 2002
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc6ce4b-0ca1-4088-a735-14e7aaab3e9e · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Parameter-efficient transfer learning for nlp
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0902e3d4-21db-44f2-9fc0-7c4032e3bb85 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models chirp" from the
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2faff45e-d3eb-48be-981a-02427295ff7f · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Vggsound: A large-scale audio-visual dataset
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ff947f8-7e51-45f2-818c-0293bc9479c2 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Transfer learning from audio-visual grounding to speech recognition
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e79ce49c-91aa-4a0b-9ee2-afc02c60a9d6 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Contrastive audio-visual masked autoencoder
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dcc3903-aa82-4f68-8eba-c72ec0734ce1 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Cav-mae sync: Improving contrastive audio-visual mask autoencoders via fine-grained alignment
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740fd3d5-3be0-41bc-a799-88fc3db0ef27 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Convolutions die hard: Open- vocabulary segmentation with single frozen convolutional clip.Advances in Neural Information Processing Systems, 36, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5844169b-cf35-4fac-99e3-629622c96a78 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Reproducible scaling laws for contrastive language-image learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d0a6d57-062f-41af-a1f4-69d217dd6f40 · outbound
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models LAIP two poolers
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.