Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:58:02.357429Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2506.06537.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:58:02.357429Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:58:02.208194Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T05:58:02.433864Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9c324de4-f448-45d4-9973-c134069f57ec · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models This capability is essential for applica- tions in video understanding, surveillance, human-computer in- teraction, and autonomous systems
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba2b96cc-0daf-4343-abb5-a88b9eeb2afc · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Many approaches leverage cross-modal attention [1, 2, 3] combined with contrastive learning [4, 5] to align audio and visual features
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b467b46-f159-4823-b8ab-669a8d04bc28 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05360814-6668-4846-b41d-7cdc1d1a6efe · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Models Below, we elaborate on the final models built for each approach
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7266a948-ce5a-4a2f-97ae-f12a5361c794 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models By systematically integrating audio, vision, and text representations, we bridge modality gaps and achieve fine-grained segmentation without requiring A VS-specific annotations
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45b681cf-81c4-47a9-8c43-0945a0339f6f · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51b6b9a9-a56e-40c4-a262-d9929e24d7fe · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning to localize sound sources in visual scenes: Analysis and applications,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbc8d8e5-6e64-4c80-844f-d0602e249a49 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea2187a2-3ab3-4113-9ef0-27d3b535398a · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Exploiting transformation in- variance and equivariance for self-supervised sound localisation,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61908570-0221-42fd-89a2-e50426373aeb · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Localizing visual sounds the easy way,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 187b50d7-1b94-47c1-a2b1-a173a5565f20 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Localizing visual sounds the hard way,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8330852f-ff89-4c02-b9ba-acd7124ae7d0 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning audio-visual source local- ization via false negative aware contrastive learning,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe06a59f-5fdb-47ab-a449-d45ab740654f · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Marginnce: Robust sound localization with a negative margin,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2d24cae-a9bd-49af-86f2-29dc969e99db · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Audio–visual segmen- tation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a97a0451-fba9-402e-b8a0-3e358bed884f · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Improving audio-visual segmentation with bidirectional generation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea52bc2d-7569-4054-80fd-e8d55ad2b4fb · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Selm: Selective mechanism based audio-visual segmentation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 110a07f7-5a14-41a5-a5a0-4d25e7cd41e2 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bavs: Bootstrapping audio-visual segmentation by integrating foundation knowledge,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0f3a086-d99a-4b58-add1-057deb845d7d · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning transferable visual models from natural language supervision,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b97ce92b-dd08-458a-a942-8ba4a2e442af · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Natural Language Supervision for General-Purpose Audio Representations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b957c4a4-d464-45c3-8a17-463ab759847e · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language mul- timodal research,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8a97d62-a16f-4e8d-9b2a-5cfb5cca0335 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models AudioCaps: Generating Captions for Audios in The Wild,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d9c8752-8a7a-40f4-a793-eda42895a494 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Microsoft coco: Common objects in context,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4704ecb-6a7c-4a03-bfd6-782595e87095 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cead3a2-851f-41a0-9133-c6a0c9ce0374 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdcf4a11-9322-438d-b2be-c3d27ffbf35f · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning to visually localize sound sources from mixtures without prior source knowl- edge,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef3a48c4-44d5-4ca7-923b-66227211d06f · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Adaptive selection based referring im- age segmentation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1d42340-6c74-4b35-a8e2-918242b51871 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Beats: Audio pre-training with acoustic tokenizers,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3955a45d-4ea2-4ede-af26-ed42f347f5f7 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Audio set: An ontology and human-labeled dataset for audio events,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e948701-6ee2-4222-99c9-14d9ced8b76b · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Clotho: An audio cap- tioning dataset,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 809d8bfc-ead5-43fe-8513-94e268eb35f4 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models spaCy 2: Natural language under- standing with Bloom embeddings, convolutional neural networks and incremental parsing,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99888b15-3398-4842-943c-7ca0b879f59c · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models High-resolution image synthesis with latent diffusion models,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406847a3-bac8-415e-b23d-bce3dce0f5e7 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Vggsound: A large-scale audio-visual dataset,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96a2c20c-d4bb-4263-b611-9c364e4e3414 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Segment anything,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3dec38ed-1032-4ad8-8887-aca36f518f87 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Unraveling instance associations: A closer look for audio-visual segmentation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 005345c1-4e4b-48e5-b82e-acfb4c9df8b5 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models A closer look at weakly-supervised audio-visual source localization,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f67a6e9-214e-4beb-8c99-694d45256cc4 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridg- ing vision and language encoders: Parameter-efficient tuning for referring image segmentation,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f16500da-de98-4eac-9740-46fc5134d510 · outbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Cris: Clip-driven referring image segmentation,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b467b46-f159-4823-b8ab-669a8d04bc28 · inbound
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.