Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:46:09.224346Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 3 inbound Pith citation observations for arXiv:2411.11278.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:46:09.224346Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:42:22.538919Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-11T13:56:18.448461Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 15ae7e47-9079-4e17-aff9-8493e5226573 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Look, listen and learn
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b654f469-11c6-46ed-bd46-4020b5b52226 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Objects that sound
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3d8a9846-c091-4486-ad9b-cde801a175a6 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Cross-modal label contrastive learning for unsupervised audio-visual event localization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 54e70ea2-a5b7-46d3-b39d-807029624e50 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization VGGSound: A large-scale audio-visual dataset
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d36fa57-ac3a-4b43-9698-acdf7ee7dc05 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Localizing visual sounds the hard way
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6232c27f-83f8-49f8-92db-e107fafa0fb8 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Cm-pie: Cross-modal per- ception for interactive-enhanced audio-visual video parsing
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8a14af03-16fa-4481-b53b-4a3d0046855c · outbound
Towards Open-Vocabulary Audio-Visual Event Localization VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e941a208-8151-4e4c-a3ef-d426edcd63ae · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual event localization via re- cursive fusion by joint co-attention
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 08ec1f37-eb68-4f14-a17f-6688462cd2ce · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Col- lecting cross-modal presence-absence evidence for weakly- supervised audio-visual event perception
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 03af3082-c28c-4052-b44d-472b72a03977 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Learning event-specific localization preferences for audio-visual event localization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d84a4fff-4d56-4430-b5cb-eddc723d8748 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Imagebind: One embedding space to bind them all
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 957f44bc-9f62-468c-b23e-716e6fea6864 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Audio-Visual Instance Segmentation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96ab497-95ca-4e75-9cff-26a6729e0e55 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Instance-level panoramic audio-visual saliency detection and ranking
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5246e0ec-f4c8-4027-a389-f584c7f24db5 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Open- vocabulary audio-visual semantic segmentation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5f1eb038-f382-4a6a-be20-a9785bfcb55b · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Unitr: A unified transformer-based framework for co-object and multi-modal saliency detection
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a346d44a-fc6d-4e6e-81f3-b145e9878fee · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Improving audio-visual segmentation with bidirectional generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 874d31ba-32d2-4cfa-8b29-edc990c1a07f · outbound
Towards Open-Vocabulary Audio-Visual Event Localization CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d306687a-1075-4646-866d-1cfdbeeb8a15 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91dfb26e-d9f5-4f0f-89a2-30f94572cd51 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Learning to answer questions in dynamic audio-visual scenarios
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 22acee5f-a790-4170-88bb-9f2f8279567e · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Progressive spatio- temporal perception for audio-visual question answering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0a9f6f5f-dd27-4a60-b15c-cca3ed389c4c · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Object-aware adaptive-positivity learning for audio- visual question answering
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e31d5b9-e864-45e8-aff4-c19f3f8681a7 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Patch-level Sounding Object Tracking for Audio-Visual Question Answering
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 589c317f-813f-4a64-a4f7-0b500e8e0b48 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Dual- modality seq2seq network for audio-visual event localiza- tion
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 948e84e3-e531-4587-b11d-f3b592fa2aaf · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Ave-clip: Audioclip-based multi-window temporal transformer for au- dio visual event localization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7b1c059b-1e82-405e-846a-55817dcaaf67 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization T-vsl: Text-guided visual sound source localization in mixtures
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9bb406f7-6335-40dd-b73d-58ac7d1a56f7 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Contrastive Conditional Latent Diffusion for Audio-visual Segmentation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e220090-092e-4cf7-a549-d71aec5aefb0 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Multimodal variational auto-encoder based audio-visual segmentation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 69ced0a5-5a5b-4617-bd3c-2074af874c0e · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Tavg- bench: Benchmarking text to audible-video generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c3a994b7-43ff-4711-a239-bb9de0336e58 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Avgzslnet: Audio-visual gen- eralized zero-shot learning by reconstructing label features from multi-modal embeddings
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d974016a-dbb9-4d62-a7da-f52c4ecae096 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Temporal and cross-modal at- tention for audio-visual zero-shot learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0d2011d4-2b03-463c-80cb-468d84ce853a · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual generalised zero-shot learning with cross-modal attention and language
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c273f183-ba0f-4b3d-b47b-84f3c7c1f979 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Localizing visual sounds the easy way
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 532280ed-dce2-424b-8d00-1fbf9c3fb8eb · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual Generalized Zero-shot Learning the Easy Way
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e1baf286-9313-4b41-8078-83b2e386eef0 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization PG-Video-LLaVA: Pixel Grounding Large Video-Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3d1c9b-ae43-4380-b520-da37344b8188 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Coordinated joint multimodal embeddings for gen- eralized audio-visual zero-shot classification and retrieval of videos
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 96feb531-bea0-4026-91b0-970b7faa4a88 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Multiple sound sources localization from coarse to fine
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b994d640-c782-4562-8410-0d4860a4332f · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Learn- ing transferable visual models from natural language super- vision
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 35c39a77-6b3a-465e-ba69-8414e423aa8f · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Fine-grained audible video description
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1c508258-4c6d-426e-bdbb-c3f94c062d87 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual event localization in unconstrained videos
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7d863dad-2e44-40a8-aab0-12e60cb51587 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5c33d0ff-547b-4703-b12e-4454b95a9959 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Attention is all you need
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ac6509aa-b06f-4c3c-a7fa-3e181c1cbf3c · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Dual attention matching for audio-visual event localization
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4f50f783-1427-4760-8706-ac9a7d3f2cea · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Span-based audio-visual localization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation df237228-da4d-46ce-acf6-e08773a08056 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9bbcc665-b21d-40be-b3de-975a8b0bbe90 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Cross-modal background suppres- sion for audio-visual event localization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 37030a44-de52-4b64-913e-cc1ee7b9df23 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Cross-modal relation-aware networks for audio-visual event localization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e8d08bf5-7c82-43b6-beed-f170b905e4a7 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Avqa: A dataset for audio- visual question answering on videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 595072d6-bec8-41d5-911a-50e0a92a590a · outbound
Towards Open-Vocabulary Audio-Visual Event Localization MM-Pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1f5ecc0c-7de5-4b10-8d1d-c3084f9907a6 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Ope- nA VE: Moving towards open set audio-visual event localiza- tion
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5bfb899d-dbe4-41ae-b09e-6ccaccffd004 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 852432ce-a2c2-4942-8fed-79a9d7ebf983 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Positive sample propagation along the audio- visual event line
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9da1f39c-47e7-491a-80d6-66f8272c44ae · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Contrastive pos- itive sample propagation along the audio-visual event line
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2c50363e-0199-41f4-b865-1a4ca2b65e4d · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Audio–visual segmentation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fdd1c146-5842-494b-8172-761b85846dd7 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Improving Audio-Visual Video Parsing with Pseudo Visual Labels
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 628c6665-f9bf-4f0b-8fe9-9feeb3d577ba · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Audio-Visual Segmentation with Semantics
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6175d65a-f33b-4108-a611-6608d1b1ccd6 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Label-anticipated event disentan- glement for audio-visual video parsing
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation afc1a1ef-2e6f-4c20-9714-4ecb58d0ea83 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Ad- vancing weakly-supervised audio-visual video parsing via segment-wise pseudo labeling
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2304d58c-8bc4-4702-93b2-f8a419010663 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f6023f4-5c01-4c81-926d-3ef258f96086 · outbound
Towards Open-Vocabulary Audio-Visual Event Localization Instruction: For the given 10-second video, divide it into 10 one-second segments. For each segment, if its audio and visual streams describe the same event, assign the label “x
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c92d3f27-fe3c-45ba-ae38-67084b2e76f7 · inbound
Patch-level Sounding Object Tracking for Audio-Visual Question Answering Towards Open-Vocabulary Audio-Visual Event Localization
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48fc6d8a-1ce9-49e8-858d-5fa97ed8ff73 · inbound
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing Towards Open-Vocabulary Audio-Visual Event Localization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 713eb6f8-0602-456b-9bf5-3a8246146c5a · inbound
Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration Towards Open-Vocabulary Audio-Visual Event Localization
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.