Pith. sign in

Paper Citation Record · LEDGER

Towards Open-Vocabulary Audio-Visual Event Localization

As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 3 inbound Pith citation observations for arXiv:2411.11278.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11278 v3

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:46:09.224346Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:42:22.538919Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T13:56:18.448461Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy46
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15ae7e47-9079-4e17-aff9-8493e5226573 · outbound

This paper cites Look, listen and learn.

Towards Open-Vocabulary Audio-Visual Event Localization Look, listen and learn

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:10.047568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:08.980979Z digest=sha256:5242f3dd5e2f54a1d41eef80018124a3e8fb1ac4ba9877ec28a106157bf61b77

Observation b654f469-11c6-46ed-bd46-4020b5b52226 · outbound

This paper cites Objects that sound.

Towards Open-Vocabulary Audio-Visual Event Localization Objects that sound

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:10.033663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:08.985768Z digest=sha256:74e32e4d1b80e5e47e419bd424ef80d6eece83949e67cb548d5774a116c52382

Observation 3d8a9846-c091-4486-ad9b-cde801a175a6 · outbound

This paper cites Cross-modal label contrastive learning for unsupervised audio-visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Cross-modal label contrastive learning for unsupervised audio-visual event localization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:10.019915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:08.990953Z digest=sha256:81ecad26da95408c35dd7896b3a2209f0ec054561e9bbb3194c253d30b068076

Observation 54e70ea2-a5b7-46d3-b39d-807029624e50 · outbound

This paper cites VGGSound: A large-scale audio-visual dataset.

Towards Open-Vocabulary Audio-Visual Event Localization VGGSound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:10.005165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:08.995675Z digest=sha256:5cbe530d439ff963139d4054140686b6fd721c816fd80cfaa28b49995739ffdb

Observation 2d36fa57-ac3a-4b43-9698-acdf7ee7dc05 · outbound

This paper cites Localizing visual sounds the hard way.

Towards Open-Vocabulary Audio-Visual Event Localization Localizing visual sounds the hard way

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.990268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:08.999948Z digest=sha256:cab83cbd44a7aaeafd4663fffcd2bef8a887b8eb31279c81c6e16f2b490faca0

Observation 6232c27f-83f8-49f8-92db-e107fafa0fb8 · outbound

This paper cites Cm-pie: Cross-modal per- ception for interactive-enhanced audio-visual video parsing.

Towards Open-Vocabulary Audio-Visual Event Localization Cm-pie: Cross-modal per- ception for interactive-enhanced audio-visual video parsing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.976520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.004285Z digest=sha256:24bda44330c31eacd24901195b633cb89ff695d10b54194a9a59f6b08adf68bb

Observation 8a14af03-16fa-4481-b53b-4a3d0046855c · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Towards Open-Vocabulary Audio-Visual Event Localization VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.008696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.008696Z digest=sha256:5b73f4e469fb29778726df9bfa9967acb9bc28c9b2c95dd7bb50a68705a55ff9

Observation e941a208-8151-4e4c-a3ef-d426edcd63ae · outbound

This paper cites Audio-visual event localization via re- cursive fusion by joint co-attention.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual event localization via re- cursive fusion by joint co-attention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.962817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.014532Z digest=sha256:dff86418c4a7003e691a427475f3a455a28302186a8e6315e6ee561a7c64dbb9

Observation 08ec1f37-eb68-4f14-a17f-6688462cd2ce · outbound

This paper cites Col- lecting cross-modal presence-absence evidence for weakly- supervised audio-visual event perception.

Towards Open-Vocabulary Audio-Visual Event Localization Col- lecting cross-modal presence-absence evidence for weakly- supervised audio-visual event perception

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.949431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.018653Z digest=sha256:3d8c64f5fdd641589b2ffd3aca145f18aa270383ad53d44ad7e1026b19febf6d

Observation 03af3082-c28c-4052-b44d-472b72a03977 · outbound

This paper cites Learning event-specific localization preferences for audio-visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Learning event-specific localization preferences for audio-visual event localization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.936627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.022625Z digest=sha256:b2bd4a3ce25c3792e915a6a4ffa653b61418a19f7d3a7fae177d1fece0ebb2a4

Observation d84a4fff-4d56-4430-b5cb-eddc723d8748 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Towards Open-Vocabulary Audio-Visual Event Localization Imagebind: One embedding space to bind them all

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.923665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.026858Z digest=sha256:dcfd0c4f6465ce9f3ecf5b72038980558481e4cb71b10cb6f5f4039335dfeaae

Observation 957f44bc-9f62-468c-b23e-716e6fea6864 · outbound

This paper cites Audio-Visual Instance Segmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-Visual Instance Segmentation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.031350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.031350Z digest=sha256:deeb346c6a15b13d27b3627ed7f80695585d93172c6167a2fd08b2dc6cac5c36

Observation a96ab497-95ca-4e75-9cff-26a6729e0e55 · outbound

This paper cites Instance-level panoramic audio-visual saliency detection and ranking.

Towards Open-Vocabulary Audio-Visual Event Localization Instance-level panoramic audio-visual saliency detection and ranking

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.910930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.035837Z digest=sha256:4054443aa53dbb7e4ee1c9c88e0c264a80218d9efd19d222069223f921d484b8

Observation 5246e0ec-f4c8-4027-a389-f584c7f24db5 · outbound

This paper cites Open- vocabulary audio-visual semantic segmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Open- vocabulary audio-visual semantic segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.897879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.039915Z digest=sha256:fa89efc6f8e8cae4fd7801da08fd93bf14094426d42d05e232723a4f258fe678

Observation 5f1eb038-f382-4a6a-be20-a9785bfcb55b · outbound

This paper cites Unitr: A unified transformer-based framework for co-object and multi-modal saliency detection.

Towards Open-Vocabulary Audio-Visual Event Localization Unitr: A unified transformer-based framework for co-object and multi-modal saliency detection

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.884566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.044185Z digest=sha256:dee4a5f9a0bdb3b7d707726aa4c411f0b1686fbde725ac6c38fb737ed7692e74

Observation a346d44a-fc6d-4e6e-81f3-b145e9878fee · outbound

This paper cites Improving audio-visual segmentation with bidirectional generation.

Towards Open-Vocabulary Audio-Visual Event Localization Improving audio-visual segmentation with bidirectional generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.870776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.048212Z digest=sha256:0fbc8e025fbf4aab312ae9125811733b804c93c36e36c1b993a9fd61c8942c01

Observation 874d31ba-32d2-4cfa-8b29-edc990c1a07f · outbound

This paper cites CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization.

Towards Open-Vocabulary Audio-Visual Event Localization CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:46:09.390007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.052170Z digest=sha256:697d62c0b0429ded4ee16d620ae1ca085a87737b2bac3398846010afe4e5cd4d

Observation d306687a-1075-4646-866d-1cfdbeeb8a15 · outbound

This paper cites Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control.

Towards Open-Vocabulary Audio-Visual Event Localization Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.056514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.056514Z digest=sha256:859d9de4ce30782e96707e4e7fe68b9a5cd785b739b5efce007fb6f023878a37

Observation 91dfb26e-d9f5-4f0f-89a2-30f94572cd51 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

Towards Open-Vocabulary Audio-Visual Event Localization Learning to answer questions in dynamic audio-visual scenarios

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.856923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.060887Z digest=sha256:a334cfb814c2c04599be1619aa6de5ca140e2cd2aa91e6d1c280d9db8779e9ce

Observation 22acee5f-a790-4170-88bb-9f2f8279567e · outbound

This paper cites Progressive spatio- temporal perception for audio-visual question answering.

Towards Open-Vocabulary Audio-Visual Event Localization Progressive spatio- temporal perception for audio-visual question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.843528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.064881Z digest=sha256:29fdceaab46c0c6c41d0c95f187051fd3193ccc778ebd9445c2f5e8983fb82ee

Observation 0a9f6f5f-dd27-4a60-b15c-cca3ed389c4c · outbound

This paper cites Object-aware adaptive-positivity learning for audio- visual question answering.

Towards Open-Vocabulary Audio-Visual Event Localization Object-aware adaptive-positivity learning for audio- visual question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.830691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.068921Z digest=sha256:ab697917d069f8fe8f21efa08f9191d9ae290753fdce72901e3c72a59f0a8752

Observation 8e31d5b9-e864-45e8-aff4-c19f3f8681a7 · outbound

This paper cites Patch-level Sounding Object Tracking for Audio-Visual Question Answering.

Towards Open-Vocabulary Audio-Visual Event Localization Patch-level Sounding Object Tracking for Audio-Visual Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.072989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.072989Z digest=sha256:184699efdf983243bb6b6e91ebf85822e5a1160080a690f827a2dc10c876b441

Observation 589c317f-813f-4a64-a4f7-0b500e8e0b48 · outbound

This paper cites Dual- modality seq2seq network for audio-visual event localiza- tion.

Towards Open-Vocabulary Audio-Visual Event Localization Dual- modality seq2seq network for audio-visual event localiza- tion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.817876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.077495Z digest=sha256:6024d7d3fcba01a4ab5b9366b27be658bffb038b42de2ab19c03f75b6c559108

Observation 948e84e3-e531-4587-b11d-f3b592fa2aaf · outbound

This paper cites Ave-clip: Audioclip-based multi-window temporal transformer for au- dio visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Ave-clip: Audioclip-based multi-window temporal transformer for au- dio visual event localization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.805042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.081450Z digest=sha256:8ac40d5b8e393ec15c6c68da453c1f79bf64a328ce775e42f4efc421d3fcf391

Observation 7b1c059b-1e82-405e-846a-55817dcaaf67 · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Towards Open-Vocabulary Audio-Visual Event Localization T-vsl: Text-guided visual sound source localization in mixtures

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.791568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.085692Z digest=sha256:d181f8388b2c4c6debc5959f01decbad1b44280039791e3a346bfee5e43b2ea7

Observation 9bb406f7-6335-40dd-b73d-58ac7d1a56f7 · outbound

This paper cites Contrastive Conditional Latent Diffusion for Audio-visual Segmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Contrastive Conditional Latent Diffusion for Audio-visual Segmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.089546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.089546Z digest=sha256:d2e1251ef396faf1592865637f771345a15450c606070a89a33bf19544bd0365

Observation 5e220090-092e-4cf7-a549-d71aec5aefb0 · outbound

This paper cites Multimodal variational auto-encoder based audio-visual segmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Multimodal variational auto-encoder based audio-visual segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.778579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.093768Z digest=sha256:3a316835796a556b487062483eb4b1bf84fcda8e2b04264ad54a0a9e4039110b

Observation 69ced0a5-5a5b-4617-bd3c-2074af874c0e · outbound

This paper cites Tavg- bench: Benchmarking text to audible-video generation.

Towards Open-Vocabulary Audio-Visual Event Localization Tavg- bench: Benchmarking text to audible-video generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.765637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.098002Z digest=sha256:3a8e0b12b44561922cc5ed1dc8aaac1f2a3fccf40938772ac63fbf641ad635a9

Observation c3a994b7-43ff-4711-a239-bb9de0336e58 · outbound

This paper cites Avgzslnet: Audio-visual gen- eralized zero-shot learning by reconstructing label features from multi-modal embeddings.

Towards Open-Vocabulary Audio-Visual Event Localization Avgzslnet: Audio-visual gen- eralized zero-shot learning by reconstructing label features from multi-modal embeddings

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.752061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.101963Z digest=sha256:440d2cc8df414881e2c7862a8c18d55a54ee1b3187ea7c52494cc936ca808209

Observation d974016a-dbb9-4d62-a7da-f52c4ecae096 · outbound

This paper cites Temporal and cross-modal at- tention for audio-visual zero-shot learning.

Towards Open-Vocabulary Audio-Visual Event Localization Temporal and cross-modal at- tention for audio-visual zero-shot learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.738739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.106230Z digest=sha256:217d4e267c04b09ac07b26c9d01508aef039da5d2a2223ac3a116182d7b06a54

Observation 0d2011d4-2b03-463c-80cb-468d84ce853a · outbound

This paper cites Audio-visual generalised zero-shot learning with cross-modal attention and language.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual generalised zero-shot learning with cross-modal attention and language

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.725021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.110153Z digest=sha256:968297c844a2ca6e83db85081f0a21651ccffc2537b2d0bb02bbcfa07ccf47db

Observation c273f183-ba0f-4b3d-b47b-84f3c7c1f979 · outbound

This paper cites Localizing visual sounds the easy way.

Towards Open-Vocabulary Audio-Visual Event Localization Localizing visual sounds the easy way

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.711381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.114184Z digest=sha256:af3d218e00eed8e408a5d24b41841cbd3d22c2f95bb3d67916e0aa42e109d93a

Observation 532280ed-dce2-424b-8d00-1fbf9c3fb8eb · outbound

This paper cites Audio-visual Generalized Zero-shot Learning the Easy Way.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual Generalized Zero-shot Learning the Easy Way

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:46:09.331158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.118070Z digest=sha256:f8df5c6965776738b630dd6459d75060a3218bfd9ec44f511b771a391c3f04c8

Observation e1baf286-9313-4b41-8078-83b2e386eef0 · outbound

This paper cites PG-Video-LLaVA: Pixel Grounding Large Video-Language Models.

Towards Open-Vocabulary Audio-Visual Event Localization PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.122320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.122320Z digest=sha256:8527f0f26a752190294040a2299bc000dc15dce1f63b4aaac2a933f864fc6d51

Observation 2b3d1c9b-ae43-4380-b520-da37344b8188 · outbound

This paper cites Coordinated joint multimodal embeddings for gen- eralized audio-visual zero-shot classification and retrieval of videos.

Towards Open-Vocabulary Audio-Visual Event Localization Coordinated joint multimodal embeddings for gen- eralized audio-visual zero-shot classification and retrieval of videos

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.697966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.126554Z digest=sha256:9bf8f0d1f8afd24e8134c6007d39fa7eb4135c6820183ba3e2b899094a112ecf

Observation 96feb531-bea0-4026-91b0-970b7faa4a88 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Towards Open-Vocabulary Audio-Visual Event Localization Multiple sound sources localization from coarse to fine

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.685165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.130581Z digest=sha256:f8817849faefe19f8da00b04c7ee9d4510cf32ab97733b3db65f212929c4b226

Observation b994d640-c782-4562-8410-0d4860a4332f · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Towards Open-Vocabulary Audio-Visual Event Localization Learn- ing transferable visual models from natural language super- vision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.672549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.134348Z digest=sha256:0a92ab62689446e679f216599098f0c710d6685686d90128cd03861d9f8b3c27

Observation 35c39a77-6b3a-465e-ba69-8414e423aa8f · outbound

This paper cites Fine-grained audible video description.

Towards Open-Vocabulary Audio-Visual Event Localization Fine-grained audible video description

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.660094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.138191Z digest=sha256:5b4210a2db7d1d0d09073be36d0ca5bed2a81e427ee8bafe15491c9c730cbaf8

Observation 1c508258-4c6d-426e-bdbb-c3f94c062d87 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual event localization in unconstrained videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.647523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.142176Z digest=sha256:19c338920f3e91885b5ca40093b389dbaff36c959a7e601b9cd7d96c84fe79f3

Observation 7d863dad-2e44-40a8-aab0-12e60cb51587 · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

Towards Open-Vocabulary Audio-Visual Event Localization Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.635048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.146224Z digest=sha256:b60707fd49b10c5338373edfc4ddd8a796372f71b4b225dc39e58ccc4546a940

Observation 5c33d0ff-547b-4703-b12e-4454b95a9959 · outbound

This paper cites Attention is all you need.

Towards Open-Vocabulary Audio-Visual Event Localization Attention is all you need

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.621497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.150182Z digest=sha256:44d92084aca9474dc43db60973e9c9913619f094847539a59ad7eca2dc690eb5

Observation ac6509aa-b06f-4c3c-a7fa-3e181c1cbf3c · outbound

This paper cites Dual attention matching for audio-visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Dual attention matching for audio-visual event localization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.608346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.154304Z digest=sha256:a9907835dc194bf90016e01952e5ea36298c4dfb6c3d0ec092973ad07e4d1f54

Observation 4f50f783-1427-4760-8706-ac9a7d3f2cea · outbound

This paper cites Span-based audio-visual localization.

Towards Open-Vocabulary Audio-Visual Event Localization Span-based audio-visual localization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.595451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.158467Z digest=sha256:d6bbb5a02dbdbdae1e39dcec8357f2f35cedbdd2d007891c03589fe9852de367

Observation df237228-da4d-46ce-acf6-e08773a08056 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.581830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.162578Z digest=sha256:837a04a17b2b9a5ef770a5a3241cee15ee00fcd929acb3daecfe9f997f06f725

Observation 9bbcc665-b21d-40be-b3de-975a8b0bbe90 · outbound

This paper cites Cross-modal background suppres- sion for audio-visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Cross-modal background suppres- sion for audio-visual event localization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.568696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.167077Z digest=sha256:5a04bed852ffd4b9419df4763f407026a466906e6ee49a04329fd8a2dd5526bb

Observation 37030a44-de52-4b64-913e-cc1ee7b9df23 · outbound

This paper cites Cross-modal relation-aware networks for audio-visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Cross-modal relation-aware networks for audio-visual event localization

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.554943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.171117Z digest=sha256:b4d71869c6f1fad12543fd5ebc0b5309d8a74684bc36ae7b7b7b1cdcca7c69d4

Observation e8d08bf5-7c82-43b6-beed-f170b905e4a7 · outbound

This paper cites Avqa: A dataset for audio- visual question answering on videos.

Towards Open-Vocabulary Audio-Visual Event Localization Avqa: A dataset for audio- visual question answering on videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.541968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.175171Z digest=sha256:1ea03c740762530a832c9218ac6bd21f5c9b8a68fe9458d157c0d7d738bbd214

Observation 595072d6-bec8-41d5-911a-50e0a92a590a · outbound

This paper cites MM-Pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing.

Towards Open-Vocabulary Audio-Visual Event Localization MM-Pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.529073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.179231Z digest=sha256:7a470796b2c6e891f2fc4b7002fc51469c72363af0759c581629709408ffce5a

Observation 1f5ecc0c-7de5-4b10-8d1d-c3084f9907a6 · outbound

This paper cites Ope- nA VE: Moving towards open set audio-visual event localiza- tion.

Towards Open-Vocabulary Audio-Visual Event Localization Ope- nA VE: Moving towards open set audio-visual event localiza- tion

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.515990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.183240Z digest=sha256:9dd8a482388abf17808ed332cd4681772422a462901738de7b949757433a0db9

Observation 5bfb899d-dbe4-41ae-b09e-6ccaccffd004 · outbound

This paper cites Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing.

Towards Open-Vocabulary Audio-Visual Event Localization Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.187257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.187257Z digest=sha256:5d8dea6f2e99441e1486d76d35c0ec7e383dc643102883efd3defa74d504bb4e

Observation 852432ce-a2c2-4942-8fed-79a9d7ebf983 · outbound

This paper cites Positive sample propagation along the audio- visual event line.

Towards Open-Vocabulary Audio-Visual Event Localization Positive sample propagation along the audio- visual event line

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.501606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.191431Z digest=sha256:d83896b6eff32a9e4c8d012960eb695d37cc338cdd8326512891fea649a979fd

Observation 9da1f39c-47e7-491a-80d6-66f8272c44ae · outbound

This paper cites Contrastive pos- itive sample propagation along the audio-visual event line.

Towards Open-Vocabulary Audio-Visual Event Localization Contrastive pos- itive sample propagation along the audio-visual event line

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.487226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.195775Z digest=sha256:6b8ed456387ecfa6e361786c86610ba6f5436399d713e7f0178e6730def6e74c

Observation 2c50363e-0199-41f4-b865-1a4ca2b65e4d · outbound

This paper cites Audio–visual segmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Audio–visual segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.473057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.199812Z digest=sha256:a47cb93628a6a767a126b663bd9b70d78ecb83dc521fc101744cfad6afb9cdcd

Observation fdd1c146-5842-494b-8172-761b85846dd7 · outbound

This paper cites Improving Audio-Visual Video Parsing with Pseudo Visual Labels.

Towards Open-Vocabulary Audio-Visual Event Localization Improving Audio-Visual Video Parsing with Pseudo Visual Labels

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.203800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.203800Z digest=sha256:d5e7c30fd011f23ba8a19c69b2754239c06c55b888e3bc3c635423d5f099d787

Observation 628c6665-f9bf-4f0b-8fe9-9feeb3d577ba · outbound

This paper cites Audio-Visual Segmentation with Semantics.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-Visual Segmentation with Semantics

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.208097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.208097Z digest=sha256:0f6b1cf85eae7929a3a3012c09d1359ed14a05825a6ae87781e9cbc796f087df

Observation 6175d65a-f33b-4108-a611-6608d1b1ccd6 · outbound

This paper cites Label-anticipated event disentan- glement for audio-visual video parsing.

Towards Open-Vocabulary Audio-Visual Event Localization Label-anticipated event disentan- glement for audio-visual video parsing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.459970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.212176Z digest=sha256:bd2edb4bcba99af4013b7e176fb1bcfd51391ee8feae1a695dce2dbf77aafc7c

Observation afc1a1ef-2e6f-4c20-9714-4ecb58d0ea83 · outbound

This paper cites Ad- vancing weakly-supervised audio-visual video parsing via segment-wise pseudo labeling.

Towards Open-Vocabulary Audio-Visual Event Localization Ad- vancing weakly-supervised audio-visual video parsing via segment-wise pseudo labeling

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.446303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.216131Z digest=sha256:dcd2e6ca9fac151c7572a26d6b2ff66919ec456ff69a85c6c617ebed48ca36be

Observation 2304d58c-8bc4-4702-93b2-f8a419010663 · outbound

This paper cites Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration.

Towards Open-Vocabulary Audio-Visual Event Localization Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.220172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.220172Z digest=sha256:fec52c890c4ceecf8cc9bfbeebdaec2f58aa3e6eacc381f03d704613ed25afe2

Observation 9f6023f4-5c01-4c81-926d-3ef258f96086 · outbound

This paper cites Instruction: For the given 10-second video, divide it into 10 one-second segments. For each segment, if its audio and visual streams describe the same event, assign the label “x.

Towards Open-Vocabulary Audio-Visual Event Localization Instruction: For the given 10-second video, divide it into 10 one-second segments. For each segment, if its audio and visual streams describe the same event, assign the label “x

Reference 2024

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T18:46:09.432304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T18:46:09.224346Z digest=sha256:ff25300906e07f00b501f95a4ed8eff6a117eb5dbb6971aafd62f85fe59752b0

Pith citing papers

Observation c92d3f27-fe3c-45ba-ae38-67084b2e76f7 · inbound

Patch-level Sounding Object Tracking for Audio-Visual Question Answering cites this paper.

Patch-level Sounding Object Tracking for Audio-Visual Question Answering Towards Open-Vocabulary Audio-Visual Event Localization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:22.538919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:42:22.538919Z digest=sha256:82bb91388f56e5ccdb8a40ee68d4dba2370bb40bd817cbb703bfbec406d9ae60

Observation 48fc6d8a-1ce9-49e8-858d-5fa97ed8ff73 · inbound

Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing cites this paper.

Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing Towards Open-Vocabulary Audio-Visual Event Localization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:13:01.650537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:13:01.650537Z digest=sha256:fb86ccd3c34336a00f8d5797a3fe72b16eb9ab9571da89eca9c9f3b0c8e03e4d

Observation 713eb6f8-0602-456b-9bf5-3a8246146c5a · inbound

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration cites this paper.

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration Towards Open-Vocabulary Audio-Visual Event Localization

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:56:18.454468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T13:56:18.376640Z digest=sha256:4c04de4b18b8333be2c307a1d7379c426a68277c437f09116dbf090469c3c580