Pith. sign in

Paper Citation Record · LEDGER

Towards Open-Vocabulary Audio-Visual Event Localization

As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 3 inbound Pith citation observations for arXiv:2411.11278.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11278 v3

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:46:09.224346Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:42:22.538919Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T13:56:18.448461Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy46
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15ae7e47-9079-4e17-aff9-8493e5226573 · outbound

This paper cites Look, listen and learn.

Towards Open-Vocabulary Audio-Visual Event Localization Look, listen and learn

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:10.047568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:08.980979Z digest=sha256:9a2fe7a685565bf2fa1756893aee98720a963b64a8c1ad9fae326c4c5cbc784e

Observation b654f469-11c6-46ed-bd46-4020b5b52226 · outbound

This paper cites Objects that sound.

Towards Open-Vocabulary Audio-Visual Event Localization Objects that sound

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:10.033663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:08.985768Z digest=sha256:5b0eebda2b2d1e76740ae578e4350e91bb27678827d31a387220bc0e7241f34d

Observation 3d8a9846-c091-4486-ad9b-cde801a175a6 · outbound

This paper cites Cross-modal label contrastive learning for unsupervised audio-visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Cross-modal label contrastive learning for unsupervised audio-visual event localization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:10.019915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:08.990953Z digest=sha256:429c2f04e2b906918617f6b312f191f88905d24e5315130eed966cd0a11d0f0f

Observation 54e70ea2-a5b7-46d3-b39d-807029624e50 · outbound

This paper cites VGGSound: A large-scale audio-visual dataset.

Towards Open-Vocabulary Audio-Visual Event Localization VGGSound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:10.005165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:08.995675Z digest=sha256:244b8555d0a7eff197b7e44ed960df1d9fb213fd065ac2d4b38dec765d45baba

Observation 2d36fa57-ac3a-4b43-9698-acdf7ee7dc05 · outbound

This paper cites Localizing visual sounds the hard way.

Towards Open-Vocabulary Audio-Visual Event Localization Localizing visual sounds the hard way

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.990268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:08.999948Z digest=sha256:141361894026afc60de65fe1da2cd3826f98809ba1edc193c4348c71c782bd6d

Observation 6232c27f-83f8-49f8-92db-e107fafa0fb8 · outbound

This paper cites Cm-pie: Cross-modal per- ception for interactive-enhanced audio-visual video parsing.

Towards Open-Vocabulary Audio-Visual Event Localization Cm-pie: Cross-modal per- ception for interactive-enhanced audio-visual video parsing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.976520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.004285Z digest=sha256:084f76dd8e80b4fb5e8d75746873eaaf156e3ceddca05aa4d8668aae3ebfac83

Observation 8a14af03-16fa-4481-b53b-4a3d0046855c · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Towards Open-Vocabulary Audio-Visual Event Localization VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.008696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.008696Z digest=sha256:5b73f4e469fb29778726df9bfa9967acb9bc28c9b2c95dd7bb50a68705a55ff9

Observation e941a208-8151-4e4c-a3ef-d426edcd63ae · outbound

This paper cites Audio-visual event localization via re- cursive fusion by joint co-attention.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual event localization via re- cursive fusion by joint co-attention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.962817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.014532Z digest=sha256:c3fae1dd23df6e2a58538aeaea279fb59b84224d6b48d8a2da137c34297b7606

Observation 08ec1f37-eb68-4f14-a17f-6688462cd2ce · outbound

This paper cites Col- lecting cross-modal presence-absence evidence for weakly- supervised audio-visual event perception.

Towards Open-Vocabulary Audio-Visual Event Localization Col- lecting cross-modal presence-absence evidence for weakly- supervised audio-visual event perception

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.949431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.018653Z digest=sha256:900138b54c25dade6250a41113afe0dd7a1b569380ecfa844e31c7ad4fcef230

Observation 03af3082-c28c-4052-b44d-472b72a03977 · outbound

This paper cites Learning event-specific localization preferences for audio-visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Learning event-specific localization preferences for audio-visual event localization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.936627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.022625Z digest=sha256:d0ade79217a10f9eb8e5286b8bcef22c4a9c195c7d0982babcae03bd0d886ea1

Observation d84a4fff-4d56-4430-b5cb-eddc723d8748 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Towards Open-Vocabulary Audio-Visual Event Localization Imagebind: One embedding space to bind them all

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.923665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.026858Z digest=sha256:a92f25c55ed92cefab2c8d0df26307321f198cfa72b73c36b4800bab12e2df90

Observation 957f44bc-9f62-468c-b23e-716e6fea6864 · outbound

This paper cites Audio-Visual Instance Segmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-Visual Instance Segmentation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.031350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.031350Z digest=sha256:deeb346c6a15b13d27b3627ed7f80695585d93172c6167a2fd08b2dc6cac5c36

Observation a96ab497-95ca-4e75-9cff-26a6729e0e55 · outbound

This paper cites Instance-level panoramic audio-visual saliency detection and ranking.

Towards Open-Vocabulary Audio-Visual Event Localization Instance-level panoramic audio-visual saliency detection and ranking

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.910930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.035837Z digest=sha256:a55585d2db44bdb8ba766d35689413bb6055c48ee12dc4a54cc172b90d8e0710

Observation 5246e0ec-f4c8-4027-a389-f584c7f24db5 · outbound

This paper cites Open- vocabulary audio-visual semantic segmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Open- vocabulary audio-visual semantic segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.897879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.039915Z digest=sha256:615bdf546282e0ca5e48e634bd384b6c7bd659cc06dc16149e25b9fac9a977cc

Observation 5f1eb038-f382-4a6a-be20-a9785bfcb55b · outbound

This paper cites Unitr: A unified transformer-based framework for co-object and multi-modal saliency detection.

Towards Open-Vocabulary Audio-Visual Event Localization Unitr: A unified transformer-based framework for co-object and multi-modal saliency detection

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.884566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.044185Z digest=sha256:4c0989300fbe072a9fa952820fbb645809f5bfa85859d3c6e6dd9e437df613b4

Observation a346d44a-fc6d-4e6e-81f3-b145e9878fee · outbound

This paper cites Improving audio-visual segmentation with bidirectional generation.

Towards Open-Vocabulary Audio-Visual Event Localization Improving audio-visual segmentation with bidirectional generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.870776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.048212Z digest=sha256:f40a132bd6e97d74d5ffc8c62996222074ec031ed3c2d0e3ab5351b6e9d7ae5c

Observation 874d31ba-32d2-4cfa-8b29-edc990c1a07f · outbound

This paper cites CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization.

Towards Open-Vocabulary Audio-Visual Event Localization CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:46:09.390007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.052170Z digest=sha256:7a7b9fbaa576de55e13d9d2934e734febb4e998e75c30178c60941b2cda349bc

Observation d306687a-1075-4646-866d-1cfdbeeb8a15 · outbound

This paper cites Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control.

Towards Open-Vocabulary Audio-Visual Event Localization Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.056514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.056514Z digest=sha256:859d9de4ce30782e96707e4e7fe68b9a5cd785b739b5efce007fb6f023878a37

Observation 91dfb26e-d9f5-4f0f-89a2-30f94572cd51 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

Towards Open-Vocabulary Audio-Visual Event Localization Learning to answer questions in dynamic audio-visual scenarios

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.856923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.060887Z digest=sha256:6b8ca5241b8d3f68531fea750352f29aacbcad2a7a9781966d27622db7ab97f7

Observation 22acee5f-a790-4170-88bb-9f2f8279567e · outbound

This paper cites Progressive spatio- temporal perception for audio-visual question answering.

Towards Open-Vocabulary Audio-Visual Event Localization Progressive spatio- temporal perception for audio-visual question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.843528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.064881Z digest=sha256:7d9ec70f31b9329e16e799bf4182dfa91a84917bb69abc6b52101d06108d4dc0

Observation 0a9f6f5f-dd27-4a60-b15c-cca3ed389c4c · outbound

This paper cites Object-aware adaptive-positivity learning for audio- visual question answering.

Towards Open-Vocabulary Audio-Visual Event Localization Object-aware adaptive-positivity learning for audio- visual question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.830691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.068921Z digest=sha256:c83e3a92c3a02e28bd549a60424042698be1004af64657577c1f27060555d1af

Observation 8e31d5b9-e864-45e8-aff4-c19f3f8681a7 · outbound

This paper cites Patch-level Sounding Object Tracking for Audio-Visual Question Answering.

Towards Open-Vocabulary Audio-Visual Event Localization Patch-level Sounding Object Tracking for Audio-Visual Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.072989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.072989Z digest=sha256:184699efdf983243bb6b6e91ebf85822e5a1160080a690f827a2dc10c876b441

Observation 589c317f-813f-4a64-a4f7-0b500e8e0b48 · outbound

This paper cites Dual- modality seq2seq network for audio-visual event localiza- tion.

Towards Open-Vocabulary Audio-Visual Event Localization Dual- modality seq2seq network for audio-visual event localiza- tion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.817876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.077495Z digest=sha256:8483f65e3b93c8548eb107624bcc5d6f6346f3178f882486e0d48a172f6ccddc

Observation 948e84e3-e531-4587-b11d-f3b592fa2aaf · outbound

This paper cites Ave-clip: Audioclip-based multi-window temporal transformer for au- dio visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Ave-clip: Audioclip-based multi-window temporal transformer for au- dio visual event localization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.805042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.081450Z digest=sha256:934c5171b9b035be4684dfd0e2101dfed573280903a64d86451df6c9f05fbb4a

Observation 7b1c059b-1e82-405e-846a-55817dcaaf67 · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Towards Open-Vocabulary Audio-Visual Event Localization T-vsl: Text-guided visual sound source localization in mixtures

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.791568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.085692Z digest=sha256:6e17b4444fe8ef2a6501329cfe97b648345038edc3b358da38bf3140390efceb

Observation 9bb406f7-6335-40dd-b73d-58ac7d1a56f7 · outbound

This paper cites Contrastive Conditional Latent Diffusion for Audio-visual Segmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Contrastive Conditional Latent Diffusion for Audio-visual Segmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.089546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.089546Z digest=sha256:d2e1251ef396faf1592865637f771345a15450c606070a89a33bf19544bd0365

Observation 5e220090-092e-4cf7-a549-d71aec5aefb0 · outbound

This paper cites Multimodal variational auto-encoder based audio-visual segmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Multimodal variational auto-encoder based audio-visual segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.778579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.093768Z digest=sha256:6247980bf1dd6196b69b48e4df418b5fa1fcaa70bdb4048d88577bd72f0f7082

Observation 69ced0a5-5a5b-4617-bd3c-2074af874c0e · outbound

This paper cites Tavg- bench: Benchmarking text to audible-video generation.

Towards Open-Vocabulary Audio-Visual Event Localization Tavg- bench: Benchmarking text to audible-video generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.765637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.098002Z digest=sha256:9851cec4bb50a47d672fbb3dfd63d06d8c3e9d82f72a72065bfad11a19ec2178

Observation c3a994b7-43ff-4711-a239-bb9de0336e58 · outbound

This paper cites Avgzslnet: Audio-visual gen- eralized zero-shot learning by reconstructing label features from multi-modal embeddings.

Towards Open-Vocabulary Audio-Visual Event Localization Avgzslnet: Audio-visual gen- eralized zero-shot learning by reconstructing label features from multi-modal embeddings

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.752061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.101963Z digest=sha256:97ac8c7ecd6f8d0f3e669763cf217634e45083e3b7100265c2d992d28789ac6d

Observation d974016a-dbb9-4d62-a7da-f52c4ecae096 · outbound

This paper cites Temporal and cross-modal at- tention for audio-visual zero-shot learning.

Towards Open-Vocabulary Audio-Visual Event Localization Temporal and cross-modal at- tention for audio-visual zero-shot learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.738739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.106230Z digest=sha256:211520388bc28590404623c17eb437ce62ba4dd10a805c5a72ba6e4e191d8624

Observation 0d2011d4-2b03-463c-80cb-468d84ce853a · outbound

This paper cites Audio-visual generalised zero-shot learning with cross-modal attention and language.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual generalised zero-shot learning with cross-modal attention and language

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.725021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.110153Z digest=sha256:638a66e5a2d451892586274a01e952ec9456a9ee082fe94147a46b9eceee0812

Observation c273f183-ba0f-4b3d-b47b-84f3c7c1f979 · outbound

This paper cites Localizing visual sounds the easy way.

Towards Open-Vocabulary Audio-Visual Event Localization Localizing visual sounds the easy way

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.711381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.114184Z digest=sha256:c26c7bc0f95db5d95c1c124f12b1e1f0bb2aadb3970fbe2d6b3d4acbeac3c93a

Observation 532280ed-dce2-424b-8d00-1fbf9c3fb8eb · outbound

This paper cites Audio-visual Generalized Zero-shot Learning the Easy Way.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual Generalized Zero-shot Learning the Easy Way

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:46:09.331158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.118070Z digest=sha256:5f73c5573ad03c05db339ce2070ba9b4bbaae30667bb7e85aa750a57bc842758

Observation e1baf286-9313-4b41-8078-83b2e386eef0 · outbound

This paper cites PG-Video-LLaVA: Pixel Grounding Large Video-Language Models.

Towards Open-Vocabulary Audio-Visual Event Localization PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.122320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.122320Z digest=sha256:8527f0f26a752190294040a2299bc000dc15dce1f63b4aaac2a933f864fc6d51

Observation 2b3d1c9b-ae43-4380-b520-da37344b8188 · outbound

This paper cites Coordinated joint multimodal embeddings for gen- eralized audio-visual zero-shot classification and retrieval of videos.

Towards Open-Vocabulary Audio-Visual Event Localization Coordinated joint multimodal embeddings for gen- eralized audio-visual zero-shot classification and retrieval of videos

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.697966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.126554Z digest=sha256:594fb825eca4d0c93a788d500083b986405819a5c0b60019ff75ac733cbb1b7e

Observation 96feb531-bea0-4026-91b0-970b7faa4a88 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Towards Open-Vocabulary Audio-Visual Event Localization Multiple sound sources localization from coarse to fine

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.685165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.130581Z digest=sha256:0176f3aaa25c8085e94d17f27c12b5f0403493fa17e4a5196eef4cf5acd1003a

Observation b994d640-c782-4562-8410-0d4860a4332f · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Towards Open-Vocabulary Audio-Visual Event Localization Learn- ing transferable visual models from natural language super- vision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.672549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.134348Z digest=sha256:69cd38eeee414a5f287ed6aa33669dc71a5b3ef015d4a061c8be05ce57cac30b

Observation 35c39a77-6b3a-465e-ba69-8414e423aa8f · outbound

This paper cites Fine-grained audible video description.

Towards Open-Vocabulary Audio-Visual Event Localization Fine-grained audible video description

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.660094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.138191Z digest=sha256:fdea1324f11e027ed1b8fdf8e79de3facf460741a73b0190217d96d63b6ef2b4

Observation 1c508258-4c6d-426e-bdbb-c3f94c062d87 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-visual event localization in unconstrained videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.647523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.142176Z digest=sha256:710567c49ce0f86a716095160843989220447e4058114000af90e527157c78b0

Observation 7d863dad-2e44-40a8-aab0-12e60cb51587 · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

Towards Open-Vocabulary Audio-Visual Event Localization Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.635048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.146224Z digest=sha256:5d85b188645358d1b00d0b26ff704e8f569cf6f036c2961fa440a7c39bd3425c

Observation 5c33d0ff-547b-4703-b12e-4454b95a9959 · outbound

This paper cites Attention is all you need.

Towards Open-Vocabulary Audio-Visual Event Localization Attention is all you need

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.621497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.150182Z digest=sha256:52dbfe0ccaef2cca39ff2736be60b3caf35a05f0f04352e67c0c0023456135c0

Observation ac6509aa-b06f-4c3c-a7fa-3e181c1cbf3c · outbound

This paper cites Dual attention matching for audio-visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Dual attention matching for audio-visual event localization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.608346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.154304Z digest=sha256:2c3d14ee82c78be0316c1e1b478f2877ea3924d548838f3341b7ad10d02d0e50

Observation 4f50f783-1427-4760-8706-ac9a7d3f2cea · outbound

This paper cites Span-based audio-visual localization.

Towards Open-Vocabulary Audio-Visual Event Localization Span-based audio-visual localization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.595451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.158467Z digest=sha256:d226c20ef3ce58387199404b61b0f4eea0f1d484f36c0935437a6e8652138274

Observation df237228-da4d-46ce-acf6-e08773a08056 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.581830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.162578Z digest=sha256:e9a868b7bac88f21988ec54fbeccec2366121522744565c5378ca80a9ce2f31a

Observation 9bbcc665-b21d-40be-b3de-975a8b0bbe90 · outbound

This paper cites Cross-modal background suppres- sion for audio-visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Cross-modal background suppres- sion for audio-visual event localization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.568696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.167077Z digest=sha256:cbab2d7b3e915a208d8fd50e5cfd05b2b04f6f42849dd0caf803426deaf7bdb8

Observation 37030a44-de52-4b64-913e-cc1ee7b9df23 · outbound

This paper cites Cross-modal relation-aware networks for audio-visual event localization.

Towards Open-Vocabulary Audio-Visual Event Localization Cross-modal relation-aware networks for audio-visual event localization

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.554943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.171117Z digest=sha256:3bb3168eab6bb50c62fa1205c189a8ed124d7e26896477ae32348f12e717e795

Observation e8d08bf5-7c82-43b6-beed-f170b905e4a7 · outbound

This paper cites Avqa: A dataset for audio- visual question answering on videos.

Towards Open-Vocabulary Audio-Visual Event Localization Avqa: A dataset for audio- visual question answering on videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.541968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.175171Z digest=sha256:a04eb1c719b4607f1f64771eb5a9a75c3cc22453b4b008c211e718983cd976f1

Observation 595072d6-bec8-41d5-911a-50e0a92a590a · outbound

This paper cites MM-Pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing.

Towards Open-Vocabulary Audio-Visual Event Localization MM-Pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.529073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.179231Z digest=sha256:e5135716f15dfa6b2ab702cb75684d7b164f5c14f8ea0dc3c369d6f928f18695

Observation 1f5ecc0c-7de5-4b10-8d1d-c3084f9907a6 · outbound

This paper cites Ope- nA VE: Moving towards open set audio-visual event localiza- tion.

Towards Open-Vocabulary Audio-Visual Event Localization Ope- nA VE: Moving towards open set audio-visual event localiza- tion

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.515990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.183240Z digest=sha256:4afd38e85468de94d889b01a65d950273b95471950f5881397936cc9363cb44d

Observation 5bfb899d-dbe4-41ae-b09e-6ccaccffd004 · outbound

This paper cites Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing.

Towards Open-Vocabulary Audio-Visual Event Localization Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.187257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.187257Z digest=sha256:5d8dea6f2e99441e1486d76d35c0ec7e383dc643102883efd3defa74d504bb4e

Observation 852432ce-a2c2-4942-8fed-79a9d7ebf983 · outbound

This paper cites Positive sample propagation along the audio- visual event line.

Towards Open-Vocabulary Audio-Visual Event Localization Positive sample propagation along the audio- visual event line

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.501606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.191431Z digest=sha256:da9c1dac649fd25d0a07dfbb42577b03a64ca20e3052ca0dcecbd1ce880d6aec

Observation 9da1f39c-47e7-491a-80d6-66f8272c44ae · outbound

This paper cites Contrastive pos- itive sample propagation along the audio-visual event line.

Towards Open-Vocabulary Audio-Visual Event Localization Contrastive pos- itive sample propagation along the audio-visual event line

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.487226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.195775Z digest=sha256:7e0f2c4b7256dbb02149ed9e9c3de73c44be934ab176e246137a5ca24462cf24

Observation 2c50363e-0199-41f4-b865-1a4ca2b65e4d · outbound

This paper cites Audio–visual segmentation.

Towards Open-Vocabulary Audio-Visual Event Localization Audio–visual segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.473057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.199812Z digest=sha256:f0832cfe80da75b04a42a0589fcab82eeea274b9a4c1b0134be7ec631026be2d

Observation fdd1c146-5842-494b-8172-761b85846dd7 · outbound

This paper cites Improving Audio-Visual Video Parsing with Pseudo Visual Labels.

Towards Open-Vocabulary Audio-Visual Event Localization Improving Audio-Visual Video Parsing with Pseudo Visual Labels

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.203800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.203800Z digest=sha256:d5e7c30fd011f23ba8a19c69b2754239c06c55b888e3bc3c635423d5f099d787

Observation 628c6665-f9bf-4f0b-8fe9-9feeb3d577ba · outbound

This paper cites Audio-Visual Segmentation with Semantics.

Towards Open-Vocabulary Audio-Visual Event Localization Audio-Visual Segmentation with Semantics

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.208097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.208097Z digest=sha256:0f6b1cf85eae7929a3a3012c09d1359ed14a05825a6ae87781e9cbc796f087df

Observation 6175d65a-f33b-4108-a611-6608d1b1ccd6 · outbound

This paper cites Label-anticipated event disentan- glement for audio-visual video parsing.

Towards Open-Vocabulary Audio-Visual Event Localization Label-anticipated event disentan- glement for audio-visual video parsing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.459970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.212176Z digest=sha256:beed1ddfd109c6620f73da8af53d2c91732df1cadda5533f64ba3e22d9295d7f

Observation afc1a1ef-2e6f-4c20-9714-4ecb58d0ea83 · outbound

This paper cites Ad- vancing weakly-supervised audio-visual video parsing via segment-wise pseudo labeling.

Towards Open-Vocabulary Audio-Visual Event Localization Ad- vancing weakly-supervised audio-visual video parsing via segment-wise pseudo labeling

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:46:09.446303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.216131Z digest=sha256:cb52ed174cda4b900734798ae19b108d1eb1edb6b6264ec2ec850a8e0ec5c46a

Observation 2304d58c-8bc4-4702-93b2-f8a419010663 · outbound

This paper cites Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration.

Towards Open-Vocabulary Audio-Visual Event Localization Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T18:46:09.220172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:46:09.220172Z digest=sha256:fec52c890c4ceecf8cc9bfbeebdaec2f58aa3e6eacc381f03d704613ed25afe2

Observation 9f6023f4-5c01-4c81-926d-3ef258f96086 · outbound

This paper cites Instruction: For the given 10-second video, divide it into 10 one-second segments. For each segment, if its audio and visual streams describe the same event, assign the label “x.

Towards Open-Vocabulary Audio-Visual Event Localization Instruction: For the given 10-second video, divide it into 10 one-second segments. For each segment, if its audio and visual streams describe the same event, assign the label “x

Reference 2024

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T18:46:09.432304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:46:09.224346Z digest=sha256:7f7857f9f3cf34818ba69d0ffaaa4d5baedabbdc4e88bda17527b004e02d78a4

Pith citing papers

Observation c92d3f27-fe3c-45ba-ae38-67084b2e76f7 · inbound

Patch-level Sounding Object Tracking for Audio-Visual Question Answering cites this paper.

Patch-level Sounding Object Tracking for Audio-Visual Question Answering Towards Open-Vocabulary Audio-Visual Event Localization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:22.538919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:42:22.538919Z digest=sha256:82bb91388f56e5ccdb8a40ee68d4dba2370bb40bd817cbb703bfbec406d9ae60

Observation 48fc6d8a-1ce9-49e8-858d-5fa97ed8ff73 · inbound

Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing cites this paper.

Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing Towards Open-Vocabulary Audio-Visual Event Localization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:13:01.650537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:13:01.650537Z digest=sha256:fb86ccd3c34336a00f8d5797a3fe72b16eb9ab9571da89eca9c9f3b0c8e03e4d

Observation 713eb6f8-0602-456b-9bf5-3a8246146c5a · inbound

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration cites this paper.

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration Towards Open-Vocabulary Audio-Visual Event Localization

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:56:18.454468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T13:56:18.376640Z digest=sha256:23e1a04dd44519ef1ef2a393efbd089b20fac672e437d2e35548b0c2977106b0