Pith. sign in

Paper Citation Record · LEDGER

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

As of 8 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2506.23196.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23196 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:53:02.484752Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:03:13.496634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:06:19.413098Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy49
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6adf58fb-1397-49f7-ac85-0b38c8e91849 · outbound

This paper cites Maas: Multi-modal assignation for active speaker detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Maas: Multi-modal assignation for active speaker detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.228540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:56.383564Z digest=sha256:a284353b4e34d53086498e3e08f7335daa5cda62ae499fe360c2389444596c9e

Observation 28ad82c1-9139-4562-a667-3d761d6b933e · outbound

This paper cites Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:56.439313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:56.439313Z digest=sha256:84bb28726756182a860d8937b65ada6ceec0be36e1b174b99d76ac101ef48153

Observation d47090ca-d83c-438b-ae89-a2525b2859f1 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.031192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:56.521994Z digest=sha256:87436dc7ba7ea398ba82461b0e7478aa164a8e6c5b58596f6f8c64e5b1681026

Observation 571de3b4-4ced-492a-bbb9-51b03117b901 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.821632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:56.624292Z digest=sha256:186618d74617ca614f51656b876cea3a261c5610cdbc84c6cda2a5ec578b22c8

Observation 7803e0e6-ca3a-4bb7-bcb6-0a03ba5ed688 · outbound

This paper cites Augmented transformer with adaptive graph for tem- poral action proposal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Augmented transformer with adaptive graph for tem- poral action proposal generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.636257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:56.734272Z digest=sha256:ce9888510aa99f31e04840db57a24144af7473716aee6d0077eb94db6ba84d6e

Observation 7951d87f-64e0-4430-b5da-69ee90e7f60b · outbound

This paper cites Re- thinking the faster r-cnn architecture for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Re- thinking the faster r-cnn architecture for temporal action localization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.462906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:56.792960Z digest=sha256:3967258f1a1a58ade73d23be675ee84e1247538fc07e863a5fad6ee33bcb6cbe

Observation e9949550-2b40-4f80-abfb-904784a08b92 · outbound

This paper cites Tallformer: Temporal ac- tion localization with a long-memory transformer.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Tallformer: Temporal ac- tion localization with a long-memory transformer

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.244143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:56.866148Z digest=sha256:7a049f77466430e6dfb6c7171854d067d046b95d3a60a6188033cbf0b0f106bd

Observation 86994997-3545-44b3-b9fa-b88eb1c620a3 · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Yolo-world: Real-time open-vocabulary object detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:56.942447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:56.942447Z digest=sha256:5e25ed64bc001e2559f951af5c7f2f986fb99b881147bf7be0319f6dc3aba479

Observation ad95e9b2-49f4-46ee-8d25-d3c8b3e77441 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.029369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:57.080544Z digest=sha256:81f383c81dac288d55ebebfee896ecaf2d7eadfd3295c8227d748fb194c36c2b

Observation 839918ca-bc76-4068-b0b0-a4420566535b · outbound

This paper cites Slowfast networks for video recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Slowfast networks for video recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.855384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:57.209727Z digest=sha256:3af8e838ac732b7e1dc109828cd3306b9cd4c7f3d8235cf5dae0b2423ac80491

Observation 4945c444-97ff-41ba-8e53-54c9f5253efe · outbound

This paper cites Multi-Resolution Audio-Visual Feature Fusion for Temporal Action Localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Multi-Resolution Audio-Visual Feature Fusion for Temporal Action Localization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:53:02.785904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:57.316945Z digest=sha256:99078dbfe3204fe2fef06d80d639bdc8ec635de6e30b579aeb6206070ec32d73

Observation 1ff4a861-5dc7-4132-94a8-9ed8a34b6d33 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio set: An ontology and human- labeled dataset for audio events

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.698271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:57.408175Z digest=sha256:4c87180e03b96bac1ab1e67fb1fa48cf801f02ed2065f77f4e66c2ad57ba3b12

Observation c02bd514-8400-48f0-aab9-74bc46808358 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.463129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:57.495631Z digest=sha256:2cefba1526d3b3f2b135b4073515d54d6cd0d911c6ecbd99583c69d97ade6567

Observation 8c37d1a1-7b8a-46d5-88bd-e30a619c88e8 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Momentum contrast for unsupervised visual rep- resentation learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.588220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.588220Z digest=sha256:16378c132553c03c4c27b3a79769b7f3f7340647f4b7652fa019247e79277612

Observation 6c7fba88-18db-429b-b253-bbc1b260e576 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Cnn archi- tectures for large-scale audio classification

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.259205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:57.718925Z digest=sha256:6be5740722c40a54725796597434ca7fac4f72368c1f7499cdc1b69ed36d9b7e

Observation 4dda55a1-49df-46a2-b808-68de94d62b04 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mix and local- ize: Localizing sound sources in mixtures

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.797874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.797874Z digest=sha256:3233ba497cffe6d0a693b0ca52d76079bf2e65b6fef0618cf90aa01d3493cae3

Observation 45643fd4-60c9-4124-aee3-f8be220466a2 · outbound

This paper cites in the wild.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.051074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:57.881569Z digest=sha256:77f1db767d3112b56574e00c78384106221ff2ea9d3daf85e25b2f5fe26e56b2

Observation 80271b71-b850-4d92-9f4c-96a1906fe8c5 · outbound

This paper cites Causal inference meets deep learning: A compre- hensive survey.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Causal inference meets deep learning: A compre- hensive survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.820174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.000025Z digest=sha256:ba7f861731f6e1e980b9da740947bf03ff602645b10abf3db20f6b3b062806eb

Observation b31a13d5-21c4-4f59-8c87-5aed19bb0808 · outbound

This paper cites Epic-fusion: Audio-visual temporal bind- ing for egocentric action recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Epic-fusion: Audio-visual temporal bind- ing for egocentric action recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.637366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.137657Z digest=sha256:9d47e50dbd963e4b3ea1f2d119be2e06e611730c25afd80608cc0cdfc35acc1e

Observation 174d3af3-f042-43d5-b345-337e0e8e60ae · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.255447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.255447Z digest=sha256:f4885a7269d724bb305c0f90ccde9833d459698d67cdd7c1d873a2fec201a9e5

Observation 67fe8343-b81f-444f-9065-b2f8139842fa · outbound

This paper cites Learning salient boundary feature for anchor- free temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Learning salient boundary feature for anchor- free temporal action localization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.492907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.345630Z digest=sha256:a0b028da63694f8352c62a0f238211624badc2a1bf03857973f98c01ba73fbee

Observation 9f0d2f44-002d-4240-addb-f9b5143a8927 · outbound

This paper cites Bsn: Boundary sensitive network for temporal action proposal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bsn: Boundary sensitive network for temporal action proposal generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.315122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.449901Z digest=sha256:e989fad72e473b805cfd1544ab751a0bd58ad14574ce54d844dbc146b22eb7c0

Observation 4f7e9a09-c5c0-4f1b-b4aa-68115365b8dc · outbound

This paper cites Bmn: Boundary-matching network for temporal action pro- posal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bmn: Boundary-matching network for temporal action pro- posal generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.542307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.542307Z digest=sha256:164e464f12bd163087e8253d2f82667c99d1c539747099f3967e4bb495d5fab7

Observation c11b35a9-a764-4db3-b9e1-42bc4f26be31 · outbound

This paper cites Progressive boundary refine- ment network for temporal action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Progressive boundary refine- ment network for temporal action detection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.184494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.669443Z digest=sha256:021a593b14952304073b9ba8127b748e3a252298de4f51b5549e4603b399a580

Observation 145a2007-89c2-451d-8ebb-17466280acb4 · outbound

This paper cites Dense modality interaction network for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dense modality interaction network for audio-visual event localization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.038197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.774936Z digest=sha256:f757cfaa8351c9aaa256e46f9712535bd485ac4c98175ceb61428bd7579c0bee

Observation 8e6cbd72-b2e8-469b-be94-c96ae7f4cf77 · outbound

This paper cites Multi-shot temporal event localization: a benchmark.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Multi-shot temporal event localization: a benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.873567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.879243Z digest=sha256:2524bfe437036adfeed81f5c3dcf2a51552b278be71c07f483cf9016b03fb67d

Observation 26109f82-ccde-42db-88d9-595505edf9da · outbound

This paper cites End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.703945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.983937Z digest=sha256:05a13bb8bfddd758a9d74dc29a93aae5df5b92fe02db80d89f3d7cc371cb8e9c

Observation 0a9edb4e-f849-4d1f-8cdc-95ea4c662f25 · outbound

This paper cites Gaussian temporal awareness networks for action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Gaussian temporal awareness networks for action localization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.524721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:59.119360Z digest=sha256:4c105aad4aa36be39395d78805d27fa26a8e89dd5bc5ea2fb7693128c0c46c68

Observation cbf46e29-3e88-4069-84ca-3f61652542e3 · outbound

This paper cites Proposal-free temporal action detection via global segmen- tation mask learning.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Proposal-free temporal action detection via global segmen- tation mask learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.374265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:59.215515Z digest=sha256:c18bf4eb752795c7c2c05f3678089cf0767f206b73f1a07c0d69c7400b0c45ea

Observation f4bc45dc-7480-4baa-8f38-360c9f7af30d · outbound

This paper cites Attention bottlenecks for multimodal fusion.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Attention bottlenecks for multimodal fusion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.185290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:59.303981Z digest=sha256:5861299a4bc72b14a22fc84f2125d01b7e2a75eb705725d2b182cd5bb27c69c4

Observation fc53abfb-da85-4cd7-a9d3-e13ef2b7475a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.408856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.408856Z digest=sha256:5a60082ffbf46debb68ce5b239d6f59b60332c9574a355161e0f52129c2d2cd9

Observation 57daa4b0-d0df-4a9f-90cc-335b1600a2e1 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual scene analysis with self-supervised multisensory features

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.046715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:59.543003Z digest=sha256:eb4bf19157053ae6e4cc12baa68b03d0902c1e48ee58b9e841fad8bac911989a

Observation ccab2e5c-8bc7-4cca-a040-28fc3e3f8836 · outbound

This paper cites A review of deep learning techniques in audio event recognition (aer) applications.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding A review of deep learning techniques in audio event recognition (aer) applications

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.877014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:59.622948Z digest=sha256:60486fd4ee1dd145529bbd93efd49915996750ddad9538ad9a0ee6fde29cd8b4

Observation 7cd8b596-8223-4b2a-a5e4-5003dea88cea · outbound

This paper cites Owl (observe, watch, listen): Audiovisual temporal context for localizing actions in ego- centric videos.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Owl (observe, watch, listen): Audiovisual temporal context for localizing actions in ego- centric videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.703910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:59.736617Z digest=sha256:d4388d98304f7de2b83f75bdb10d675c8a373f710a5c40e54eaa3df857c5f94c

Observation d160e83d-f8e2-4f1a-887e-ef62cb617543 · outbound

This paper cites You only look once: Unified, real-time object de- tection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding You only look once: Unified, real-time object de- tection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.832187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.832187Z digest=sha256:f7519d979c52e726ea08c4d29a26a338f2a933644c355cd3aa30fdbc13bbf69f

Observation 9b5c229e-e276-4ca3-91a4-bba32b5d3587 · outbound

This paper cites Action sensitivity learning for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Action sensitivity learning for temporal action localization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.922852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.922852Z digest=sha256:0d7f7864b3ae7c200784c94187c2fb1fba2d784046f2a280a613c152966187d0

Observation e4f8db01-5425-4c93-b34b-101ce7945aec · outbound

This paper cites Temporal Action Localization with Enhanced Instant Discriminability.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal Action Localization with Enhanced Instant Discriminability

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.048092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.048092Z digest=sha256:03295ea1730a2d0115396c1023bb37f3aa2ee9938ac66856f6fa7c2c4e66f2eb

Observation 6f57e020-44bd-4943-8662-b77ad44fd57b · outbound

This paper cites Tridet: Temporal action detection with relative boundary modeling.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Tridet: Temporal action detection with relative boundary modeling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.463989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.153296Z digest=sha256:fc170deb377a57b33a2eacb95cdfe7717ebc33d25c04a0d19e1d41512fc46617

Observation 022467a2-d6ab-4a35-9887-f852a27e4e6e · outbound

This paper cites Re- laxed transformer decoders for direct action proposal gener- ation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Re- laxed transformer decoders for direct action proposal gener- ation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.134101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.248074Z digest=sha256:90df3b48eed50626d89c1b7a9f4a633dd6443e5b0e83f0c1e284aa510972a004

Observation 3a408326-4418-43bf-9b67-396df641f1d3 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual event localization in unconstrained videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.791765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.340814Z digest=sha256:a3ba75d452aef988a32bd6b482bd0cff995295ccad4fad51175207c7b7171c6f

Observation 0dd7ef7a-81b6-4b5e-a317-d1eabb452cf0 · outbound

This paper cites Deep learning-based action detection in untrimmed videos: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(4):4302– 4320, 2022.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Deep learning-based action detection in untrimmed videos: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(4):4302– 4320, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.518398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.491979Z digest=sha256:83c98ca6b034f5870565060b963909065c8c3e9926d783d5dd48d1edc7141450

Observation 98b04cc2-17fc-4d29-882b-6bf07547d5c8 · outbound

This paper cites You only hear once: a yolo-like algorithm for audio segmentation and sound event detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding You only hear once: a yolo-like algorithm for audio segmentation and sound event detection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.182668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.581690Z digest=sha256:8f3589d42a1e2ced8cbbf832da19d2091a3cb0587e25c6a699665b8d6e429ab9

Observation 7d18d4f0-4ac5-40a9-bb98-6fa3bfae5bf0 · outbound

This paper cites Temporal Action Proposal Generation with Transformers.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal Action Proposal Generation with Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.679239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.679239Z digest=sha256:29716ef788eb9a45eddb8fcf00617e89eaf7a07625bb51fdff82c102b8cbcdae

Observation f8281b68-9c52-4f7c-b270-0664967836d9 · outbound

This paper cites Rcl: Recurrent continuous localization for temporal action detec- tion.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Rcl: Recurrent continuous localization for temporal action detec- tion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.894285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.753442Z digest=sha256:ee922714502fdb997e2d2f36bbad65e9b69f82a2a6d76d4d9516cf2b6cb4d9bf

Observation 782a3c0a-83ca-4af2-a2aa-a91be51477db · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.706511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.861896Z digest=sha256:9e5dcb9f71b5752c962928a943a4e08c87792176514483c7ee172caaa814abb9

Observation 75eb99e0-c396-48bb-8d86-7d12f6842739 · outbound

This paper cites An efficient spatio-temporal pyramid transformer for action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding An efficient spatio-temporal pyramid transformer for action detection

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.526823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.939127Z digest=sha256:ac09f6d023d59d227cd337fae26148daed9c9d4754221d4ab0237d636f0ed85d

Observation e2ea7e57-8730-4fd1-b28f-78e1db40ad0e · outbound

This paper cites Dual attention matching for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dual attention matching for audio-visual event localization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.011214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.011214Z digest=sha256:d5d7ab841443211d27fd095de21d41fd854d5b2879ca2261abda2cbc17d7a0ed

Observation ac19a1c4-9689-4c5a-9d93-259124fe18b1 · outbound

This paper cites Dual relation network for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dual relation network for temporal action localization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.267675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.089484Z digest=sha256:4ad025ca055ac72911dc2e0b0e2661981c9a2f9f27a303dddd3a10130e76242d

Observation 91a4d0c6-7b42-4adf-85ea-1b8476481198 · outbound

This paper cites Learning to refactor action and co-occurrence fea- tures for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Learning to refactor action and co-occurrence fea- tures for temporal action localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.930022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.205808Z digest=sha256:f8f76f96885364adfdc4a57b5feb276b5b9db3096ae5fe44716eda12bc555ccb

Observation eeeb6297-1c94-4403-b82a-165461958cee · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audiovisual SlowFast Networks for Video Recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.288755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.288755Z digest=sha256:651a653fcf29428f8553a116cd8f4aa4316f739a8b2e9f50897facabb04d0102

Observation f2d36044-0dc2-4d7f-bd5c-48f4b73bdce9 · outbound

This paper cites G-tad: Sub-graph localization for tempo- ral action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding G-tad: Sub-graph localization for tempo- ral action detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.649130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.409621Z digest=sha256:30e00a9fcba05ed851520274320d4beefad6086d131b3f62426707ab59bc43c7

Observation ace9aa87-5799-4b7e-a854-2a1df9253e52 · outbound

This paper cites Audio-visual event localization by learning spatial and semantic co-attention.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual event localization by learning spatial and semantic co-attention

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.300311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.497261Z digest=sha256:4b79e3a3873c4084a3f9a5819818961291d69c9c200756d0c0ac382e2a8338bf

Observation 98644f61-cfbb-4f39-b9fd-1f9d39f146d3 · outbound

This paper cites Temporal pyramid network for action recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal pyramid network for action recognition

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.087532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.572693Z digest=sha256:463c12822670daca00c0ae3f3a9720146083530bb54f2c7b03b942da7256c53c

Observation 3e4f63c9-fc4c-4de1-b7bc-31f74b47d100 · outbound

This paper cites Revisiting anchor mechanisms for temporal ac- tion localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Revisiting anchor mechanisms for temporal ac- tion localization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.961697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.667259Z digest=sha256:f64e3181f711ae09b0e59afdce7ddcb7208d223eec87a6dcf18f5c064c62df44

Observation da56df42-9e95-4be0-ae0e-8dd3e2543d4c · outbound

This paper cites Mpn: Multimodal parallel network for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mpn: Multimodal parallel network for audio-visual event localization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.839549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.739029Z digest=sha256:b8fc0197561e9755c27b33aaed954e7b05b971cc42fcf5bacc163208526a7983

Observation eec95379-296b-495c-a05a-27b0704c6cfc · outbound

This paper cites Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.739506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.822512Z digest=sha256:eef95a0dff94d61c055fe01c63bc618bd5bdb91ab6c4156df19d589bf212f019

Observation d9106c28-a30f-4748-a608-e214a731aa28 · outbound

This paper cites Graph con- volutional networks for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Graph con- volutional networks for temporal action localization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.628158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.887537Z digest=sha256:4f48b9d66c9acf5a202c05df527924a122466d6ead598f4c0292656be1fe9f1b

Observation ee01785a-809a-4097-9844-f5cb33925766 · outbound

This paper cites Actionformer: Lo- calizing moments of actions with transformers.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Actionformer: Lo- calizing moments of actions with transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.984274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.984274Z digest=sha256:27dddf9f631d3634f309f2319ee9b8d8359caaf31edc26fc8a64e2c251821b83

Observation d8015066-5343-4a61-a764-7c223f5fe782 · outbound

This paper cites Video self- stitching graph network for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Video self- stitching graph network for temporal action localization

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.522826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.073232Z digest=sha256:ff95794a35045ef63a165adb254abea7e87a4bff83829d509de09fc96fe6aed8

Observation b311c305-b820-4278-8bce-a976fe78e285 · outbound

This paper cites Bottom-up temporal action localization with mutual regularization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bottom-up temporal action localization with mutual regularization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.412100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.139404Z digest=sha256:f8a665a14f5daea4f9c74c555445eddcef966eda3c2b43fa048b09c952fac9b7

Observation 25fc7dc9-2928-4cbc-b6dd-d57993c85d7b · outbound

This paper cites Enriching local and global contexts for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Enriching local and global contexts for temporal action localization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.286231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.210024Z digest=sha256:dcff1c1612ba887b7d9c4993409db0e5353f53ffc9b236b23e70b4c74d3bd461

Observation 9b8f3925-1ac2-4f58-af5d-74b122c37715 · outbound

This paper cites Our code provides further information.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Our code provides further information

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.160024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.296518Z digest=sha256:fb13d352624e070b5dd4decac94a4da92cc6ad704ca8412af7dc652c65c070f0

Observation bbc44088-7e4e-420f-a1c0-02c4353f1649 · outbound

This paper cites Performance was measured us- ing mAP@[0.5:0.05:0.95], along with the average mAP.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Performance was measured us- ing mAP@[0.5:0.05:0.95], along with the average mAP

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.035271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.400314Z digest=sha256:91c91650d0d5e3816fd59c418ef1500105a52efcdf6eddd2dd5269c63e2506e5

Observation cda74189-f55c-4927-a1f0-11c380a57b4c · outbound

This paper cites This approach ensures greater feature consis- tency across different temporal resolutions, ultimately im- proving the regression of an event’s location.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding This approach ensures greater feature consis- tency across different temporal resolutions, ultimately im- proving the regression of an event’s location

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:02.924236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.484752Z digest=sha256:67752ca0f28ce5bc1b33a9c84fc4eeb587ac176c0ef3cee4426c708438bd0096

Pith citing papers

Observation 47b14a0d-d100-4117-a088-aa437b827869 · inbound

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing cites this paper.

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:19.415238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:03:13.496634Z digest=sha256:8c078d708e46bdc495d80100c60f81875a9089d9a3ea165b04ac3a3311684da1