Pith. sign in

Paper Citation Record · LEDGER

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

As of 22 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2506.23196.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23196 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:53:02.484752Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:03:13.496634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:06:19.413098Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy49
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6adf58fb-1397-49f7-ac85-0b38c8e91849 · outbound

This paper cites Maas: Multi-modal assignation for active speaker detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Maas: Multi-modal assignation for active speaker detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.228540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:56.383564Z digest=sha256:6f35926952914502e98ae4006f6686968718276e71466c5d20a8a7b26bd90ff6

Observation 28ad82c1-9139-4562-a667-3d761d6b933e · outbound

This paper cites Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:56.439313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:56.439313Z digest=sha256:7234f21d2e5bb5c6f4e929ba6344cde203e00a6e7f6b7c5d228c28b2a7c9225f

Observation d47090ca-d83c-438b-ae89-a2525b2859f1 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.031192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:56.521994Z digest=sha256:7184f74f603154835d13f3c8d618d2f1c6758d47eaea7668408ab563b8eb4d06

Observation 571de3b4-4ced-492a-bbb9-51b03117b901 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.821632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:56.624292Z digest=sha256:ae78daa818f19dc2f4a039c7f17b35b682187172a6c0a098c057e5b24699d8af

Observation 7803e0e6-ca3a-4bb7-bcb6-0a03ba5ed688 · outbound

This paper cites Augmented transformer with adaptive graph for tem- poral action proposal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Augmented transformer with adaptive graph for tem- poral action proposal generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.636257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:56.734272Z digest=sha256:114fdfce3f31c72c30736c8c40417a5253e5389dba3b0dfda9243fd4fae53052

Observation 7951d87f-64e0-4430-b5da-69ee90e7f60b · outbound

This paper cites Re- thinking the faster r-cnn architecture for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Re- thinking the faster r-cnn architecture for temporal action localization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.462906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:56.792960Z digest=sha256:59286183bf054244e992bb4e41997c371c084011db4b4f77674d7f76f4e72569

Observation e9949550-2b40-4f80-abfb-904784a08b92 · outbound

This paper cites Tallformer: Temporal ac- tion localization with a long-memory transformer.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Tallformer: Temporal ac- tion localization with a long-memory transformer

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.244143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:56.866148Z digest=sha256:1d366866953ce037ffb8b8129640399da2805a31d2af447f7ada5626d8041a9f

Observation 86994997-3545-44b3-b9fa-b88eb1c620a3 · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Yolo-world: Real-time open-vocabulary object detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:56.942447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:56.942447Z digest=sha256:ba0a5202ff2906d58fba3667af615ecc11ed65c7bce2dc201564bec15bc834fd

Observation ad95e9b2-49f4-46ee-8d25-d3c8b3e77441 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.029369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:57.080544Z digest=sha256:a408e84f6fa6dd3edf9be0dee182d4e2bb5fc8c5718fdf7300a2d6fec7e4861e

Observation 839918ca-bc76-4068-b0b0-a4420566535b · outbound

This paper cites Slowfast networks for video recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Slowfast networks for video recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.855384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:57.209727Z digest=sha256:809335e554fe5230cb004a7044dbbe702c045a67625b62e5e4bb15c7376c28f1

Observation 4945c444-97ff-41ba-8e53-54c9f5253efe · outbound

This paper cites Multi-Resolution Audio-Visual Feature Fusion for Temporal Action Localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Multi-Resolution Audio-Visual Feature Fusion for Temporal Action Localization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:53:02.785904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:57.316945Z digest=sha256:a557ad4ef86a407a78732c3165088e48a69a138f03118a11f0239191b7c50f31

Observation 1ff4a861-5dc7-4132-94a8-9ed8a34b6d33 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio set: An ontology and human- labeled dataset for audio events

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.698271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:57.408175Z digest=sha256:958e05ce043e8e87dc0d14e455aab3a0ee245d74dba834d99dc3d2be99dfdf32

Observation c02bd514-8400-48f0-aab9-74bc46808358 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.463129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:57.495631Z digest=sha256:589502c99a5d0ce25f6d537090eb9c6abb8385880d7491a2fc3116d29bd4f7f2

Observation 8c37d1a1-7b8a-46d5-88bd-e30a619c88e8 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Momentum contrast for unsupervised visual rep- resentation learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.588220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.588220Z digest=sha256:12c9c9dc259820945c766791b111afae27621340787337ae81db636cb77c7de1

Observation 6c7fba88-18db-429b-b253-bbc1b260e576 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Cnn archi- tectures for large-scale audio classification

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.259205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:57.718925Z digest=sha256:a0b341108577a720b411ef4d313b4f71589a7dd6754a90bc20d7e2e4dcb3ddae

Observation 4dda55a1-49df-46a2-b808-68de94d62b04 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mix and local- ize: Localizing sound sources in mixtures

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.797874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.797874Z digest=sha256:ed733633471bb37e77eaa98220b89f4c79be26b8c585db36f00469cda10096db

Observation 45643fd4-60c9-4124-aee3-f8be220466a2 · outbound

This paper cites in the wild.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.051074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:57.881569Z digest=sha256:b96996d04798ec250ed724717d6d3233c23fcea22de924dbd31fa8f76652f18c

Observation 80271b71-b850-4d92-9f4c-96a1906fe8c5 · outbound

This paper cites Causal inference meets deep learning: A compre- hensive survey.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Causal inference meets deep learning: A compre- hensive survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.820174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:58.000025Z digest=sha256:6aadd2fc75c0415f0794c4804e61faea55f0f0e66ea4af940a35baa8eded8f30

Observation b31a13d5-21c4-4f59-8c87-5aed19bb0808 · outbound

This paper cites Epic-fusion: Audio-visual temporal bind- ing for egocentric action recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Epic-fusion: Audio-visual temporal bind- ing for egocentric action recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.637366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:58.137657Z digest=sha256:ec5965d3cf5f5a994d163db1b9b5a379a993d1544668bf02369dc6ee3cc5aeee

Observation 174d3af3-f042-43d5-b345-337e0e8e60ae · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.255447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.255447Z digest=sha256:cbdbc526209581ae59c6ee5828236af1f5d98269b516e0df6b456a14ed68d91b

Observation 67fe8343-b81f-444f-9065-b2f8139842fa · outbound

This paper cites Learning salient boundary feature for anchor- free temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Learning salient boundary feature for anchor- free temporal action localization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.492907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:58.345630Z digest=sha256:64e277118ab7948c9719beeddabc9191e35d1d163b7fffcdedf015302e6a905b

Observation 9f0d2f44-002d-4240-addb-f9b5143a8927 · outbound

This paper cites Bsn: Boundary sensitive network for temporal action proposal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bsn: Boundary sensitive network for temporal action proposal generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.315122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:58.449901Z digest=sha256:4a97dcc595e69347f4a268bbe3de74f560997e273dee8fb2ce652862e218e144

Observation 4f7e9a09-c5c0-4f1b-b4aa-68115365b8dc · outbound

This paper cites Bmn: Boundary-matching network for temporal action pro- posal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bmn: Boundary-matching network for temporal action pro- posal generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.542307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.542307Z digest=sha256:f6ffead4cb22ecf5208de2f6e385b22cab687cb40efe9b88431096a787d52af3

Observation c11b35a9-a764-4db3-b9e1-42bc4f26be31 · outbound

This paper cites Progressive boundary refine- ment network for temporal action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Progressive boundary refine- ment network for temporal action detection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.184494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:58.669443Z digest=sha256:e6400905808a1fd7482bfe9fdeeed3fddf9a358552c7ff12af32014840c2069d

Observation 145a2007-89c2-451d-8ebb-17466280acb4 · outbound

This paper cites Dense modality interaction network for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dense modality interaction network for audio-visual event localization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.038197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:58.774936Z digest=sha256:9a13f22e9e69be5fee05a0842a99d4ea1b867f200e306ed8346fcbde3b43cbc1

Observation 8e6cbd72-b2e8-469b-be94-c96ae7f4cf77 · outbound

This paper cites Multi-shot temporal event localization: a benchmark.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Multi-shot temporal event localization: a benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.873567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:58.879243Z digest=sha256:d088f4c46ae4bd89ec2a603b18a8eee819af32b911e0ba60513459389c38d5bb

Observation 26109f82-ccde-42db-88d9-595505edf9da · outbound

This paper cites End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.703945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:58.983937Z digest=sha256:450688151a313b9362ce9c09311c3e325e5308465e7ff2652c6c119d36f34ee1

Observation 0a9edb4e-f849-4d1f-8cdc-95ea4c662f25 · outbound

This paper cites Gaussian temporal awareness networks for action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Gaussian temporal awareness networks for action localization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.524721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:59.119360Z digest=sha256:a478df45c12bb92019f19a4fd1eec22dead50f4fd4d96705c4a68a2097119ed3

Observation cbf46e29-3e88-4069-84ca-3f61652542e3 · outbound

This paper cites Proposal-free temporal action detection via global segmen- tation mask learning.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Proposal-free temporal action detection via global segmen- tation mask learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.374265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:59.215515Z digest=sha256:7233ffef78e89ace8a8c0bc815dda17bdeee6f4e3345657a47cc409bb93e0480

Observation f4bc45dc-7480-4baa-8f38-360c9f7af30d · outbound

This paper cites Attention bottlenecks for multimodal fusion.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Attention bottlenecks for multimodal fusion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.185290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:59.303981Z digest=sha256:25cd28db5b98ad52a1440bfdd6b9006d1a9f11a3cc67b6b6e91494e203ae3877

Observation fc53abfb-da85-4cd7-a9d3-e13ef2b7475a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.408856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.408856Z digest=sha256:151b0903124bbd147e74a2f591a149047c281269bcb0b23424bab8bd25d9816f

Observation 57daa4b0-d0df-4a9f-90cc-335b1600a2e1 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual scene analysis with self-supervised multisensory features

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.046715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:59.543003Z digest=sha256:2b50e8c36ea42287d33b92db577bc60d675cc36dea8f424b8eef8d1f918c4e38

Observation ccab2e5c-8bc7-4cca-a040-28fc3e3f8836 · outbound

This paper cites A review of deep learning techniques in audio event recognition (aer) applications.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding A review of deep learning techniques in audio event recognition (aer) applications

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.877014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:59.622948Z digest=sha256:ccd97a634d88d70a1eaa1ccf00fcdf5f9b3efb154421b3f41ad28fa42f8a2094

Observation 7cd8b596-8223-4b2a-a5e4-5003dea88cea · outbound

This paper cites Owl (observe, watch, listen): Audiovisual temporal context for localizing actions in ego- centric videos.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Owl (observe, watch, listen): Audiovisual temporal context for localizing actions in ego- centric videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.703910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:52:59.736617Z digest=sha256:b223fc3eb07b4b04852e02e784f7b131b095a90033c0b7bdc820fd75b3122e55

Observation d160e83d-f8e2-4f1a-887e-ef62cb617543 · outbound

This paper cites You only look once: Unified, real-time object de- tection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding You only look once: Unified, real-time object de- tection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.832187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.832187Z digest=sha256:5ebf312291cb4fd3d709e2d89ed4081473ab2e8ad08e99a7a2eb0cf7d17e5fe6

Observation 9b5c229e-e276-4ca3-91a4-bba32b5d3587 · outbound

This paper cites Action sensitivity learning for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Action sensitivity learning for temporal action localization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.922852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.922852Z digest=sha256:9dea46e81a2a133669b7fe5904a1500a455eac6318697c50cc225e9c6ea17f82

Observation e4f8db01-5425-4c93-b34b-101ce7945aec · outbound

This paper cites Temporal Action Localization with Enhanced Instant Discriminability.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal Action Localization with Enhanced Instant Discriminability

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.048092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.048092Z digest=sha256:24aa7705aa66826fbd71e3a6d27e930ccdbc767605cd12391425204fb7b48e3f

Observation 6f57e020-44bd-4943-8662-b77ad44fd57b · outbound

This paper cites Tridet: Temporal action detection with relative boundary modeling.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Tridet: Temporal action detection with relative boundary modeling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.463989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:00.153296Z digest=sha256:90f9739929f036882a784193d583e8792f4ea224c392b66246a7ad32836504bd

Observation 022467a2-d6ab-4a35-9887-f852a27e4e6e · outbound

This paper cites Re- laxed transformer decoders for direct action proposal gener- ation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Re- laxed transformer decoders for direct action proposal gener- ation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.134101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:00.248074Z digest=sha256:680da3a14f0006d56363f185d9f6f840cb9faa919c7b9ce0b2982c40aaf95d16

Observation 3a408326-4418-43bf-9b67-396df641f1d3 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual event localization in unconstrained videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.791765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:00.340814Z digest=sha256:fe7af85e7698d9314390ad9f9fb648ffc56d1f70a608bdb7aaef703d9908d407

Observation 0dd7ef7a-81b6-4b5e-a317-d1eabb452cf0 · outbound

This paper cites Deep learning-based action detection in untrimmed videos: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(4):4302– 4320, 2022.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Deep learning-based action detection in untrimmed videos: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(4):4302– 4320, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.518398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:00.491979Z digest=sha256:2a3b9260e95ff60e150a8745bb6eabca7b0f14598f61b31b505f70794a6ea275

Observation 98b04cc2-17fc-4d29-882b-6bf07547d5c8 · outbound

This paper cites You only hear once: a yolo-like algorithm for audio segmentation and sound event detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding You only hear once: a yolo-like algorithm for audio segmentation and sound event detection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.182668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:00.581690Z digest=sha256:8285263544cb1ab3b1c1e29a29499e17422b4194c62cc0fe4930e55e2be24ba8

Observation 7d18d4f0-4ac5-40a9-bb98-6fa3bfae5bf0 · outbound

This paper cites Temporal Action Proposal Generation with Transformers.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal Action Proposal Generation with Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.679239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.679239Z digest=sha256:9d460205a88ccfce84b79a1b5101bda41a493a777c3f45c80948477efaba0885

Observation f8281b68-9c52-4f7c-b270-0664967836d9 · outbound

This paper cites Rcl: Recurrent continuous localization for temporal action detec- tion.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Rcl: Recurrent continuous localization for temporal action detec- tion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.894285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:00.753442Z digest=sha256:b11d711b9257a4f7f5e4bb9bdf1a49ee344f36af256da580943d049f93822deb

Observation 782a3c0a-83ca-4af2-a2aa-a91be51477db · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.706511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:00.861896Z digest=sha256:29ce5957c43e7fd37533c395a083fa3ac78a989a16828b0d0ec6cf78061c46c9

Observation 75eb99e0-c396-48bb-8d86-7d12f6842739 · outbound

This paper cites An efficient spatio-temporal pyramid transformer for action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding An efficient spatio-temporal pyramid transformer for action detection

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.526823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:00.939127Z digest=sha256:e0ce70f034683fadd521e4ba773fde898b1221717456a717a201eb4a824dbb40

Observation e2ea7e57-8730-4fd1-b28f-78e1db40ad0e · outbound

This paper cites Dual attention matching for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dual attention matching for audio-visual event localization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.011214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.011214Z digest=sha256:ea3aca440047dfef4e567cc1ebf4e2ad5ce69fe2e393c05708989aff051d9748

Observation ac19a1c4-9689-4c5a-9d93-259124fe18b1 · outbound

This paper cites Dual relation network for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dual relation network for temporal action localization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.267675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:01.089484Z digest=sha256:5e8af7cb61643978126f74bc639ca3f997d881bc3e4c54b662d34a2a21a35a51

Observation 91a4d0c6-7b42-4adf-85ea-1b8476481198 · outbound

This paper cites Learning to refactor action and co-occurrence fea- tures for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Learning to refactor action and co-occurrence fea- tures for temporal action localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.930022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:01.205808Z digest=sha256:7812c3a739ae148eae81454a3606e1449e86d31d3772f63d82e4c96375eff51a

Observation eeeb6297-1c94-4403-b82a-165461958cee · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audiovisual SlowFast Networks for Video Recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.288755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.288755Z digest=sha256:c6ef277b2d6b708fd10c3ba73b932fa75a14d4ab72af96496c960eeac7c8fb3f

Observation f2d36044-0dc2-4d7f-bd5c-48f4b73bdce9 · outbound

This paper cites G-tad: Sub-graph localization for tempo- ral action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding G-tad: Sub-graph localization for tempo- ral action detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.649130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:01.409621Z digest=sha256:f2a5be9c64e77a27f539ce1aa3c19e4113a314cdd282f4b9b422b9f4d5ad3fd5

Observation ace9aa87-5799-4b7e-a854-2a1df9253e52 · outbound

This paper cites Audio-visual event localization by learning spatial and semantic co-attention.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual event localization by learning spatial and semantic co-attention

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.300311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:01.497261Z digest=sha256:e60b7c162c5c0fe3debe9c26b9260265a1136874f26f9184f14d3b1cac6755c3

Observation 98644f61-cfbb-4f39-b9fd-1f9d39f146d3 · outbound

This paper cites Temporal pyramid network for action recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal pyramid network for action recognition

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.087532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:01.572693Z digest=sha256:5b850c17915509409c3f987ebf7ab2d37a6e43668be1902ac844870a9b0f1020

Observation 3e4f63c9-fc4c-4de1-b7bc-31f74b47d100 · outbound

This paper cites Revisiting anchor mechanisms for temporal ac- tion localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Revisiting anchor mechanisms for temporal ac- tion localization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.961697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:01.667259Z digest=sha256:f8fd8cf71e58b0edb9fa130e41e5b18fe7092d727385117ceba1b1bfa04ecc27

Observation da56df42-9e95-4be0-ae0e-8dd3e2543d4c · outbound

This paper cites Mpn: Multimodal parallel network for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mpn: Multimodal parallel network for audio-visual event localization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.839549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:01.739029Z digest=sha256:73e5c208c0e6d151f69a74fcf580c66ceb8aebb083d28327542fea66f96e22bb

Observation eec95379-296b-495c-a05a-27b0704c6cfc · outbound

This paper cites Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.739506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:01.822512Z digest=sha256:9a4cc55ee8f99d5c651dae2277042f1b3763cd8579a9185db5d449589abd9fc1

Observation d9106c28-a30f-4748-a608-e214a731aa28 · outbound

This paper cites Graph con- volutional networks for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Graph con- volutional networks for temporal action localization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.628158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:01.887537Z digest=sha256:afb2dea29f4a19bb98e4efea7cdf8b72e0aac51482dbcfa58b5cb6ff9a505632

Observation ee01785a-809a-4097-9844-f5cb33925766 · outbound

This paper cites Actionformer: Lo- calizing moments of actions with transformers.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Actionformer: Lo- calizing moments of actions with transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.984274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.984274Z digest=sha256:3a762e1adc0131378aa0431edfa5f12319e63a2cdeefb12f9e5fe91f4c3b0065

Observation d8015066-5343-4a61-a764-7c223f5fe782 · outbound

This paper cites Video self- stitching graph network for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Video self- stitching graph network for temporal action localization

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.522826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:02.073232Z digest=sha256:1269dd24aca35ab31da8ae7ee1d1c6da295103a10ddac072a5f28f2adff695d6

Observation b311c305-b820-4278-8bce-a976fe78e285 · outbound

This paper cites Bottom-up temporal action localization with mutual regularization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bottom-up temporal action localization with mutual regularization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.412100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:02.139404Z digest=sha256:46d966f4bbbea1466b0c988042bbed1cbe13e1b669b5190a819005320244a02a

Observation 25fc7dc9-2928-4cbc-b6dd-d57993c85d7b · outbound

This paper cites Enriching local and global contexts for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Enriching local and global contexts for temporal action localization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.286231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:02.210024Z digest=sha256:c3946df76c0e18f78b468c9a5109cf9504b6f0885e03058cff19fefa52f21f43

Observation 9b8f3925-1ac2-4f58-af5d-74b122c37715 · outbound

This paper cites Our code provides further information.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Our code provides further information

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.160024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:02.296518Z digest=sha256:30363b8ce02c00c5b059e591b7b62d0bea7ed3877b0da2a97bcc383ad81cb477

Observation bbc44088-7e4e-420f-a1c0-02c4353f1649 · outbound

This paper cites Performance was measured us- ing mAP@[0.5:0.05:0.95], along with the average mAP.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Performance was measured us- ing mAP@[0.5:0.05:0.95], along with the average mAP

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.035271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:02.400314Z digest=sha256:05d10e69aacd47d0bd9db77106caf33a221fb2e9ebd38b8ac36cd15050ab9126

Observation cda74189-f55c-4927-a1f0-11c380a57b4c · outbound

This paper cites This approach ensures greater feature consis- tency across different temporal resolutions, ultimately im- proving the regression of an event’s location.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding This approach ensures greater feature consis- tency across different temporal resolutions, ultimately im- proving the regression of an event’s location

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:02.924236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:53:02.484752Z digest=sha256:29a4ca07013a684470fb677cf5c9ea0bb9890f3c886708328e86b886df10094a

Pith citing papers

Observation 47b14a0d-d100-4117-a088-aa437b827869 · inbound

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing cites this paper.

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:19.415238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:03:13.496634Z digest=sha256:e440dbe3cf850e90b8e86989ce36b0b345b9103d2289bb8a19b875ef9857c66a