Pith. sign in

Paper Citation Record · LEDGER

NowYouSee Me: Context-Aware Automatic Audio Description

As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2412.10002.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10002 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:31:32.064560Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:48:07.172477Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:20:18.426408Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eda46f5a-119d-481d-a301-582e08fd6b04 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

NowYouSee Me: Context-Aware Automatic Audio Description Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.782434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.782434Z digest=sha256:101637bcceeadd589d7f901411188c3963c0fdb4b813fce216993eb23672ac2c

Observation 3f3899a5-9941-45e9-9eeb-4e9bf0c616a5 · outbound

This paper cites Localizing mo- ments in video with natural language.

NowYouSee Me: Context-Aware Automatic Audio Description Localizing mo- ments in video with natural language

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.965621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.788137Z digest=sha256:f0c4853f3148f74f529c5b93fdd0717275d74b762f7e207871159786ebba3423

Observation c765cc33-2ac1-41f4-b558-c18c74d051ee · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

NowYouSee Me: Context-Aware Automatic Audio Description Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.794184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.794184Z digest=sha256:fe326600d7f400ae4caa2ec0714298ac6f0e1e830720923127dc718920e8a595

Observation 02f7c4d0-4544-4a94-a2ea-96c0c2e72f86 · outbound

This paper cites an unresolved cited work.

NowYouSee Me: Context-Aware Automatic Audio Description Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:31:32.934436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.799426Z digest=sha256:dbd7a63714d737870090937fe8726b78431c9d670f0aa3e4ac4f32f09c09a8d7

Observation 169d9801-b770-4451-9d7c-bab01b8e3258 · outbound

This paper cites Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning.

NowYouSee Me: Context-Aware Automatic Audio Description Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.915145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.804778Z digest=sha256:a3770ab3773b1555873db5247e3c1d669d500ff37c37df1df90247ddb39c8014

Observation 4c08fb09-3196-4d62-b6e1-5eb913f5dc34 · outbound

This paper cites Sketch, ground, and refine: Top-down dense video caption- ing.

NowYouSee Me: Context-Aware Automatic Audio Description Sketch, ground, and refine: Top-down dense video caption- ing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.810431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.810431Z digest=sha256:8c6acc1b2468c1dc1c3e8b1899dc722edf5b01df14fc988162d2562c558b8d16

Observation 5c14ec62-e9b2-40ac-954a-c5d715bb7523 · outbound

This paper cites Tall: Temporal activity localization via language query.

NowYouSee Me: Context-Aware Automatic Audio Description Tall: Temporal activity localization via language query

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.815949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.815949Z digest=sha256:75a83096c4345c8a2dc802f4df089d35c6b2baf248442217fa14ad557d2c1721

Observation d4c04edd-c36c-4fc0-9e54-697c9a73d88f · outbound

This paper cites A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation.

NowYouSee Me: Context-Aware Automatic Audio Description A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.821234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.821234Z digest=sha256:eb1b60d723240b1acd79d3aeb246f8e1bf7ec10df99cc05841aa4f28cfa6e8ea

Observation 3d3d1017-4228-4522-a3dc-5f8d6f44e4ab · outbound

This paper cites Hippo: Recurrent memory with optimal polynomial projections.

NowYouSee Me: Context-Aware Automatic Audio Description Hippo: Recurrent memory with optimal polynomial projections

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.873275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.828209Z digest=sha256:bdd071067cc518bf8ed7cb4c11109e7c64db41971f93aa26919ddbc750528534

Observation 18533dc9-b5e8-4699-aeb0-f9feef93be7f · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

NowYouSee Me: Context-Aware Automatic Audio Description Efficiently Modeling Long Sequences with Structured State Spaces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.833285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.833285Z digest=sha256:842c728c8566465e6d6bda078bba2b1801fa83077992a57097508f351dc0fda4

Observation d5af77a2-bcde-49a4-aa34-509141aa3ccc · outbound

This paper cites Efficiently mod- eling long sequences with structured state spaces.

NowYouSee Me: Context-Aware Automatic Audio Description Efficiently mod- eling long sequences with structured state spaces

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.838588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.838588Z digest=sha256:242fca83f779763c5a58126c4bf99d9bb098eca37d607980762e93edcf1c5835

Observation 73b22e65-67f6-4dc2-9a06-98e17e3e95c4 · outbound

This paper cites Combining recurrent, convolutional, and continuous-time models with linear state space layers.

NowYouSee Me: Context-Aware Automatic Audio Description Combining recurrent, convolutional, and continuous-time models with linear state space layers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.842644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.843985Z digest=sha256:c4baec33d9c717ee16032038feeee7e582368c965a24abd24d18974dd6bf6b00

Observation 537f0914-9eca-46ea-9c11-37ba63a7f9fc · outbound

This paper cites AutoAD II: The sequel-who, when, and what in movie audio description.

NowYouSee Me: Context-Aware Automatic Audio Description AutoAD II: The sequel-who, when, and what in movie audio description

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.824270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.848871Z digest=sha256:b8dea3ff00f7f855f60e5a0931053b7b6bb54f32e9f2bb042a62dde171ec2717

Observation a81da768-95c7-455f-8dbe-5f09ec4aa10b · outbound

This paper cites AutoAD: Movie description in context.

NowYouSee Me: Context-Aware Automatic Audio Description AutoAD: Movie description in context

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.806794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.854303Z digest=sha256:ce1dcb0610bef09140b6cbf33b43e41b18cbd1fd12a799bb6c709af407cfe8f5

Observation 11854838-f118-4b0d-9477-1c574aa14c5a · outbound

This paper cites Multi-modal dense video captioning.

NowYouSee Me: Context-Aware Automatic Audio Description Multi-modal dense video captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.787787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.861362Z digest=sha256:eb1586a1c310b1aeca1e9e1288fcdf550a1bc831f9cd637e76bb2e00dce22667

Observation ccb431ae-cd61-4084-afa5-dc628872ec3c · outbound

This paper cites an unresolved cited work.

NowYouSee Me: Context-Aware Automatic Audio Description Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:31:32.769276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.866469Z digest=sha256:20ae285d794f2c7cf1375285fd36a2fc42112e7675bb3b279f9e1061cb7c8dab

Observation a29f79fc-4049-4c9e-b54f-9617b3885c52 · outbound

This paper cites Long movie clip classification with state-space video models.

NowYouSee Me: Context-Aware Automatic Audio Description Long movie clip classification with state-space video models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.750694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.872080Z digest=sha256:5f97f9c713f9b12a6baea84d1b131c07939f5bc56906c89f5f117d968f330ac7

Observation 6e79a8ef-5187-48cc-9ace-0f170a6af2d1 · outbound

This paper cites Dense-captioning events in videos.

NowYouSee Me: Context-Aware Automatic Audio Description Dense-captioning events in videos

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.733025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.878506Z digest=sha256:60c57b2c3ea35d8b7cfaf7e28b6d07a0bbae513806f53287b70f8dd609e820f9

Observation 6e676614-bf02-465a-8561-b8e4911afd53 · outbound

This paper cites Beam search algorithms for multilabel learning.

NowYouSee Me: Context-Aware Automatic Audio Description Beam search algorithms for multilabel learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.713737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.883391Z digest=sha256:126901ce9908367c5a9889d7407c24e48f08fc34a199d5d46e70e31c13a29efe

Observation b5f4b709-bab1-4ced-b35c-cb1e89fa7a61 · outbound

This paper cites Video event detection by inferring temporal instance labels.

NowYouSee Me: Context-Aware Automatic Audio Description Video event detection by inferring temporal instance labels

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.696465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.888879Z digest=sha256:df3644ba985249b760bc35fbbc4bb4ea90b01a735ec9ac2c5375339d938e72c9

Observation fac09a50-22b5-450b-9a92-95513f3b1c0e · outbound

This paper cites Jointly localizing and describing events for dense video captioning.

NowYouSee Me: Context-Aware Automatic Audio Description Jointly localizing and describing events for dense video captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.894695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.894695Z digest=sha256:bfaaf70e5ad28708804d9ad72d3d2af46dc7ce06a057268f4c6312eb76335c03

Observation 3fcc5074-191c-42bb-bbb9-2ee7f6eb2692 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

NowYouSee Me: Context-Aware Automatic Audio Description Rouge: A package for automatic evaluation of summaries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.900399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.900399Z digest=sha256:42c3b7c79ae7e3ad2542d8faaced3e03441ab413611a09eb4b4170cdd5449d4c

Observation 47fb8825-74ae-4876-a2a0-c59280555ed7 · outbound

This paper cites SwinBERT: End-to-end transformers with sparse attention for video cap- tioning.

NowYouSee Me: Context-Aware Automatic Audio Description SwinBERT: End-to-end transformers with sparse attention for video cap- tioning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.655705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.906066Z digest=sha256:6cf41fb0d2078b134c6cd3630c6a15ecb530bb24ff2304ab57dac57f5e56f715

Observation da0889ed-c841-45bf-bc19-22a6464cc6aa · outbound

This paper cites Focal loss for dense object detection.

NowYouSee Me: Context-Aware Automatic Audio Description Focal loss for dense object detection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.911151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.911151Z digest=sha256:7f051ea3f6eb969b0adb2f3668e0888753d2ca44f818a112627eb39b3e922b08

Observation b72f0e64-bf75-4713-b4e7-49cc90a03896 · outbound

This paper cites Decoupled Weight Decay Regularization.

NowYouSee Me: Context-Aware Automatic Audio Description Decoupled Weight Decay Regularization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.917182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.917182Z digest=sha256:8b8dd598674910755cef3c9879881baf0c1a2b2295aaa8dc2a2376d673330202

Observation f62521b1-3deb-4b8b-ade5-5aaa826e5411 · outbound

This paper cites Event detection and anal- ysis from video streams.

NowYouSee Me: Context-Aware Automatic Audio Description Event detection and anal- ysis from video streams

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.624752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.922778Z digest=sha256:104b28e9afb87559e3087d8597a52cd8d86e2368237c4b8f582cdcc350ff08ba

Observation e04431e0-1ec3-45ed-9993-0b22c43e3bb5 · outbound

This paper cites Local- global video-text interactions for temporal grounding.

NowYouSee Me: Context-Aware Automatic Audio Description Local- global video-text interactions for temporal grounding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.606098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.928061Z digest=sha256:01af5920dfc33bf7920b4504d0080b7921446ae356316aac250042260b3c195e

Observation 50ba587f-e0d5-4c08-aeb7-af9ede21b4dc · outbound

This paper cites Streamlined dense video captioning.

NowYouSee Me: Context-Aware Automatic Audio Description Streamlined dense video captioning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.588579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.933669Z digest=sha256:0b27450811d57faab8b1b47e7124ded60c689fd8f37afdb24f23b20f1eb84c63

Observation b3275a07-561d-430d-84da-416197736ec0 · outbound

This paper cites S4nd: Modeling images and videos as multidimensional signals using state spaces.

NowYouSee Me: Context-Aware Automatic Audio Description S4nd: Modeling images and videos as multidimensional signals using state spaces

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.571010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.938867Z digest=sha256:9e28b7104f441f16f7705141391857b14b86d7deac7944ec9518a2b8723f692a

Observation b69d6272-838d-4dc3-a4cc-85d0eac3cd6e · outbound

This paper cites Queryd: A video dataset with high-quality text and audio narrations.

NowYouSee Me: Context-Aware Automatic Audio Description Queryd: A video dataset with high-quality text and audio narrations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.552422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.944035Z digest=sha256:82940c96fb5c45809f26b3e20d721b83567f206e1d6f0c694cb93feecd8d97ab

Observation 7cd6cdb1-6e75-44e0-bcc6-fe6c82c5f2be · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

NowYouSee Me: Context-Aware Automatic Audio Description Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.949734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.949734Z digest=sha256:87681288684f233b96f8b44c1a3b756dcd56774b3e70a9a5d4ac78bde49b67aa

Observation 3a334275-9cbb-4b8c-a9b0-3f66772102f5 · outbound

This paper cites Learning transferable visual models from natural language supervision.

NowYouSee Me: Context-Aware Automatic Audio Description Learning transferable visual models from natural language supervision

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.521971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.954497Z digest=sha256:557ea14d4312dce8e5cd38d4b6a9ed19132c21fa31eb4f04d0c0aa19e99cd922

Observation d2cfb3a8-c265-478e-a923-db0c5a0fa2d6 · outbound

This paper cites Language models are unsu- pervised multitask learners.

NowYouSee Me: Context-Aware Automatic Audio Description Language models are unsu- pervised multitask learners

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.959527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.959527Z digest=sha256:8e3cc2f2b324eb2d64aaa19a99236e7541031483f99193a56f08e145fd80980a

Observation 53d76604-1957-428d-9fd7-1ef7dd03657d · outbound

This paper cites Broaden Your Views for Self-Supervised Video Learning.

NowYouSee Me: Context-Aware Automatic Audio Description Broaden Your Views for Self-Supervised Video Learning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T16:31:32.133567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.964409Z digest=sha256:c0024d61121a41f620a3053540abe686ad1ff816f4c194fac796615b487cf246

Observation 81b73fe0-3e15-413f-a9c4-46cee8de74b8 · outbound

This paper cites A dataset for movie description.

NowYouSee Me: Context-Aware Automatic Audio Description A dataset for movie description

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.489683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.970324Z digest=sha256:c1f9f18c6c20c30d4debe3ef6037fdfa726ca4f6e7862945cf105f27e5e05948

Observation 9a9ad50c-5d15-49ad-a128-f696fadad2a6 · outbound

This paper cites Movie description.

NowYouSee Me: Context-Aware Automatic Audio Description Movie description

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:31.975611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:31.975611Z digest=sha256:43cf61cd1001580f76e5de891c679a2a7f62e9dfa819692ef0ccca7ac05cad67

Observation fb1e2c7b-b802-4388-aa21-b96ce06e5be1 · outbound

This paper cites Fine-grained audible video description.

NowYouSee Me: Context-Aware Automatic Audio Description Fine-grained audible video description

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.457423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.980889Z digest=sha256:fb451a6c47d742391acc29099d57485e8c05aa027d1d2da68a175f766ea889b6

Observation 927ea1f4-af12-4fcd-9360-f60aae6e43a9 · outbound

This paper cites TriDet: Temporal action detection with rela- tive boundary modeling.

NowYouSee Me: Context-Aware Automatic Audio Description TriDet: Temporal action detection with rela- tive boundary modeling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.440532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.987128Z digest=sha256:3a91cc8e2b39d8d25bf295bf02e48b14d6ce954e4e5a6ce5faa501a1a5f276df

Observation eb10aeaa-11ac-4118-851f-a9224cd8f932 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions.

NowYouSee Me: Context-Aware Automatic Audio Description Mad: A scalable dataset for language grounding in videos from movie audio descriptions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.422531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.992465Z digest=sha256:a1fdf113cef667194174638a194cce17dc9f8f41f182c104fafdd189ddca6500

Observation 486cf176-0beb-4fd9-8f09-33e227efbe79 · outbound

This paper cites You need to read again: Multi-granularity perception net- work for moment retrieval in videos.

NowYouSee Me: Context-Aware Automatic Audio Description You need to read again: Multi-granularity perception net- work for moment retrieval in videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.403804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:31.998203Z digest=sha256:c09a9a355956ded0bc452d699df04f711f1278b85c1e80942075f59065c170f1

Observation 485b0d9e-d9d9-4680-89da-ef280fb73b68 · outbound

This paper cites Long-form video-language pre- training with multimodal temporal contrastive learning.

NowYouSee Me: Context-Aware Automatic Audio Description Long-form video-language pre- training with multimodal temporal contrastive learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.385866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:32.003529Z digest=sha256:054aa0af428f4f04763557804ec6d350d8dd400d95b1fa72af8f63723290c3aa

Observation f0d367c3-6e80-40e9-adea-8c14e2281714 · outbound

This paper cites Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research.

NowYouSee Me: Context-Aware Automatic Audio Description Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:32.009315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:32.009315Z digest=sha256:db0a25b946d448f685be90686e5fbb162185f0689a2d594e0ea2fef0e55fd410

Observation ff990f01-f205-48cb-861f-edc65d84603d · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

NowYouSee Me: Context-Aware Automatic Audio Description Cider: Consensus-based image description evalua- tion

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:31:32.014750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:31:32.014750Z digest=sha256:269da2a696523f8b1bf318b492d6e7abc800adf3bdfcee889e7a25bebf89c348

Observation aa0ca21c-b335-4d6a-a058-8f9c9bf53c71 · outbound

This paper cites Long-short temporal contrastive learning of video transform- ers.

NowYouSee Me: Context-Aware Automatic Audio Description Long-short temporal contrastive learning of video transform- ers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.354715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:32.019666Z digest=sha256:3eb5fe05f3d5a52dc6af1bc6063aa06de905e777f9ad801803173f1751d5b0ad

Observation b5d1468b-48dd-4d9e-bf26-363b95cb3706 · outbound

This paper cites Bidirectional attentive fusion with context gating for dense video captioning.

NowYouSee Me: Context-Aware Automatic Audio Description Bidirectional attentive fusion with context gating for dense video captioning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.337202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:32.025835Z digest=sha256:4c5bcf9bf70a926580c036f6cd0490614d381b8c70e7469e7fbdec3fb2290b67

Observation df329e2e-6c3e-4fe9-9422-da8397a91ea6 · outbound

This paper cites Deformable video trans- former.

NowYouSee Me: Context-Aware Automatic Audio Description Deformable video trans- former

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.319802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:32.031676Z digest=sha256:93c2406a28a72cb0f8395dd7110ff2ce1a8a579e0522e31c1611f899d6df18cc

Observation f6f6fedd-f415-45b7-8e80-79a91501a290 · outbound

This paper cites Selective structured state-spaces for long-form video understanding.

NowYouSee Me: Context-Aware Automatic Audio Description Selective structured state-spaces for long-form video understanding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.300982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:32.036922Z digest=sha256:6aa0bb795b5dd561c03aa165640f8d87116289282ad8735e669d7351021e49ac

Observation f48afc03-2d05-48c8-9030-df35109607ef · outbound

This paper cites Event-centric hierarchical representation for dense video captioning.

NowYouSee Me: Context-Aware Automatic Audio Description Event-centric hierarchical representation for dense video captioning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.280879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:32.042244Z digest=sha256:5e43f2bf53385734186813e853e99a8c22a7a65d2c7d182febba2233bc84f070

Observation ec5e20e3-8229-4082-bda5-a8f555803553 · outbound

This paper cites Negative sample matters: A renaissance of met- ric learning for temporal grounding.

NowYouSee Me: Context-Aware Automatic Audio Description Negative sample matters: A renaissance of met- ric learning for temporal grounding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.262014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:32.047984Z digest=sha256:5d17da4f342653cdfbfa68a3c5eb5614f0ecb72e7211ec85369260d787c35f15

Observation 6901906a-6d1e-42f5-a402-b3917616a379 · outbound

This paper cites Towards Long- Form Video Understanding.

NowYouSee Me: Context-Aware Automatic Audio Description Towards Long- Form Video Understanding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.242680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:32.053690Z digest=sha256:9d12020d1af9331c1b44b4bd7698dcbd2cac89e4b9b6a392ba4ee8e8c187faed

Observation 60f6bb55-bf39-48c7-a287-f3ad340c48ec · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

NowYouSee Me: Context-Aware Automatic Audio Description Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.223303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:32.059518Z digest=sha256:ec2fe56807131569d1edfbf8f5f1f893ef007fdcb6bd9ae38984399c17106cc6

Observation d7435eb9-9278-4dc6-9b36-31098aae2845 · outbound

This paper cites Multi-modal interaction graph convolu- tional network for temporal language localization in videos.

NowYouSee Me: Context-Aware Automatic Audio Description Multi-modal interaction graph convolu- tional network for temporal language localization in videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:31:32.205302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:31:32.064560Z digest=sha256:c49f3bab6d334cb4ac444b561518c9a8b8b6cbff780b7b9a98c3c4f4767a6412

Pith citing papers

Observation 57a20458-e932-4b76-93c1-4b894ad55745 · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text NowYouSee Me: Context-Aware Automatic Audio Description

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.172477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.172477Z digest=sha256:62e451da45646d567fe571854df6106f87a02919c8be3006171f16b762019b09

Observation e616a5a0-59cd-4e20-a11e-9eba3d580b2b · inbound

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning cites this paper.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning NowYouSee Me: Context-Aware Automatic Audio Description

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:20:18.566942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:20:15.541739Z digest=sha256:679c456e41e77048660bc67548d046b1843492dcf267a4fff9a2c091f1b59912