Pith. sign in

Paper Citation Record · LEDGER

Towards Long Video Understanding via Fine-detailed Video Story Generation

As of 12 August 2026, this Paper Citation Record lists 100 of 108 outbound references and 0 inbound Pith citation observations for arXiv:2412.06182.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06182 v2

Coverage vector

measured 100 of 108 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:59:13.619204Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 108 outbound references displayed

  • verified exact0
  • verified fuzzy57
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b76aa785-a8f7-40b9-9d24-9b0847fe086a · outbound

This paper cites Intermediary-guided bidi- rectional spatial-temporal aggregation network for video-based visible- infrared person re-identification,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Intermediary-guided bidi- rectional spatial-temporal aggregation network for video-based visible- infrared person re-identification,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.299656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.299656Z digest=sha256:9c66a9e272a1315ed83340172d050a92a1edfefee9ed68c6f0ed775d43bd345e

Observation 03b84715-37a0-4f07-9597-baef3232868b · outbound

This paper cites Video moment re- trieval via comprehensive relation-aware network,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video moment re- trieval via comprehensive relation-aware network,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.304058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.304058Z digest=sha256:01166e962fdcab241d4e2dc06ad7f16d286a03e555c2c87c86780c258f04488a

Observation 3c3e072d-7723-4ee6-b14c-14fd0680d94e · outbound

This paper cites Self-supervised adversarial video summarizer with context latent sequence learning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Self-supervised adversarial video summarizer with context latent sequence learning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.308146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.308146Z digest=sha256:900e5dfc643bd20ede5170b14d0a8fe3203ef53e8e2d79c058b51f58c4dee295

Observation 2c88bb87-0543-4ad9-8de1-7925763f3a63 · outbound

This paper cites Complementarity- aware space learning for video-text retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Complementarity- aware space learning for video-text retrieval,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.312043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.312043Z digest=sha256:6f0bb37e17981444b2ff0884a951e9d7692f921a704e0d236c1bc63fa08fa59f

Observation 157f858f-1c4c-4644-9cf2-07453d74dfcb · outbound

This paper cites Moma-lrg: Language-refined graphs for multi-object multi-actor activity parsing,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Moma-lrg: Language-refined graphs for multi-object multi-actor activity parsing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.315626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.315626Z digest=sha256:b5d2ccf8ef9ad875d8b198a8e35a6855089bdd0effce5543f9a4f2384d34a6ca

Observation 3c5484a6-9cf2-4b62-8e17-7f1bf7c1b273 · outbound

This paper cites Videoclip: Contrastive pre- training for zero-shot video-text understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Videoclip: Contrastive pre- training for zero-shot video-text understanding,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.319397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.319397Z digest=sha256:a761f8a5404a42ecdc05320b347939c0d4b11dceb956816f3d380d79c28db276

Observation 6ddc86b7-04cf-4c20-a1b1-d33e8b7c07ea · outbound

This paper cites Graph convolutional module for temporal action localization in videos,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Graph convolutional module for temporal action localization in videos,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.323074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.323074Z digest=sha256:339fdd32eb285dcb6fb09344b7b1ac19e59a26293ac065749e36f7722355cee2

Observation 304fefc4-3b33-4afd-a4bb-1de82d73b17c · outbound

This paper cites Evcap: Element- aware video captioning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Evcap: Element- aware video captioning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.326324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.326324Z digest=sha256:92072d3f9c57ce8d2bec991dcdc293a7ae6282e1b541872f0d4dfa3f8f41d01d

Observation 246b13d7-365c-4a74-bc4c-c0679c3f3802 · outbound

This paper cites Multi-granularity interaction and integration network for video question answering,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Multi-granularity interaction and integration network for video question answering,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.329312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.329312Z digest=sha256:7add429f1c036243fbf19237739413783dad6a76c491887ccca825cc7e9a74f6

Observation 106cebde-0324-45a0-808f-ff7265d0a44d · outbound

This paper cites Video question answering with semantic disentanglement and reasoning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video question answering with semantic disentanglement and reasoning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.332350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.332350Z digest=sha256:3a260b2832f97700e9b7aa51c6920191d937b3c127c102f8af78779cc2b28e7b

Observation 548d2b2a-2e8e-473e-85f3-74925b113941 · outbound

This paper cites Multilevel semantic interaction alignment for video–text cross-modal retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Multilevel semantic interaction alignment for video–text cross-modal retrieval,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.335894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.335894Z digest=sha256:5151e8d21ddf5d4bf6c8ce0deb8aee33e4e453cbcf1c521b64e45549fc38f24e

Observation 37202589-7b99-4879-8fac-b0d933076c10 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Towards Long Video Understanding via Fine-detailed Video Story Generation VideoChat: Chat-Centric Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.339068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.339068Z digest=sha256:3d7984240e55c8ca5e1bdac3c9038a60bb26c4a027e59f55ebb2ef9ad586476f

Observation bc909dc9-ba3a-4750-bbe9-39769e529306 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video-llama: An instruction-tuned audio-visual language model for video understanding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.342311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.342311Z digest=sha256:9161bb43683657b82aa474275c5103787d6ef552d8a8e33ce08166856f60cb6f

Observation f77b4e5e-7d18-4849-a028-b1c48563771f · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Towards Long Video Understanding via Fine-detailed Video Story Generation MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.345138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.345138Z digest=sha256:b3f7047a71886affde990659f9683862a1f3132850275ba38ace46ba64494aef

Observation 831d8942-a346-4f40-9a3b-25cb57a1004c · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.348682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.348682Z digest=sha256:4cbc6e29a2bc1a78887cffbf5adeb3247d393b94f73fed0aa7f3f3e1d208bf24

Observation 89fee407-103d-46cb-831e-8e19fb6d605f · outbound

This paper cites Language models with image descriptors are strong few-shot video-language learners,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Language models with image descriptors are strong few-shot video-language learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.352109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.352109Z digest=sha256:32051816384db48322998dc0460da8f3af6693cf87554f336a6e88f59ee52611

Observation 68e75ac0-42de-4cbc-8e57-69a8d3917f5a · outbound

This paper cites Compressed video action recognition with dual-stream and dual-modal transformer,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Compressed video action recognition with dual-stream and dual-modal transformer,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.355176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.355176Z digest=sha256:91c726db862bf606c23130e8f2ac8d89d4b2958b6af616879d10f3ddc444156f

Observation ccab0104-9ecf-4506-8be1-da3f20434797 · outbound

This paper cites Dynamic spatial focus for efficient compressed video action recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Dynamic spatial focus for efficient compressed video action recognition,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.358243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.358243Z digest=sha256:e231fd61a573ad1a87af71b7bb304af19ba8fb8796f792826227c37d09802ec0

Observation 5f1d5e46-a24d-44ee-88fa-eee4f7a9d064 · outbound

This paper cites Alignment-guided temporal atten- tion for video action recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Alignment-guided temporal atten- tion for video action recognition,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.361520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.361520Z digest=sha256:91f23fdd24ccbe9f6cb313533aae2f09840a17ed0e285620f7f95ba509b50609

Observation 787aab86-06d4-43a0-befd-c74f18c3aa3b · outbound

This paper cites Slowfast networks for video recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Slowfast networks for video recognition,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.364500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.364500Z digest=sha256:59f10abe61a45fe2bc2378378d4beef1bd595d1cc6b1fbc329a8321b7e18fd56

Observation 2071a66f-0e0e-4850-9335-33bc33a806d5 · outbound

This paper cites Temporal distinct representation learning for action recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Temporal distinct representation learning for action recognition,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.367934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.367934Z digest=sha256:98ffabd36caf955199db3178ce424707abd699f82210c6ce8efe566c3ce94dcd

Observation 3d99e65b-f375-4487-aafe-5d7a2e96741b · outbound

This paper cites Truncate-split-contrast: a framework for learning from mislabeled videos,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Truncate-split-contrast: a framework for learning from mislabeled videos,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.371085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.371085Z digest=sha256:834120784dd342930193b6e72f72874d45a62e1053b0e75ad93f54fe27fbc356

Observation de6591e7-2ab0-41c1-a586-ed130cdf71c1 · outbound

This paper cites Reading-strategy inspired visual representation learning for text-to- video retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Reading-strategy inspired visual representation learning for text-to- video retrieval,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.374125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.374125Z digest=sha256:8c149ef6e8b630fa1a3d9a0a08a763e39d78dcc6722d7d085b9945d60942fa94

Observation 903fe238-f9d2-46cb-89ec-7c40649e1aff · outbound

This paper cites Use what you have: Video retrieval using representations from collaborative experts,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Use what you have: Video retrieval using representations from collaborative experts,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.377335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.377335Z digest=sha256:96074f8a9c9d55048811c4d4d1fe67f9025cc7dec8bdcb9e3c3884577ac70252

Observation 881de220-95ef-413a-8ee8-773a53b9718a · outbound

This paper cites Dual encoding for zero-example video retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Dual encoding for zero-example video retrieval,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.380823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.380823Z digest=sha256:45e5dbfd9fdf3d2d8e1314a51504ef9d0f08aebe51ec3e01d5bb886b6fbe97c4

Observation 61a2a580-296b-4df4-bf20-aff3036720a2 · outbound

This paper cites Locvtp: Video-text pre-training for temporal localization,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Locvtp: Video-text pre-training for temporal localization,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.383984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.383984Z digest=sha256:d1b68f8be3bb1de98643a6eb3b0575230889fbfa3cec04e819336e5b719a303c

Observation adad5951-16a8-4bfb-83c9-20e48977a0a0 · outbound

This paper cites Breaking winner-takes-all: Iterative-winners-out networks for weakly supervised temporal action localization,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Breaking winner-takes-all: Iterative-winners-out networks for weakly supervised temporal action localization,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.387374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.387374Z digest=sha256:2db25e16b1974540aa61a5f717b8bfa749aa411a08a96b47158c98a0cd582e85

Observation 8c562c82-db62-4be2-bd76-c3c12052f130 · outbound

This paper cites Cross time-frequency transformer for temporal action localization,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Cross time-frequency transformer for temporal action localization,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.390767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.390767Z digest=sha256:5bf6c05f9ec72330552e67560ae5fa3a189836a439a145869d52660256bd0022

Observation 1740bebf-8852-4dcf-87ce-0ac1d1bf3b23 · outbound

This paper cites Slow motion matters: A slow motion enhanced network for weakly supervised temporal action localization,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Slow motion matters: A slow motion enhanced network for weakly supervised temporal action localization,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.394100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.394100Z digest=sha256:f4e2237c82aef59b5229d6a037cfece40948af1ef5cb49c948ecc82db7f72091

Observation ce42a0f6-b6d2-4169-a8d3-cbee63e055e4 · outbound

This paper cites Long-form video- language pre-training with multimodal temporal contrastive learning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Long-form video- language pre-training with multimodal temporal contrastive learning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.397280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.397280Z digest=sha256:40049262b8394697f7cc8fecf00ebd193f2249265ab3648a0304175a5208bad8

Observation 4320506b-7c55-417b-ad3d-d9a5b657c202 · outbound

This paper cites VideoGraph: Recognizing Minutes-Long Human Activities in Videos.

Towards Long Video Understanding via Fine-detailed Video Story Generation VideoGraph: Recognizing Minutes-Long Human Activities in Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.400389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.400389Z digest=sha256:96416521ada50752a5a634bd3bfe0eb1ee1e93ddfa45adaa4271cb498fc5edfd

Observation be66552c-aac9-465d-a353-7201b08941d3 · outbound

This paper cites Supervoxel attention graphs for long-range video modeling,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Supervoxel attention graphs for long-range video modeling,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.404329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.404329Z digest=sha256:bac8d73570bbcd08adcbc6aeaf2b18aca9459243213ce52c1e40f7e709c464f1

Observation 0c772c00-a4df-49fe-a93a-02cc99cf828f · outbound

This paper cites Long movie clip classification with state-space video models,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Long movie clip classification with state-space video models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.539583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.407515Z digest=sha256:180b7b203ba44c0d76289432764e9d5bbb7ff8d4d63d581069a9f171fc4daa9e

Observation be7e0557-c7f6-41d0-93a9-e7b5db8f2549 · outbound

This paper cites S4nd: Modeling images and videos as multidimensional signals with state spaces,.

Towards Long Video Understanding via Fine-detailed Video Story Generation S4nd: Modeling images and videos as multidimensional signals with state spaces,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.528911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.410492Z digest=sha256:1e346497cd03d6c9cf4032a013307c6e230e76299d29f0d2810efb8a81c140fe

Observation 6a883042-e723-4ce8-8dba-bfb316fce0f6 · outbound

This paper cites Selective structured state-spaces for long-form video understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Selective structured state-spaces for long-form video understanding,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.518072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.413725Z digest=sha256:7ab33f3041cbab7a4b17b4d879401a4bf2ba6f74e3f8e9c8049df8bc808c63e7

Observation 3c55eb5c-84ae-440d-ba0d-78efe043e69b · outbound

This paper cites Efficiently modeling long sequences with structured state spaces,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Efficiently modeling long sequences with structured state spaces,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.507046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.416851Z digest=sha256:cf2638df8e6aa5d4390d2a3fea4abf3edabe427d1bb00308c2c63c8c7ce30589

Observation b707d111-080e-4719-841a-d0e10ae281ac · outbound

This paper cites Mgsampler: An explainable sampling strategy for video action recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Mgsampler: An explainable sampling strategy for video action recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.495316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.419845Z digest=sha256:80a12358bc65d3c6b164baa783e0a83e5908eb62e0bdb8f4751ed002bf41a62e

Observation ec1f12c2-12fd-4792-8e34-5910efa98bbd · outbound

This paper cites Adaframe: Adaptive frame selection for fast video recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Adaframe: Adaptive frame selection for fast video recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.484590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.422737Z digest=sha256:3c4c903453928640b4e18e8bd208fc96ce1c25920ffe89acb3a75d56bb5ce0a2

Observation 7e4b57c6-eef4-4842-883e-935ad5c79ac4 · outbound

This paper cites Localizing moments in long video via multimodal guidance,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Localizing moments in long video via multimodal guidance,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.474385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.425583Z digest=sha256:bea52de8f6bd709d18e947aa5123c77059256cdea944df5625d73dac7b7f70f5

Observation 00a45254-eb61-4d1d-bec5-d4f60dcc6e25 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Mad: A scalable dataset for language grounding in videos from movie audio descriptions,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.463310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.428711Z digest=sha256:6967ffe579025474dd5bea7621c1c012f031e735780a8ce857c0f1141cdc1bad

Observation 97ae6fcd-5ebe-4097-9584-57e18ee9b50e · outbound

This paper cites End-to-end learning of visual representations from uncurated instructional videos,.

Towards Long Video Understanding via Fine-detailed Video Story Generation End-to-end learning of visual representations from uncurated instructional videos,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.452622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.431718Z digest=sha256:f43f520c13c28a759c4d0b0904b94080c5e392d165644491484f29aa2f4d9272

Observation bf90c54e-8f4c-4c1e-81b9-45c03751ff83 · outbound

This paper cites Merlot: Multimodal neural script knowledge models,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Merlot: Multimodal neural script knowledge models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.441521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.434656Z digest=sha256:a28d3cac79eb01a4ef49faec7364a819443758689eef6c18374d5de496cbeaa4

Observation 8a745a5d-ad95-4236-9c48-2f872ed5b492 · outbound

This paper cites Scaling up vision-language pre-training for image captioning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Scaling up vision-language pre-training for image captioning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.431524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.437890Z digest=sha256:f8a18b364d0481f0735ea7efb28ed64847ad0e4437ec600bfc27bb04633ef317

Observation 6ded8614-ee68-4727-827a-9f20bc76284e · outbound

This paper cites Gbc: Guided alignment and adaptive boosting clip bridging vision and language for robust action recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Gbc: Guided alignment and adaptive boosting clip bridging vision and language for robust action recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.420598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.440965Z digest=sha256:cb259db9257424abd2b9481765951dc7eae393625ce712707e5bfb4c3365ad7c

Observation 8170842f-9f6f-46c9-b9a9-151258bcb288 · outbound

This paper cites Unsu- pervised pre-training for temporal action localization tasks,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Unsu- pervised pre-training for temporal action localization tasks,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.409107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.444160Z digest=sha256:c482ed0ab7530cdec6390aa8979cd3aa65c3596d85f06c91057e3c60fe6cff80

Observation 548090d7-eebb-48e5-8077-948385896a68 · outbound

This paper cites Clip4clip: An empirical study of CLIP for end to end video clip retrieval and captioning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Clip4clip: An empirical study of CLIP for end to end video clip retrieval and captioning,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.397729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.447486Z digest=sha256:122cd2e8cb475e0094d881c645711a9d610465f3d5e4fcc2c9a4130f5a412749

Observation 7d6a0576-d691-432a-9b02-92d48310de48 · outbound

This paper cites Human action recognition and prediction: A survey,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Human action recognition and prediction: A survey,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.450578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.450578Z digest=sha256:6ee9b3ffab84e59e48bdbf90c690ba72cdeb4da8258f6933a24eeef83338925d

Observation 986671f5-f4f9-4321-aa71-3a98f247ffa4 · outbound

This paper cites A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,.

Towards Long Video Understanding via Fine-detailed Video Story Generation A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.379512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.453369Z digest=sha256:fa6e19e29b65b4a32745a79a71c3baeb0e034fd3dc5805209274406c2c468d11

Observation 2081cbe2-ef2e-4b4b-80e2-26cf5f84a614 · outbound

This paper cites Lavender: Unifying video-language understanding as masked lan- guage modeling,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Lavender: Unifying video-language understanding as masked lan- guage modeling,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.368806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.456277Z digest=sha256:aea5134adcbcbe44ddc8765040153c78e1f6acec34889188667337648c954b62

Observation be59d67c-e8a0-4cc3-b078-0cf00140ee16 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.358415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.459592Z digest=sha256:49f185690f8d5708794504beeabb3d5ce16dbcc3583b4f1330c89a29394cb4b4

Observation a6411d7e-81c2-46b2-9290-6155fde763eb · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Videomae v2: Scaling video masked autoencoders with dual masking,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.347902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.463104Z digest=sha256:55a84cf9d21ff7603d4c6fe4f409c828cc722fc86e3c3c8d8568035120421822

Observation a7316983-e468-4d51-9030-b1592160bf1b · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

Towards Long Video Understanding via Fine-detailed Video Story Generation VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.466341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.466341Z digest=sha256:3c2ddcc24bcb0ffff3c5ebc4e91bdce8bdfbf1e591afa985083cce028e028837

Observation ba31b6c5-0eb5-4468-a395-30e98376d61e · outbound

This paper cites All in one: Exploring unified video-language pre-training,.

Towards Long Video Understanding via Fine-detailed Video Story Generation All in one: Exploring unified video-language pre-training,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.336103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.470172Z digest=sha256:d874ed040f4e98ef3f71dea3c104c779c2a6667344f9f829509849f240a3101a

Observation 69b91bbb-8d7a-4fc6-9e55-b0d424f983e9 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Towards Long Video Understanding via Fine-detailed Video Story Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.473248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.473248Z digest=sha256:e8b6c2f3e74c8d0a0d86fa8a00ea04ea1b80114b98f8dcab7e3e652dde493dad

Observation 03d1f286-a7fd-485c-9b5e-b3d8dbeb4c64 · outbound

This paper cites Language mod- els are few-shot learners,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Language mod- els are few-shot learners,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.325979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.476828Z digest=sha256:b30213d1dae3f4c82e38d6e0dca9fbecf0f79c16f4d874f87d10d110de85ed8c

Observation 98458942-a666-484f-8f2d-3d651a394c1b · outbound

This paper cites GLM: general language model pretraining with autoregressive blank infilling,.

Towards Long Video Understanding via Fine-detailed Video Story Generation GLM: general language model pretraining with autoregressive blank infilling,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.315557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.480137Z digest=sha256:aac027e3758c8fc9cd510960fa526d71ade8425bcec7691c9eda0cf6a49973a9

Observation c8db19b9-f461-438d-9453-3f5b81894773 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Towards Long Video Understanding via Fine-detailed Video Story Generation LLaMA: Open and Efficient Foundation Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.483476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.483476Z digest=sha256:07eb59c17f321f20e1fda5443a8c40b57c5ec3d3d2ee2d02335091fd88361954

Observation b41f2851-7240-4be8-bf5a-73bca0b77d97 · outbound

This paper cites MIST : Multi-modal iterative spatial-temporal transformer for long-form video question answering,.

Towards Long Video Understanding via Fine-detailed Video Story Generation MIST : Multi-modal iterative spatial-temporal transformer for long-form video question answering,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.303595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.487191Z digest=sha256:c66d11fff69da19a6f7624e3ffe5cdd5b6f14cec950d00acdf820a28082fa213

Observation cd159143-971a-4beb-bada-fac3eb9a60c5 · outbound

This paper cites A joint sequence fusion model for video question answering and retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation A joint sequence fusion model for video question answering and retrieval,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.291561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.490546Z digest=sha256:27a7352b8996f5ee3f93c9b713de33a07baa7a17acd4199577690d300bb8ae55

Observation d2387de6-656c-40f1-ab0f-432ba582b063 · outbound

This paper cites Video question answering via gradually refined attention over appear- ance and motion,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video question answering via gradually refined attention over appear- ance and motion,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.279595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.493599Z digest=sha256:7f005f568671fb7c8f67b34536894b3f8dcafefca03fe707f4397e80797747ab

Observation bff493ec-b099-46b2-b64f-2e3d8539f224 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Activitynet-qa: A dataset for understanding complex web videos via question answering,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.268566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.496737Z digest=sha256:e4602fb8e201ef2002907fbc284b9b4b5c82a7aeade843e37c8cee3f8512be62

Observation 522cb6ed-6d9f-4a78-8400-76f8bf6ab18e · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Swinbert: End-to-end transformers with sparse attention for video captioning,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.256936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.499776Z digest=sha256:17fbc27dd807cda86c3b37e05d0a4aa6d2e51b10e653e14731d117438a33e989

Observation 08ad18dd-0240-4738-ae53-f7bf796daec9 · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.245299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.502597Z digest=sha256:29630cd0bb505451ae0f93b22b39589d05242d3210ce05c49c02901b6462c488

Observation 7cadfdc3-52c4-4d92-8868-78f9cb26701b · outbound

This paper cites Dense- captioning events in videos,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Dense- captioning events in videos,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.233671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.506128Z digest=sha256:a94b3f8ae2fc8008829889b5d4f31beefc5aaa478bcd36fed63f242c9c9f05a5

Observation df2e5754-c242-4eb4-9377-621b106bdf0f · outbound

This paper cites BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Towards Long Video Understanding via Fine-detailed Video Story Generation BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.221672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.509171Z digest=sha256:18ccfa382a7ce17dafea8c6e5dd2adeda06bbb78df5af487f2536a3ac00e4149

Observation 0f9707f9-ad5b-47f5-9be7-fc6146ad7ac3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Towards Long Video Understanding via Fine-detailed Video Story Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.512195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.512195Z digest=sha256:3687d0d90d2e29906ab1715cca720c609d93998876abb3f4c9298f4118f5e25f

Observation 447ac7d9-1ccd-48aa-8c0b-d5ab82159b1f · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Towards Long Video Understanding via Fine-detailed Video Story Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.516220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.516220Z digest=sha256:20731cfea42b959aa87fcf539035334a50734902372a273484061090659db8fc

Observation ebac29eb-40df-406a-84f7-88fdeb1499aa · outbound

This paper cites Pyscenedetect,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Pyscenedetect,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.209964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.519600Z digest=sha256:9e5754d4ce02dad78d47d0c29fa3b8cdad0bd033cff180b01ddc2ccd37f56d1e

Observation 4b1ddd92-00f5-44b7-9f41-36cc95970907 · outbound

This paper cites Decord: An efficient video loader for deep learning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Decord: An efficient video loader for deep learning,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.197450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.522639Z digest=sha256:5d33ad009ab2c13f3cad4c4725364a1c0eb4f2d738f4af2683d27f4310166120

Observation f11bf45d-6f12-488f-b4fd-81bd7de680d8 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Learning transferable visual models from natural language supervision,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.525872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.525872Z digest=sha256:edb9d7a1190a3fd40603d417bae0f4a1e8ac9f12e4dfdfe0199b5464a35344d6

Observation d5849239-8def-49a6-a4c3-cd867a0ed0b8 · outbound

This paper cites DINO: DETR with improved denoising anchor boxes for end-to-end object detection,.

Towards Long Video Understanding via Fine-detailed Video Story Generation DINO: DETR with improved denoising anchor boxes for end-to-end object detection,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.179339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.529027Z digest=sha256:796fb1e358545320b597e654467ca807456b5b63ab1d1f7d2b24ec316ae3317f

Observation ee5186e6-8829-4917-80a0-799e771344e1 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Sentence-bert: Sentence embeddings using siamese bert-networks,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.168721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.532056Z digest=sha256:7ab3c1cfd1eb26c53fda60873102e0db7503cb0488c281b994a2fb2f1a7a9795

Observation 11655f67-fe3d-4ae2-a455-a70c149ca17b · outbound

This paper cites Partially relevant video retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Partially relevant video retrieval,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.158363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.535135Z digest=sha256:c601e9b5ba4ca8e37f844e747f54e525c26082455c3af106076bebe30ace8308

Observation 5a466a5f-d901-4c80-88a4-acaad7416c99 · outbound

This paper cites Joint searching and grounding: Multi-granularity video content retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Joint searching and grounding: Multi-granularity video content retrieval,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.147009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.538757Z digest=sha256:2829e8a8a671b0ffe9d86c6cde07c66dd5f772187584355e4742abc59bf1eeec

Observation bef96942-deb3-4b17-ad8e-b537412a17ba · outbound

This paper cites Hit: Hierar- chical transformer with momentum contrast for video-text retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Hit: Hierar- chical transformer with momentum contrast for video-text retrieval,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.136686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.542232Z digest=sha256:2f3f72699abd20b6d1767b145a41595a2c2095cd741f2acd25fa463e3bcadf45

Observation 619faa62-256d-4b58-b41f-2d23bc20fbff · outbound

This paper cites Multi-modal trans- former for video retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Multi-modal trans- former for video retrieval,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.126362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.545156Z digest=sha256:9297720e8261dffc0e90337cb8df435f46b587e72716f951b3743dbca7938624

Observation 818e37cf-5c0d-4737-b57a-81c6e1766cc0 · outbound

This paper cites Cross-modal and hierarchical modeling of video and text,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Cross-modal and hierarchical modeling of video and text,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.115696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.548293Z digest=sha256:9053e51e98cb7ffd004697975cc49e9274817f623664346810daf7ff4533bfd3

Observation 05a59bbb-5df8-486d-91ba-046535002610 · outbound

This paper cites Eclipse: Efficient long- range video retrieval using sight and sound,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Eclipse: Efficient long- range video retrieval using sight and sound,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.105621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.551159Z digest=sha256:3715f9ea065015d66113193ddfe1f6afa39b0ed9335c588f44ea92573169de6e

Observation 7750a39d-77a5-481a-9aa4-1712f209efa4 · outbound

This paper cites TALL: temporal activity localization via language query,.

Towards Long Video Understanding via Fine-detailed Video Story Generation TALL: temporal activity localization via language query,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.095823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.554216Z digest=sha256:d9647fe134c08e9767f336fcec52762655dc9ffa3cbcc1e1109a98242c8844ca

Observation f586d4ae-e6cc-47e5-8c52-ee0c3c2a2acb · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Hollywood in homes: Crowdsourcing data collection for activity understanding,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.085989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.557624Z digest=sha256:1b3fe9f05244f3b3b6bf92fa554cc0db66b6ff7e7d2554c98d8244056d0e606f

Observation 2c271b89-00d5-48d8-ad34-618c428d11a8 · outbound

This paper cites MSR-VTT: A large video descrip- tion dataset for bridging video and language,.

Towards Long Video Understanding via Fine-detailed Video Story Generation MSR-VTT: A large video descrip- tion dataset for bridging video and language,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.076134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.560460Z digest=sha256:5ea8e908915417c6c478df811f45d72b8d3e236a575bfcc670ca8c4d42b4110b

Observation f42f821b-3b50-472a-875e-c61b44477382 · outbound

This paper cites Question generation via overgenerating transformations and ranking,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Question generation via overgenerating transformations and ranking,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.065916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.563116Z digest=sha256:f0638b0dd661493f7a51ac6314efffbe0df79d17d5914991b45b3fc6d3660a14

Observation b12f594c-664e-4e6a-b75c-8f8707ec5244 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Egoschema: A diagnostic benchmark for very long-form video language understanding,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.056879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.566380Z digest=sha256:9515301fe5538335a1ca4c32b1d53b97ebea9eae77713b4174e6e7a16552335b

Observation 4a7c3694-66ff-4de1-bb2f-002d1f1b1d72 · outbound

This paper cites Ego4d: Around the world in 3, 000 hours of egocentric video,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Ego4d: Around the world in 3, 000 hours of egocentric video,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.047172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.569192Z digest=sha256:9e445314434de04e5618550e8742c89b4c4ec6ace6304a0863c64b771711f370

Observation 884ac27b-36fd-4ab8-9bcd-0063a955b8f8 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.036272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.572112Z digest=sha256:33c6f8b3ce8f487874f2293e18fbf950ae108b61254684924a415f6cce5314b4

Observation 22780f74-4a8e-43b9-8b2e-19825e63da4e · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.026777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.574991Z digest=sha256:df3f51437844187dcaacd53f82fc1e051870d804777249d404638d705041ba8f

Observation de50226d-8903-4b14-bd41-818fd009af55 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Towards Long Video Understanding via Fine-detailed Video Story Generation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.578033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.578033Z digest=sha256:8cae2316198bea227f329d5a0b70d994044bf775b8e57d7af521f8b589a4afd5

Observation 3b77bdc2-774f-4302-b3fb-3a29b29d2bd5 · outbound

This paper cites Sharegpt: Share your wildest chatgpt conversations with one click,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Sharegpt: Share your wildest chatgpt conversations with one click,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.017304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.581465Z digest=sha256:02c944beaf41413ec0511a6bec01a0537e380269c9ace330c477a99022ca6a48

Observation 4e65467a-9dc8-4212-b7e0-f05e6b940993 · outbound

This paper cites AnglE-optimized Text Embeddings.

Towards Long Video Understanding via Fine-detailed Video Story Generation AnglE-optimized Text Embeddings

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.585145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.585145Z digest=sha256:dddabe7b11557cb2522a52ad304ffd75fd1f7c5d30b064acc9d24aa32a721a54

Observation 7be46fc4-02ac-4d22-af5d-02c447adc113 · outbound

This paper cites Video corpus moment retrieval with contrastive learning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video corpus moment retrieval with contrastive learning,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.007278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.588390Z digest=sha256:1331e42f0329ee2fb862be2f8ce045edd156d2eb5f1d7f5e954dceb65353d1a4

Observation 8875e9a0-e7ed-4f2b-9067-a3ebc767d51c · outbound

This paper cites TVR: A large-scale dataset for video-subtitle moment retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation TVR: A large-scale dataset for video-subtitle moment retrieval,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.997127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.591759Z digest=sha256:8c1d032c27fa2576ca113c2b56cd478bfe4f8ebf870ded1cd6c06a9d27e66316

Observation 3aeba3ae-12b5-4ea8-96d3-233932ff55b5 · outbound

This paper cites Dual learning with dynamic knowledge distillation for partially relevant video retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Dual learning with dynamic knowledge distillation for partially relevant video retrieval,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.986868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.594774Z digest=sha256:de327562a0d3f53def4c92f55bc4e4307f57c3295d573d70e13523f9ac34f0ff

Observation f92a385a-4186-4f2f-af3c-81ae0c0f3a94 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.597852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.597852Z digest=sha256:4b565df7a9b9d8b6b328686005863a45f53b3fc53b43dbb15753f4d61ec0e5aa

Observation ab79c81a-42d1-4f97-aa46-be999033b701 · outbound

This paper cites Just ask: Learning to answer questions from millions of narrated videos,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Just ask: Learning to answer questions from millions of narrated videos,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.976799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.601448Z digest=sha256:7a1d29a078bdd9bf41a52366dbbe3d3882f977e220ad669e517991b976d7f385

Observation 2528ebbc-87f5-4143-8735-33589eb2ea56 · outbound

This paper cites MERLOT RESERVE: neural script knowledge through vision and language and sound,.

Towards Long Video Understanding via Fine-detailed Video Story Generation MERLOT RESERVE: neural script knowledge through vision and language and sound,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.965769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.604459Z digest=sha256:2c4efe7468d8b4f6654541122ba0960b74530038b061824c58c2b9dc586e10a4

Observation cf7169bb-7386-4afe-a9d6-dd790f1767c7 · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Zero-shot video question answering via frozen bidirectional language models,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.955780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.607548Z digest=sha256:396e98d1f83c157d4fe4ff0d213981bf58993194b0bf2b91ac541862fa0d7020

Observation 921d8ba2-5ff9-4525-8263-44a3a5f22613 · outbound

This paper cites Hitea: Hierarchical temporal-aware video-language pre-training,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Hitea: Hierarchical temporal-aware video-language pre-training,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.944568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.610461Z digest=sha256:914cb392727f766de7726bf1b263836ad532f478d8f727564b1b80d426445235

Observation 0e4515bd-6b0e-47f2-80a4-eff5f7432f4b · outbound

This paper cites Self-chained image-language model for video localization and question answering,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Self-chained image-language model for video localization and question answering,

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.933048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.613053Z digest=sha256:45768ffa47728978ee42194fe1f3d28ab7562c6e164a3ece75a38e71456b55d8

Observation 22431014-9daa-4649-902f-4214ea607eb1 · outbound

This paper cites Memory consolidation enables long-context video understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Memory consolidation enables long-context video understanding,

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.922071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:59:13.616348Z digest=sha256:d7aefbb27de0bc2ceb6d085bfac7adbafc3e7aa7ecbb5d8c74380f3c745c11a7

Observation 8f5f7710-bbd0-4204-ab66-f4012eaf31af · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Towards Long Video Understanding via Fine-detailed Video Story Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.619204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.619204Z digest=sha256:a940430ff255bbdf135ddd04c67d8aaae0f922dda36a9bbdbbc91bd609e34c57

Pith citing papers

No inbound Pith citation observations are available.