Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:59:13.619204Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 100 of 108 outbound references and 0 inbound Pith citation observations for arXiv:2412.06182.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:59:13.619204Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 108 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b76aa785-a8f7-40b9-9d24-9b0847fe086a · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Intermediary-guided bidi- rectional spatial-temporal aggregation network for video-based visible- infrared person re-identification,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03b84715-37a0-4f07-9597-baef3232868b · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Video moment re- trieval via comprehensive relation-aware network,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c3e072d-7723-4ee6-b14c-14fd0680d94e · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Self-supervised adversarial video summarizer with context latent sequence learning,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c88bb87-0543-4ad9-8de1-7925763f3a63 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Complementarity- aware space learning for video-text retrieval,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 157f858f-1c4c-4644-9cf2-07453d74dfcb · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Moma-lrg: Language-refined graphs for multi-object multi-actor activity parsing,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c5484a6-9cf2-4b62-8e17-7f1bf7c1b273 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Videoclip: Contrastive pre- training for zero-shot video-text understanding,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ddc86b7-04cf-4c20-a1b1-d33e8b7c07ea · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Graph convolutional module for temporal action localization in videos,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 304fefc4-3b33-4afd-a4bb-1de82d73b17c · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Evcap: Element- aware video captioning,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246b13d7-365c-4a74-bc4c-c0679c3f3802 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Multi-granularity interaction and integration network for video question answering,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106cebde-0324-45a0-808f-ff7265d0a44d · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Video question answering with semantic disentanglement and reasoning,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 548d2b2a-2e8e-473e-85f3-74925b113941 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Multilevel semantic interaction alignment for video–text cross-modal retrieval,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37202589-7b99-4879-8fac-b0d933076c10 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation VideoChat: Chat-Centric Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc909dc9-ba3a-4750-bbe9-39769e529306 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Video-llama: An instruction-tuned audio-visual language model for video understanding,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f77b4e5e-7d18-4849-a028-b1c48563771f · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831d8942-a346-4f40-9a3b-25cb57a1004c · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89fee407-103d-46cb-831e-8e19fb6d605f · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Language models with image descriptors are strong few-shot video-language learners,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e75ac0-42de-4cbc-8e57-69a8d3917f5a · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Compressed video action recognition with dual-stream and dual-modal transformer,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccab0104-9ecf-4506-8be1-da3f20434797 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Dynamic spatial focus for efficient compressed video action recognition,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f1d5e46-a24d-44ee-88fa-eee4f7a9d064 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Alignment-guided temporal atten- tion for video action recognition,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 787aab86-06d4-43a0-befd-c74f18c3aa3b · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Slowfast networks for video recognition,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2071a66f-0e0e-4850-9335-33bc33a806d5 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Temporal distinct representation learning for action recognition,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d99e65b-f375-4487-aafe-5d7a2e96741b · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Truncate-split-contrast: a framework for learning from mislabeled videos,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de6591e7-2ab0-41c1-a586-ed130cdf71c1 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Reading-strategy inspired visual representation learning for text-to- video retrieval,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 903fe238-f9d2-46cb-89ec-7c40649e1aff · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Use what you have: Video retrieval using representations from collaborative experts,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 881de220-95ef-413a-8ee8-773a53b9718a · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Dual encoding for zero-example video retrieval,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a2a580-296b-4df4-bf20-aff3036720a2 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Locvtp: Video-text pre-training for temporal localization,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adad5951-16a8-4bfb-83c9-20e48977a0a0 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Breaking winner-takes-all: Iterative-winners-out networks for weakly supervised temporal action localization,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c562c82-db62-4be2-bd76-c3c12052f130 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Cross time-frequency transformer for temporal action localization,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1740bebf-8852-4dcf-87ce-0ac1d1bf3b23 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Slow motion matters: A slow motion enhanced network for weakly supervised temporal action localization,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce42a0f6-b6d2-4169-a8d3-cbee63e055e4 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Long-form video- language pre-training with multimodal temporal contrastive learning,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4320506b-7c55-417b-ad3d-d9a5b657c202 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation VideoGraph: Recognizing Minutes-Long Human Activities in Videos
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be66552c-aac9-465d-a353-7201b08941d3 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Supervoxel attention graphs for long-range video modeling,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c772c00-a4df-49fe-a93a-02cc99cf828f · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Long movie clip classification with state-space video models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be7e0557-c7f6-41d0-93a9-e7b5db8f2549 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation S4nd: Modeling images and videos as multidimensional signals with state spaces,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6a883042-e723-4ce8-8dba-bfb316fce0f6 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Selective structured state-spaces for long-form video understanding,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3c55eb5c-84ae-440d-ba0d-78efe043e69b · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Efficiently modeling long sequences with structured state spaces,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b707d111-080e-4719-841a-d0e10ae281ac · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Mgsampler: An explainable sampling strategy for video action recognition,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ec1f12c2-12fd-4792-8e34-5910efa98bbd · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Adaframe: Adaptive frame selection for fast video recognition,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7e4b57c6-eef4-4842-883e-935ad5c79ac4 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Localizing moments in long video via multimodal guidance,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 00a45254-eb61-4d1d-bec5-d4f60dcc6e25 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Mad: A scalable dataset for language grounding in videos from movie audio descriptions,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 97ae6fcd-5ebe-4097-9584-57e18ee9b50e · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation End-to-end learning of visual representations from uncurated instructional videos,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bf90c54e-8f4c-4c1e-81b9-45c03751ff83 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Merlot: Multimodal neural script knowledge models,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8a745a5d-ad95-4236-9c48-2f872ed5b492 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Scaling up vision-language pre-training for image captioning,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ded8614-ee68-4727-827a-9f20bc76284e · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Gbc: Guided alignment and adaptive boosting clip bridging vision and language for robust action recognition,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8170842f-9f6f-46c9-b9a9-151258bcb288 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Unsu- pervised pre-training for temporal action localization tasks,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 548090d7-eebb-48e5-8077-948385896a68 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Clip4clip: An empirical study of CLIP for end to end video clip retrieval and captioning,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7d6a0576-d691-432a-9b02-92d48310de48 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Human action recognition and prediction: A survey,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986671f5-f4f9-4321-aa71-3a98f247ffa4 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2081cbe2-ef2e-4b4b-80e2-26cf5f84a614 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Lavender: Unifying video-language understanding as masked lan- guage modeling,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be59d67c-e8a0-4cc3-b078-0cf00140ee16 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a6411d7e-81c2-46b2-9290-6155fde763eb · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Videomae v2: Scaling video masked autoencoders with dual masking,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a7316983-e468-4d51-9030-b1592160bf1b · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba31b6c5-0eb5-4468-a395-30e98376d61e · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation All in one: Exploring unified video-language pre-training,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 69b91bbb-8d7a-4fc6-9e55-b0d424f983e9 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d1f286-a7fd-485c-9b5e-b3d8dbeb4c64 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Language mod- els are few-shot learners,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 98458942-a666-484f-8f2d-3d651a394c1b · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation GLM: general language model pretraining with autoregressive blank infilling,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8db19b9-f461-438d-9453-3f5b81894773 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation LLaMA: Open and Efficient Foundation Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b41f2851-7240-4be8-bf5a-73bca0b77d97 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation MIST : Multi-modal iterative spatial-temporal transformer for long-form video question answering,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cd159143-971a-4beb-bada-fac3eb9a60c5 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation A joint sequence fusion model for video question answering and retrieval,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2387de6-656c-40f1-ab0f-432ba582b063 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Video question answering via gradually refined attention over appear- ance and motion,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bff493ec-b099-46b2-b64f-2e3d8539f224 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Activitynet-qa: A dataset for understanding complex web videos via question answering,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 522cb6ed-6d9f-4a78-8400-76f8bf6ab18e · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Swinbert: End-to-end transformers with sparse attention for video captioning,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 08ad18dd-0240-4738-ae53-f7bf796daec9 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7cadfdc3-52c4-4d92-8868-78f9cb26701b · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Dense- captioning events in videos,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation df2e5754-c242-4eb4-9377-621b106bdf0f · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0f9707f9-ad5b-47f5-9be7-fc6146ad7ac3 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 447ac7d9-1ccd-48aa-8c0b-d5ab82159b1f · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebac29eb-40df-406a-84f7-88fdeb1499aa · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Pyscenedetect,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4b1ddd92-00f5-44b7-9f41-36cc95970907 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Decord: An efficient video loader for deep learning,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f11bf45d-6f12-488f-b4fd-81bd7de680d8 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Learning transferable visual models from natural language supervision,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5849239-8def-49a6-a4c3-cd867a0ed0b8 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation DINO: DETR with improved denoising anchor boxes for end-to-end object detection,
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee5186e6-8829-4917-80a0-799e771344e1 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Sentence-bert: Sentence embeddings using siamese bert-networks,
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 11655f67-fe3d-4ae2-a455-a70c149ca17b · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Partially relevant video retrieval,
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5a466a5f-d901-4c80-88a4-acaad7416c99 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Joint searching and grounding: Multi-granularity video content retrieval,
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bef96942-deb3-4b17-ad8e-b537412a17ba · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Hit: Hierar- chical transformer with momentum contrast for video-text retrieval,
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 619faa62-256d-4b58-b41f-2d23bc20fbff · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Multi-modal trans- former for video retrieval,
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 818e37cf-5c0d-4737-b57a-81c6e1766cc0 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Cross-modal and hierarchical modeling of video and text,
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 05a59bbb-5df8-486d-91ba-046535002610 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Eclipse: Efficient long- range video retrieval using sight and sound,
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7750a39d-77a5-481a-9aa4-1712f209efa4 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation TALL: temporal activity localization via language query,
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f586d4ae-e6cc-47e5-8c52-ee0c3c2a2acb · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Hollywood in homes: Crowdsourcing data collection for activity understanding,
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2c271b89-00d5-48d8-ad34-618c428d11a8 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation MSR-VTT: A large video descrip- tion dataset for bridging video and language,
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f42f821b-3b50-472a-875e-c61b44477382 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Question generation via overgenerating transformations and ranking,
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b12f594c-664e-4e6a-b75c-8f8707ec5244 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Egoschema: A diagnostic benchmark for very long-form video language understanding,
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4a7c3694-66ff-4de1-bb2f-002d1f1b1d72 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Ego4d: Around the world in 3, 000 hours of egocentric video,
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 884ac27b-36fd-4ab8-9bcd-0063a955b8f8 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Next-qa: Next phase of question-answering to explaining temporal actions,
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 22780f74-4a8e-43b9-8b2e-19825e63da4e · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Quo vadis, action recognition? a new model and the kinetics dataset,
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation de50226d-8903-4b14-bd41-818fd009af55 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b77bdc2-774f-4302-b3fb-3a29b29d2bd5 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Sharegpt: Share your wildest chatgpt conversations with one click,
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4e65467a-9dc8-4212-b7e0-f05e6b940993 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation AnglE-optimized Text Embeddings
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be46fc4-02ac-4d22-af5d-02c447adc113 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Video corpus moment retrieval with contrastive learning,
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8875e9a0-e7ed-4f2b-9067-a3ebc767d51c · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation TVR: A large-scale dataset for video-subtitle moment retrieval,
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3aeba3ae-12b5-4ea8-96d3-233932ff55b5 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Dual learning with dynamic knowledge distillation for partially relevant video retrieval,
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f92a385a-4186-4f2f-af3c-81ae0c0f3a94 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab79c81a-42d1-4f97-aa46-be999033b701 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Just ask: Learning to answer questions from millions of narrated videos,
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2528ebbc-87f5-4143-8735-33589eb2ea56 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation MERLOT RESERVE: neural script knowledge through vision and language and sound,
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cf7169bb-7386-4afe-a9d6-dd790f1767c7 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Zero-shot video question answering via frozen bidirectional language models,
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 921d8ba2-5ff9-4525-8263-44a3a5f22613 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Hitea: Hierarchical temporal-aware video-language pre-training,
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0e4515bd-6b0e-47f2-80a4-eff5f7432f4b · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Self-chained image-language model for video localization and question answering,
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 22431014-9daa-4649-902f-4214ea607eb1 · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation Memory consolidation enables long-context video understanding,
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8f5f7710-bbd0-4204-ab66-f4012eaf31af · outbound
Towards Long Video Understanding via Fine-detailed Video Story Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.