Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:24.588694Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 2 inbound Pith citation observations for arXiv:2505.24158.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:24.588694Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:27:56.359929Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T03:06:18.765685Z
84 of 84 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0b0f98ab-5ecc-4387-99bc-7583abcaa79a · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a25c717-9def-4308-8ff1-c6049a818e78 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68155e93-d9bf-49c6-b5da-6cad2496cd05 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vqa: Visual question answering
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bc3e1f1-f3a4-4f92-a82e-04cd9341d2bc · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c4e58b-8675-472d-a7ed-ac3f59df9c45 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Solving mixed-integer quadratic programming problems with ibm-cplex: a progress report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f8b0ea-6d17-4e2d-9aba-f79e9be26d52 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders On the Opportunities and Risks of Foundation Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b1037ce-82e4-47c4-a89d-2e18f8978df0 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Language models are few-shot learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 896ac7d1-7c57-4717-a50c-28a592b161b9 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Hourvideo: 1-hour video-language understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c3568a3a-b413-41e3-88e1-e0f954ffb8ae · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Sharegpt4video: Improving video understanding and generation with better captions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 27e4ed29-6118-42b5-b76a-559978453cd2 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3b44931-b510-4d07-8757-d9f2446ea122 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 762e141e-19cd-44e0-bc51-3a40d50f8291 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d484031f-54be-415b-940f-bb6fb78340da · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 376adc86-d204-4089-8f9e-cecd44d61290 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders An image is worth 16x16 words: Transformers for image recognition at scale
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2676de86-cf63-4741-8c40-e4de5cae84b2 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vlmevalkit: An open-source toolkit for evaluating large multi- modality models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f21be05-08c7-49b4-9be9-780019e38bd2 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Slowfast networks for video recognition
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 809b72ba-d3a4-41c5-8700-937a966fc819 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 47dd4bc0-84f5-4db0-b975-1d6a06d65372 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2381eef4-f583-4f76-94e7-c73d273512bf · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders M-llm based video frame selection for efficient video understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3750f7a5-fdfc-4ed9-aafa-7789b8c5546b · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Chat-univi: Unified visual representation empowers large language models with image and video understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa668bc0-cee4-43e7-b805-46e964487e50 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Language repository for long video understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a8d39fa3-b1cf-4f27-b2a2-2587ffad92a7 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders An image grid can be worth a video: Zero-shot video question answering using a vlm
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4241b8f0-2af6-46c9-9a51-1bf2b12aaab9 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Lmms-eval: Accelerating the development of large multimoal models, March 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 291baed7-2303-4cf2-977a-841f9a3c3bf3 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LLaVA-OneVision: Easy Visual Task Transfer
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d203498-19e8-4011-bb17-7995405bec65 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 60b66c9c-95ff-4c99-aa4f-806c7132fb93 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders VideoChat: Chat-Centric Video Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01bfaaa7-1cec-48f4-8706-3dc35bce83d9 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llama-vid: An image is worth 2 tokens in large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d5a836a-f36a-4515-98b8-f0983000e696 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-llava: Learning united visual representation by alignment before projection
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 638b84e3-ccde-4770-8b05-ab1b71fd4652 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vila: On pre- training for visual language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c00adb3d-44ed-426f-a8e7-1acd59c1535a · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e16f8ea2-1456-4418-80ed-0254c24804d7 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Visual instruction tuning.Advances in neural information processing systems (NeurIPS), 36:34892–34916, 2023
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6954443f-d472-471a-ba74-4e194b2f62cf · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b426fb8e-9c4a-44d3-9d07-e7078607e366 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Lost in the middle: How language models use long contexts
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2f4cc9f1-282b-41ed-9b30-625553421422 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders St-llm: Large language models are effective temporal learners
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 81f55e11-6d72-4da8-a94c-73e27f7c50ec · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Bolt: Boost large vision-language model without training for long-form video understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5342c0cd-7585-46c2-a601-c2740e0a62c8 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Drvideo: Document retrieval based long video understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4d72409-db44-44cc-b222-457dff76e831 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffbf24f6-cd17-462d-8bb0-3c82bb673924 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ae471a22-25b4-492d-a176-bca78cc195e3 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Morevqa: Exploring modular reasoning models for video question answering
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0abc113c-9916-4985-825b-dc805945947b · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Branch-and-bound algorithms: A survey of recent advances in searching, branching, and pruning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d1c22e62-5148-4f3c-bd01-6c185a24dec6 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Chatgpt: Optimizing language models for dialogue, 2023
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 02da2d3a-496d-4709-bcc1-4a52e6737496 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Too many frames, not all useful: Efficient strategies for long-form video qa
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation deb2fcb6-9589-4fa4-b275-b8777c51689d · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Momentor: Advancing video large language model with fine-grained temporal reasoning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9347cb9c-82d3-4a9e-bf0e-53f2dcf2f7f3 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Learning transferable visual models from natural language supervision
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bc64dc1-212e-4d68-9cf6-0e9ba1c2a852 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders The knapsack problem: a survey
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312e0497-9293-483e-9372-348a673c1466 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f19b5cbc-5429-425a-8010-6d84e9d4ed61 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbd4598c-ac5d-4e0c-aeba-4bd4cc9806cf · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Two-stream convolutional networks for action recognition in videos
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d462dcf3-d607-4bff-9bed-dd3901c0c289 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Moviechat: From dense token to sparse memory for long video understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7b755651-2206-48b2-9d89-86b81dba6344 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Mdp3: A training-free approach for list-wise frame selection in video-llms
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 542d090a-bde5-476e-93c8-631aaba47c51 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Adaptive keyframe sampling for long video understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb5be959-6296-4046-9e8b-69de40a997a2 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Gemma: Open Models Based on Gemini Research and Technology
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bb7053e-8653-4f74-a8ed-f522d94005a6 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Cambrian-1: A fully open, vision- centric exploration of multimodal llms
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a09a5b92-a3c3-46d7-9c28-29122d677e3c · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46bcfb17-ae86-4fea-9e96-e026cddb3bfc · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Attention is all you need
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a9736535-42eb-4e5f-8817-6603449fa2e5 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Show and tell: A neural image caption generator
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cf88f0d5-6501-4dae-b17b-c9deec7298aa · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Efficient large language models: A survey.Transactions on Machine Learning Research (TMLR), 2024
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8e98f985-ca8e-4fc2-a759-67f3c756a9be · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Weakly supervised gaussian contrastive grounding with large multimodal models for video question answering
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88731b3a-d85e-4308-ab9d-076a50bdcc2a · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26926a99-6fe9-4763-ae56-44f168a64c7d · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae110e84-5baf-4ba8-94c9-2236d33e803d · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Videoagent: Long-form video under- standing with large language model as agent
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eb9742ac-5b10-4d64-a3ed-357f9bcc1d79 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Videotree: Adaptive tree-based video representation for llm reasoning on long videos
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7cd90885-7735-4b77-8764-6ba6759e772e · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 061e4745-6e27-4634-a3ca-076def2997d1 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7998ee8f-53e2-406a-944f-7d9d6e446f8b · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Next-qa: Next phase of question-answering to explaining temporal actions
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 519ac6c8-ad9d-415d-bb49-44f01783acce · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Can i trust your answer? visually grounded video question answering
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c3b083f2-ee51-4cb1-a99d-b3b5c2ef56c0 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Effective long-context scaling of foundation models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a990e067-485f-4cf9-852f-ee77af800c35 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cba55f33-db7f-408d-99fa-78317f5baaca · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb3d9dd3-06aa-45e8-bdb6-fdb75c4eda2e · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Zero-shot video question answering via frozen bidirectional language models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 51652350-edb5-4a59-8448-5a2ab454b129 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9298836e-13e0-461f-ab6c-1abfb7ea459e · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Dense connector for mllms
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ef5db10-d2dc-4306-898f-c0a74ce8d362 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Generative Frame Sampler for Long Video Understanding
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e9d53f51-3628-4ccd-8161-bafbb69f4d9a · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1388fada-bbd9-4c81-83ce-e5b474ed503f · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Self-chained image-language model for video localization and question answering
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a2d295cf-2c3e-4d3b-864f-226d08d397da · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Frame-voyager: Learning to query frames for video large language models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 49e98b79-f970-452f-9e47-f1468db5e4da · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Sigmoid loss for language image pre-training
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 842a74e2-9b22-453c-b188-65da3f3cc9e1 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders A simple llm framework for long-range video question-answering
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0f16f33-24b5-4af9-ad4e-4b8e07355c6d · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Long Context Transfer from Language to Vision
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79b9db04-5fc5-48b4-aadf-697adeb8f07e · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llava-next: A strong zero-shot video understanding model, April 2024
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b30ebedc-b77a-4463-b3a6-ef8c7a773f6d · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b028f069-8365-4ff5-903f-0d9928d4f313 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MLVU: Benchmarking Multi-task Long Video Understanding
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4793def1-2e8a-4082-a776-2a2c5df55d59 · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 447bbede-b0be-4fbc-9feb-22827ba5a7bd · outbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Apollo: An Exploration of Video Understanding in Large Multimodal Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76dbb3a0-84e8-434c-9c6f-48b3029920af · inbound
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 767c7c7e-b77e-4aeb-8c50-e2dfe77869ac · inbound
CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.