Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T23:04:21.463842Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2605.26014.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T23:04:21.463842Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8ed27c2f-2875-4f32-8273-05c7e910d6a8 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Temporal chain of thought: Long-video understanding by thinking in frames.Advances in Neural Information Processing Systems, 38:143018–143046, 2026
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf37c793-386e-4fea-b83a-403cc00ca777 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0c1e0b0f-925a-4f91-ab34-0410f6dc74e1 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Perception tokens enhance visual reasoning in multimodal language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b85823d2-4229-4e84-be95-744a3e30e6d6 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Spatialdreamer: Incentivizing spatial reasoning via active mental imagery.arXiv preprint arXiv:2512.07733
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e002077-db44-4bd0-b0da-84e45fcd9f89 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Eagle 2.5: Boosting long-context post-training for frontier vision-language models.Advances in Neural Information Processing Systems, 38: 91077–91100, 2026
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e283d9-d204-441a-8743-5cce1de4df21 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d72b92-c5eb-4d43-9246-24b882d50954 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b86c6586-2eb6-4d46-bb9a-5bbd73d411ec · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 840b73b0-a587-498c-af76-e3de316843fc · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Don’t look only once: Towards multimodal interactive reasoning with selective visual revisitation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a945a249-5493-4350-9b7c-7b8188b729c7 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 433ae2e0-4eaa-49c3-a4f7-36f6b24270f2 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e9e022b0-0a57-49d8-8b19-8f615fb6a7ff · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-r1: Reinforcing video reasoning in mllms.Advances in Neural Information Processing Systems, 38:99114–99137, 2026
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6da11c3-3a5f-456f-97e3-52fb17e41955 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Mme: A comprehensive evaluation benchmark for multimodal large language models.Advances in Neural Information Processing Systems, 38, 2026
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff3ce45b-1a70-4467-8747-bcdd131ac89b · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Blink: Multimodal large language models can see but not perceive
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f0b526-4d84-449c-935b-dc8265e2006d · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3e5b7e4b-668c-4e5d-8a96-d687ad83697e · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 073c46b2-b814-4303-beb5-490553676874 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c599d7-8362-4486-ad1b-21a4216a95c6 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Training Large Language Models to Reason in a Continuous Latent Space
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 42cb633d-a214-4ce6-8196-39810dab466d · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models GPT-4o System Card
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1ef76558-25a3-49ae-9f1e-11e85ca42551 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Visioncoach: Reinforcing grounded video reasoning via visual-perception prompting
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5405cfb-5c63-4d27-9add-07cf6554e9df · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Latent Visual Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d710614a-dad6-47ca-946a-6920d15a98f8 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 645d133c-f232-407d-ba77-d3223737a921 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02789d4-5c79-4d0c-a9e8-0c09eab6fb0c · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Llama-vid: An image is worth 2 tokens in large language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edd91668-ed11-4a1e-ae09-d4c4bc76fcc5 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Univer- sal video temporal grounding with generative multi-modal large language models.Advances in Neural Information Processing Systems, 38:64426–64455, 2026
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8274f1bf-9d38-489c-b2d1-edbcafc5c3fb · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-llava: Learning united visual representation by alignment before projection
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d6f08e3-5618-4a3b-aefe-2db9a49e015e · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Kangaroo: A powerful video-language model supporting long-context video input: J
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1dff096-465c-41be-8138-8c7d113fa18c · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models St-llm: Large language models are effective temporal learners
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d1a293-f65c-42d7-9f0b-7ca651a906cf · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics: ACL 2024, pages 8731–8772, 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332ffb27-fe4e-4210-90ae-dd15b2a82d4d · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Sat: Spa- tial aptitude training for multimodal language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 41d7a7ef-7ca3-465a-86f5-e44bb5d11d83 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Mull-Tokens: Modality-Agnostic Latent Thinking
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6c74e4bb-7fa0-49cc-8719-d908b62f7470 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Zoomeye: Enhancing multimodal llms with human-like zooming capabilities through tree-based image exploration
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb8d4dc2-4bbb-4abd-86f2-da6f4089aa39 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Codi: Com- pressing chain-of-thought into continuous space via self-distillation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff17e02-9429-4880-ae5f-d28a82840e38 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d315af14-af7f-41a9-bce4-ba2f5d1b03f6 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Wan: Open and Advanced Large-Scale Video Generative Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1951b92a-731d-4d5f-a925-63c395ab0937 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-thinker: Sparking” thinking with videos” via reinforcement learning.arXiv preprint arXiv:2510.23473
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ddbca35b-dfde-45b2-972f-4b90e20b0f71 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 50133aa3-2b10-46ab-a869-f38dc07a5318 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation df04e209-d8c1-4065-b552-8939cfdf5654 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Time-r1: Post-training large vision language model for temporal video grounding.Advances in Neural Information Processing Systems, 38:83330– 83364, 2026
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f1fa6a-fd3d-433a-b2fe-8e7a4b3444c4 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d1fd8fbc-7905-4148-a208-f46299698392 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9c8bb1ec-0283-48c4-9bae-2a968b2deed9 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 314c3d7f-4428-486a-a4bf-aa89d081ac11 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Visual planning: Let’s think only with images
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 53f5bd53-bb66-4f09-8a9f-ce94ab1daed7 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead86fe9-3639-44a2-b16f-34c923b51c50 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0263200f-4e1f-4d9b-ada5-12bf700c67d1 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f2a19d64-227e-4966-b2e5-461947f50484 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 38255bd4-f7fe-477d-9ada-8fbd9b4543a5 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95dfcfab-c8d0-40bd-b52d-892e0696a4a1 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69750faa-9f81-4477-af30-ff24fbcc021e · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Cmmcot: Enhancing complex multi-image comprehension via multi-modal chain-of-thought and memory augmentation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50b57f9b-3c25-41b3-b925-5308d6dd85f8 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Long Context Transfer from Language to Vision
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0c8ae7cb-59e4-4584-96dd-08dfbfe59926 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl.arXiv e-prints, pages arXiv–2505, 2025
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12f59734-c3e1-49ce-8718-91decbbbff57 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models LLaV A-NeXT: A strong zero-shot video understanding model, 2024
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c177a1e-5681-4bd1-9c71-71475485eaec · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4e7eda27-9a17-4d93-9a2d-9fef7b857d41 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Mmvu: Measuring expert-level multi-discipline video understanding
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5b7ce6d-fd73-4a12-bb8b-fd4f47701776 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Reagent-v: A reward-driven multi-agent framework for video understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a589d5-7cdd-473f-a8e3-412f1d2b8dfd · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Emergence of superposition: Unveiling the training dynamics of chain of continuous thought.arXiv preprint arXiv:2509.23365, 2025a
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e282e05-9ddb-4dc4-a4db-147600d69eb9 · outbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
No inbound Pith citation observations are available.