Pith. sign in

Paper Citation Record · LEDGER

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models

As of 4 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2605.26014.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26014 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T23:04:21.463842Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact32
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ed27c2f-2875-4f32-8273-05c7e910d6a8 · outbound

This paper cites Temporal chain of thought: Long-video understanding by thinking in frames.Advances in Neural Information Processing Systems, 38:143018–143046, 2026.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Temporal chain of thought: Long-video understanding by thinking in frames.Advances in Neural Information Processing Systems, 38:143018–143046, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:9300c6331bb6aaac30be8fa37f5121b0358fa136f85fd454e163198cfd6d4724

Observation bf37c793-386e-4fea-b83a-403cc00ca777 · outbound

This paper cites Qwen2.5-VL Technical Report.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Qwen2.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.149530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:a8c57b02b16ff4bbfd21a84d7cdf4207793fccb16c804ff3932d8297c7f7543b

Observation 0c1e0b0f-925a-4f91-ab34-0410f6dc74e1 · outbound

This paper cites Perception tokens enhance visual reasoning in multimodal language models.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Perception tokens enhance visual reasoning in multimodal language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:dc9985a103630b26b32780d00160894373a4413fd84f824103c358a5db32b766

Observation b85823d2-4229-4e84-be95-744a3e30e6d6 · outbound

This paper cites Spatialdreamer: Incentivizing spatial reasoning via active mental imagery.arXiv preprint arXiv:2512.07733.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Spatialdreamer: Incentivizing spatial reasoning via active mental imagery.arXiv preprint arXiv:2512.07733

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.145445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:c63d061085eb7ee21beabb57cde30924cd6f82c9d923b372d5c71365494c555a

Observation 5e002077-db44-4bd0-b0da-84e45fcd9f89 · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.Advances in Neural Information Processing Systems, 38: 91077–91100, 2026.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Eagle 2.5: Boosting long-context post-training for frontier vision-language models.Advances in Neural Information Processing Systems, 38: 91077–91100, 2026

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:ed9379340921d3e5f86280a313883a3e238008c5b9e50ceaf8fa95f81553e70e

Observation 97e283d9-d204-441a-8743-5cce1de4df21 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:d6ac82fa9e2a397cb394aee84fb7dc2d12bee08419cf230c54cf5efeb0c8efbe

Observation 99d72b92-c5eb-4d43-9246-24b882d50954 · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.175936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:b4b43b20be5c50b2fa9ccde133332bbaadaa14368da44720e946863ba8780ef4

Observation b86c6586-2eb6-4d46-bb9a-5bbd73d411ec · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.194371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:103cb72e2919bab67ea3b1f15c7481b5df8fe7df0e4f86417a7850f9655c3624

Observation 840b73b0-a587-498c-af76-e3de316843fc · outbound

This paper cites Don’t look only once: Towards multimodal interactive reasoning with selective visual revisitation.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Don’t look only once: Towards multimodal interactive reasoning with selective visual revisitation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:9fe4b871e8b7f155cc6e613e57d42cefa8a720460cda0073330fc2fec58915ca

Observation a945a249-5493-4350-9b7c-7b8188b729c7 · outbound

This paper cites From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.196716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:957e28b11b61865d6adc0f2ebfbdad7e6637dcbec5bb5f6f285124550773dd06

Observation 433ae2e0-4eaa-49c3-a4f7-36f6b24270f2 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.202123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:25f3bd81a4af4fa0f446fd5b9a5421a58c599326bc7600046c901c1f95e6d8dd

Observation e9e022b0-0a57-49d8-8b19-8f615fb6a7ff · outbound

This paper cites Video-r1: Reinforcing video reasoning in mllms.Advances in Neural Information Processing Systems, 38:99114–99137, 2026.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-r1: Reinforcing video reasoning in mllms.Advances in Neural Information Processing Systems, 38:99114–99137, 2026

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:dc829a1b47f7338f8861bf1173212b85fefd7912b16d91bfef52d025987ca094

Observation f6da11c3-3a5f-456f-97e3-52fb17e41955 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.Advances in Neural Information Processing Systems, 38, 2026.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Mme: A comprehensive evaluation benchmark for multimodal large language models.Advances in Neural Information Processing Systems, 38, 2026

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:e7d333115184d705d16e0d03cd25ccacf1e9acd57f4df27487f7fa91c6f5b8ba

Observation ff3ce45b-1a70-4467-8747-bcdd131ac89b · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Blink: Multimodal large language models can see but not perceive

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:8ae5b3db015152e5e915be7a6ed08b53201cee910a23fe55ce54e34cf7e862ef

Observation c0f0b526-4d84-449c-935b-dc8265e2006d · outbound

This paper cites ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.204792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:863fd4a4491bd3b09b284d69f0df717bf99993fe7a8cd41a2629f62fc4b9bdcf

Observation 3e5b7e4b-668c-4e5d-8a96-d687ad83697e · outbound

This paper cites Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.209910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:9c45032a2a89b39e0348f19e5257a6d83cfd774a7eaea1cda163febe36f2707b

Observation 073c46b2-b814-4303-beb5-490553676874 · outbound

This paper cites Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:58bc1173e34493a9fd95677ce41ad7fc8ff81191b16fb62ffa10830dcd87ac1b

Observation 93c599d7-8362-4486-ad1b-21a4216a95c6 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Training Large Language Models to Reason in a Continuous Latent Space

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.212178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:701dc8ac7aaa799c1f161c66cb0f461029004b6b6f506eb9c533777f047ff23f

Observation 42cb633d-a214-4ce6-8196-39810dab466d · outbound

This paper cites GPT-4o System Card.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models GPT-4o System Card

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.173451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:65d05e9c72c6dd6921a6c820a511660b2fe16f63cf97ad5a4a9ab329dc765748

Observation 1ef76558-25a3-49ae-9f1e-11e85ca42551 · outbound

This paper cites Visioncoach: Reinforcing grounded video reasoning via visual-perception prompting.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Visioncoach: Reinforcing grounded video reasoning via visual-perception prompting

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.217378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:6f96692ef2c2951b0b3c39d88c1525de99e58049eaf81035d9ee0277eb04c1e2

Observation d5405cfb-5c63-4d27-9add-07cf6554e9df · outbound

This paper cites Latent Visual Reasoning.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Latent Visual Reasoning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.147369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:27c617c3fde5ba1df2ee996d20b7efb882b7e35c2e0bb12b96c1bb8841c3b5ab

Observation d710614a-dad6-47ca-946a-6920d15a98f8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.219580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:830b1686944639c3d3ba5536e57fdd4e6119456f96cd7451764f59a7be0d3d36

Observation 645d133c-f232-407d-ba77-d3223737a921 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:94d1f15c00c79ff1bba69a9411614f63af34e209fac98e8a10243ea8aaa2c053

Observation f02789d4-5c79-4d0c-a9e8-0c09eab6fb0c · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:6c722d1c46810807baf7de7a08ec83b9c42d142e9bb8ce5ab1ef39f90a866213

Observation edd91668-ed11-4a1e-ae09-d4c4bc76fcc5 · outbound

This paper cites Univer- sal video temporal grounding with generative multi-modal large language models.Advances in Neural Information Processing Systems, 38:64426–64455, 2026.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Univer- sal video temporal grounding with generative multi-modal large language models.Advances in Neural Information Processing Systems, 38:64426–64455, 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:2fca3c5b5a08c60dd07972e54a9f9f1bf439b3095277ce209ee8290b65abb23c

Observation 8274f1bf-9d38-489c-b2d1-edbcafc5c3fb · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-llava: Learning united visual representation by alignment before projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:9020c8c9d784605ada02f87f280edf8b5ea2c536633c48c87727b547344b8848

Observation 1d6f08e3-5618-4a3b-aefe-2db9a49e015e · outbound

This paper cites Kangaroo: A powerful video-language model supporting long-context video input: J.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Kangaroo: A powerful video-language model supporting long-context video input: J

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:5d2439c674582994dd1a97cfa079167ef87db494a522aed2b5896876061084b9

Observation b1dff096-465c-41be-8138-8c7d113fa18c · outbound

This paper cites St-llm: Large language models are effective temporal learners.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models St-llm: Large language models are effective temporal learners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:2558e5ae8587b211d9fd614b2c0c5b7cafb541acc76c7fa0821c3a4451f0c612

Observation 18d1a293-f65c-42d7-9f0b-7ca651a906cf · outbound

This paper cites Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics: ACL 2024, pages 8731–8772, 2024.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics: ACL 2024, pages 8731–8772, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:449b1d48dc06d0d65e966813366bcac87fb81014f1147714a3f709dc529398f0

Observation 332ffb27-fe4e-4210-90ae-dd15b2a82d4d · outbound

This paper cites Sat: Spa- tial aptitude training for multimodal language models.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Sat: Spa- tial aptitude training for multimodal language models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.222058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:20aefeb40f1fd93f3ed3f3896ac5cee5f6eb1bee84d2636f1fbb41dbca1f1294

Observation 41d7a7ef-7ca3-465a-86f5-e44bb5d11d83 · outbound

This paper cites Mull-Tokens: Modality-Agnostic Latent Thinking.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Mull-Tokens: Modality-Agnostic Latent Thinking

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.171278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:351e9bd18506ecd996175a6013fef2a9d966df499b1f99f66b8ac0a83fb80afc

Observation 6c74e4bb-7fa0-49cc-8719-d908b62f7470 · outbound

This paper cites Zoomeye: Enhancing multimodal llms with human-like zooming capabilities through tree-based image exploration.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Zoomeye: Enhancing multimodal llms with human-like zooming capabilities through tree-based image exploration

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:a6f75d87aead731360440e96a5660014bfd99dc9dedcf14c62103b3a39843624

Observation cb8d4dc2-4bbb-4abd-86f2-da6f4089aa39 · outbound

This paper cites Codi: Com- pressing chain-of-thought into continuous space via self-distillation.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Codi: Com- pressing chain-of-thought into continuous space via self-distillation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:3d7b6523532093440a01a2c2da32ff2e6cc2c6ce7547bc673f3a8bc47ed39d20

Observation dff17e02-9429-4880-ae5f-d28a82840e38 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.183555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:448fa3df686c432d6af8628dae9b982895afb3440899f3e7cffa60a68187a028

Observation d315af14-af7f-41a9-bce4-ba2f5d1b03f6 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.164050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:db3d29162c8d152e2baea08b29b975c9e7e1e52ecafd3020bcce161a97e2096d

Observation 1951b92a-731d-4d5f-a925-63c395ab0937 · outbound

This paper cites Video-thinker: Sparking” thinking with videos” via reinforcement learning.arXiv preprint arXiv:2510.23473.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-thinker: Sparking” thinking with videos” via reinforcement learning.arXiv preprint arXiv:2510.23473

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:14:02.169937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:c567565c129310f9f3a9b11dc1c060735ccc3b09dabd1a8788c9e6b396400617

Observation ddbca35b-dfde-45b2-972f-4b90e20b0f71 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.182899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:0c2f967887aa806bb2fa8971d9e45d9f0366995beb9332768a18724b1241751b

Observation 50133aa3-2b10-46ab-a869-f38dc07a5318 · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.180351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:f974d618afd4ff3715a0847c59cd3b40f3c3cf8368928a9e7a7ee6df366c643d

Observation df04e209-d8c1-4065-b552-8939cfdf5654 · outbound

This paper cites Time-r1: Post-training large vision language model for temporal video grounding.Advances in Neural Information Processing Systems, 38:83330– 83364, 2026.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Time-r1: Post-training large vision language model for temporal video grounding.Advances in Neural Information Processing Systems, 38:83330– 83364, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:cfb6d35f005ef72114c1e93c445615ef474413cafb82505cde033f56b2198481

Observation 01f1fa6a-fd3d-433a-b2fe-8e7a4b3444c4 · outbound

This paper cites VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.214718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:c3a0cdb4d0704799d6b512e0c11068909974485e72f5902ccdedb1b9e48f7a53

Observation d1fd8fbc-7905-4148-a208-f46299698392 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.214187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:87d1aac08a9b298ceccead13b38c58f3c9b77c87a5eebb949b52e211edf0093e

Observation 9c8bb1ec-0283-48c4-9bae-2a968b2deed9 · outbound

This paper cites SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.177596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:d08c7f70a16b27b1f3f849416fd2378c29bd165609d3644dd99829ff37b550ee

Observation 314c3d7f-4428-486a-a4bf-aa89d081ac11 · outbound

This paper cites Visual planning: Let’s think only with images.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Visual planning: Let’s think only with images

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.201167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:a120e4ae11188bd0d352201f307c6c5a151f12e82a6e4175da6f733ca4e226cd

Observation 53f5bd53-bb66-4f09-8a9f-ce94ab1daed7 · outbound

This paper cites an unresolved cited work.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:b662f898017777c31d72dc643392a1bec55f35604ec64f90685d80ae9ef8fe49

Observation ead86fe9-3639-44a2-b16f-34c923b51c50 · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.161748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:86c3e1eacc27d5fdb200eb60999c03d4bdde1b822181608d0c8ca05f5788dbfe

Observation 0263200f-4e1f-4d9b-ada5-12bf700c67d1 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.218728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:fadb965443f922bcc2f83a2e3a232731f8fd85caf6b96c97b3fb45b49376564a

Observation f2a19d64-227e-4966-b2e5-461947f50484 · outbound

This paper cites LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.223873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:de6e8f87f1fff1a7385626962a41470c047e193f7f60bd8957be76ac9e005616

Observation 38255bd4-f7fe-477d-9ada-8fbd9b4543a5 · outbound

This paper cites When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.224535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:8bd1d155a551e7526d673b789c3289516e57b0484e38d5d129cf9b5438043c2a

Observation 95dfcfab-c8d0-40bd-b52d-892e0696a4a1 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.207358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:5e522d58a44ac109d12ff7a08001cb818821fd5a2557815d98ab4243e359b2b8

Observation 69750faa-9f81-4477-af30-ff24fbcc021e · outbound

This paper cites Cmmcot: Enhancing complex multi-image comprehension via multi-modal chain-of-thought and memory augmentation.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Cmmcot: Enhancing complex multi-image comprehension via multi-modal chain-of-thought and memory augmentation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:37d0e7b64d92db552694f81f2c33f0659081641cebc703c0be7dcf30522e0b57

Observation 50b57f9b-3c25-41b3-b925-5308d6dd85f8 · outbound

This paper cites Long Context Transfer from Language to Vision.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Long Context Transfer from Language to Vision

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.186158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:d334f0af415a6f0c93c50cb53f8fd481f054de24c3c0657b08fdf70cb5164d5d

Observation 0c8ae7cb-59e4-4584-96dd-08dfbfe59926 · outbound

This paper cites Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl.arXiv e-prints, pages arXiv–2505, 2025.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl.arXiv e-prints, pages arXiv–2505, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:8a940adfcbb565d104df42207823a99ba7a971d18b730356dc0efcc5bf9eacc0

Observation 12f59734-c3e1-49ce-8718-91decbbbff57 · outbound

This paper cites LLaV A-NeXT: A strong zero-shot video understanding model, 2024.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models LLaV A-NeXT: A strong zero-shot video understanding model, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:ade9aa5b5f2074b63daf95a07d0ec3001643bbc7efe0433b22c319a5104878eb

Observation 2c177a1e-5681-4bd1-9c71-71475485eaec · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.193248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:6180f3c86ab6b2e2e88204698a5853dfa65babca28d4183da5b5622224832101

Observation 4e7eda27-9a17-4d93-9a2d-9fef7b857d41 · outbound

This paper cites Mmvu: Measuring expert-level multi-discipline video understanding.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Mmvu: Measuring expert-level multi-discipline video understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:a666c122925195fb310a6d51967498a70a01b38e3a5f8c8865002141d4386419

Observation e5b7ce6d-fd73-4a12-bb8b-fd4f47701776 · outbound

This paper cites Reagent-v: A reward-driven multi-agent framework for video understanding.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Reagent-v: A reward-driven multi-agent framework for video understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-29T23:04:21.463842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:fda4b54253dcc0159a5eee9d427f864653e13a1ea164633a522168587e6df562

Observation 20a589d5-7cdd-473f-a8e3-412f1d2b8dfd · outbound

This paper cites Emergence of superposition: Unveiling the training dynamics of chain of continuous thought.arXiv preprint arXiv:2509.23365, 2025a.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Emergence of superposition: Unveiling the training dynamics of chain of continuous thought.arXiv preprint arXiv:2509.23365, 2025a

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.172467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:7968ddf9194f2b33d05b98f164b133342f0ad5a9e320baca3e745f293761b171

Observation 5e282e05-9ddb-4dc4-a4db-147600d69eb9 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.178567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:152256717e677b3ee209bfa5a6ee0d8a42932c0ccbda1ab991f687e3acfe2c00

Pith citing papers

No inbound Pith citation observations are available.