Pith. sign in

Paper Citation Record · LEDGER

Demystifying Video Reasoning

As of 10 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 7 inbound Pith citation observations for arXiv:2603.16870.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.16870 v3

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T02:34:01.740152Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:24:28.926512Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T22:26:17.260644Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved81
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f37a201-64e9-40b5-8115-2a32baea8c3d · outbound

This paper cites Advances in neural information processing systems 35, 23716–23736 (2022).

Demystifying Video Reasoning Advances in neural information processing systems 35, 23716–23736 (2022)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:52.921494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:52.921494Z digest=sha256:df0f43dd54ee1d74959fbf61b269d886d718969f3e9947df0c3fba9f357908c4

Observation eaff6a74-4695-4296-8363-0a1950ad5250 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Demystifying Video Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:52.964918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:52.964918Z digest=sha256:23069572f89dc381a1efd94192b9a4124810a1ad1e11ceedbcd1b12c0a6a164d

Observation f4820378-d8fd-49c6-b65c-495db2a7bb93 · outbound

This paper cites Neuron 100(2), 490–509 (2018).

Demystifying Video Reasoning Neuron 100(2), 490–509 (2018)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.095672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.095672Z digest=sha256:007963d4f56b60c741e472d50ae8972c8b94145f1098036d89648ecd011e778a

Observation 097e7648-d32c-4efd-8214-3d0523b5fa13 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Demystifying Video Reasoning BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.187763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.187763Z digest=sha256:9b086e4198cb1edfa0ecc90354f7d09440953017d4b6590081bcf905ddc6d3a5

Observation 9dd078b8-1151-42d5-85c8-cc2c9a113731 · outbound

This paper cites arXiv preprint arXiv:2602.12279 (2026).

Demystifying Video Reasoning arXiv preprint arXiv:2602.12279 (2026)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.311838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.311838Z digest=sha256:1fcaaa7d567b4f804aed51b49322ba0df877855e61a2da925ccb997deb67b53a

Observation 39eb3e3e-cf45-4af9-b7e1-0628fe0c2829 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Demystifying Video Reasoning In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.448633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.448633Z digest=sha256:59c2d6699b3d8f40e8189cdf79973117d12d8fcab1b62c200ba6f449330201e4

Observation b2339c8b-b782-4231-8d29-031878628030 · outbound

This paper cites Thinking with Generated Images.

Demystifying Video Reasoning Thinking with Generated Images

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.570433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.570433Z digest=sha256:0fd57846011f8fd88f93ee8dc71412663601d9081ad89feb82877253aef2ae4b

Observation fd3ad8d6-45e0-42fb-8ad1-61a466a4cdcd · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Demystifying Video Reasoning Emerging Properties in Unified Multimodal Pretraining

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.617312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.617312Z digest=sha256:097f3f10dad6b67049235729cad63bd6722050395d983592e6285857253f461e

Observation f70ffa70-5579-41ee-a828-8fb71f5b77ba · outbound

This paper cites In: International conference on machine learning.

Demystifying Video Reasoning In: International conference on machine learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.674465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.674465Z digest=sha256:b9b6e206d503f07ca586a63dd1f0f9c9a42e7c7a4965012051fe04205baf0d83

Observation 516faf33-70ee-4909-9683-b156ffc9db19 · outbound

This paper cites GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning.

Demystifying Video Reasoning GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.733424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.733424Z digest=sha256:885f7e6140afafeacc9973f457db5ed043f6b4bfac298f95e58828e22633fce4

Observation 11b88191-a3d4-4dfa-985f-636b32b79294 · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

Demystifying Video Reasoning Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.809611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.809611Z digest=sha256:ef8c348ec96836b105b37f60cf0b5dac164f4e267624f28693dedca755ae3402

Observation ffd9ecd0-296e-4f19-8bc1-37937673d194 · outbound

This paper cites arXiv preprint arXiv:2510.27684 (2025).

Demystifying Video Reasoning arXiv preprint arXiv:2510.27684 (2025)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.870811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.870811Z digest=sha256:264616609a7e16490eb0bf2a458b791a36f4c0cf1e0febe366050cea4e203b40

Observation 3766ccda-d2e8-4b89-a5ea-a612eab982a5 · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

Demystifying Video Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.930732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.930732Z digest=sha256:7e8ab0382cf1e54c02cbe87ece62ed3d409e831bb9818fc54aeebb105159b1b3

Observation c9ae9374-51f2-4a29-943f-6e9afd7eb16a · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Demystifying Video Reasoning SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.970685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.970685Z digest=sha256:d09916dc99e5432137a965427758afeee312ca1eff88de19651ed94a6d528fdf

Observation ad251a02-08dc-4a6a-8bc9-d18bc9b24c14 · outbound

This paper cites Technical Report Veo 3.1, Google DeepMind (January 2026), https: //blog.google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/ , released January 13, 2026.

Demystifying Video Reasoning Technical Report Veo 3.1, Google DeepMind (January 2026), https: //blog.google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/ , released January 13, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.074142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.074142Z digest=sha256:f79e606c0bb25bc01cb8ca351762739d9f05d0135abcd484323b5383d55cb16d

Observation 17c05c7b-cb86-4bb6-a657-19a2bcbeb692 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

Demystifying Video Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.197765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.197765Z digest=sha256:f25bff56af91021fd403217e42dd378c1e775fd9af1c0dc567c22f1c44c51d85

Observation 2b14d321-d112-415f-a26f-9371d35a4e4a · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

Demystifying Video Reasoning LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.270513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.270513Z digest=sha256:fd3a110edbbfbdcfcc1a3d62292a63b026c410ec62beebab94d281644bac0538

Observation e7546311-49ee-460a-88b2-9e59e75daff6 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Demystifying Video Reasoning Training Large Language Models to Reason in a Continuous Latent Space

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.345493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.345493Z digest=sha256:a2ac8fc42e5a9d9e11a6c58d47babbbc53f4b9256b934da6efa54af6f5b7fa0d

Observation 01f0d2eb-bc92-44c4-b0f5-3acaccabad6b · outbound

This paper cites arXiv preprint arXiv:2512.02622 (2025).

Demystifying Video Reasoning arXiv preprint arXiv:2512.02622 (2025)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.416375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.416375Z digest=sha256:f7223f08f8ab4d61146099eda49437628206c2b6cf813c30aaa67cba3220de73

Observation 5763056b-9a86-49f8-b9b8-c70d511e0706 · outbound

This paper cites In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H.

Demystifying Video Reasoning In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.489463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.489463Z digest=sha256:e6d34bfafd00b4e6fb94fba21fa116821228f90481454056207ba5250b7bf628

Observation dae0de17-10b7-422b-885e-d6148340e0fc · outbound

This paper cites In: Proceedings of the 2023 conference on empirical methods in natural language processing.

Demystifying Video Reasoning In: Proceedings of the 2023 conference on empirical methods in natural language processing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.573616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.573616Z digest=sha256:bd31cd158d2502e2f8991a14e77afac173bc27447a2e7e37d450d4798a13f01b

Observation e7ce3380-e8bc-4589-92ef-baf681f3e1f3 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024).

Demystifying Video Reasoning In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024)

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.644722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.644722Z digest=sha256:e7fbf5404f5e178d8e6c5b1afc2a9fb2c20d8adfcb4930038b92a8cd09aee832

Observation 9036b58d-d12f-4cf9-ae27-f4630177d170 · outbound

This paper cites VChain: Chain-of-Visual-Thought for Reasoning in Video Generation.

Demystifying Video Reasoning VChain: Chain-of-Visual-Thought for Reasoning in Video Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.768104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.768104Z digest=sha256:e27a42439ec5443157462d6adc049e11375365e643e8c0b4a58d020963e57131

Observation 37feeb0c-ab25-4076-be2b-32d2b6ff2ec8 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence (2025).

Demystifying Video Reasoning IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.904741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.904741Z digest=sha256:0257dfc666dcbad46c674dd7c2f37a42d49451afcec0181f67949d6d073f0ccd

Observation 0e924498-084f-4499-a413-563005529d2f · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Demystifying Video Reasoning T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.036331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.036331Z digest=sha256:02dcf34f7e078c34c24138da2b125421698d8b490800d274544db4637f9f31c0

Observation a35fe085-c00a-4c49-a1ad-29dee50fef5a · outbound

This paper cites Auto-Encoding Variational Bayes.

Demystifying Video Reasoning Auto-Encoding Variational Bayes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.156467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.156467Z digest=sha256:3b6f3dd318bcc930f0303856805aade9aa23921b667daa07e4d235b495c1ca0b

Observation af5925bf-b8b4-4ed8-a628-f116730c4493 · outbound

This paper cites NeurIPS (2023).

Demystifying Video Reasoning NeurIPS (2023)

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.236134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.236134Z digest=sha256:093dba36767ffe5e870a387282b716e8bccbfd535068e5731854b9bc3c866ead

Observation 1b4af0e4-6043-42d9-99c4-a6f6edff39b5 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Demystifying Video Reasoning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.289369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.289369Z digest=sha256:380bae79971c7c4327d66d239ae08fd6777b8bd899432b1d6580adecd29ac8ab

Observation d0278788-d2a9-490f-8f71-bfbe3a30a828 · outbound

This paper cites In: International conference on machine learning.

Demystifying Video Reasoning In: International conference on machine learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.356358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.356358Z digest=sha256:1e0647a9d1589ba79ecf354fe7f670d1168c5ac13b7fbd09c213416dd9c10710

Observation 70886151-777a-419f-a5b6-14b5dd8acc6e · outbound

This paper cites simultaneous audio-visual generation.

Demystifying Video Reasoning simultaneous audio-visual generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.397252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.397252Z digest=sha256:e641825786791e46c0a4bed6aff9938e0e59dd0b1bd1cd99a9316e959f9eb072

Observation e17a3ed1-3199-4e5e-92d3-77967eb97d3b · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

Demystifying Video Reasoning Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.476154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.476154Z digest=sha256:6c5182989dc35e4261794c984ca0fbc7960ee92854adee737ca6254eb00be770

Observation 96511202-ec72-4d57-9479-e51dd39acf60 · outbound

This paper cites In: International conference on machine learning.

Demystifying Video Reasoning In: International conference on machine learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.605877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.605877Z digest=sha256:54e4c022d12fb822bf1126aa077d86c46f154b38a969fadb0bce553dda8c8ecb

Observation 21d33951-39b9-4f9b-a8c4-8db51d955fc8 · outbound

This paper cites UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation.

Demystifying Video Reasoning UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.705653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.705653Z digest=sha256:8f3a35ce737089a1fa9d88a775712613c8056537db0f72601aa32f19294cedb5

Observation e78c0c14-6d6a-4dd8-9330-ccad39e82f4f · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

Demystifying Video Reasoning Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.844314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.844314Z digest=sha256:3b35d9f2ccfc9b642def364fa353bba668eb62bf3fece4451a23ac1f1f31445d

Observation bb69a5f8-d580-49f2-ac6d-90bc26e5857e · outbound

This paper cites arXiv preprint arXiv:2512.11464 (2025).

Demystifying Video Reasoning arXiv preprint arXiv:2512.11464 (2025)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:55.981219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:55.981219Z digest=sha256:ae0c2acc4d3a665f298a0f4154e54bbbbc4b7cf420f419fe67ff6e600c9eb3d9

Observation 5584e8f9-1c0f-4618-9749-6390c1de95e7 · outbound

This paper cites an unresolved cited work.

Demystifying Video Reasoning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:56.134175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:56.134175Z digest=sha256:67d2a44359b8d83abb326fb3b778a0f48f8437620d9be860256d94cca67474a8

Observation 264f299b-d416-4343-aaba-85d72ae9d937 · outbound

This paper cites an unresolved cited work.

Demystifying Video Reasoning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:56.279752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:56.279752Z digest=sha256:2551331dd4d871f88f3af6dd41f1db83450aaaa1ef6301e72e5f4bec73d960b4

Observation e63bac1c-283b-44a5-b421-70e161b0795b · outbound

This paper cites arXiv preprint arXiv:2511.16668 (2025).

Demystifying Video Reasoning arXiv preprint arXiv:2511.16668 (2025)

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:56.359412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:56.359412Z digest=sha256:0c6118cb1e0a706119343dc70e4c201af5824c9e5c8fe60b0ca98e3ac4c7e598

Observation 89997916-1492-496e-ae4e-ab84913000cd · outbound

This paper cites an unresolved cited work.

Demystifying Video Reasoning Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:56.495190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:56.495190Z digest=sha256:7b01801994feacfa9831dc2197775c01bf31d8c326494f0ec4b0dd6171b20ab0

Observation 1163cc56-2a48-4352-904f-9974574aa65f · outbound

This paper cites Neuron 110(6), 914–934 (2022).

Demystifying Video Reasoning Neuron 110(6), 914–934 (2022)

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:56.656117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:56.656117Z digest=sha256:456d1c89afb083bd01e0b83fca311ec6d26aea469f079098abafb0d4c0e6feac

Observation f6c8bd32-0cdd-45c3-84d6-221d60f0de36 · outbound

This paper cites an unresolved cited work.

Demystifying Video Reasoning Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:56.760929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:56.760929Z digest=sha256:db19feb140c46a0db4e36a5aca19d4ddcde0605c847311c3bdf9ed5baf1e338b

Observation 2947cdd5-766c-4992-b2cb-55f7e3e03625 · outbound

This paper cites Transfer between Modalities with MetaQueries.

Demystifying Video Reasoning Transfer between Modalities with MetaQueries

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:56.874814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:56.874814Z digest=sha256:10b9ee9389495037cdbc22f26f1daf7af68c92a80d368b0aae4cc29312b5c1ce

Observation 5b3b89fb-608a-4898-9555-3de5ba8fe668 · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

Demystifying Video Reasoning In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:56.976042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:56.976042Z digest=sha256:e3035a2277b2d59d4f9b16dbfb7b1bc14877354171b8a1a32dfb7b9bfbe89453

Observation df04f4f8-7114-49d9-b9a6-f30f0b6197b6 · outbound

This paper cites Nature 497(7447), 74–79 (2013).

Demystifying Video Reasoning Nature 497(7447), 74–79 (2013)

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.091383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.091383Z digest=sha256:37d3f3fda98b65a7b2f71fc3317860b6422426a0bbd2a2864c77c751bcf3cbec

Observation 12acc4b0-9372-4e7f-8a69-0d04b3b6be29 · outbound

This paper cites arXiv preprint arXiv:2508.05606 (2025).

Demystifying Video Reasoning arXiv preprint arXiv:2508.05606 (2025)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.214974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.214974Z digest=sha256:17278404ac98bc95233b7c387e5d3708f8ad19dc20aa21b282401cca35a35a9b

Observation 37ff14c8-08d2-4a33-988a-48a3b7e55eba · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Demystifying Video Reasoning In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.366863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.366863Z digest=sha256:7f7223057171b068cca1c467ac5154f20d2b356c1889b5d62bfd9fe3be303558

Observation 3fb10181-f827-4d19-a156-0722e0220711 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

Demystifying Video Reasoning In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.492008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.492008Z digest=sha256:f5d99111d72085ff52d606b16a4a7fd0dafb7d98f0bacf013aac0d834bbf0fb7

Observation 8c5a1038-46f4-4a9f-9d37-48d223a30356 · outbound

This paper cites an unresolved cited work.

Demystifying Video Reasoning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.616115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.616115Z digest=sha256:696d8b4d2babfa8dc4937f7e39b1b101918090218fb29b5cc19ad5119a17a7cf

Observation 83fc4cf3-1d08-439c-acfc-809263045dce · outbound

This paper cites arXiv preprint arXiv:2509.24791 (2025).

Demystifying Video Reasoning arXiv preprint arXiv:2509.24791 (2025)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.733051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.733051Z digest=sha256:fc5da2d9e7e0cf226a98ca1133055e38a29db525791eb1520ab34bae3b22b10f

Observation 8f54f030-7751-43f7-9d1a-645e1db41827 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

Demystifying Video Reasoning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.876583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.876583Z digest=sha256:4dd7c0475bfa3b83d11ccd853dfcf241c7a739cdc3b2f5145a4b6f2ec6289698

Observation 42abb59e-6135-4e0b-97bd-1be2c051949d · outbound

This paper cites arXiv preprint arXiv:2510.14958 (2025).

Demystifying Video Reasoning arXiv preprint arXiv:2510.14958 (2025)

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.975767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.975767Z digest=sha256:dba5ff900a9e088bafe06f90a44008bf919e41876321e32632d3ff940b3dbd9b

Observation d11b36a1-2c3f-4a61-af49-4c4ce76fc6da · outbound

This paper cites an unresolved cited work.

Demystifying Video Reasoning Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:58.119826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:58.119826Z digest=sha256:801254e2013af7e340540424ded6f115a73ec6373970d5b829837fee4c01a9f6

Observation a1a96eba-cd25-4557-b2ee-d3583052f493 · outbound

This paper cites arXiv preprint arXiv:2507.06119 (2025).

Demystifying Video Reasoning arXiv preprint arXiv:2507.06119 (2025)

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:58.225199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:58.225199Z digest=sha256:e786ac831911320a1715a48fcdf0ef6d5f4d926e3506f6264e312b948c9ebb5a

Observation 670869ea-7046-448a-a5a5-ed58481567aa · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Demystifying Video Reasoning Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:58.341491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:58.341491Z digest=sha256:f7899c339deebb5783ed67d6335724a912cfa9e5a41bb3c2f7dcdd2d7099d231

Observation f471e0dc-1dea-4e03-9dc0-e2fe01729b3d · outbound

This paper cites Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm.

Demystifying Video Reasoning Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:58.465302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:58.465302Z digest=sha256:4d243d9aa1e1727998bba669e65a87f0e125d636fd08d4fe9a3162e7098df413

Observation 25c15500-b632-4d7c-ac3a-9bdf58c806f8 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

Demystifying Video Reasoning In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:58.622416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:58.622416Z digest=sha256:cbc40eb31cb57a2993eb14b9eedbf571e07092bf0c4c60c09452c019fcfe3dd0

Observation d9051897-54b9-4e15-a5ba-b5f46a8e0f90 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Demystifying Video Reasoning Wan: Open and Advanced Large-Scale Video Generative Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:58.782463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:58.782463Z digest=sha256:f094a554453c7132c35fc909dc3240285063091b4813066e266a69647bb4d7db

Observation 11eb632c-a796-451b-8407-2cc90c0e87b7 · outbound

This paper cites arXiv preprint arXiv:2602.20159 (2026), https://arxiv.org/abs/2602.20159.

Demystifying Video Reasoning arXiv preprint arXiv:2602.20159 (2026), https://arxiv.org/abs/2602.20159

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:58.952437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:58.952437Z digest=sha256:f1fcec9ee64a3a356e2beea7caa2d2a3fce618ceb64c2d7199a63b599a8f95cb

Observation af1c5d49-5e6d-4a16-9327-8967dda424b5 · outbound

This paper cites International Journal of Computer Vision 133(5), 3059–3078 (2025).

Demystifying Video Reasoning International Journal of Computer Vision 133(5), 3059–3078 (2025)

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:59.115726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:59.115726Z digest=sha256:dda6b258317a7d678562265d46c34d5d3d3fc476971918a96d4485f38c0e5020

Observation 1102ae90-26a7-48bd-b9d1-0cf1f919283e · outbound

This paper cites Emergent Abilities of Large Language Models.

Demystifying Video Reasoning Emergent Abilities of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:59.253418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:59.253418Z digest=sha256:4a8b5d2478a8f87139c633ea49f64a3cbf595a04041b69e3d080c1dc8b7eb4ac

Observation 4d1ab11c-f965-4e39-8430-6c37eb09ee6e · outbound

This paper cites Advances in neural information processing systems 35, 24824–24837 (2022).

Demystifying Video Reasoning Advances in neural information processing systems 35, 24824–24837 (2022)

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:59.385565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:59.385565Z digest=sha256:2e9904584cb061cf7de46b81bbd3ddfe094a0f632afa788722207a5de7339dd7

Observation 235c9488-a005-442c-aae0-b67a4bbaab26 · outbound

This paper cites Video models are zero-shot learners and reasoners.

Demystifying Video Reasoning Video models are zero-shot learners and reasoners

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:59.549335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:59.549335Z digest=sha256:84b9078fd41d9618dd37ba8100bac78c350caab2a0a8d865964db2b20aee9fe8

Observation 45021f51-674c-4b4d-8f3d-d5b1ba22c13d · outbound

This paper cites In: International conference on machine learning.

Demystifying Video Reasoning In: International conference on machine learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:59.709374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:59.709374Z digest=sha256:c4875ad225e9e83b99a96a28c0e4b992133bf4fe78112976026ca2cfc188f96d

Observation 187208af-d9ec-47d9-9672-b487604587ea · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Demystifying Video Reasoning Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:59.819309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:59.819309Z digest=sha256:493eaa9c930a6023d50846410cc86f96e337926e91f36e17d731255c9132bda0

Observation 9df42cb4-f0c1-4cdf-9ea2-3e0d77a4cf38 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Demystifying Video Reasoning OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:59.925218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:59.925218Z digest=sha256:519e4ac7944be6546857c67ecd696f01d6b229dff05cc8f297fff654148d5098

Observation 12d425a0-b7e2-4e3f-8cf5-3146297b1b50 · outbound

This paper cites arXiv preprint arXiv:2510.04290 (2025).

Demystifying Video Reasoning arXiv preprint arXiv:2510.04290 (2025)

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:59.981987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:59.981987Z digest=sha256:772fdcdf0847b8f99b890e75a8dff056336b91ac3b650a714869470d4467bee9

Observation a0967b82-5fb2-430d-ae00-a9f2fe1aabca · outbound

This paper cites MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO.

Demystifying Video Reasoning MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.137676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.137676Z digest=sha256:6a2903578ac45a87b552c5cdd5dc56110ebd54e24d9c36f39c058d053cc658b8

Observation bf877170-9014-46ea-b23a-59efc1ded141 · outbound

This paper cites an unresolved cited work.

Demystifying Video Reasoning Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.247038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.247038Z digest=sha256:15a9d6f0a598924102495d2b3ce59f720a83c30494eaf9b5164515ba2da5224c

Observation 7cbc9ef9-354b-4349-8e04-04e7a71963fc · outbound

This paper cites Understanding Aha Moments: from External Observations to Internal Mechanisms.

Demystifying Video Reasoning Understanding Aha Moments: from External Observations to Internal Mechanisms

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.332846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.332846Z digest=sha256:5c51b4ed2f734dbf67fcd2a3cb6db32cb48fa43dca327bdf232882280e14aee4

Observation 4a963288-41e8-4616-ab0d-fe9e3d9022d2 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Demystifying Video Reasoning CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.443293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.443293Z digest=sha256:5671c9c4db2791f3c2f1e3b8609add1076a7ddad6762eab294ff852c65a08f6d

Observation cf894df8-e668-4f5a-9d44-adf6bd97ecc9 · outbound

This paper cites In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S.

Demystifying Video Reasoning In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.579291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.579291Z digest=sha256:c530b54a45d37db865c8606c81bb7d2de866869de04e01bdb39f0040bbefc588

Observation 18660236-bce0-4e81-8c7b-69634b7a9b28 · outbound

This paper cites In: International Conference on Learning Representations (ICLR) (2023).

Demystifying Video Reasoning In: International Conference on Learning Representations (ICLR) (2023)

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.709082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.709082Z digest=sha256:a9c6f92d03bebffd04282e81bec86d83b65db1020773eeefa27d177ff7108b97

Observation 5e93bda6-b364-42ba-8712-53a4c8278e1f · outbound

This paper cites arXiv preprint arXiv:2511.08585 (2025).

Demystifying Video Reasoning arXiv preprint arXiv:2511.08585 (2025)

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.817156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.817156Z digest=sha256:00549f452cf5bc537834f8821754e1179ee63d137dc335be582464c7fe886744

Observation 1fd5c3ea-c979-4587-a900-29ed70942d0a · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Demystifying Video Reasoning Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.921622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.921622Z digest=sha256:a695953fe74f331d70504c9ce8d175db2bb421a702e26232f47e3c1c5cbd9db3

Observation 10b44e85-d8ec-4886-aa4e-cb42eca79b8c · outbound

This paper cites FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving.

Demystifying Video Reasoning FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:01.085561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:01.085561Z digest=sha256:c8c8bdba7897720a7d8fad5175e3183d0bf78f98d39a98ab700045c207fd7599

Observation 33b67750-7ebb-4e76-9b76-f56f83e82902 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Demystifying Video Reasoning In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:01.178591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:01.178591Z digest=sha256:076d8cd444b4bf72e3f702313458fbaf41e307e4e4b39761c5849df6cd6c9adc

Observation a3d61ab9-1ca5-40dd-8f67-f100d397e96e · outbound

This paper cites Advances in Neural Information Processing Systems 37, 12847–12871 (2024).

Demystifying Video Reasoning Advances in Neural Information Processing Systems 37, 12847–12871 (2024)

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:01.245035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:01.245035Z digest=sha256:252c9e7bbd25cdb82df07e945d1be96a6bb5885e0effe33e00dae38c485f35ea

Observation 0871f3ab-e617-4dea-8835-4fdf9977c74b · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

Demystifying Video Reasoning VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:01.381315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:01.381315Z digest=sha256:08210934eec9b70e7debe40f0de160df36012f4cd77a050aaf8abd91be3575a1

Observation 62405ab9-1e9b-4b96-8ca9-86eeba6799b6 · outbound

This paper cites an unresolved cited work.

Demystifying Video Reasoning Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:01.488539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:01.488539Z digest=sha256:33fd6c7039667ca70fb2a3e0b29ff509cebf66712ef53863b07991d04d478ac8

Observation e988bc26-b644-4d96-aa89-0e5d9a344453 · outbound

This paper cites VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model.

Demystifying Video Reasoning VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:01.603385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:01.603385Z digest=sha256:760f8c6877d30a9f14cd2148f8a82edc111f77453a8c4862bc76eeb7078d910a

Observation db262d51-c76b-4fd2-b7c1-d22937148fd1 · outbound

This paper cites Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark.

Demystifying Video Reasoning Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:01.740152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:01.740152Z digest=sha256:955e12952ad751702052a09069924f6c9de3463252a05211b2dc88f1b25fa4fd

Pith citing papers

Observation 41a988b5-909d-4a01-87d5-ba45aa6f9e3f · inbound

Do multimodal models imagine electric sheep? cites this paper.

Do multimodal models imagine electric sheep? Demystifying Video Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-27T02:05:13.880714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:24:01.933339Z digest=sha256:7be922ca82b3e5909cea61c1bcebd5907d9a8246f9544dc6b55d278b8729a51a

Observation ae335b51-2285-4a55-8cf5-7468b845d319 · inbound

Video Models Can Reason with Verifiable Rewards cites this paper.

Video Models Can Reason with Verifiable Rewards Demystifying Video Reasoning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-27T02:05:13.880714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T15:03:14.894952Z digest=sha256:23cfade9473a4223462ffd9f8ba271e6d7ecd4be63a77a2d60f74598ee12f423

Observation bdde8ebd-63b9-4e90-a23c-a3ec92fa1b4e · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization Demystifying Video Reasoning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.262860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:26:21.284810Z digest=sha256:3dee4d6d5ef5f69cba20ceb62ada1c75ce96523c4cb84c33cf85a6cad0caf5b0

Observation 522a6245-66ba-4160-b948-fb59ee29d99c · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization Demystifying Video Reasoning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:44:36.848444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T10:38:22.619277Z digest=sha256:140f316a6f4d9da526ccfb1ffda49c0fbc10ace1e6e4ab5474d528a97ad8ce2a

Observation 92abe9cd-18cb-4901-9e05-57beaba17782 · inbound

Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction cites this paper.

Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction Demystifying Video Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:24:28.926512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:24:28.926512Z digest=sha256:66f4e104cb03db6b6837bc97ede9dd2f65d607c0ef2285601aaa92c2a7a1cce4

Observation baafe26f-2357-4050-b7fb-99c5e9acca6e · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Demystifying Video Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:03.269966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:03.269966Z digest=sha256:a2f337d2b0b20afb4bd6a5353d2f1c25dbba8dc19553e0c29943906a022b4d1e

Observation a1e35699-3f7b-4289-b025-7f4af4bed506 · inbound

Visual prompt engineering for video models cites this paper.

Visual prompt engineering for video models Demystifying Video Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T02:13:11.772187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:13:11.772187Z digest=sha256:36f892a6ecf2c7fa3d6c63de8d75445b199960579f10d4cface5db7662322ab2