Pith. sign in

Paper Citation Record · LEDGER

Thinking in Video: Can Video Generators Really Reason About the Real World?

As of 6 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2607.17523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17523 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:47:13.750039Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3b7b776-f939-4895-a031-237815ada913 · outbound

This paper cites Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:04.983265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:04.983265Z digest=sha256:27e2055596cce2c246b7a42c7bd9fcde3877562f7b63243b7e5293d32695cfe2

Observation b268883f-6059-47e9-814b-78ba70662e6f · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Thinking in Video: Can Video Generators Really Reason About the Real World? Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.062133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.062133Z digest=sha256:8c8ffff23b76c43fa5f3bd07fe5e9e83967f728edef0acb95ea98544f089d2b5

Observation 3362a094-430d-475b-b4a9-73d0ce640b1d · outbound

This paper cites Sora 2.https://openai.com/index/sora-2, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Sora 2.https://openai.com/index/sora-2, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.183713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.183713Z digest=sha256:e02231a515e9a35e4f241a227f8e153ead5debfa0a1f91e4cd3854fcd6b01785

Observation 92675ba5-d455-41dd-a856-07ad00e3b75d · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

Thinking in Video: Can Video Generators Really Reason About the Real World? HunyuanVideo 1.5 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.356951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.356951Z digest=sha256:3e20ba81f7c18dfbb52a650ae4fd63a7f9d03e1606f273863a85a8eb88cc73a7

Observation 6c2d970b-d489-4ddd-ba65-e5a47f9ce6df · outbound

This paper cites Veo 3.https://aistudio.google.com/models/veo-3, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Veo 3.https://aistudio.google.com/models/veo-3, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.530332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.530332Z digest=sha256:eb5c516bc4a89c90360ca30fd0b07b45492c2f940930a6baf278fdb079e114e5

Observation ea11f366-8280-42aa-bc23-0f5cda6c1acc · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Wan: Open and Advanced Large-Scale Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.639163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.639163Z digest=sha256:2446c8a7315f2e23e5a4aab6d6bf7283862a6f5278306f3c12a44afb4f51b0ba

Observation 41acdf30-f207-439f-a381-4e6b87a1b8b2 · outbound

This paper cites From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI.

Thinking in Video: Can Video Generators Really Reason About the Real World? From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.751020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.751020Z digest=sha256:15737f12eca073bef8c25f13c2cc7c62424a810b8e8f9c3357b2d6fc18641f15

Observation 6f71d0e0-f3e8-46bc-a99b-beb811933f59 · outbound

This paper cites World Simulation with Video Foundation Models for Physical AI.

Thinking in Video: Can Video Generators Really Reason About the Real World? World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.927225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.927225Z digest=sha256:b2a05d374c9b8f5781f50cb78176c4eb1407855d3d01cd363e42afc6e1dc284f

Observation 183adc3c-406c-406f-9150-098284b56556 · outbound

This paper cites Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.134736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.134736Z digest=sha256:6bcb5ccf25252a214b436cf2fc35b6e1443a4feacaec1a99b8b5555d4564bd38

Observation 5920ead6-c2f8-48fb-9f7d-3462801e86f6 · outbound

This paper cites VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.272600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.272600Z digest=sha256:6bd8fd117005d012a84743bce6ba4e466934ccb5e17d31dda9a4f01b70d12b9a

Observation 28b0cf78-ac9b-4bb2-842d-17a99a13261c · outbound

This paper cites The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking.

Thinking in Video: Can Video Generators Really Reason About the Real World? The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.472741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.472741Z digest=sha256:b3fd75fb2fb0d066e36643d9b38af1050a1630d30908f5eca6ced8a8e3efe04f

Observation bcf34f7e-8d07-4f11-b174-56cac699311d · outbound

This paper cites Think- ing with video: Video generation as a promising multimodal reasoning paradigm.

Thinking in Video: Can Video Generators Really Reason About the Real World? Think- ing with video: Video generation as a promising multimodal reasoning paradigm

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.688624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.688624Z digest=sha256:3412f63edfaf25bf39844fdb7858bb81e9de04ce0c4cf625c98164a40d4281ea

Observation 79b302f7-46fa-476c-9608-8aaf018b30bb · outbound

This paper cites Reasoning via video: The first evaluation of video models’ reasoning abilities through maze-solving tasks.arXiv preprint arXiv:2511.15065, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Reasoning via video: The first evaluation of video models’ reasoning abilities through maze-solving tasks.arXiv preprint arXiv:2511.15065, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.902542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.902542Z digest=sha256:77480901ac8f41684aa6dbde264b6fa31f8721aeaa39c6639ba83e75c2134e5e

Observation 2849c63f-b004-45be-a2a4-aea3c427aa1b · outbound

This paper cites Weave: Unleashing and benchmarking the in-context interleaved comprehension and generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? Weave: Unleashing and benchmarking the in-context interleaved comprehension and generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.063535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.063535Z digest=sha256:142b3da0f5724b88e0fc1a38453c0cf21993189c32318a9f737cd2c95c51549a

Observation eb068e54-0837-4871-aaad-0c0c9c255731 · outbound

This paper cites Tivibench: Benchmarking think-in-video reasoning for video generative models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Tivibench: Benchmarking think-in-video reasoning for video generative models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.188868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.188868Z digest=sha256:68d651cf1b3f4b15614d99610e56246f43cd699686f916f520e79090bb3db0e4

Observation df8b54b6-e33e-4890-b847-28dae3b270fa · outbound

This paper cites Evaluating gemini robotics policies in a veo world simulator.arXiv preprint arXiv:2512.10675, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Evaluating gemini robotics policies in a veo world simulator.arXiv preprint arXiv:2512.10675, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.292420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.292420Z digest=sha256:296003fe8ec58136248d7269c29acca3f7644cfc79062dd811142a11cfb07ccb

Observation 340e3464-76b5-4ebb-8910-eeda2658bb8a · outbound

This paper cites Ggbench: A geometric generative reasoning benchmark for unified multimodal models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Ggbench: A geometric generative reasoning benchmark for unified multimodal models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.399742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.399742Z digest=sha256:6783135ba71516815877e2584ae0ce3931aa7cb26545081a3d79cd744a833935

Observation c588afa7-58f2-49db-bfbf-e07be89a9d08 · outbound

This paper cites Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts.

Thinking in Video: Can Video Generators Really Reason About the Real World? Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.550254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.550254Z digest=sha256:05a9fde621e64563b884f4088d62635858f86a915e852a37c40dcb1828284007

Observation 7c54f164-3993-4179-9b48-475a6b2ee650 · outbound

This paper cites Physgen: Rigid-body physics-grounded image-to-video generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? Physgen: Rigid-body physics-grounded image-to-video generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.733004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.733004Z digest=sha256:36a7ffd21c231859081ff2d1fc70c9ae8d6a7c4fb20b5e3499f5090aa9d460b5

Observation 4e7da424-a68d-4c27-8ee9-6419c9091ffb · outbound

This paper cites Exploring the Evolution of Physics Cognition in Video Generation: A Survey.

Thinking in Video: Can Video Generators Really Reason About the Real World? Exploring the Evolution of Physics Cognition in Video Generation: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.906970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.906970Z digest=sha256:f3ebf91165c14a7f19b6c325eba8b91e130b46b87350e9ab7168172a3b62eaea

Observation 2d7942be-06a5-4a8f-8624-74c3ee6f3621 · outbound

This paper cites Improved techniques for training gans.Advances in neural information processing systems, 29, 2016.

Thinking in Video: Can Video Generators Really Reason About the Real World? Improved techniques for training gans.Advances in neural information processing systems, 29, 2016

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.045401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.045401Z digest=sha256:dfee4ca79c8c593e05e2dd8db0ce81198c0760c41881088c77d5809933e0c92b

Observation 4384fc51-e1a3-477b-8302-9fe4f393075b · outbound

This paper cites FVD: A new metric for video generation,.

Thinking in Video: Can Video Generators Really Reason About the Real World? FVD: A new metric for video generation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.215955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.215955Z digest=sha256:b2e61c447536e2f981a99ba2f5a13d60ee1e93ba69c14712674618a14c35231d

Observation 4bbd32dc-1b98-4c46-bbe5-53e1ca6f2694 · outbound

This paper cites On the content bias in fréchet video distance.

Thinking in Video: Can Video Generators Really Reason About the Real World? On the content bias in fréchet video distance

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.494735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.494735Z digest=sha256:ee4d5708475b07654bfb2a38784cb8ef068dc01c8fac5f2f2b8f7339d18477df

Observation 0993dd37-f5d0-4ae7-a738-2c0c41d2a414 · outbound

This paper cites The Essential Role of Causality in Foundation World Models for Embodied AI.

Thinking in Video: Can Video Generators Really Reason About the Real World? The Essential Role of Causality in Foundation World Models for Embodied AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.632064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.632064Z digest=sha256:d587833ffad8f7d4fe8af2ed800c15e4870aa6aae06100a39d9952e129b8f9dd

Observation f4d139b7-71ad-4e7d-a290-2d3d1a41e8fd · outbound

This paper cites Diffusion art or digital forgery? investigating data replication in diffusion models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Diffusion art or digital forgery? investigating data replication in diffusion models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.796963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.796963Z digest=sha256:876b81221d05cfeb11c1e3346c8cb7aa1dc32175bb43a6273cd70093a4d5e460

Observation 8d3fbe95-ac2c-4864-9de3-afdfa3d18455 · outbound

This paper cites an unresolved cited work.

Thinking in Video: Can Video Generators Really Reason About the Real World? Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.955102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.955102Z digest=sha256:7b35c59f08d23794519b0794da7b1b010d1e7fb91babe59d45b31421157d497f

Observation 346eedb5-5d52-47ae-8da8-8821aa4496aa · outbound

This paper cites Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.107864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.107864Z digest=sha256:d79c677310f0a9580e4984d8c6286afe3117120c0c2f22d362f393310ed03a44

Observation 681bab52-da7f-4f80-8208-920d355eba84 · outbound

This paper cites RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models.

Thinking in Video: Can Video Generators Really Reason About the Real World? RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.268892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.268892Z digest=sha256:54d05ba978cb2bb1f0e572e806de57464374fce26eba9faa9a154ebd18dc8577

Observation 7cfdd990-47cb-419e-9141-1f8aba064fdb · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Vbench: Comprehensive benchmark suite for video generative models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.423659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.423659Z digest=sha256:72b0ebdd7cc77cbeed5ad54eaac52c61d6050ce2e371b421a51a67d26878ac1d

Observation b03da153-1e28-4dde-80d4-d80e60feac85 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Evalcrafter: Benchmarking and evaluating large video generation models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.568545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.568545Z digest=sha256:d184a48f7f9cbfff374c0840a2a3e3f13d21d3c36426d47ac90b4dbc26eb4b92

Observation d16d4994-edc7-486f-8ca0-c70b18120e6d · outbound

This paper cites Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024.

Thinking in Video: Can Video Generators Really Reason About the Real World? Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.717609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.717609Z digest=sha256:c895f0299a0405c9c2f37c5765531a6821a2f56e1008017cbe8131ae78e42e9a

Observation cfbbbf81-aae1-4bb4-8f0c-b48111c36c11 · outbound

This paper cites What about gravity in video generation? post-training newton’s laws with verifiable rewards.arXiv preprint arXiv:2512.00425, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? What about gravity in video generation? post-training newton’s laws with verifiable rewards.arXiv preprint arXiv:2512.00425, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.883305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.883305Z digest=sha256:b50d32418e5e080000db9c64040f04114fce23a6b762d91109156928874c8c73

Observation f1606849-161b-4dde-92ac-bcd5521a617c · outbound

This paper cites T2v-compbench: A comprehensive benchmark for compositional text-to- video generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? T2v-compbench: A comprehensive benchmark for compositional text-to- video generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.028648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.028648Z digest=sha256:d1681bf7ca5f4b0fc7aed5bb50ab19620e535f4690872779d526fe310ffeb683

Observation 81b672e5-e006-454e-9c39-3baadfdbd46b · outbound

This paper cites Vibe: A text-to-video benchmark for evaluating hal- lucination in large multimodal models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Vibe: A text-to-video benchmark for evaluating hal- lucination in large multimodal models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.192580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.192580Z digest=sha256:b656609ea7f86d2ab50e41826723d61d32131f36291f96f53a24a55791300607

Observation 6005acd2-3838-4655-b5b2-f94ef9c5749b · outbound

This paper cites World Models.

Thinking in Video: Can Video Generators Really Reason About the Real World? World Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.369130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.369130Z digest=sha256:2f4ff9b7aecd6461a8f538d2a6254eeeb65f89c1a07ea3f029aa7a74dfe4d7e1

Observation 865b148a-0639-4b83-bed5-70275e86a37e · outbound

This paper cites A path towards autonomous machine intelligence version 0.9.

Thinking in Video: Can Video Generators Really Reason About the Real World? A path towards autonomous machine intelligence version 0.9

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.488113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.488113Z digest=sha256:054237912301b485ade4ef25b6f9b6949795a998a92e523e1d9d08b686b012ec

Observation 1fe3a9cc-1e28-4c26-bd81-91db74b9caa6 · outbound

This paper cites Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability.

Thinking in Video: Can Video Generators Really Reason About the Real World? Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.650083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.650083Z digest=sha256:d27e7b6f3928fafebe42ebd8400dc7b55d89d26ee9883658b1c92208fb7fb841

Observation bdc30d75-403e-4271-bb63-5728bbb293b0 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Thinking in Video: Can Video Generators Really Reason About the Real World? Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.778104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.778104Z digest=sha256:1d03bca4286e617a6fe1568aedc7167362f764557f29f7dad5ffa2e21cc1c0d7

Observation 6d7ff8a3-ca4f-4be4-a1bd-f9fb1fdcd440 · outbound

This paper cites The sound of water: Inferring physical properties from pouring liquids.

Thinking in Video: Can Video Generators Really Reason About the Real World? The sound of water: Inferring physical properties from pouring liquids

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.959513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.959513Z digest=sha256:2409e0325ada2980df81133bc03392b9dc251d99f3af34171577331f50657371

Observation e1672087-e3a2-4a30-8032-89efe9002863 · outbound

This paper cites Physion: Evaluating Physical Prediction from Vision in Humans and Machines.

Thinking in Video: Can Video Generators Really Reason About the Real World? Physion: Evaluating Physical Prediction from Vision in Humans and Machines

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.071407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.071407Z digest=sha256:df790656cb4e6a9b4db8958899195927d43bf68d31e49e0e564b4609b55508cc

Observation 56746adb-4ee7-424c-b99f-fe235f2ce35d · outbound

This paper cites TLD: A Vehicle Tail Light signal Dataset and Benchmark.

Thinking in Video: Can Video Generators Really Reason About the Real World? TLD: A Vehicle Tail Light signal Dataset and Benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.177637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.177637Z digest=sha256:0ef49dea6ea7c7385551163fdec93572d2819d4b264ea42e4b1896c4ca56a922

Observation 3119fe02-6076-421d-8623-8fb47675dd61 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Thinking in Video: Can Video Generators Really Reason About the Real World? The Kinetics Human Action Video Dataset

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.299311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.299311Z digest=sha256:c462f72a30e64166620c729c808cde7176d9c931f78c2177b0eee57e0c79ae73

Observation 0f149436-6454-4854-ba04-5fbd4a2ec778 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Thinking in Video: Can Video Generators Really Reason About the Real World? Ego4d: Around the world in 3,000 hours of egocentric video

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.456060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.456060Z digest=sha256:69d700a785141add463b9749d67f9e1004766a8c6cc8aa18cfecce2b2c6a71db

Observation 0c291af9-b8e3-4bc6-b573-09e9e32aa427 · outbound

This paper cites Gemma 3 Technical Report.

Thinking in Video: Can Video Generators Really Reason About the Real World? Gemma 3 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.624753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.624753Z digest=sha256:12f36e0df8f448cb6b5087b894fe7e801cd130277d9c6f073c3a05980f73af41

Observation 7c66c50e-c190-47bc-94e1-ae62a2776045 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Thinking in Video: Can Video Generators Really Reason About the Real World? Robust speech recognition via large-scale weak supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.737381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.737381Z digest=sha256:243e584ec8b3ed8265b4771dc00cf047943e193f0563710c449cd9fac7416758

Observation bb04f007-3916-4d36-9401-c8711314876a · outbound

This paper cites Qwen3-VL Technical Report.

Thinking in Video: Can Video Generators Really Reason About the Real World? Qwen3-VL Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.848598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.848598Z digest=sha256:fc4c604daae67080f49a9c74a4fc72c494ae113011d2fd1888655d8df1dab5a1

Observation 289c7a03-9513-4594-ae01-e88413b45fc5 · outbound

This paper cites OpenAI GPT-5 System Card.

Thinking in Video: Can Video Generators Really Reason About the Real World? OpenAI GPT-5 System Card

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.966954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.966954Z digest=sha256:3d3d1b41e67c0ef6cdf906ad5b5790cceb732734ce2cd676f309609ca474a63a

Observation a6f56fdd-2510-4055-8df2-90eceae09a42 · outbound

This paper cites Scaling zero-shot reference-to-video generation.arXiv preprint, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Scaling zero-shot reference-to-video generation.arXiv preprint, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.088675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.088675Z digest=sha256:9a7b94a29ad9aacebf88bbf2337915ef85543a8a33684365ce644e1596818bac

Observation 4917b50f-6f35-4b29-a0d0-58d904216787 · outbound

This paper cites Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation.

Thinking in Video: Can Video Generators Really Reason About the Real World? Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.237328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.237328Z digest=sha256:dd980bbf331c0b87ee2cb5457f3a9974ba744ce0445b4dc274c10a5555a01a8c

Observation f2523d10-3c97-4cc1-bb65-31af0cb879ce · outbound

This paper cites Stable video infinity: Infinite-length video generation with error recycling.arXiv preprint arXiv:2510.09212, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Stable video infinity: Infinite-length video generation with error recycling.arXiv preprint arXiv:2510.09212, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.350191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.350191Z digest=sha256:8dfbf42b298e33f49cfd3b366d23a93888d96b5410c5651cfae7cec21312b7db

Observation 5f57e935-e817-4ec8-8ad7-79753866fcaa · outbound

This paper cites LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling.

Thinking in Video: Can Video Generators Really Reason About the Real World? LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.458101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.458101Z digest=sha256:0290fc4e1ec9ad51610689f8ca91e949de8295605f1ec62489d9c7227a7dda7d

Observation 349359b0-731f-4220-8f57-5b05ee3c02d4 · outbound

This paper cites V-reasonbench: Toward unified reasoning benchmark suite for video generation models, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? V-reasonbench: Toward unified reasoning benchmark suite for video generation models, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.611214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.611214Z digest=sha256:5641d458d6f92cd21dd3cd4eb1ff24fe5885edbb0c4bb8520e583880acd8f7f7

Observation 27ab5293-97ed-488f-9cb7-7bcc0a468d3a · outbound

This paper cites Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.825883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.825883Z digest=sha256:0a71c7912d0a8b0d6211c77bf72152cb3935d158df2ff7437ae9694cbde2ef7c

Observation 2fc1c363-0dc2-4bfa-a372-cdecc44e767f · outbound

This paper cites Ruler-bench: Probing rule-based reasoning abilities of next-level video generation models for vision foundation intelligence, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Ruler-bench: Probing rule-based reasoning abilities of next-level video generation models for vision foundation intelligence, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.956508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.956508Z digest=sha256:c31d9348346dd1a7d522720032605a1e6fc41e4974d2a7c2f71a1566ede3f39a

Observation 542dec33-ec6a-4828-b0a5-1c969dd8ed89 · outbound

This paper cites Can world simulators reason? gen-vire: A generative visual reasoning benchmark,.

Thinking in Video: Can Video Generators Really Reason About the Real World? Can world simulators reason? gen-vire: A generative visual reasoning benchmark,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.071487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.071487Z digest=sha256:9fcb06cb5b6a0e891e1e870ac4e66f6d521fbafddae52f1711061f616f673d32

Observation 8d8fd854-c9e1-49ed-a6c3-6e3595abf3dc · outbound

This paper cites CCHall: A novel benchmark for joint cross-lingual and cross- modal hallucinations detection in large language models.

Thinking in Video: Can Video Generators Really Reason About the Real World? CCHall: A novel benchmark for joint cross-lingual and cross- modal hallucinations detection in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.292641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.292641Z digest=sha256:910a60adad01aee4e2b504f78ba26651b78624fbcfac2b8fcb7940b31ca18d6f

Observation 6d0f5da2-5c50-4887-a9a1-c7fd7163c42a · outbound

This paper cites Large language models meet nlp: A survey.Frontiers of Computer Science, 20(11):2011361, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Large language models meet nlp: A survey.Frontiers of Computer Science, 20(11):2011361, 2026

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.421131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.421131Z digest=sha256:77eecc1841e3e0f528cf56234c974fb397ab03eba324a29428f2a5cc005dbe8f

Observation dcc1aacd-a901-45d8-a871-ea3b0a1777c7 · outbound

This paper cites Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark.arXiv preprint arXiv:2510.26802, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark.arXiv preprint arXiv:2510.26802, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.478604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.478604Z digest=sha256:c36c3713302d0a25e81249a4be3fe534ae7126f965b4afd808df3dad81b975a4

Observation ba003690-3d5e-4339-81a0-d116a936937c · outbound

This paper cites ISBN 979-8-89176-251-0.

Thinking in Video: Can Video Generators Really Reason About the Real World? ISBN 979-8-89176-251-0

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.365314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.365314Z digest=sha256:7c8a28d5ab1c73badc7c816a652d8c06e131da0ae8b4346b96b7bf0248ec1ddd

Observation 30cf12d7-beed-4d12-9caa-a1ba8ba98eb2 · outbound

This paper cites Latent Visual Cache for Video Reasoning.

Thinking in Video: Can Video Generators Really Reason About the Real World? Latent Visual Cache for Video Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.597841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.597841Z digest=sha256:76c593d61f4d4a40346c079c1bf42d87890e31235b8d7939ae77a78e892b1071

Observation e985d63c-cca1-4f9b-9f41-16b1dcd9e458 · outbound

This paper cites Mmgr: Multi-modal generative reasoning.arXiv preprint arXiv:2512.14691, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Mmgr: Multi-modal generative reasoning.arXiv preprint arXiv:2512.14691, 2025

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.651779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.651779Z digest=sha256:1938f5e13f1381a89cf92f7fd0dc92848a07b74591c1462eab56db57e2378b78

Observation 38e2166c-e519-4414-a023-a8ccac217909 · outbound

This paper cites Video models start to solve chess, maze, sudoku, mental rotation, and raven’matrices.arXiv preprint arXiv:2512.05969, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Video models start to solve chess, maze, sudoku, mental rotation, and raven’matrices.arXiv preprint arXiv:2512.05969, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.540223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.540223Z digest=sha256:747a551bece9ab21745f02779ec1d9d4e82ff9d5b1b9e8d2391faa384ba8c43c

Observation 307fd135-8114-41b9-a092-782e9b338d94 · outbound

This paper cites Beyond the last frame: Process-aware evaluation for generative video reasoning, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Beyond the last frame: Process-aware evaluation for generative video reasoning, 2026

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.750039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.750039Z digest=sha256:9074fc8a7a4c50ac4300fcc9e8000f0ef959af8c3f083ec683af41cfba8fc784

Observation 605f4c84-aa7a-48b6-9c5d-73b0fd3f4a94 · outbound

This paper cites an unresolved cited work.

Thinking in Video: Can Video Generators Really Reason About the Real World? Unresolved cited work

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.331772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.331772Z digest=sha256:f9450c138140ed03f537611af62f7da301af84c80f8a97ff3a64d626cc930bb9

Observation 179ec37e-e83d-4301-9e90-1933f2fb7440 · outbound

This paper cites an unresolved cited work.

Thinking in Video: Can Video Generators Really Reason About the Real World? Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.188710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.188710Z digest=sha256:1e123ef1ce8a39459b691a8ad5622dcff37802df80d185fa6a89da3ec9338a81

Pith citing papers

No inbound Pith citation observations are available.