Pith. sign in

Paper Citation Record · LEDGER

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

As of 6 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 13 inbound Pith citation observations for arXiv:2511.04570.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.04570 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:54:24.641649Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:23:36.183222Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:46:40.976623Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact28
  • verified fuzzy20
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a6210ad-2d6c-4050-91db-2043878113d4 · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm System card: Claude opus 4 & claude sonnet 4

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:32:18.381268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:e7f79efaf2173cabca964ef7b868f883071c00dbf70910e9831c11e42af46172

Observation 24783198-2c0f-4560-aa44-422cd7016f1b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.174854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:ed17f0b2929f637694ebcd9b39435e0ee45bb9b377b5fb8423b96bf37b6f841f

Observation 8d1e9b02-affb-4ddc-9554-3c31e578e769 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.171420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:6b20aaf4c4486361f91adb6216226eeca9f22eb72db0a50f3ede87424caa4d18

Observation 01759ace-d87f-411a-ba98-e7310f835613 · outbound

This paper cites PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:55:35.167332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:e43132a4bdcf45776c9854e25fbeb0795c2ab0802b742554da324b0721800ce8

Observation d50a2f0e-551c-4a75-9b42-45221f4d661d · outbound

This paper cites ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.178851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:78e313ac752bc423db9bab07d317449c8d88ff9c99479cc51dc52ee68fb6f5f4

Observation 797019a0-089f-45b2-b3e0-066747a726d6 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.153536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:edae5f073acce2c45cc2b57d75cfc6e8632cdfb36245f1603e2fc2247ad5ef2e

Observation 1d2ae126-c49f-4a38-a5e6-079127cc34eb · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.148843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:0867fa0bd5fc1b9da16853b4078e2b0f7ba9902cd90d9f0a67fa340817c5e9bd

Observation f6d047db-2a61-435b-af07-280e9e5a2503 · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Emu3.5: Native Multimodal Models are World Learners

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:43bf800d6a45297360947ce8206df95323fb3e6c53908d950a0ea4f77489fd3f

Observation ae65fa33-1014-463c-adb2-793fa5cdd2e0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.144057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:108dceb0436d547037c392f166fa8c0e2da10665ded6dbf7265e6751ad56888f

Observation af316306-02c3-46c5-a372-a213b2a8501a · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Emerging Properties in Unified Multimodal Pretraining

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.127900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:04930d6ea1f9814621338db37f7182f1fb16b5795c01d92f949d566a55f82651

Observation b22ae060-4b90-4ef5-b9fd-aa062bf65f31 · outbound

This paper cites A Survey on In-context Learning.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm A Survey on In-context Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.139311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:5fcbfbb755c3219524d04419b624cd6f7f9af8db941dfb443f2ab17dac1466a4

Observation 03d3891d-1d02-4249-9528-092d4c542964 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.122041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:6de588b449819f37a21e5e611b6ce625fcf19016f62744eb6e3fb763fcde9b12

Observation c7de72b6-4ee3-48d4-aad6-5762bac31158 · outbound

This paper cites Veo 3.https://aistudio.google.com/models/veo-3.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Veo 3.https://aistudio.google.com/models/veo-3

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.468963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:9d3bbbb9c400418528295232a9a000383b580656bbdebc168ad8cb3ac2ed3dbe

Observation e5474820-96a6-42dd-842a-4c82e00e5cfd · outbound

This paper cites Gemini 2.5 flash & 2.5 flash image model card.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Gemini 2.5 flash & 2.5 flash image model card

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.446959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:bc60af8a91c35ce2d17d6afb9d63f597880dc54fbd9ba15999b5e82c9a789dd0

Observation ba2d6648-2481-4cb3-bac8-cc9625714b58 · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:05:35.493339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:47f00b8b6573ff3bd36f4f74ee4bfbb4ff49201e901bac1f6f015708b2e1b379

Observation f6cf9f60-b7b2-409c-8e87-cf770d961d8d · outbound

This paper cites Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:55:35.134168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:2735ef987f2e04861a381e4dc5dbc791c4caf46e698ba85fdedef9043049d7c4

Observation 67898fd8-ab45-406a-ab6d-8d6362d280fa · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Measuring Massive Multitask Language Understanding

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.105346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:16f06f9fa3ccebad3dc38d19e0e13ae9be8d1837dfb913c05c08f0df7e55aadd

Observation b282db82-6de7-4f38-abd8-4a7a0b29f2a9 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Measuring Mathematical Problem Solving With the MATH Dataset

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.109122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:e9ebcd75087eb0b64bd6e72f357f24689975707177e360ec544dd6a4e8f79935

Observation 33abd66d-501a-4660-9dca-bb67f03953cf · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.113178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:31f773ba3e1cddde0d2bee278ba3a50e5778270eb152cec64373e1f4e1efe3f3

Observation 28bb8881-8720-433a-a626-32391014ba2d · outbound

This paper cites Enhancing Advanced Visual Reasoning Ability of Large Language Models.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Enhancing Advanced Visual Reasoning Ability of Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:55:35.117377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:731778818c21b472d0cf8fe98ae25f3b82c19f8ebe0a675efb755f9115f840b7

Observation 67697a59-fd70-4836-8005-c5cddc0cfa60 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computervision, pages 216–233.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computervision, pages 216–233

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:32:18.378628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:59ebb99b0720c837d924684f2e2513ff7dfcd9d250e2bf0747ac4b4f0f00e986

Observation 2d594898-c52e-47af-b569-cfc6d31da99e · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.182626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:3337edbc170ee4952919206265b62f3b86df16399843689684553f1558a16f80

Observation 374e3105-e63e-49be-ab2d-a804c600d60d · outbound

This paper cites GPT-4o System Card.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm GPT-4o System Card

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.158165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:3c0346c04f5788b2b3e295e3394a0e8f329b1cdfa7ebbba0e48ad98a375cfe54

Observation 7d929fe1-a5c8-4e47-a0a1-d331604adebf · outbound

This paper cites Gpt-5 system card.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Gpt-5 system card

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:32:18.375199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:a3495e824d8f23777f9a2c4b23977f5873505dbf7b489a265d07b21ea4a0b1f1

Observation 32784cca-2eaf-4d46-8e4a-75cd43405d74 · outbound

This paper cites Learning to reason with llms.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Learning to reason with llms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:32:18.397886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:6dd3d42791dcf8a37d5eae3c355326a2b7b33fde6583a9e088aa2ecc0fdafb9a

Observation 35b2e7e6-79d7-483c-b10a-d3f7cf646338 · outbound

This paper cites OpenAI o3 and o4-mini System Card.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm OpenAI o3 and o4-mini System Card

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.520694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:72e5b4faf163425f242cde1e7d1ff67ca5ee29d792daf48e26acbeffa393df16

Observation 6ea2b7f8-9b9e-4384-9659-5ac7b35aa2d2 · outbound

This paper cites Improving language understanding by generative pre-training.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Improving language understanding by generative pre-training

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.472735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:c1cfdbaf4a8bbbf656c63ee6d58fa761ba2fbb4b706c4f493650c4ad1c0bed2f

Observation f8d2e659-612c-439e-9c5b-0130df993cef · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Gpqa: A graduate-level google-proof q&a benchmark

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.461522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:b8ca7b0623b2fad63624f55014c903afabc187f3509bbade46b499e8f458382e

Observation d002bcd8-fa71-4753-9ad8-a7ac67e46cd0 · outbound

This paper cites Introducing Gen-3 Alpha: A New Frontier for Video Generation.https://runwayml.com/ research/introducing-gen-3-alpha, June 2024.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Introducing Gen-3 Alpha: A New Frontier for Video Generation.https://runwayml.com/ research/introducing-gen-3-alpha, June 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.503402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:d7510bbe476d42cfa406180de62273d05d3682555705f8e8a7247700ffaf256c

Observation 667298d5-fa3d-498c-a207-4e89d4a69b8c · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.097599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:eb9adee9a7c89c624fffa9b822a52b18d2beef648296ea895cd4d1f4d4a84934

Observation 2f351b37-97b5-4816-b2c7-3da16616038b · outbound

This paper cites Code2logic: Game-code-driven data synthesis for enhancing vlms general reasoning.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Code2logic: Game-code-driven data synthesis for enhancing vlms general reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:55:35.102017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:109a3c50122969ffd5eedee27282efc57f7f843bf62c4b96c9027e1b8023aff5

Observation 8a2e46f6-9f6c-42a0-ba6f-499a75b6b0a0 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Wan: Open and Advanced Large-Scale Video Generative Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.092815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:e140643a42087977572394b3b6cbf86a1903473b3dad06eb541ca7f9d4d64d90

Observation bc26ea48-6c7a-4a7e-a2be-a2538bf22d65 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.Advancesin Neural Information ProcessingSystems, 37:95095–95169.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Measuring multimodal mathematical reasoning with math-vision dataset.Advancesin Neural Information ProcessingSystems, 37:95095–95169

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.513972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:439698f5fe7a0e93a9a7d5c50ed0dd4f4acf984639798b020f3bb6c7b6703ce3

Observation ec7f6cd8-3745-49f9-bb91-c0d1176833f8 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.081325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:fc212ef5aebff6aa77faad444fe7ea8bd1f6b0942a3ab134eb2f7dbf28252507

Observation 26640546-4595-49d2-8227-8db44df323fb · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.490061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:3822ffb54cf46cc184e17c52e8dd28b2d6faf82f3b1bb43bbe18e40bf45fba24

Observation 4b80c960-58df-4b07-9aca-2dc8a0465b32 · outbound

This paper cites Chain-of-thoughtpromptingelicitsreasoninginlargelanguagemodels.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Chain-of-thoughtpromptingelicitsreasoninginlargelanguagemodels

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.517363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:7acaf28f6dc0ad58e94703e851ee7f99de497a17903ed8443212fdb3c5447c10

Observation 9980435e-eb6a-4904-aba3-28082a0df574 · outbound

This paper cites Video models are zero-shot learners and reasoners.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Video models are zero-shot learners and reasoners

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.077440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:3d1372dc647a68fec848cadf2c1d21b6fb9af74ba2a1e2c4f156c87aca2d192c

Observation 9b2cffb6-d8f0-48ed-9deb-35bdb42c302b · outbound

This paper cites Qwen3 Technical Report.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Qwen3 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.088958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:b643b8c0aea5f87f2ec743a6dfecf1ecc9a414286691407ce710ad3193c3f519

Observation 33445499-1eab-4f7e-b8f4-84c94001d3ce · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:29:59.463518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:bd7bcf235a22394d90d544da8341030696d2da1c3172450e3d98ae70562c44e4

Observation 3b10c898-003a-447d-a413-8ccdbb4580b3 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.454507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:1dd24853f563ca4ed9d71a03f793986100b7d84934fffec23cf8779f79bfdd79

Observation 2818f3ee-7a98-4abb-84e0-01309a1319fc · outbound

This paper cites Rewatch-r1: Boosting complex video reasoning in large vision-language models through agentic data synthesis.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Rewatch-r1: Boosting complex video reasoning in large vision-language models through agentic data synthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:32:18.392093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:e6bab864bb384ad8e7aeb938096c7fe61c427d0d88d4894a730e67a013c94254

Observation ff49280c-142a-447b-8902-af0fa6d7ca9a · outbound

This paper cites Rewatch-r1: Boosting complex video reasoning in large vision-language models through agentic data synthesis.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Rewatch-r1: Boosting complex video reasoning in large vision-language models through agentic data synthesis

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:55:35.073729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:f922cae76dff886a4b511d28663b11902118008a0600afc2fc84140fa7ecd491

Observation 558a108b-cffe-40fe-a71d-db26eebee2c9 · outbound

This paper cites ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:55:35.070077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:f74977f5eea8f56dd9b324865117b6b83bde3f475cad8424ed9b710886db068f

Observation 95dbf2cf-d48c-4ec7-bdbc-ebde4de5fcc8 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Multimodal Chain-of-Thought Reasoning in Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.066474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:62716566a2f91cdbbaa1dd847b81744c54d5eaf7de382cf75733a2bcc89db15c

Observation 3002ca2b-7775-4fcb-ba72-258f490afabe · outbound

This paper cites Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:55:35.063035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:231a175d80db069f0d0f814f45395a5ecf335d16aef3f0dd469ee1bbc5ff3537

Observation 8296ca9e-a8a8-42fe-a1db-930756ac1115 · outbound

This paper cites coverage difference.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm coverage difference

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T01:05:35.464918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:808d0f306c9f5d82c7ea324deea1552a8eed80c9534c5e9aaf0f6f4c64a937c0

Observation 5c117c6a-be6a-40fc-9b44-e2b5f517b2c2 · outbound

This paper cites The answer is.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm The answer is

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:32:18.394798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:3b46a40c642d3f2b8d1bda5006078065cda2656f409d7a652d063f891440dc88

Observation f5a2f150-603e-4cf9-8172-d63b039bdc1b · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:32:18.402684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:837731159485a7d7f136dfd34f244b126644ba879b1d7cbf68a475211da457d5

Observation b2518fb1-db59-4b4b-96f2-1f43497b120e · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:32:18.400238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:b650b75283bc115db502c4244d8d2c363a701481eef09aa0f0de1e08aaf740c5

Observation bf7726de-3f47-492d-8390-7ab393eb13f1 · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:05:35.476186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:507bc7ecd07b8d99d6b4e078da5eec0685961384346f65ac7dd6c81abbdbd585

Observation 78c6c0a6-8bcb-4dde-b20c-df2ab079af1d · outbound

This paper cites 4” vs “4.0.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm 4” vs “4.0

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.496576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:fecf3034e565e178f58d25abb1f38db8c18ff5273bb8a4c0ebcf6abc50da32da

Observation 7bd62a01-49e3-4670-b73e-d8b531d5822a · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:05:35.450763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:36932d340f220c811ae17a49db4455a8a68c94243a966819c76245ace1e19f30

Observation 1f3bfedc-c78a-4592-aa0e-e8e7171a56b1 · outbound

This paper cites Your task is to determine if an audio transcript from a solution video contains the correct answer to a given question.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Your task is to determine if an audio transcript from a solution video contains the correct answer to a given question

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.506874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:9608bb95e92728dcedb8b22f012aedb5fdaf331216e0e3dbdf90b7e451f10b7a

Observation a97d2ecb-46f9-48e3-aa32-968e18595bb3 · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:32:18.383694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:f67a0fe495b152247cb1ece73f0160ceffd3da626a920739fc188799edaba1a9

Observation e5025b50-3d1a-4bb7-9701-ae8e44c39ed3 · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:05:35.527297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:51bb92f82f9b0817c1dcfb0bdbdf5c5b836de6e0587f07b1e97c0478e2d60288

Observation 21884392-a320-449e-a8b3-dc1a9a5c9650 · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:05:35.486572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:74b9d8343999c0466d4f48145a423ee8a354cae89d2b9d4ad9edf446663dbc32

Observation 3c55ef35-7654-474b-884b-747217e1e351 · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:05:35.499973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:d58c70babf8b68f52345fa20163e2d3b861c93e08abec7f1f459942ff2f3adf3

Observation 59a965d3-dca6-4c83-a6d0-669c0da2bda4 · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:05:35.457927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:7b3ed4583b983c3db36971ae1a4c097d7b1bee96234222f08d6de3a167c576e3

Observation 43197749-3c3d-437b-9b45-434bd6b8e7c5 · outbound

This paper cites the correct answer is.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm the correct answer is

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.479760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:7cfc307f088999b14a4e46fd46f3fcf361742e38023232f74f39903845300814

Observation 1f83a8ec-a6f3-46f6-9b51-f12fc8f9b502 · outbound

This paper cites A” to “E.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm A” to “E

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:05:35.523988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:ae21930c4ab9099aafc37cafd723eeb4f68434634188fe68324200783298fa13

Observation 92b680eb-d926-4adb-9ed1-b7c3f2cfaa3f · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:05:35.443280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:298e608ade642910ea5346ded6d99422483fd60c91907323f07421718243a0bc

Observation d99680c1-0226-4d58-8356-f9bd6b40a448 · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:05:35.510263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:c4f5eeba5bebb89a46eb9bec4e878321cc6d02f967b99bb1926960e45cc10ccf

Observation 142902e0-7aa7-47b7-b232-ab122af26538 · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:32:18.389228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:ced8a212c6f65ef897df2292593e7145135d909b8dac1f7db4dc478d34176efc

Observation b2a13b9f-e45f-4c46-b994-454835469a8e · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:32:18.386887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:e4d18270fa0f4508f5848bb8360c5b28b26da0bfb211ee8caaf43156473372b0

Observation c779ccf8-d6b1-4c1a-af8d-dc2172977c6f · outbound

This paper cites an unresolved cited work.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:05:35.483051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:dfd8308725c3caa8fc31186c6d772b08cbc617c46838f3fd280a02f8aaf0df7c

Pith citing papers

Observation bd2e331e-ca62-4140-ac2b-16b475336611 · inbound

Kling-Omni Technical Report cites this paper.

Kling-Omni Technical Report Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:00:58.574782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T21:00:58.473043Z digest=sha256:44a09f9b0121c005538ffbda90e1d2c6acd4e35f9bab5aaa838f9d593dc35365

Observation 25739d30-0711-4cc9-840c-bd19894f9ed8 · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:36.183222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:36.183222Z digest=sha256:9e8be079e3363ec14a950f9438f40d7dc97c9bb08aa40493c5019feb94696a86

Observation eda22b86-aea4-40d9-b931-ce5b33cb746b · inbound

EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models cites this paper.

EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T22:25:04.080629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:25:04.080629Z digest=sha256:f5b4822afbf6ca7d7304ef0106ff9dce7508e2124a62387c379a9200c01727e4

Observation 7e65bc2e-827e-4b94-bee0-2a48c58c49b4 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:0d1cd4960f354ce21761c0689aecd2b15fd591a6ed60ebee47e0dfd803af444a

Observation f471e0dc-1dea-4e03-9dc0-e2fe01729b3d · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:58.465302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:58.465302Z digest=sha256:4d243d9aa1e1727998bba669e65a87f0e125d636fd08d4fe9a3162e7098df413

Observation 06207086-defd-46b3-971c-ed06e3d5b810 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:03.125700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:cb983560c90fc92dad5aba536dc4ff8ac89db96d22467f9ebe5cf0acf8cd6c03

Observation 05861950-70fd-49ee-ac44-c85e1a1a0057 · inbound

Video Models Can Reason with Verifiable Rewards cites this paper.

Video Models Can Reason with Verifiable Rewards Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:07:37.190619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T15:03:14.894952Z digest=sha256:9b2a0dbdcfea9a756a44698c568bf2faf69650fbbed4af039f1b544e128a2c61

Observation 6e84f652-d9c5-4d06-90ed-905ab170667c · inbound

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation cites this paper.

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.880681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T18:17:52.284353Z digest=sha256:16c2f8d15d982b2516601cac8dbae1f88d73d979f7e3e53f98bf55f03a984e1d

Observation ed1158d7-a5ac-48a2-bceb-0f6ba35a35ea · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.349322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:26:21.284810Z digest=sha256:2af0ba454cc61313935366e6d29a1f4b393747f2eea3b95786dd188d875c0832

Observation e574146b-899b-42f2-95b0-b6c82a161f70 · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:44:36.835287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T10:38:22.619277Z digest=sha256:f722cf9a9318e1e91019ecb0670c097dc9986b4cfe4e8581739bb652a2ed91db

Observation 8d548877-7965-45a0-9a44-69bf59a5576c · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:05:40.236863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:ab5544fd09b7a4156cf053fec4a3df9f83046764b9db772693b88968826ffacd

Observation d93dd5b9-05b3-40a7-b701-81dc87bb4625 · inbound

OpenCoF: Learning to Reason Through Video Generation cites this paper.

OpenCoF: Learning to Reason Through Video Generation Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.977806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:bde8a35c8fd15dca053c85bb535d8a4d7e07acdcb4c73a0c704a6dd1031132ef

Observation caece559-b57d-4b2e-988e-553c55d87fae · inbound

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence cites this paper.

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:27.491168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:27.491168Z digest=sha256:1ee96dba11c84e5add18c19ab3aecf7e4683d9e0876ccc5361610d4c8ec2c495