Pith. sign in

Paper Citation Record · LEDGER

OpenCoF: Learning to Reason Through Video Generation

As of 16 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2607.08763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08763 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T01:43:37.265723Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact31
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9023a7b8-c75c-4c9e-a37d-4b6b650524a3 · outbound

This paper cites Lumiere: A space-time diffusion model for video generation.

OpenCoF: Learning to Reason Through Video Generation Lumiere: A space-time diffusion model for video generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.710227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:64e835b3b751299da3eec1270038b966fb26719a5de94adbf745e2ab70afdf71

Observation b0ff4bdc-f080-49a9-bb31-5b9db8fb612d · outbound

This paper cites Mmgr: Multi-modal generative reasoning.

OpenCoF: Learning to Reason Through Video Generation Mmgr: Multi-modal generative reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.951462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:290224965f4c452bfb75617e35c401d143f600c715e45648bec5acde255a2f50

Observation 6891cb1f-3420-475d-9ca0-96b142251f54 · outbound

This paper cites arXiv preprint arXiv:2511.13704 , year=.

OpenCoF: Learning to Reason Through Video Generation arXiv preprint arXiv:2511.13704 , year=

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T01:46:40.962721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:6c58bd8774e92398552ac2a86f5e2d9ea2fe8e0cdea617911d0072c07cc84a85

Observation af841a16-d433-4220-9a54-c8c40f4ca6d0 · outbound

This paper cites MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning.

OpenCoF: Learning to Reason Through Video Generation MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.976230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:698cfae486872dadbdeb5d84911aa7bec261984a9a0b17539d9fe6efb14480bc

Observation 188fa137-da5c-4ef5-8765-2935a5d615e1 · outbound

This paper cites UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining.

OpenCoF: Learning to Reason Through Video Generation UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.957102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:5eb07ef5f1ec90b7823c8b4650acb61fe289ecd1589384763fd2f6511f0bd1c8

Observation b6fb3e8c-3d45-4037-9f75-6ff610ad38fd · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

OpenCoF: Learning to Reason Through Video Generation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.941636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:8b0e12fc9c770384f2662bfd4651a63fdadbb25744f4a86a37e2ed248c6deef7

Observation 9c5d7403-a04f-43d1-b567-b3c9a1805600 · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

OpenCoF: Learning to Reason Through Video Generation G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.940154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:9856383d05ccf2eb8124d9edeb2a231f500d2cd40478e84a656a9c85b2ee4c4d

Observation f0a263af-4767-4c36-b6eb-9cd479634827 · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

OpenCoF: Learning to Reason Through Video Generation Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.934444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:0c76ce1c38e66586c96ec72a1a065d2f5337ec33623ab9e762580b9b75adcde2

Observation b2a5a924-b357-4e98-8ab7-379e66038bf5 · outbound

This paper cites Veo-3 technical report.

OpenCoF: Learning to Reason Through Video Generation Veo-3 technical report

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.708214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:c215d6e97fc2f177de0511ba4abfd97dc1bbf49577565ac0dbf8478704182d5e

Observation 6463ea4a-889b-4a75-ad3f-e0f17b8412e2 · outbound

This paper cites Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark.

OpenCoF: Learning to Reason Through Video Generation Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.959738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:77f14f6b1196d3c223848c05a998c8f9f8b30c3786e67dc4b1c187345a06ae93

Observation cd97e1ec-a4f5-4930-a23c-4cc791835aca · outbound

This paper cites ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both.

OpenCoF: Learning to Reason Through Video Generation ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.964934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:c77219355bd67dca5f1f81ac61fb18475d6be0a7243ca6db033f688ad9359b4f

Observation 1a65e0b6-ca78-4ed1-a16d-6e96e54fe19b · outbound

This paper cites Ruler-bench: Probing rule-based reasoning abilities of next-level video generation models for vision foundation intelligence.

OpenCoF: Learning to Reason Through Video Generation Ruler-bench: Probing rule-based reasoning abilities of next-level video generation models for vision foundation intelligence

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.957301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:ab32b19517f4580c898b244e32db97c55778155e3e354b499e8075136adee646

Observation e9e10b7e-e588-4610-9a61-d22cbc2a186f · outbound

This paper cites DeepEyesV2: Toward Agentic Multimodal Model.

OpenCoF: Learning to Reason Through Video Generation DeepEyesV2: Toward Agentic Multimodal Model

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.983972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:1e01418a74cc04eb7d00336a3e4c21d85ac0e958c5c870b58e72839b413faa48

Observation 11df88a7-3bc2-483d-9b0d-f1daedfc5519 · outbound

This paper cites Lora: Low-rank adaptation of large language models.Iclr, 1(2):3.

OpenCoF: Learning to Reason Through Video Generation Lora: Low-rank adaptation of large language models.Iclr, 1(2):3

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.700940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:c5f0d4128ad45d61eb89e3a3dd76a5029bc562014a45a4698f60416f09a6fc37

Observation 5d7a631a-470a-42c3-aee0-a34a45a791fe · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Information Processing Systems, 37:139348–139379.

OpenCoF: Learning to Reason Through Video Generation Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Information Processing Systems, 37:139348–139379

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.697171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:4d6e2cbc2748a4fa77fdc0acf7b3dd6876219fb54a60c4da970120616594ba1a

Observation 7a8a1784-94f6-4d85-9bb5-47133c928073 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

OpenCoF: Learning to Reason Through Video Generation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.945808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:11fd7f22ba821fc8e51f5441a74f8260d5ddbd4ed236007b8e9e5eeebf1ae6e5

Observation 062dbfca-809a-4311-a779-d6b9b3c251ce · outbound

This paper cites Large language models are zero-shot reasoners.Advancesin neural information processing systems, 35:22199–22213, 2022.

OpenCoF: Learning to Reason Through Video Generation Large language models are zero-shot reasoners.Advancesin neural information processing systems, 35:22199–22213, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.695293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:72ad76b7577bd01d7be42fad3a6b5db60375c146bf0b41cd6470dc7bdb800e2a

Observation c3b4b50d-260e-4c33-8979-ba81089ceac8 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

OpenCoF: Learning to Reason Through Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.978808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:66907e2b976e959e498bd4cbcfee72ce8b4f3d6fcbc43a8678321ea77b9eb550

Observation 39325a9b-6503-4a4c-bd97-ab03340701e3 · outbound

This paper cites Kling ai: Next-generation ai creative studio.https://klingai.com/, June 2024.

OpenCoF: Learning to Reason Through Video Generation Kling ai: Next-generation ai creative studio.https://klingai.com/, June 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.702717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:f0f611f4930ba7dace4be037c9197a291529b585469bd4fa5dd6abe37ada58e0

Observation d877ac51-8731-44a6-80f7-fdadc40eac68 · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

OpenCoF: Learning to Reason Through Video Generation Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.925629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:627ca41e830a4e273c915aebdd49ad17a57eea966414d19691ec37e5ca2b482d

Observation c45ac3be-9ebc-4caf-8d91-1f5861551ca8 · outbound

This paper cites Thinking in frames: How visual context and test-time scaling empower video reasoning.

OpenCoF: Learning to Reason Through Video Generation Thinking in frames: How visual context and test-time scaling empower video reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.989376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:3ba4a52a4baeb7520ab5c719a05bd5f18668bd813ba8501e2f54cd83ea1ddbaa

Observation c86affcc-5c45-4d49-9652-8851960cbdd9 · outbound

This paper cites Beyond the last frame: Process-aware evaluation for generative video reasoning.

OpenCoF: Learning to Reason Through Video Generation Beyond the last frame: Process-aware evaluation for generative video reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.946670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:cad8c69adeb77f2488cb37b3873cdc4693f70b1b32f5744b3e0cc2e8c7a1331f

Observation a2378171-af96-4dfb-a16e-4a3455e26c58 · outbound

This paper cites Can world simulators reason? Gen-ViRe: A generative visual reasoning benchmark.

OpenCoF: Learning to Reason Through Video Generation Can world simulators reason? Gen-ViRe: A generative visual reasoning benchmark

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.966106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:d5d182e20132276d9cc280392544680c70f899ab99a36b8682237f9fc31707f5

Observation 5ae965ed-4668-48cd-8f50-42c07938ef76 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

OpenCoF: Learning to Reason Through Video Generation MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.922981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:7bd08c0cedcc7195d8af94b1591083fd33f1f13bcfa6326bb496111b840e11d9

Observation 1c5f7a11-3783-4f99-b1e5-8d4ff7594325 · outbound

This paper cites V-reasonbench: Toward unified reasoning benchmark suite for video generation models.

OpenCoF: Learning to Reason Through Video Generation V-reasonbench: Toward unified reasoning benchmark suite for video generation models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.954809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:c6f81163c8bfddecd4a0c4e6515866a9276f7670aec561fd7cdca5bb4ad25fd5

Observation 5f7cfce7-c762-4002-8a03-3de86e4b2680 · outbound

This paper cites Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers.

OpenCoF: Learning to Reason Through Video Generation Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.693365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:cfcf256503a4f4009169bc1ffee7a14bfc699d36cc78c7f8f1c72bec663d8617

Observation b600dee4-f4a6-4258-9d8c-f2576be339ce · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

OpenCoF: Learning to Reason Through Video Generation MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.973659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:064a80dcedd5d1bf6251186c6110ff720882e22669e7936784aacb380df5319d

Observation d97fcd91-3b88-481a-81b3-23eac9c1fee7 · outbound

This paper cites Sora 2 system card.

OpenCoF: Learning to Reason Through Video Generation Sora 2 system card

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.689671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:fa255b7f514eb38a490f21eed7ba4e5cddc9f31e61dbfabbc5075a0c2159bcc5

Observation 3080deb3-be58-474c-8580-5bc03a4d6a93 · outbound

This paper cites Scalable diffusion models with transformers.

OpenCoF: Learning to Reason Through Video Generation Scalable diffusion models with transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.691493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:859ead8ee028493f59d4c50e4bbd75b25711055e5550369ed89422176895584c

Observation c8f7bd50-224c-4bc8-9a57-79e870edfdf0 · outbound

This paper cites Mme-cof-pro: Evaluating reasoning coherence in video generative models with text and visual hints.arXiv preprint arXiv:2603.20194, 2026.

OpenCoF: Learning to Reason Through Video Generation Mme-cof-pro: Evaluating reasoning coherence in video generative models with text and visual hints.arXiv preprint arXiv:2603.20194, 2026

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.981614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:652592c56dbbc845b1a37ea6cccba57bdb7ab1f24d58488fb490f49a270cc67f

Observation c554f99c-a9a5-499b-8b9d-93629e585ffe · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

OpenCoF: Learning to Reason Through Video Generation U-net: Convolutional networks for biomedical image segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.706478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:78a2f1ffe9851b1a0589a45fdce5da304d38f50ae39304db8e3e4f1e1b25a21c

Observation 414590e5-777f-4818-ab82-213f3920114d · outbound

This paper cites Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model.

OpenCoF: Learning to Reason Through Video Generation Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.968677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:0e8348713af68ec03fea031d140d1969d8f646c0299fa96ddc276fdede5a476d

Observation 5dd83c35-a621-42e2-991b-2ddf1ddccd6e · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

OpenCoF: Learning to Reason Through Video Generation Seedance 2.0: Advancing Video Generation for World Complexity

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.986572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:949d7ed06c65831fb4b969fbd412e1876c97fc90cbef8ef60b485a6dc29b5e1b

Observation 754f53a1-1095-4b26-9443-797670817931 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning.

OpenCoF: Learning to Reason Through Video Generation Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.687949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:7741a41a0d7fc3492a2f6699aaa93a2003b1402d4a4ed64ef895bbccf1bbd817

Observation 7a687809-ffa1-465a-a467-5f64c2433ae0 · outbound

This paper cites Image editing in gemini just got a major upgrade.

OpenCoF: Learning to Reason Through Video Generation Image editing in gemini just got a major upgrade

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.686062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:3f809f155a2d03b028505ec3d8e1436907a8ded3a43e5465563ce3365740467a

Observation 7a5f0106-3f86-4f6b-bd99-dac096f1d38f · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

OpenCoF: Learning to Reason Through Video Generation Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.944092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:8a36bd0a2dc50cac3cf4eb7bec7560cf7ae4c94a983170412b685877501efa18

Observation bb2c96df-62fb-4a1f-9510-3665177844d3 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

OpenCoF: Learning to Reason Through Video Generation Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.718006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:7923dcd5ae315e329b4844815401c4fe1f3d949f766db5699dd845acd79f00df

Observation d93dd5b9-05b3-40a7-b701-81dc87bb4625 · outbound

This paper cites Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm.

OpenCoF: Learning to Reason Through Video Generation Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.977806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:df34d5bae74404ec0c3348c9904c6c9bbcebd85d053ef2dc6dc8772c37e93fb9

Observation 4320165e-6448-49aa-b662-beb1d6a27e8e · outbound

This paper cites Attention is all you need.Advancesin neural information processing systems, 30.

OpenCoF: Learning to Reason Through Video Generation Attention is all you need.Advancesin neural information processing systems, 30

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.712067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:60be56f3719e4a5b6027256e189cd7acdd166bb811090df0eeb2d69a86d42f91

Observation 15fceebf-d8b6-4570-8021-90f0cf75f9c5 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

OpenCoF: Learning to Reason Through Video Generation Bridgedata v2: A dataset for robot learning at scale

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.713777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:5ecc9b9df642f1f59c8b2716a4230d6c1294473ee36ff33691706771a80e9a3a

Observation 9e5b6d57-8719-429f-8121-96e685de7bcd · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

OpenCoF: Learning to Reason Through Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.980174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:f3dded12d8fc1b9cc19ebb12bed61784b794a28478d45cf4ed58c513e2a4d717

Observation dc76d491-4e59-4cb6-baca-d3a1fae1a3e1 · outbound

This paper cites A very big video reasoning suite.

OpenCoF: Learning to Reason Through Video Generation A very big video reasoning suite

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.919795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:cbb1f6e23bf27cb7b9b27412d6ce780516dadad88d971dc4be54a3e3082b499e

Observation 2b3fc277-3f0b-4234-93cf-8483a272a386 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advancesin neural information processing systems, 35:24824–24837.

OpenCoF: Learning to Reason Through Video Generation Chain-of-thought prompting elicits reasoning in large language models.Advancesin neural information processing systems, 35:24824–24837

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.715805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:6f9d26b11abb9bc699f19fc84756c3196c756635c117dcfc229a7a4195866225

Observation 8657b9ef-67e8-4500-a5f8-38582f8953d2 · outbound

This paper cites Video models are zero-shot learners and reasoners.

OpenCoF: Learning to Reason Through Video Generation Video models are zero-shot learners and reasoners

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.972750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:95adb0d2e9117a4ffee9be883851eced8ff3994633a147ea9aa8f3635c53f66d

Observation c0047a30-0832-4b0f-8df6-e7a5a46116af · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

OpenCoF: Learning to Reason Through Video Generation HunyuanVideo 1.5 Technical Report

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.933515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:0dd3cff75694041aef6085c0d9c3af16be009ddb5f06093ca64559d28581fd95

Observation 885deb5a-6120-4ae8-906b-f0a07f982aaa · outbound

This paper cites Visual planning: Let’s think only with images.

OpenCoF: Learning to Reason Through Video Generation Visual planning: Let’s think only with images

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.960034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:8d0826555396f52219e71fe0afa0f073d6df4ca156ff900d39bc09234958f253

Observation cba367d6-33e6-403f-9bf3-ca666b0c4672 · outbound

This paper cites Reasoning via video: The first evaluation of video models’ reasoning abilities through maze-solving tasks.

OpenCoF: Learning to Reason Through Video Generation Reasoning via video: The first evaluation of video models’ reasoning abilities through maze-solving tasks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:46:40.949249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:6dc70a353315d3b234f12033a9d0669734ca0f8d518e6f6e7e534c483a32f8f7

Observation fe13e993-bead-430a-8d8a-d06b6f3d0e08 · outbound

This paper cites R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization.

OpenCoF: Learning to Reason Through Video Generation R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.699135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:4857851cb8a4ff2d80d9397c4343e2429b2c7d9c0895081f188be3578b0a0298

Observation 4627b822-39cc-4b87-984d-7413f66e04e2 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? ECCV 2024, 2024.

OpenCoF: Learning to Reason Through Video Generation Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? ECCV 2024, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T01:46:41.704599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:7b1390f64f6a71036d8136e0cdc1bb4263c83c6fac723eee9266e2a89fe2e69e

Observation a231b1ce-35ce-4484-963e-8c973cf20d25 · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

OpenCoF: Learning to Reason Through Video Generation DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.971205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:671f36bedfc08ca607a43e479af4c034cbc738cfd5f00bf60b03da9073ab2f5d

Pith citing papers

No inbound Pith citation observations are available.