Pith. sign in

Paper Citation Record · LEDGER

Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2601.23224.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.23224 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T18:23:40.234499Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T17:27:15.492915Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation befe0f0b-9765-41fa-bce1-312fc6b69e6e · inbound

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously cites this paper.

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T18:23:40.234499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:23:40.234499Z digest=sha256:097ec31c0cd785d13d9581e92c5d956855e6736bf02776c53fab0695a6a966a6

Observation 72739661-56ac-40cb-a1f6-6f65e9e570cf · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:54.352773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T14:06:27.953376Z digest=sha256:374c637bfab966d25abcafdf296e5c5ff12cb5fdd450f57db907288c77b82946

Observation fcb3e49e-c940-4a4f-a6d7-3077eace1f4e · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:54.352773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:49:41.654207Z digest=sha256:1b703c7dca331b637c6f6a5393c2bf732fe28d244788819a53b15678f36dec6c

Observation eceb86bc-57f8-4394-8298-faa5faf0b4d1 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:54.352773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:35:59.553683Z digest=sha256:50ce87291abe1dd763852e8db6134fcc6e31faea46af8266d638533d86577d89

Observation 94c1d8ba-cfc7-4428-ba33-1c0e1619b2d2 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:10:23.940870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:08:19.956833Z digest=sha256:52cc344dfa002cb19ba64a43758c07f16d5ac180a152713e872b172759905b6b

Observation 670d6987-35b0-40e8-a420-64e9aa06f73a · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:54.352773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T07:32:12.180233Z digest=sha256:2b2644aebd1835406c5f3593aa997fe1fc477682bcd90b8447c9c42962f909cd

Observation bd90550d-0b58-44b7-9105-d3fe25a305a5 · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.474877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:b835e7ae3ddb9380fe46cb99e762b3b315822e6effb8c3ec58ef88b52bd620db

Observation f1224666-04e5-4bcb-be0a-0ba2d4c7743c · inbound

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding cites this paper.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.121262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:1a6d31e71f5d52bdaa00b03b3db066b49ceb1d05be3b2a689282c9241fb801f2

Observation 86e4d4c6-5e9c-4cd8-9236-65ce6d196aa6 · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:56:55.396325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:5cf0447cdc4949b52b6f8dc25cdc2bd1b5a578ac420d1f5af7b899b99114a15a

Observation 7b88abd7-62ea-4c28-84a8-edaa87a03c80 · inbound

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering cites this paper.

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:46:56.871066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T01:52:33.768494Z digest=sha256:d6573b15e83c1b9df29839c8459cece8ecf8dea4261ce8fffb29a319df58c218

Observation b9cb58ef-339d-40ea-bf26-e2da5cad6ca0 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.494462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:7f53b840efd8de91b3c025723170192f81ef1cf52642f00b235f075cdbac840c

Observation 7eb20e68-e4ab-472d-83e9-bc2fcebe7712 · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.221547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.221547Z digest=sha256:31d6818869ad75b6668c065eed2403cf02e0bc43b653863c00db6de97520bae1

Observation 2d04a5fb-0e4a-4430-bed8-93c0705a9293 · inbound

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA cites this paper.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.593997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.593997Z digest=sha256:b948c36f0b7826593779cc1d2676fa3e25f2c11c67bef12ad40b06d631d8ff4a

Observation f154b320-e218-433b-92d4-f285507ebcde · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.780478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.780478Z digest=sha256:7352799974596360a7fc9c8a5aded9fcc1b104bac576e03672f0420b39f1022b