Pith. sign in

Paper Citation Record · LEDGER

Video as the New Language for Real-World Decision Making

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2402.17139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.17139 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:57.292927Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:27.914181Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cb68f435-9ef9-496f-9a99-9d4e3f324fae · inbound

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation cites this paper.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Video as the New Language for Real-World Decision Making

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.397994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:a30c1ae688fcf117377552ce3febfc62e775143c68a5e82f1d5ff9a30998967d

Observation d9e99fea-0c5a-4df0-bae7-857ae8ae0f22 · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models Video as the New Language for Real-World Decision Making

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:21:45.130034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:3349a663a19f94b66c5ececdca9ed260a15be969a948913e03698adee5264e44

Observation b9323015-c13c-4360-b99d-9cbe805562ee · inbound

Humanoid World Models: Open World Foundation Models for Humanoid Robotics cites this paper.

Humanoid World Models: Open World Foundation Models for Humanoid Robotics Video as the New Language for Real-World Decision Making

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:57.292927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:57.292927Z digest=sha256:7158577abf0df910081fb61dd6fc9388d914c8ff0243e18d1c099efcb2b50d2f

Observation 2ada284e-c6b7-4b1d-9580-983545486f75 · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents Video as the New Language for Real-World Decision Making

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:57.884964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:57.884964Z digest=sha256:86c82cbf19176dfc4336d9382c6c4f90962f4b3ae2e662a202f6163061cb4788

Observation bd6a97b2-880f-4f4a-81a5-1f4dba080336 · inbound

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics cites this paper.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video as the New Language for Real-World Decision Making

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.250329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.250329Z digest=sha256:a54a7b44972060b1ed0a6f043f435326e704624d6fa5c988546bdef1cd742f83

Observation 37625214-519f-46b9-9d50-251d091c359f · inbound

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models cites this paper.

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models Video as the New Language for Real-World Decision Making

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:31.134881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:31.134881Z digest=sha256:18aa76d4f438a5e0e8b5d7d0e96dc2765a46dd60287e51268dfac6be4e7463a2

Observation ebdd81f7-2e05-4cc2-aa21-62ef55f20a77 · inbound

Whole-Body Conditioned Egocentric Video Prediction cites this paper.

Whole-Body Conditioned Egocentric Video Prediction Video as the New Language for Real-World Decision Making

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:27:13.090738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:27:13.090738Z digest=sha256:871ef58b7f5d84578be1bb49e4f0e27e08a0d55c8c8e162d33197bf1ba2a2508

Observation 9d3b7601-69a2-4004-9bce-c21bbd9453f9 · inbound

T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation cites this paper.

T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation Video as the New Language for Real-World Decision Making

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:28.927615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:28.927615Z digest=sha256:e322d06e818cfdd9cad6bad7e88d54f7d0b8f025cbaee04c05042e65c1933998

Observation 4a4b10f8-090a-45da-8050-70aa98552be8 · inbound

Video models are zero-shot learners and reasoners cites this paper.

Video models are zero-shot learners and reasoners Video as the New Language for Real-World Decision Making

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.722401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:0e7428c7bcd5118b2c13af1b703cab6ad4b9427e63cd2a0d5b47cf9f6c9e2cde

Observation 87e01a35-60f0-4852-9adf-c989b5237486 · inbound

Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning cites this paper.

Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning Video as the New Language for Real-World Decision Making

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:53:54.888033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:52:22.919967Z digest=sha256:44881a8efea5e58f31d51c422c23641aedada9791a0cede1bea7d744dcc6a392

Observation 52267ed1-33aa-4896-8491-9ed18aef500c · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Video as the New Language for Real-World Decision Making

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:36.964436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:36.964436Z digest=sha256:8c91304eaa03284dd5d15d19768060aef97e56f0cf9c63c294200576e450a3d7

Observation a645df80-184e-4802-96b5-9920ab6b6856 · inbound

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models cites this paper.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Video as the New Language for Real-World Decision Making

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:46:02.852584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:51:12.604102Z digest=sha256:0d55d6fb16693b23908e4beda8a7b2b61376911bcc260542a2886e194b012a50

Observation 16b3b3ec-13db-46c2-b9a2-84c7af644737 · inbound

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models cites this paper.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Video as the New Language for Real-World Decision Making

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:27.915621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:efe1d23bc8997b402833225b803255ee0ad064866da0c2685c3f3269e1be1d34

Observation a8759541-caa2-4281-a86f-58fc7310dc89 · inbound

Visual prompt engineering for video models cites this paper.

Visual prompt engineering for video models Video as the New Language for Real-World Decision Making

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:13:11.587979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:13:11.587979Z digest=sha256:f9425e004397b34a81e305cc578f633500603fc2d13470a9c5a10eb42219881d