Pith. sign in

Paper Citation Record · LEDGER

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

As of 6 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 22 inbound Pith citation observations for arXiv:2510.27607.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.27607 v3

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:00:48.172438Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T22:22:57.515817Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T04:09:35.372388Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6df8e2d7-5b34-4efc-bbd4-3b7d601b151c · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:45.901236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:45.901236Z digest=sha256:35a214d4694cd373e23820b9ad3e63ee0aa869138a75382c51515c1ae001e909

Observation a0cefecb-96fe-4570-833e-a39278f44d64 · outbound

This paper cites Enerverse: Envisioning embodied future space for robotics manipulation.arXiv preprint arXiv:2501.01895,.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Enerverse: Envisioning embodied future space for robotics manipulation.arXiv preprint arXiv:2501.01895,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.060097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.060097Z digest=sha256:690330f05e475578c506269198ac178d30c07394686cf4e3d2b3ae26d2ef0060

Observation e1b77faa-ffaf-447f-86f0-50c9ae4f7756 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.231397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.231397Z digest=sha256:6c446ff7c008eaeb1e143df81479d9d0c4e776f20167ab4fce9d503c2b8bf745

Observation 32ec8f49-49da-4f6e-8e9d-b268b49409eb · outbound

This paper cites an unresolved cited work.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.963775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.963775Z digest=sha256:a0a1bb6dc9501d26b4dd65ce8a3f36550f6a9ba12b123f5c4175c58c1a596ae8

Observation fd1067cb-fee6-469e-a104-84c13e6da7a7 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.695032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.695032Z digest=sha256:4851723cb2e6d3f66f81a0e6cbb1d412510422c3fd7e0f94356fa0de2a5de236

Observation 3f52d05e-4059-46fa-9a42-e8e09518940d · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model DINOv2: Learning Robust Visual Features without Supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.799936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.799936Z digest=sha256:573fb52d2ac9f056cde64fbf2f4a875b5f700f83f75b1e7b5dd0294d8d3c3fb2

Observation c6bbb80e-455e-4326-b6c8-5b277c9c80eb · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.893433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.893433Z digest=sha256:0373bfbc7c388fda464992cdfdbd03cf41209a18ff412cfc10b12c87f920d2f1

Observation 3bddf48f-d85e-4084-a283-4d9b984e5ff2 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.062153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.062153Z digest=sha256:9548b41e5b7af4c5ce81971ba2827e7dc17ed0a9923527e3ed0e136d7d74d66b

Observation dbf338ef-eadf-4f53-90f2-7ef1640b41f8 · outbound

This paper cites Genie centurion: Accelerating scalable real-world robot training with human rewind-and-refine guidance.arXiv preprint arXiv:2505.18793,.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Genie centurion: Accelerating scalable real-world robot training with human rewind-and-refine guidance.arXiv preprint arXiv:2505.18793,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.375007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.375007Z digest=sha256:55987d3e92ee38efe08f343984791b2a2815d346a72f555842ca338974233381

Observation fcaf861b-f407-475b-98a7-80f134dbf97c · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.457905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.457905Z digest=sha256:825b0d9ee14c5fe1c3f73ce4f4fc733cfca008d2906e10be8c03324c4a834bc8

Observation 1f89afb5-6bc5-496d-a056-6dc01b81b31c · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.591114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.591114Z digest=sha256:ffde83a69c0948d69bcb83c52503a4103ebc4e434e09d531be2e726f0847b8ee

Observation 16f87478-21fb-4e43-baed-96286917918a · outbound

This paper cites FLARE: Robot Learning with Implicit World Modeling.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model FLARE: Robot Learning with Implicit World Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.749003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.749003Z digest=sha256:e3322ad702ed46d0c9cd8a97f967f254e398fb966db03796fd84748a13a44071

Observation cb2975b3-4393-4e7f-bb87-a867b40fcd6c · outbound

This paper cites DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.904165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.904165Z digest=sha256:547fe3a7bf577cec22203e6beee6d96913cc7cfcd604d385ea93493182bcb388

Observation 6e5be1c8-6589-4dd6-835e-2f62be821ac1 · outbound

This paper cites Image observations include 3 viewpoints from the left, right, and wrist.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Image observations include 3 viewpoints from the left, right, and wrist

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:48.083925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:48.083925Z digest=sha256:c9530b90d17b0034f6affeb202a478641c5e65f9003d97b2fe2b1092e7785ea0

Observation 67dcb501-5c7f-46b8-b48c-f4c26ca48b8d · outbound

This paper cites The simulated robot is a GR-1 humanoid robot with Fourier dexterous hands, enabling fine-grained grasping and manipulation.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model The simulated robot is a GR-1 humanoid robot with Fourier dexterous hands, enabling fine-grained grasping and manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:48.172438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:48.172438Z digest=sha256:66283dcd064e1fcd13357b911c91af31ee41dfbb503aeebeba419841df3bb3e6

Observation 4d37e046-644e-4868-8bf3-105848765b48 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.154975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.154975Z digest=sha256:1c855352dfb1877973a91d581ae755a95e6cd6d7f3745241ace8a9c71a3555f2

Observation 8ea14750-e58d-4048-b56d-b8b74d8d3eb6 · outbound

This paper cites Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.531852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.531852Z digest=sha256:b82fa4159f3b284cade702070076463794aeb34bd79b435af476e870b87a51e7

Observation 2e2927a6-ed44-434f-9df3-2c89bc35768f · outbound

This paper cites This&That: Language-Gesture Controlled Video Generation for Robot Planning.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model This&That: Language-Gesture Controlled Video Generation for Robot Planning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.260541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.260541Z digest=sha256:421070b0425e48c4b97db7a63d7bf0272266bdb90237fe3a11f94bc0fb1d729d

Observation 6b2874ea-6c8b-49ce-8efa-7104d79a790e · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.400140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.400140Z digest=sha256:a14daa3e7742e2d03090d0235e42dd9e837e92b4a4e37922a3bf00edc5d7890c

Observation 90529829-446a-46ba-b581-2aed5cffbc7e · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:45.804069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:45.804069Z digest=sha256:bf9641e77dbbc9bf97b89594a51d73c3dbb89be843d67f71513e0a805354eee5

Pith citing papers

Observation a1b749c4-fc25-452f-abfa-6aecfe15734e · inbound

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body cites this paper.

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T22:04:07.403410Z digest=sha256:2ab62222d988f4cfe8940944cba15588386d7aea0f9010674477fa1b74c04401

Observation 88b916c2-e2cc-44d5-a27c-2c4b8992c4bb · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:d3d885179e1537e6c593c379b18a08662faadc0891ddaa035086730107614c0e

Observation b5c46c76-bf57-419e-a747-94e7625a8a7f · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:f84378346f2e2dd1a06a4fa2aa08b4790c135add599555e7fa5c97f63740064e

Observation 96b6dee1-172f-43e1-94f9-b09c71c70ee1 · inbound

Fast-WAM: Do World Action Models Need Test-time Future Imagination? cites this paper.

Fast-WAM: Do World Action Models Need Test-time Future Imagination? Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T01:57:29.753735Z digest=sha256:3421bfb5a5fa17a7bd00a559d2defa7feed43a44325595f023b89d608788acda

Observation a0729ffa-837f-4672-bc4f-6d0e7cf66b5c · inbound

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis cites this paper.

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:16:31.378588Z digest=sha256:210d63e0d36c3ef3dd02a1de0d87bfcca1d507536dc17ade6c233df6aa121afb

Observation 0f281ee1-c735-4eb8-bae8-d163e34277d2 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T16:06:10.595164Z digest=sha256:2e00f4bb34aee48b47195d2c9f82800b49b4232bd32cd22f6c34e0dbb645725a

Observation 3a164519-be3e-4db6-aed4-cf123fe18cc1 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T19:02:19.756092Z digest=sha256:b0c480ada2584e631461c2a2350cca2a0b1429b22d2c93e2e0f29ee60e057892

Observation c7321da2-609e-46e7-91df-c5fbc8ddaa54 · inbound

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models cites this paper.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:c579dae69b230cb0e4a6a6547664f143622a97bd33a95a011bde11109ebf7d26

Observation 6f68adc5-ad17-4443-a28f-c243e88d430f · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:ea4c6dd561d320cb724cffc86cbc06926e720cbc1d34b42db7feba2526469108

Observation d1756185-b4f1-4a41-8fdd-bea0b2f35d18 · inbound

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction cites this paper.

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T04:50:07.792757Z digest=sha256:f31d26ae9d31c5d98ee75f2701f5221a85c429c8f4d3490b40765d68eb51f6af

Observation d9a28dc9-a6ed-4b2a-a428-e0d1bd845a0d · inbound

Point Tracking Improves World Action Models cites this paper.

Point Tracking Improves World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T03:51:10.921839Z digest=sha256:aa2fb3d87192db636ed8a7990f93e3970f8cc055511e96f3efe4ee7b4be82bf5

Observation 2ec89e32-5612-46ea-bdf7-5cf57c296e4e · inbound

World Models for Robotic Manipulation: A Survey cites this paper.

World Models for Robotic Manipulation: A Survey Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:33:25.043319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T12:24:18.025364Z digest=sha256:548c96d34516522ef13d27a879b09a76a54bbb53b5e6dcf66b84834262da30a0

Observation aaaf0dfb-a2e0-4c18-9125-e261cee57598 · inbound

WALL-WM: Carving World Action Modeling at the Event Joints cites this paper.

WALL-WM: Carving World Action Modeling at the Event Joints Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:36:22.390546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T14:15:14.454649Z digest=sha256:4d4adea777f3936a7d40717085090f2a50449e6b07278cbeaa757d3a1a6c455c

Observation 7a54e59b-80dc-4fd0-aae0-964e75e15e52 · inbound

Next Forcing: Causal World Modeling with Multi-Chunk Prediction cites this paper.

Next Forcing: Causal World Modeling with Multi-Chunk Prediction Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.466847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T13:31:53.905704Z digest=sha256:7c1e006e91de255c4447e6b5da5783c066b77d9a31becc2d0fe547b717538643

Observation 99f314d7-3ead-4642-aa27-b6946760aa0d · inbound

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models cites this paper.

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:08:03.388045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:40:54.963759Z digest=sha256:07990af562c142b40f4714032350025834be97984185e7b43e6518c16995caba

Observation d07178fc-7b55-4ba4-b038-3c950f608f1c · inbound

World Pilot: Steering Vision-Language-Action Models with World-Action Priors cites this paper.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:18:03.351218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:40:02.137152Z digest=sha256:f456eb7d519796c7a5178552dc81f71b430c48a7071d9e292064ffc731f49259

Observation cc3bc517-242b-425e-b212-a123e059d14d · inbound

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models cites this paper.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.842966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:782ed92683d265f1111230db4aeec0749578320415f5bed4ed34278edecf9789

Observation 9d166564-449c-40cf-9737-b0f787255849 · inbound

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI cites this paper.

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 156

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:28:44.992670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T04:02:53.110012Z digest=sha256:2b135a7155d14856425d2e8f1d9d5d82bab2828d2eb7b7ba349252b4c90cafec

Observation 7edcb786-9da0-4025-b9f2-220881f800bd · inbound

World Action Models: A Survey cites this paper.

World Action Models: A Survey Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 171

Resolution
verified exact
local_arxiv, observed 2026-07-04T04:09:35.374151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T17:11:12.686936Z digest=sha256:b4c8c3f7d2f528ea2268b0ad5e9f90bab650da8ee82ea735ff252d51ee860393

Observation 694347b4-43b9-4633-a04e-80a7d8ec6aeb · inbound

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control cites this paper.

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T10:36:56.968032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:36:56.968032Z digest=sha256:9b7fef83645c9ecf92a2e946d11a1d1d4b1bddea94094eac3f1c5a8cb465ee18

Observation 3d2c49e2-5f48-40c0-993c-b9314e9ac0ae · inbound

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation cites this paper.

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T08:19:04.131379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:19:04.131379Z digest=sha256:87eaaa56536885490507855e17ee2dc186d4289cd29e2838c260e9e69ba86f9a

Observation 6eecc954-fad2-4ecf-b563-8cbb4475a9dc · inbound

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning cites this paper.

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T22:22:57.515817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:22:57.515817Z digest=sha256:9f2e984f4bca7d869b91d146be64b9ab8ebdcf5e18f9028aac486c5615513e2f