Pith. sign in

Paper Citation Record · LEDGER

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

As of 19 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 24 inbound Pith citation observations for arXiv:2510.27607.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.27607 v3

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:00:48.172438Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:42:19.328479Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T04:09:35.372388Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6df8e2d7-5b34-4efc-bbd4-3b7d601b151c · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:45.901236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:45.901236Z digest=sha256:8c0c6a5e157baf91347171714b956d46b07e02de60941aea50101983d5102a60

Observation a0cefecb-96fe-4570-833e-a39278f44d64 · outbound

This paper cites Enerverse: Envisioning embodied future space for robotics manipulation.arXiv preprint arXiv:2501.01895,.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Enerverse: Envisioning embodied future space for robotics manipulation.arXiv preprint arXiv:2501.01895,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.060097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.060097Z digest=sha256:cdfd084d8f286aab1ff2ac60c05c9aa12a4a254aa2790a31f7728bcf1100dcd7

Observation e1b77faa-ffaf-447f-86f0-50c9ae4f7756 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.231397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.231397Z digest=sha256:2ec94bd44fc96952a3f19f8083f5f3cdcbb1e2a55c9f4c862e981b77f4b49521

Observation 32ec8f49-49da-4f6e-8e9d-b268b49409eb · outbound

This paper cites an unresolved cited work.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.963775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.963775Z digest=sha256:d3f27330cea55e56602e0ec84edd05b54fbe1598c4911891aceea5b4e7355285

Observation fd1067cb-fee6-469e-a104-84c13e6da7a7 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.695032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.695032Z digest=sha256:1877b7dc8dc9c10762ac5649c82eea650b5aff79160da36a1263f38c68af475d

Observation 3f52d05e-4059-46fa-9a42-e8e09518940d · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model DINOv2: Learning Robust Visual Features without Supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.799936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.799936Z digest=sha256:3aa095d399297aff61d6cb2f0c8354036d6769a996df217c4fa84d42fe4e0962

Observation c6bbb80e-455e-4326-b6c8-5b277c9c80eb · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.893433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.893433Z digest=sha256:ac63b2a8aa4c37d75dbc2ecd2d4261cf3be4e024763994cfba316a6c06c19424

Observation 3bddf48f-d85e-4084-a283-4d9b984e5ff2 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.062153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.062153Z digest=sha256:3b2f822dc253fcc2a2bfd3cf7cf0970911dd4fa48ac415f8fa1a4a4983c29b37

Observation dbf338ef-eadf-4f53-90f2-7ef1640b41f8 · outbound

This paper cites Genie centurion: Accelerating scalable real-world robot training with human rewind-and-refine guidance.arXiv preprint arXiv:2505.18793,.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Genie centurion: Accelerating scalable real-world robot training with human rewind-and-refine guidance.arXiv preprint arXiv:2505.18793,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.375007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.375007Z digest=sha256:24a8d743c7610980504bb4f1ac8447e3f7604ceb3fa3a1f01b268a3349f4db21

Observation fcaf861b-f407-475b-98a7-80f134dbf97c · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.457905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.457905Z digest=sha256:fd5c10f7de625641f81621387c23d3f46b8677339f3222a06ee4deb037dad1a4

Observation 1f89afb5-6bc5-496d-a056-6dc01b81b31c · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.591114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.591114Z digest=sha256:f4473ec64121700f9efea9098e9f2e6ce9888f8ddbf82d465a06182be3e83dad

Observation 16f87478-21fb-4e43-baed-96286917918a · outbound

This paper cites FLARE: Robot Learning with Implicit World Modeling.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model FLARE: Robot Learning with Implicit World Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.749003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.749003Z digest=sha256:e8374720b79b1918a4da5acdd581788b60677abde1221ffa5753ecb243027bce

Observation cb2975b3-4393-4e7f-bb87-a867b40fcd6c · outbound

This paper cites DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.904165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.904165Z digest=sha256:a1bcf1fabdf118d85f534fa0b4e0c3ab07d557608c01c4bf5fb08dcbbc7e36cc

Observation 6e5be1c8-6589-4dd6-835e-2f62be821ac1 · outbound

This paper cites Image observations include 3 viewpoints from the left, right, and wrist.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Image observations include 3 viewpoints from the left, right, and wrist

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:48.083925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:48.083925Z digest=sha256:04acaf4f3d292ff16a70c764af6d9ffc8a758e6bf3baee25feff8dd57654d2a9

Observation 67dcb501-5c7f-46b8-b48c-f4c26ca48b8d · outbound

This paper cites The simulated robot is a GR-1 humanoid robot with Fourier dexterous hands, enabling fine-grained grasping and manipulation.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model The simulated robot is a GR-1 humanoid robot with Fourier dexterous hands, enabling fine-grained grasping and manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:48.172438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:48.172438Z digest=sha256:32e182d50a9d5321e54d716c83f1f94870877d4681ef7714a149e15b76dd0e7b

Observation 4d37e046-644e-4868-8bf3-105848765b48 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.154975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.154975Z digest=sha256:e59f1ac699aac9ee6284e03d89252d24355732738bf36ab2960d009498737fc8

Observation 8ea14750-e58d-4048-b56d-b8b74d8d3eb6 · outbound

This paper cites Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.531852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.531852Z digest=sha256:c2ddc6a78df8ca53a4324f04ff2a059ef3ac8e2768c3c7cedfb55b4d7f7b12e5

Observation 2e2927a6-ed44-434f-9df3-2c89bc35768f · outbound

This paper cites This&That: Language-Gesture Controlled Video Generation for Robot Planning.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model This&That: Language-Gesture Controlled Video Generation for Robot Planning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:47.260541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:47.260541Z digest=sha256:59b726868bb5218eb761bb0bf2a708c201f82f3d7a4ff4f662c8f2615ac059d0

Observation 6b2874ea-6c8b-49ce-8efa-7104d79a790e · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.400140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.400140Z digest=sha256:77956cc736cae97e8beceee07af2dda1cae2da2672623b3f011d14ee47a08383

Observation 90529829-446a-46ba-b581-2aed5cffbc7e · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:45.804069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:45.804069Z digest=sha256:ffe903abf0a0532bdac8a8d14e06ac15be4f2f0623e4f876c0b800cae33d8e9b

Pith citing papers

Observation a1b749c4-fc25-452f-abfa-6aecfe15734e · inbound

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body cites this paper.

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T22:04:07.403410Z digest=sha256:4b17f4b071c6ce548a7fad880a2101d042bac0f9f1a02668b4a45ca818b541e2

Observation 88b916c2-e2cc-44d5-a27c-2c4b8992c4bb · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:7becdae44c788887d644330bfc6d08435b4ef7978538d1ab94d1de3ff274b0ce

Observation b5c46c76-bf57-419e-a747-94e7625a8a7f · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:71e4f56a8b03d4c809a70043be00cb8301c8fd02d46261560c765db3713b1fbc

Observation 96b6dee1-172f-43e1-94f9-b09c71c70ee1 · inbound

Fast-WAM: Do World Action Models Need Test-time Future Imagination? cites this paper.

Fast-WAM: Do World Action Models Need Test-time Future Imagination? Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T01:57:29.753735Z digest=sha256:234e1e2c416e0b983ca206fe8f5139f65ba5a2e56ecc4b117d7518e646abc55f

Observation a0729ffa-837f-4672-bc4f-6d0e7cf66b5c · inbound

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis cites this paper.

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:16:31.378588Z digest=sha256:85ee96ae9178a2ad455b7317cd6511918fa1598ffb9120748d8faebfaaaee91a

Observation 0f281ee1-c735-4eb8-bae8-d163e34277d2 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T16:06:10.595164Z digest=sha256:716cbc08fc3e7f49e212f1f9740f254b9c1aa775e662633612c37f53e6b80f61

Observation 3a164519-be3e-4db6-aed4-cf123fe18cc1 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T19:02:19.756092Z digest=sha256:2c1afd12be001935ea7e856a68ddf187289455a8201d5375270c68c5d917179c

Observation c7321da2-609e-46e7-91df-c5fbc8ddaa54 · inbound

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models cites this paper.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:e773a72967cbaf7e9d5b3284f26d5a1a22c8a030e06f0acadd5e8715861a236c

Observation 6f68adc5-ad17-4443-a28f-c243e88d430f · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:57452606da378ac4332f9b935a400e592478a0040e0cc640cb1c218ae4cd5cce

Observation d1756185-b4f1-4a41-8fdd-bea0b2f35d18 · inbound

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction cites this paper.

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T04:50:07.792757Z digest=sha256:5ee780d1ef34554fe451f2d71ebf1da9a4edab0173860b82a9175ff43b9fc339

Observation d9a28dc9-a6ed-4b2a-a428-e0d1bd845a0d · inbound

Point Tracking Improves World Action Models cites this paper.

Point Tracking Improves World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T03:51:10.921839Z digest=sha256:ee4b4b213c687e9ae535ddf1798220fcf11f1273fc90573f0ee6c543269565aa

Observation 2ec89e32-5612-46ea-bdf7-5cf57c296e4e · inbound

World Models for Robotic Manipulation: A Survey cites this paper.

World Models for Robotic Manipulation: A Survey Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:33:25.043319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T12:24:18.025364Z digest=sha256:06a3b1b2614f6b87802491eaed3ce2897ff854f186ed97716e898af9326aecde

Observation aaaf0dfb-a2e0-4c18-9125-e261cee57598 · inbound

WALL-WM: Carving World Action Modeling at the Event Joints cites this paper.

WALL-WM: Carving World Action Modeling at the Event Joints Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:36:22.390546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T14:15:14.454649Z digest=sha256:8ce4a77a4511d508143b297e692b2273eb0d8492339b7bab6acd93cdde9be39d

Observation 7a54e59b-80dc-4fd0-aae0-964e75e15e52 · inbound

Next Forcing: Causal World Modeling with Multi-Chunk Prediction cites this paper.

Next Forcing: Causal World Modeling with Multi-Chunk Prediction Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.466847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T13:31:53.905704Z digest=sha256:415486d36ce914c19d23769bf66b170808a929caf9639f790fdcf2225f4ebf2d

Observation 99f314d7-3ead-4642-aa27-b6946760aa0d · inbound

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models cites this paper.

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:08:03.388045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T09:40:54.963759Z digest=sha256:5d7249d717b961c380fb138ea9a421273ecc561bf3c395c862b591a4ae5a155a

Observation d07178fc-7b55-4ba4-b038-3c950f608f1c · inbound

World Pilot: Steering Vision-Language-Action Models with World-Action Priors cites this paper.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:18:03.351218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T09:40:02.137152Z digest=sha256:27aaf87f1bc5365f49ba1b09fb3a5ad274133c886fcdeddd2b561d6d587db7aa

Observation cc3bc517-242b-425e-b212-a123e059d14d · inbound

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models cites this paper.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.842966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:31306f5bcdc32a5360d47c42561b81c4d6c9cc3f64b5072e4a072124e8dd9453

Observation 9d166564-449c-40cf-9737-b0f787255849 · inbound

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI cites this paper.

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 156

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:28:44.992670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T04:02:53.110012Z digest=sha256:9fce8f931478367bb1f501de757d4340a75d01ffe68c47f6db49f6db30734972

Observation 7edcb786-9da0-4025-b9f2-220881f800bd · inbound

World Action Models: A Survey cites this paper.

World Action Models: A Survey Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 171

Resolution
verified exact
local_arxiv, observed 2026-07-04T04:09:35.374151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T17:11:12.686936Z digest=sha256:e1d059e129abafd84f70f20f96856c09e4eec73d0900e5b7355264308cd25c0d

Observation 694347b4-43b9-4633-a04e-80a7d8ec6aeb · inbound

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control cites this paper.

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T10:36:56.968032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:36:56.968032Z digest=sha256:343fd39e4cbfa8dd474fb92d32cd35fe600528ef7f75f5d00ecfa1f17d93afe6

Observation 3d2c49e2-5f48-40c0-993c-b9314e9ac0ae · inbound

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation cites this paper.

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T08:19:04.131379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:19:04.131379Z digest=sha256:b9606ab2566821cfbc31b4396e25d8f2d270f3a6c6eb1d50d330c2ec0d96c1cd

Observation 6eecc954-fad2-4ecf-b563-8cbb4475a9dc · inbound

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning cites this paper.

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T22:22:57.515817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:22:57.515817Z digest=sha256:2ba9a0eb9da6e3214d183c88890c382be23a5d6370baf7fc918a7948b6f38986

Observation 1ec08bbd-b908-4a7d-ba17-80f545df5fc8 · inbound

Keep the Future, Drop the Rollout: RIFT for World Action Models cites this paper.

Keep the Future, Drop the Rollout: RIFT for World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:19.328479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:42:19.328479Z digest=sha256:d1edf9e654be1cafd4ecd83486820abdcab9d47c13bf52f23780301265c6a2ee

Observation 4658eaad-6897-4de7-9859-ac80fed2cb66 · inbound

Foresight Without Seeing: Latent Futures for World Action Models cites this paper.

Foresight Without Seeing: Latent Futures for World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:26.422950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:39:26.422950Z digest=sha256:ca304182f407c4f03d827a62c2df9180c98b471dd2b5f38ad64b35d2166006c0