Pith. sign in

Paper Citation Record · LEDGER

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 7 inbound Pith citation observations for arXiv:2603.29844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.29844 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T23:21:13.658826Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:18:12.785448Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T17:58:47.598667Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact12
  • verified fuzzy29
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c39a5be-70c7-44ee-bcf6-eaef48518a25 · outbound

This paper cites Paligemma: A versatile 3b vlm for transfer.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Paligemma: A versatile 3b vlm for transfer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.233536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:e2a58ec0a433a2dd7ddce3358e605f613bb8f5f31846e4a09ff2acfc1f1a2742

Observation 477cf41a-6477-4e92-8357-7dabb987e61b · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Eagle 2.5: Boosting long-context post-training for frontier vision-language models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.239454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:883353568e82388cbbb57453d46d9835873fc3a0599d1183a27ae0fb6df494c3

Observation c97a815f-26f2-41a9-b817-7d10daff85dd · outbound

This paper cites Qwen2.5-VL Technical Report.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.359919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:dc7e1e93e77969016f1f5e23b509a324bf067373a1bc0020d23f4b025f5765fe

Observation 10662f23-2d2e-4719-a889-c0717e91495d · outbound

This paper cites Qwen3-VL Technical Report.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Qwen3-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.373508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:fcff51db4ddcd854d4853cc813d6e667e5314374714714b3eb3e8681c6398aac

Observation 41ee3560-7973-4805-a6b6-f52f964f8130 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.316201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:fc4ac466a35c1ed8e1e730961688f26607d48d927da52b3ab83b775a88bf6dea

Observation 130f669f-1b53-4008-beee-e2cccb4c89d9 · outbound

This paper cites an unresolved cited work.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-13T23:23:27.288239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:d8350d4683046a276882333d11bfedd772e1a9a04ca3f53ae99c53123b1cc723

Observation 320752ee-44ae-4abb-8b11-f91afe4f7793 · outbound

This paper cites Openvla: An open-source vision-language-action model.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Openvla: An open-source vision-language-action model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.284060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:f2d6e5d855ad6fda94f93a87553aab17e4c2c4b91c819428e9836a35ee1ac347

Observation 08df3935-4f4c-4522-8cc2-8107d7641b6d · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.338436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:d2dab542eba78ea7d8e7ae91ac9b8e408084b1c9033b0f2679959e1e0de19723

Observation aa030518-5b9b-43c2-a442-e72197ecf0e8 · outbound

This paper cites Hi robot: Open-ended instruction following with hierarchical vision-language-action models.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Hi robot: Open-ended instruction following with hierarchical vision-language-action models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.344894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:a149d33b6e95f52ffd3be776766e9065455463eab024a4d67e275a97c68285d6

Observation bcc66cfe-005a-403f-922f-d303a32b7a03 · outbound

This paper cites Rekep: Spatio- temporal reasoning of relational keypoint constraints for robotic manipulation.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Rekep: Spatio- temporal reasoning of relational keypoint constraints for robotic manipulation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:28fb62c7c39ea2ee8472bf95a9181ea9793149388c7efa2b74fa75c03312589b

Observation 6459839d-a1d2-4aa6-91e5-b25a7af9973f · outbound

This paper cites GR00T N1: An open foundation model for generalist humanoid robots.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA GR00T N1: An open foundation model for generalist humanoid robots

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.258802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:08828409666b9eaf0c41e3e4a11e434559b4775f121b0b0a640bfb12de4a8786

Observation 82abb517-cc54-4d53-a026-cb52d080aa6c · outbound

This paper cites an unresolved cited work.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-13T23:23:27.356722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:d936e221c79d383e82d9ae72c070d4af5301147e464ee91bdb05abdbab0fdeb7

Observation 1975b9a0-d6a5-4297-bf65-91f775fd62ad · outbound

This paper cites FLARE: Robot Learning with Implicit World Modeling.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA FLARE: Robot Learning with Implicit World Modeling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:59:09.111136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:a62e36cc216c599a4325308678b51c3f53934e052f6d11f24a1b0c225faaeb35

Observation ef6a66b3-6b61-4990-9a82-7f86f09ded74 · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language- action models.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Cot-vla: Visual chain-of-thought reasoning for vision-language- action models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.292108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:eeb166a8e7d81603ed35f6fc626fe3aeab01ddad411f27023c6d262d73619f2e

Observation ef532a97-aa2a-48c0-8b6d-bb3e6de5c121 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:38:25.165602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:a2feceee148a53ef5c781426b3650801ec679cdaa0eb52ebe0ef1ca0cace10e5

Observation 2dbaf191-8cfb-491f-b7ff-47d83fdb3c2a · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Llama 2: Open foundation and fine-tuned chat models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.333055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:409964a2a14be8c3037518790d2649e4e798246c96ac9460e2253ba28de626d0

Observation feca6a88-0b21-4716-9a40-5d2e81364b2b · outbound

This paper cites Visual instruction tuning.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Visual instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.262142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:7a840f9a88dcbcb9f237f25dce343135e22c3d3ad68b4c18dd806f29cc68fe82

Observation ab1e0bf0-efac-4dbb-af5c-4a66de53cb65 · outbound

This paper cites Egoplan-bench: Benchmarking multimodal large language models for human-level planning.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Egoplan-bench: Benchmarking multimodal large language models for human-level planning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.267171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:ed4fc28c397185e45b5e6a058f3bc74a35bfe1c9fbea1c1816ba783e695fee28

Observation 09802b67-2016-49bf-8eb9-7a55f4edd7a9 · outbound

This paper cites Code as policies: Language model programs for embodied control.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Code as policies: Language model programs for embodied control

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.322027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:c53c37ee27e235fac3cd6368a609483a0b073fe4358036cc44230b2e8417f762

Observation c0a6ab37-e5c3-49ce-8082-37ccd02adaad · outbound

This paper cites an unresolved cited work.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-13T23:23:27.318283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:dc57b53000a5e21e2a3adf58bbb349fea7a4db78437415223a22590d6afc4d3a

Observation 6b64f0c0-890a-4ebf-bfb4-0f0c843a94bc · outbound

This paper cites Tenenbaum, Dale Schuurmans, and Pieter Abbeel.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Tenenbaum, Dale Schuurmans, and Pieter Abbeel

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.314392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:93eb6dc653ef00f25de23cc56e28e85ab9824470ad6b2f967e9c679d36655d02

Observation e024f619-ddaa-41a3-a3c6-f57b528a210e · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.310458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:0d550bf8c241b8b055c71e88fda759766672f7b32e11e6aa21e40325956b7a0c

Observation a8907712-3fa3-46f5-8067-1cc56465e088 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.366771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:17344aa8075f37035ff580cf1a3b22f0da3a4d3c3ac62364f910b6d676aadec1

Observation 6b057b30-2481-42f0-a3fe-187af6413d85 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.353798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:2693206e4a4f91ce014829f2268401fa8368dde354e8f4765ea97371e8d5aa7f

Observation ba92cf27-12f7-454c-b563-6e50ba677222 · outbound

This paper cites Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.242994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:91723a938b14110ab0c1f7979a2a6b83f959cfd6cc5dbd48a859e5fa8bddb4de

Observation 73017f00-d469-4d35-a890-b7d40381b75a · outbound

This paper cites Gr00t n1.6: An im- proved open foundation model for generalist humanoid robots.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Gr00t n1.6: An im- proved open foundation model for generalist humanoid robots

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.300114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:854dd979638292dbe1b6d02c076a5ddcfd86a32f2d6e8e0af30f919ee043062c

Observation 27da5bd5-abe3-4225-984f-1a0e694a3056 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.297923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:de3c07cd6a6772c6a2e694102745b05e6785e92c00e008b7cbd45eddb3981dd3

Observation 72301bd2-c9e4-4fbc-a4e9-0a7d41641a42 · outbound

This paper cites Gr-3 technical report.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Gr-3 technical report

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.250618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:431714ea298bb493089e413a3a53fe68226d5e6ea9032883d96fd71af6f35d1d

Observation 625e70ec-5d19-4d0f-8cc4-9281cd0f4629 · outbound

This paper cites Igniting vlms toward the embodied space.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Igniting vlms toward the embodied space

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.271890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:647cc0607f0cf0ec1d2acdcbbaf5df305b49b4b67632230852a0146bf1dd7327

Observation af34a71e-f9c8-4be9-9936-de7a55153b92 · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Robotic control via embodied chain-of-thought reasoning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.279812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:a8f199533000ecb0f6f3b039305ccb9dcf05aed22590b7dcc15c5c6ab57b2a33

Observation 326a1e34-73bd-44ad-a249-5eb2a894d3d2 · outbound

This paper cites Molmoact: Action reasoning models that can reason in space.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Molmoact: Action reasoning models that can reason in space

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.296012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:941844c7def24a11ae86205932551a47ae2420d8e12d88040d289eb7702648c9

Observation e0ef737a-7dde-40a3-b87f-e8e0c6f86add · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipulation.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Unleashing large-scale video generative pre-training for visual robot manipulation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.329101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:6f9ccead983d0b932a1d7b5338c1927cc6d32d95e8d6ba169b998ffce04567cd

Observation 492a9166-0654-48c0-a2a4-fcb43b5c76eb · outbound

This paper cites Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.340623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:1c908f0e3bc3520b033f38c8473a6df650d2d116bc3f9bafdaf1a7e6d38bd7fa

Observation 2cd4a1cc-27dc-4930-9df8-a3148d86e8bb · outbound

This paper cites Unified vision-language-action model.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Unified vision-language-action model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.352874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:7fa5eeb23d1859d087af49060601d5240e43b7e0ed0fac9fc35dfeb4a13f5f5c

Observation 8af9dca0-a714-4bb6-b95f-4081082cb8c1 · outbound

This paper cites Worldvla: Towards autoregressive action world model.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Worldvla: Towards autoregressive action world model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.349029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:da0ec68ca8511633dfc37b4374310a3dda7c7fd019649d516d571060695c738e

Observation d664ce26-54b0-4117-a20c-906aa1370dea · outbound

This paper cites Latent action pretraining from videos.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Latent action pretraining from videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.325486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:915dc699f5834aaf312d70ac58709f48963817b3b988eb016be33c8166ccaa88

Observation d32d7c11-c71e-4a70-86a5-ec165b3956ed · outbound

This paper cites Moto: Latent motion token as the bridging language for learning robot manipulation from videos.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Moto: Latent motion token as the bridging language for learning robot manipulation from videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.337010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:2576a73f9b8cf671cf0af0b206e8c791beb207a438797d78eaa254d3d1ab29b9

Observation 1bf737ad-d08e-43fa-a1b0-71572a753e9c · outbound

This paper cites villa- x: Enhancing latent action modeling in vision-language-action models.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA villa- x: Enhancing latent action modeling in vision-language-action models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.254966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:3f16ad58fd7538b9ba221b235952c9b11a26e5873ea1ebe9a542d148efe5bb12

Observation 8d38d9e6-9be3-4e04-bdde-a4d777f4ef97 · outbound

This paper cites Unicod: Enhancing robot policy via unified continuous and discrete representation learning.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Unicod: Enhancing robot policy via unified continuous and discrete representation learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.246738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:fd3aef6b432b86a63832a6d1cf327daa395fc63bf31050a59ca1b45110a114fe

Observation aafb4aae-2b89-4ff4-8d87-188f2f7d6a72 · outbound

This paper cites EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:40:29.471447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:24753add9d03f65dab2fd06ed93357cfd1fa46b662ad88591969f23e129eddd4

Observation 3a3941dc-c1a2-46ae-903a-895fdf060210 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Diffusion policy: Visuomotor policy learning via action diffusion

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.310904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:b6a1b133a9fdb1ec56bd176c4f905c22f4367992dda211e6a68574ecf2794b86

Observation 66bd4548-52a5-4c9f-bfcc-b7db01df7cb7 · outbound

This paper cites Unified Video Action Model.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Unified Video Action Model

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.321209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:89ecfa02f01e0aab366a17e2bee38393fe06d41d709bef858ff297e0adfcbf15

Observation c8d7c8f8-64b2-4fa1-af31-fd53ca519555 · outbound

This paper cites Starvla: A lego-like codebase for vision-language-action model develop- ing.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Starvla: A lego-like codebase for vision-language-action model develop- ing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.303706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:a2795c7a9617073f23b19fe5cf9fa8113ed362cd7991bdb5abfe57593a0721d8

Observation 164e0c9b-221d-4c91-a9f2-8ce7d5fb794f · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Dinov2: Learning robust visual features without supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:23:27.307206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:7b477e1faf672141d49481e9bb94288563d8749a858d8da6d83995a5028cc562

Pith citing papers

Observation 05b73946-9726-4fb7-a1a7-53f13247b360 · inbound

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation cites this paper.

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:35:46.767295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:55:56.318157Z digest=sha256:5f5542671f8334aed53feb8e6531d1e6e7e573c1fa0c8c9f9b6079224a5e7ee0

Observation c4f45136-24d0-49db-b4b2-a63f41f1dd0b · inbound

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation cites this paper.

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T18:59:49.358909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:59:49.358909Z digest=sha256:adede4744d58b407771b1f9fe34b91742222e7ac7088269d1d8e4816b36c5033

Observation 569809f2-468a-4b66-bc89-99f587b32288 · inbound

World Models for Robotic Manipulation: A Survey cites this paper.

World Models for Robotic Manipulation: A Survey DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:33:25.001319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:24:18.025364Z digest=sha256:26b7ca3b8ac6704bd3f81482f33d58e13d2efa1b9874bd96fa0258d27124f49f

Observation 3674c4fb-e992-4a3b-88d4-22353dcda17e · inbound

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models cites this paper.

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:08:03.376362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:40:54.963759Z digest=sha256:ad44c182f18ae0dbbe809ee32c7018b80825c89feace75f3ba8d152c67b27686

Observation f7914f00-e4d1-4c9e-8a25-2fdb018b5f55 · inbound

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining cites this paper.

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:58:47.599986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:25:39.450667Z digest=sha256:fc1e49263277a2d051263402f3357e261969536167d98be4c0faba77b4e971db

Observation f7bb466b-d7c2-499c-8d49-404f97cb8a5f · inbound

WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos cites this paper.

WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T05:52:34.589171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:52:34.589171Z digest=sha256:a9a0c48a68bdc0322561995afbea1b5517782160a642738bae86162adef54efe

Observation 27bdb77f-5671-436e-ab04-abf083406e04 · inbound

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space cites this paper.

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:12.785448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:12.785448Z digest=sha256:edf43b8b27264a31393b459a0671bd1c8672ba0e13523eef86c1623bbce54e69