Pith. sign in

Paper Citation Record · LEDGER

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

As of 21 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 45 inbound Pith citation observations for arXiv:2412.03293.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03293 v3

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:38:33.644750Z

measured 117 of 117 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:46:41.112893Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T20:46:33.613525Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c5b610e7-3050-4ca6-b99e-e6c9d8e90ca9 · outbound

This paper cites write newline.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:30.677122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:30.677122Z digest=sha256:d6999d376db4b9eaa2a924404bb2f0d734e73568c612befd3516b4ba94b28176

Observation ce9a55c8-ad49-4df3-9f67-45d492791481 · outbound

This paper cites GPT-4 Technical Report.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:30.780478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:30.780478Z digest=sha256:31853a3bf5e015d998838f3579819227ac6fb5b4c62a93ddff7c72d3a7654c01

Observation 8988af72-8f17-4956-ac21-e4a0369244b4 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning RT-H: Action Hierarchies Using Language

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:30.790927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:30.790927Z digest=sha256:9941a109aa2eccd79c79b95c859429dfa3ecb71c5fba0667b6f9a5025009ea84

Observation 3b40a6d0-53fa-410d-b026-b300f2b34edf · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:30.804831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:30.804831Z digest=sha256:642fa59b7c05a2ffb8aea9feef1f4eaab58cf80481f12f6e848414f0a95309a6

Observation f7238fea-80b6-48ba-9b36-354643bcb841 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Training Diffusion Models with Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:30.833183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:30.833183Z digest=sha256:d7ebaa66cc57bb88881ad47f74d49822708925b38b6455b97feb8c9a2552337a

Observation 23ae1d4e-49e5-45dc-9e27-42b8c999e053 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:30.854034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:30.854034Z digest=sha256:6aeaeacb2667e95b4e3ef336c1b26863c6760e2c7dc4c4f45a72fe960cef3363

Observation 35e41f25-e935-4706-8940-c9fb962d7cfa · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:30.884902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:30.884902Z digest=sha256:0437bbd28f58ccacb12a39a0c0d8cc0910c5bfe71654fd2dcab1696352fc1e47

Observation 411c4de9-b80d-4448-872c-95df520308b6 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:30.927195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:30.927195Z digest=sha256:9bb0891682de45e9dce4471b2b4fecf83de2231d54ec21efa56971e9a9497121

Observation d68b3fe0-05ac-4fab-8d7e-8ada58bf5b1b · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:30.997068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:30.997068Z digest=sha256:831e43f9f357b72c8f7724a5df8bc415bb13c2049501dc00a73ccd44c2d1c769

Observation 47eed8d6-771b-4ac5-9787-1e1476d55680 · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.106963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.106963Z digest=sha256:67fce35235af457caac938859dbb37a00bb648fc757c127587ccf281c232dbee

Observation 9673ca89-e7cd-428f-9f80-c7345940c207 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.192316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.192316Z digest=sha256:b53c9456e383c6f962e8686a7c7c84e8181a29b6338f2201cd1c5f7fb3373a6b

Observation 8654041c-5c7d-4c37-9c60-69beb8382a0e · outbound

This paper cites The Ingredients for Robotic Diffusion Transformers.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning The Ingredients for Robotic Diffusion Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.198772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.198772Z digest=sha256:1a0c6f2fd0477c003e3201d7424d9673ae97d4a53b38c8d67b302f7acce1211f

Observation 623c01a6-f16f-46ce-9487-72eee06c3eae · outbound

This paper cites Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.204174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.204174Z digest=sha256:9c94a09656d8643b607dcc4fb27058d7175e64da015df04c08713948219ab3fa

Observation b8eb4d62-afb3-4f65-aa95-4ab6f8efb091 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.228000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.228000Z digest=sha256:98fea80f66678db3b670a724e377e3e473746abc6bfe8394b7473fc9849fa394

Observation 729ec95a-3061-4e01-b25e-5522133fa68b · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.238409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.238409Z digest=sha256:0570097eeb5e88395c0b6ee84bec0efd58cfe26b872735a12775743762def70a

Observation d0f9babd-24e9-4389-9d6a-4355895f37f1 · outbound

This paper cites Denoising diffusion probabilistic models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Denoising diffusion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.249751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.249751Z digest=sha256:c0ce94a8121dd590951b877077e4d5e430d8f93d6552b9587240714bf570fa53

Observation 05ff5274-167d-419d-a797-5fe3016b3c56 · outbound

This paper cites an unresolved cited work.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.283882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.283882Z digest=sha256:1732e51bc15703760248265fbb5d69a186da6165d9b25eb460526621a5cc31eb

Observation 964c2fcc-c986-4efa-b198-81242e1d7217 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.298023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.298023Z digest=sha256:a351ef98d1c1615d8af45505aeb3238a9e1f45535947fdf5ee1eac344eddc9d1

Observation c08b0fd2-20f8-4d08-811a-73c50dc8f6c3 · outbound

This paper cites An embodied generalist agent in 3d world.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning An embodied generalist agent in 3d world

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:38:36.794770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T22:38:31.306907Z digest=sha256:bae637ae2e86c28a53ce6b13221c3bd586af13a81aba2522d71f5235136f25ce

Observation bf710902-d1d2-4757-9f14-db4d63a31fe9 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.377236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.377236Z digest=sha256:f45a97a609477065ea203baa1a762de2a2be4b15bc2eb1199edcf89b121567e1

Observation cd54b650-a9a3-425b-a5f5-d1666ebb84ba · outbound

This paper cites Mail: Improving imitation learning with selective state space models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Mail: Improving imitation learning with selective state space models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:38:36.630936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T22:38:31.496666Z digest=sha256:04edac25a58e9126d7fae029a32beaaa3cb642d2ba4b0ccb80b271391e2bd595

Observation 25b8a85c-1538-4133-ad85-961a276b187e · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning VIMA: General Robot Manipulation with Multimodal Prompts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.617037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.617037Z digest=sha256:31e153c5462dee69e88c65d00e4626240efc1ee5c3ff05c3bd4913520223cdbc

Observation a119d026-f345-4dc6-893d-48081e5e0530 · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.632303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.632303Z digest=sha256:49a085d31b0bb605035b8c9b65a25e840ff608d07d30ddbfba7816f07e4ea0b5

Observation 6e24741e-ea5a-4186-bf5d-19ea7199672a · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.658250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.658250Z digest=sha256:bcc91405a1c939053eacabad8220014ca3d8dc2c5146fce78292200433561bf7

Observation 1f2a28fa-9b09-4a85-b7cf-71318b4067c3 · outbound

This paper cites J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., et al.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., et al

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:38:36.584821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T22:38:31.672368Z digest=sha256:913ca3727da0364be58c05b61495e1b445bb60eb3d469fa75a6f9e5c64a8a715

Observation d32c8843-2d7d-400c-9ccb-e39b87bb0c2b · outbound

This paper cites H., Gonzalez, J.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning H., Gonzalez, J

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.725170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.725170Z digest=sha256:c91cb67da95142f64b7e90dc86a3a0b9524965e52b7191b006071258b4f4b010

Observation 78c6b8e2-337b-4e65-ae44-19436bcf22b3 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.816900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.816900Z digest=sha256:e8c59f9d11c7b93ef0cd384c1fd9470db53e150b179e8f86b30e3add5add12ab

Observation f0b01fca-0b7e-41a6-924d-06e2317b4e49 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Autoregressive Image Generation without Vector Quantization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.911890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.911890Z digest=sha256:be25891b342618217e60ae0f6fee2b1e48dc143aa3fcd1e061f99275fea33aa3

Observation b8749297-084a-438f-8c1f-37bc88f77115 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.939665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.939665Z digest=sha256:3c9c73c799db6aa34c9bf0a42e221553e38c07cedcc35922a3f1041717ea3e22

Observation 57d67beb-2b20-48d3-9fc5-fac67a0ce124 · outbound

This paper cites Code as policies: Language model programs for embodied control.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Code as policies: Language model programs for embodied control

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.958422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.958422Z digest=sha256:f793c3f02d89f29c1ab57ecdd9a56aa848dec84ec5fc450ff34c5b48b0c40ca3

Observation 574eb622-8a13-44cd-b9f1-184eec9a40c3 · outbound

This paper cites Data Scaling Laws in Imitation Learning for Robotic Manipulation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Data Scaling Laws in Imitation Learning for Robotic Manipulation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:31.978533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:31.978533Z digest=sha256:78799271f194ed507455e8fe1d9a153fe4e283060a8c2df6ab813c6d04c72e43

Observation b90df923-f651-4c32-af31-887263ad173b · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Improved Baselines with Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.001601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.001601Z digest=sha256:e67e4686b0f477fb3416c7039e0f1c50eac7eb623524b6d09fea3f9b19785d4c

Observation 5bed63dd-9dda-4841-96c2-a033fb293541 · outbound

This paper cites an unresolved cited work.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:38:36.390407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T22:38:32.017289Z digest=sha256:7a1c8d741bdda142252112d9c3eae413c11d8d565fd186d692bf0108edaa1f76

Observation 36796e7f-eca3-4d49-8091-0668e16b3ff7 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.131516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.131516Z digest=sha256:aba580d197da065f5025cda4471ab9d78494d45b9243a93a01d5f458fa3f8492

Observation 7d16bbeb-4fed-4692-b251-d78ef42bd36c · outbound

This paper cites N., Zhu, S.-C., and Gao, J.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning N., Zhu, S.-C., and Gao, J

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.189701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.189701Z digest=sha256:1d4a078450e3fd861d2524b6f749fc6a3da8828fe4b6f0b617546842427457f8

Observation 39c46c55-c681-42a3-af14-3de9732f17b1 · outbound

This paper cites Y., Sanketi, P., Vuong, Q., Xiao, T., Sadigh, D., Finn, C., and Levine, S.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Y., Sanketi, P., Vuong, Q., Xiao, T., Sadigh, D., Finn, C., and Levine, S

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.200125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.200125Z digest=sha256:4c5e2ab1394101dd579a4d2bd203bb572a423bdbeb7aa00e41e2974cb45f508e

Observation 41a92526-83c9-4872-8aa4-498af23715c5 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.234754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.234754Z digest=sha256:41e2864ebd6d80b8252824a35bb97c8a618e695e41e4c2d1f6ea36c29f379565

Observation 55809c25-28cf-4000-bcf3-7b3f96836adc · outbound

This paper cites and Xie, S.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning and Xie, S

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.284752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.284752Z digest=sha256:ef22f932f5537a5c1a16a0a8f482762c01c5d771086960686d6f70dbe23106ca

Observation 4f816274-f3f3-484c-a2d6-b6e43a843cb7 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Film: Visual reasoning with a general conditioning layer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.319035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.319035Z digest=sha256:d9d39b245f2574ba354694fa37e3e9f4dac62a3bdfe82a6ace50549cccf436ee

Observation d1de3bde-3460-42c2-bcbc-12872c3c58c5 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.344745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.344745Z digest=sha256:1e05f4585501caf7c676b3243047cb38775dc4de3b7b7dc98e6c10ff061cd4aa

Observation f61ff552-8b34-40d9-bba2-222cee5191ff · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.442886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.442886Z digest=sha256:3b379c8e4a4aee588d7a017bdec15bbf9fb6e07f26d4bd197fa2ea3f4cfa49f9

Observation 916c612b-97a2-4976-810c-6842b1377aef · outbound

This paper cites Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.583543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.583543Z digest=sha256:fc5d098a43322e70bd7ae6d10df582c636e3a580bd6b1a4ab9795680c3e4819a

Observation 8714406a-b6bf-4b22-801a-89beb6803205 · outbound

This paper cites E., Wenzel, F., and Lioutikov, R.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning E., Wenzel, F., and Lioutikov, R

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:38:36.245525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T22:38:32.645819Z digest=sha256:53a0bfc60ba96c95d677330b68ac6dba02ba3a92da1c2dcfa37347b9330f22de

Observation cd6de92d-f06e-4ef0-a863-ba345ebf67cc · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning High-resolution image synthesis with latent diffusion models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.663912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.663912Z digest=sha256:f8cee7251d32f975cab64b7be1e3485ce065590355a0e0a3e2f6f6be06117dc3

Observation f5d58f4a-9bc9-4950-8a4e-1632028ae561 · outbound

This paper cites Yell At Your Robot: Improving On-the-Fly from Language Corrections.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Yell At Your Robot: Improving On-the-Fly from Language Corrections

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.675853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.675853Z digest=sha256:dd9af5337e7b9466632498fe8077cdf3aa6192b100968fd60d6967e64e48837a

Observation db9df151-978d-49b7-9617-02bd0b28fde2 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.693520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.693520Z digest=sha256:abb018ed27a929456c7420ba6c354c6d5653709a40e1a7be83fd3f4b648c208d

Observation 488b31cd-c809-40dd-9f10-9f37ae859311 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.711089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.711089Z digest=sha256:81fdaef830aab7ffa21b2119064048d7f951c7b9ead2ab2446a9d1c43112c839

Observation f18134a4-7b39-41be-9825-d72caaa54dc3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.724782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.724782Z digest=sha256:94a6f24427178137dd2aadfe8f2c8bb9cbfb27349a3a14e07f048706b8598f72

Observation b0bf8ac0-ff76-475f-8ffc-1f571d04549f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.736703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.736703Z digest=sha256:7baad73b90aa220ebb6a9bdfd52bb6d794934c2d18c1f0abf710f0787aacecc4

Observation df6887dc-8750-4568-a4cd-1e143c92e2e0 · outbound

This paper cites Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.747952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.747952Z digest=sha256:3e6324f9a3b8dc13e7e4f3936e05a75839a160746e952137a52c39e74874289d

Observation 34b7d7b6-4968-4407-88e5-aca06de7a728 · outbound

This paper cites Feedback Efficient Online Fine-Tuning of Diffusion Models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Feedback Efficient Online Fine-Tuning of Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.763447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.763447Z digest=sha256:be6b18363b8bbafdd17b2b5d24f100aed1e4d4864dc507affd622f9b06730b17

Observation c4662172-1294-4182-9055-36e4792b23c4 · outbound

This paper cites Equivariant Diffusion Policy.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Equivariant Diffusion Policy

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.856741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.856741Z digest=sha256:18a627862b91a2f417cf359848cd98010dd72b5e16765bd524c8c44dbca961bb

Observation 2346503c-9fd5-4f7a-9c83-69900fa373b1 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:32.964702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:32.964702Z digest=sha256:0d8f38ec1fa086d33b940f61db505bf34bf91f96de15e1c1b476537b8287d8b1

Observation 9e683b69-97cc-406d-8b30-eb47bdb315b9 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Emu3: Next-Token Prediction is All You Need

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.037973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.037973Z digest=sha256:06a686fef3b036e2904b74e94a57f03a5687e25f61104437fb39d3508267d052

Observation c52154f4-3eb1-433f-889b-9d861161e450 · outbound

This paper cites Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.057294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.057294Z digest=sha256:f3983404d0699ed83b4c50b2695f2986ea3ec3fea7b46c157615acc4c3243750

Observation 7f209aeb-04a0-4d37-bd5f-97b90186db77 · outbound

This paper cites One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.067036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.067036Z digest=sha256:f180d31cfc25e3c068e496381176faaa0efca7ebad99b286ee9fd7feacd61350

Observation 08132ec9-081a-49e1-945e-bb61426d9ae3 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.084810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.084810Z digest=sha256:2247feb2edb902fd4ff321185f100db0d120fd8715f5ae3a621f581fd415b227

Observation 482c0842-30e5-42dd-b13b-d4c2ea8d7a12 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.162000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.162000Z digest=sha256:c52764078c9726aebb61b73d4ba180ffc7262daa52d9e96baefd649dc7ba6b98

Observation 4bccf384-dd94-4b41-96d2-0b78fb3acc59 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.263595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.263595Z digest=sha256:980e23cc38d4dc0c3479f959df692d800b032815eef4698f536a1517d06046fc

Observation b223495b-6eb9-4353-8289-ee0d35baeb16 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.276125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.276125Z digest=sha256:df600976e301ba5cb9ed31e94c165124581c7f66fee9d2f4e3562ed78d15dd68

Observation 4cf806c9-31f0-40a4-a1fb-bfea83757858 · outbound

This paper cites DNAct: Diffusion Guided Multi-Task 3D Policy Learning.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning DNAct: Diffusion Guided Multi-Task 3D Policy Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.292167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.292167Z digest=sha256:78cf6a904245800b8ffa1330bdcdc4da2849cdbc1c76531968cc08abf8a6a98e

Observation e425115c-9b20-484f-b73e-69965c97da97 · outbound

This paper cites Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.305761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.305761Z digest=sha256:6e8109c599eb1022096370a7033491cad593b394fabfcaecf816591fdd3bc2f4

Observation 62c0834a-6852-4d4d-a3b6-1e7d257da7a8 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.311794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.311794Z digest=sha256:1168d9447fe40d39c996bf89534c82524aaac7df71d2ac00e360892e9413bdbd

Observation c19f2071-5ad0-4411-89f7-3ea4765360d0 · outbound

This paper cites Generalizable Humanoid Manipulation with 3D Diffusion Policies.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Generalizable Humanoid Manipulation with 3D Diffusion Policies

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.319212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.319212Z digest=sha256:0c9e31340146091a2b88447b9cfaf82d050c5bc9683f04615d36f8fd5d817999

Observation cc3f6a21-794c-443d-80db-8c870a4fdca3 · outbound

This paper cites 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:38:36.100844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T22:38:33.335420Z digest=sha256:01c75fe501f2ad048ad52a81b97995fe9fba9f525d17c31b223ba7e2f98bf97a

Observation 3d2d6e36-aafd-4c04-93f0-f236c19e2e5c · outbound

This paper cites Sigmoid loss for language image pre-training.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Sigmoid loss for language image pre-training

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.373560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.373560Z digest=sha256:9c864797cbb40f4944cc36d27c1c46417090dbe2e2010498c44ce206b51fbeb8

Observation 83c7d9b3-d2d0-41d6-bf33-e4f2503b24f9 · outbound

This paper cites GRAPE: Generalizing Robot Policy via Preference Alignment.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.391932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.391932Z digest=sha256:d4b9f56c22984fce985971c003f8efaddb92def7aac6a16babcfa74e5ff916b3

Observation 4b1abf3e-2b42-4f9a-84d7-5855e4419e12 · outbound

This paper cites MonoFormer: One Transformer for Both Diffusion and Autoregression.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning MonoFormer: One Transformer for Both Diffusion and Autoregression

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.407074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.407074Z digest=sha256:93d08ac17f4a80db837ade91eb7ed5d6e9fc0ede0655c6a945072548be4c8ca4

Observation 485b140b-6911-4e63-bb00-6b4b2a2b7266 · outbound

This paper cites Z., Tompson, J., Driess, D., Florence, P., Ghasemipour, S.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Z., Tompson, J., Driess, D., Florence, P., Ghasemipour, S

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:38:35.887897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T22:38:33.455476Z digest=sha256:eb90413e3d3c160339927cd783cce2057f17cb85b5764d5c6328d0018411d1bf

Observation 953d36f1-aa1a-472b-8fdf-077e28448eec · outbound

This paper cites Universal Actions for Enhanced Embodied Foundation Models.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Universal Actions for Enhanced Embodied Foundation Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.482273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.482273Z digest=sha256:2853215071472727bfcc707965124741b831276471bc9b92b5115ef6e8efef34

Observation a913c615-beab-43be-b4f5-77e54065cd16 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.554833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.554833Z digest=sha256:d4cea36bbc9bcff913fde2f8fc0a3f07e3a4ed8e96cb10461dd9147ef15016bc

Observation 52850c6a-4a89-44d5-a47e-d46e82a0af22 · outbound

This paper cites Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:33.644750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:33.644750Z digest=sha256:36f7169eba4478364f539383fedb5b14de9b27ede961650b05f5dc75806f5834

Pith citing papers

Observation b144605a-384f-491c-9321-e6499167c0a6 · inbound

Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation cites this paper.

Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T10:56:57.864336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:56:57.864336Z digest=sha256:6b3d0c3c31b1df4ca706973a624ebe422de7618c203b5398451e0fae9a1ab197

Observation 3d779480-121c-4ae5-bfe9-f95b0201292d · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:48:48.910288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:dcc9a0cf869f1b96e6e6fe069c3a33850a3423018687352c06a11f60c668d61b

Observation 38067976-547d-4fc8-8882-fc257db786e3 · inbound

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning cites this paper.

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:06:27.250851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T20:06:27.136345Z digest=sha256:4abb92bae779c754949d145b751b81e817ba04cad0b5f182fdef62df5c718bc5

Observation 5305cc76-1475-4aa7-8624-537f17e4a68a · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:00:48.879916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:01f0a3b5634451f6a20b664a4a5e4a600993a8979779f6f06a0c5a058320d92b

Observation f2c9b90a-b4cf-4de9-8722-36443e179810 · inbound

A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI cites this paper.

A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 171

Resolution
unresolved
no resolver link, observed 2026-08-16T04:46:41.112893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:46:41.112893Z digest=sha256:c5e2de1949c2a97d9d5418f848bb652000a3e6d97afa5424b3712cf61a878a64

Observation 795c6801-66d5-4d2a-8c02-557e8742822e · inbound

Conditioning Matters: Training Diffusion Policies is Faster Than You Think cites this paper.

Conditioning Matters: Training Diffusion Policies is Faster Than You Think Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:53.625518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:53.625518Z digest=sha256:a2efd05f3f22a40569dd4ef1e3321b24e81978c7d9b138444b82064e55762f43

Observation 21a7a0a9-4637-46e2-9892-ee2983bb4cba · inbound

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions cites this paper.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.377189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.377189Z digest=sha256:541d731c6aa8dc66c73c950cacb9bb51d80c1193e4b1ed24df5942299ee798da

Observation c555925c-c09e-4422-b2d3-13797fab3728 · inbound

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion cites this paper.

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:52:17.848780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T12:50:17.902979Z digest=sha256:7e26d0b4dcc3d04d033c6dbf831c447fb61dc8f3aaa3fcada1030529e04b0497

Observation d9b01b24-1d17-4365-b681-446034d39b78 · inbound

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models cites this paper.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.455619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.455619Z digest=sha256:5c50f348c685581f7466f0db9854353a6cea6f0bc79c635cd5a4f55cb2328480

Observation e2f290be-f8a1-46c1-8e1c-654bbba7905d · inbound

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge cites this paper.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.902047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.902047Z digest=sha256:d7fa861b35d0315217febd9df64928d558c36ac8a8dc6d18b260ee7f1bab5acb

Observation e1d03a9a-3438-4690-a6a9-9da73d0da9e6 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.449497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.449497Z digest=sha256:47ad00eee0fe038035c6daa75ead2e3a568021721b5d64e5215fb119744b1954

Observation 7ecbe182-66c7-4eb9-b7b9-233a0db0521f · inbound

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks cites this paper.

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:53.176992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:53.176992Z digest=sha256:b7d07db7977fca13a7de5c971f826d196d472a2567733bbeddad578d832837d2

Observation 086f3048-ec0e-44a0-a2fb-13ec3e1bccb0 · inbound

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models cites this paper.

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:22.391752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:22.391752Z digest=sha256:7ea7026edcdd7007d2b34523b3941fcad5a1e0841c125c96dd795b764abaf228

Observation 3c9762e3-034f-4bc4-9d83-b3e51d650d2b · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.614110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.614110Z digest=sha256:ecdb2a577cc968838a558c553500ef94bdfb0ba2933999edfe2e20e6c393520e

Observation aae6aa57-c1ab-4931-96f6-b093dcdad595 · inbound

Block-wise Adaptive Caching for Accelerating Diffusion Policy cites this paper.

Block-wise Adaptive Caching for Accelerating Diffusion Policy Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:14.009891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T09:36:09.790248Z digest=sha256:0a1b31ae76f45fc5f6afc7fa83e4a312813699c258c84bd13eb2d58902b7296f

Observation aa52847d-d72f-4257-a581-a7e6d0f8791a · inbound

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models cites this paper.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.540220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.540220Z digest=sha256:e90b9ca6dceb55746c4bdcb7700c97d572c031d17fbe1fe86a793a1bfd4f7f70

Observation e9708aa9-d99f-4b06-b3be-99c2970666f6 · inbound

WorldVLA: Towards Autoregressive Action World Model cites this paper.

WorldVLA: Towards Autoregressive Action World Model Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:57:08.303232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T22:57:07.883617Z digest=sha256:3e47904237824ded29a69db4bbab770535a47901e86b78e99a325dd377094038

Observation af514658-c5c6-4e6b-b026-931977c0abe3 · inbound

A Survey on Vision-Language-Action Models for Autonomous Driving cites this paper.

A Survey on Vision-Language-Action Models for Autonomous Driving Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 130

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:04.778023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:31:04.778023Z digest=sha256:d46abeed94e733b6933df1474b05b18ed18d78b5403bf054844dd062494b0674

Observation d23c1da5-a936-43db-a192-fa8f83e95956 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 290

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:08:35.427922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:4200aea4f6823e094cd9360607309c51ef89c076b1f66ff807217ff8f319a38e

Observation 6118946c-2d91-49bd-abf4-c75788a0bb9e · inbound

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation cites this paper.

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:30.607811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:30.607811Z digest=sha256:85dc3fa5a64b6d41d142297e095a555c419c12c8e588f0b8d4b78c3a8687384d

Observation 8a3068a2-2413-4b17-893b-59a5f87d8200 · inbound

GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions cites this paper.

GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T21:59:03.506693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:59:03.506693Z digest=sha256:c9b4443af9d5942104ddd72a8965483fae7d14dc82685382c2aac7df378b9b19

Observation f71475a0-c916-477a-9efa-81b420fe9b5b · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:46.932257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:46.932257Z digest=sha256:8b38603d47b9bb357e1a412945e670ed0c2adecdad9e1fb219c86ce17b940997

Observation 1bd6b839-6ce0-4d8e-92d6-093f9ec6d4b6 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:44.250854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:44.250854Z digest=sha256:9682ecd7b82e8f2c22ef83c073fc5bd8318f0b564d3d1f7d2e091db503073170

Observation fc906212-3aac-42d2-9d8e-004341d43b30 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 202

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:54.311759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:54.311759Z digest=sha256:12864e30cd4c9da70612848cce09124aeaf2309566223588efdc7d1a4192d677

Observation 57d86c72-065c-4732-9470-a42dcad4b8be · inbound

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy cites this paper.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.951601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.951601Z digest=sha256:cf6774dd46c07f735f85aa5a631a3be015409f5d17ff70e65cebd8cccd91264f

Observation 91f54363-3107-412c-bd71-703477719c40 · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.085685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.085685Z digest=sha256:2ddd981e0072b8205f2fe22d0f3fcb933bdb5883a62c9d1c54f2ecbf6406ae4d

Observation 0df86736-05d4-46ee-ac7d-78bfe38ee4d9 · inbound

DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation cites this paper.

DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T17:00:45.116074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:00:45.116074Z digest=sha256:7e1b89b1be510251e8f259a547ea1e5c0976da0141c83d25be69282da37bd20e

Observation e5e2c21e-164c-4a1f-9e3c-9451662013ef · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:30:18.333293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:9bc605abdcedf31bf96669ab9897410bd80fafce782bcfd7e1a5cb7ec16c1b84

Observation 7d87aa4e-6461-42e9-9273-272d44758ff9 · inbound

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance cites this paper.

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:07:47.830757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T11:06:23.555610Z digest=sha256:f2024428892c6553d8497e1c4295c3e48901a06ae54f22dfae58bc64ab967c76

Observation 23b91943-9cd2-4185-9a9d-0f1853a914b5 · inbound

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation cites this paper.

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.882539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T13:13:53.818915Z digest=sha256:d57ef42441f6131e5877a41cfc019eb60f8f4bd0e367b2ffadb836105af07d75

Observation ff9e5053-b10b-4e09-afd9-62bd655ea456 · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:30.125617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T05:47:17.494531Z digest=sha256:c31bae4e4e159e61eeadf6ba28bf70ddb2a1217c1b5dd51af0b75fea678d01ce

Observation 5853d067-9cc7-4071-8963-0320d87bb4b7 · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:17:03.723898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T03:00:26.352130Z digest=sha256:88c04af107ec48061cb9b8845c22d09a0b7d349d3ecc2676bab5ce27516bfef7

Observation a26384c0-ff05-4852-bd17-6c093afdb86c · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:25.650502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:08:43.222818Z digest=sha256:52816db52217ee3ff830e4fc9fb3275f8600cf32eb762f652f6d3fc39ca6e8e6

Observation 889590f4-6d3c-4e1b-b2a3-3c0cdee70821 · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T14:26:02.341925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:26:02.341925Z digest=sha256:be5a77cd0df53b4ee721432e4ce6b4361d553b29634a669365303c26594759fb

Observation 975ceea2-8c42-4770-8a5e-878814d0c7ff · inbound

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models cites this paper.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.315613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:aa3279f6ff07ea4e1492e950ac21bf262bec008145c83aa23c366b945c803d55

Observation 865d3931-e91c-49d1-9d11-336a5f39757b · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 269

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:18.129736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:de6a3016e2379db5935dcf5b50a2da6a79311574c2b62a1316198568d7d8e086

Observation edab43d3-6724-4e1e-8d86-c5198978556e · inbound

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data cites this paper.

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T17:57:33.179561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T17:54:50.325820Z digest=sha256:ed84b0608d681df939993086529985166aef2daf6f931da961b38bc91bfd296b

Observation 143f74b6-e385-4553-b953-ccb3bc3f422c · inbound

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding cites this paper.

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:59.255359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T01:23:02.576098Z digest=sha256:17f8e2581de9fcce7bfd2ac0c0a8e626f8c81cb52c6b0cdc5a42d28bf3d81d5e

Observation cb00d7cb-1c93-4ede-bd6e-2617738df0b7 · inbound

Diffusion Forcing Planner: History-Annealed Planning with Time-Dependent Guidance for Autonomous Driving cites this paper.

Diffusion Forcing Planner: History-Annealed Planning with Time-Dependent Guidance for Autonomous Driving Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:27:40.190980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T13:14:28.378548Z digest=sha256:f9ff1d606b5ed5c2a63e0ba191c7ce0e70374f5dd952d355b3eef82320e24b4f

Observation 74a36b4d-891e-4346-9e6c-f687633bd2f7 · inbound

MV-WAM: Manifold-Aware World Action Model with Value Augmentation cites this paper.

MV-WAM: Manifold-Aware World Action Model with Value Augmentation Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:38.129392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T14:36:42.051049Z digest=sha256:fee9378fc644f46100b8caf7228e3e580720fc8904c5111bf9c00b0c1105b56d

Observation 266c8a36-4ae2-4add-b35d-0d376db6bc53 · inbound

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models cites this paper.

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:19:47.613178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T08:57:01.861091Z digest=sha256:2ca90eb672e76ceedb369f15dae244e2c642872197773977138b79c00867d847

Observation e7f4277e-4c42-4440-8bf2-80fdf02602d3 · inbound

LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation cites this paper.

LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:29:50.765620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T07:58:17.225491Z digest=sha256:f679694b8a8ed62f3e4e14029c244a059a18458460beee6980cf0c30d81c0969

Observation 08fc0924-eb0a-4e2f-bf69-e288d02483ab · inbound

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning cites this paper.

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.797302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T06:56:45.197407Z digest=sha256:48b90b9da1f3c5b82cde78aa7b3fff9c53c11a7712d990f1a7075dd8478a4b15

Observation e6bf9257-e101-4579-9ee1-11f270853ae2 · inbound

Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation cites this paper.

Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:44:26.114231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:41:34.889711Z digest=sha256:c70e9e4397cee860f1457703319bd783accedc3a75462aed1d2447b6c6474acf

Observation ebaeb5c1-0c0e-475e-a0b0-6f9c90292ac6 · inbound

PriGo: Test-Time Primitive Guidance to Diffusion and Flow Policies for Adaptive Robotic Manipulation cites this paper.

PriGo: Test-Time Primitive Guidance to Diffusion and Flow Policies for Adaptive Robotic Manipulation Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T20:46:33.614832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T20:38:38.404267Z digest=sha256:6d0b8ccd38b4ac55c88333e92e3f958d4268ca1d981b2072faf748340568144b