Pith. sign in

Paper Citation Record · LEDGER

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

As of 5 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 100 inbound Pith citation observations for arXiv:2504.16054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16054 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T18:02:23.305313Z

measured 193 of 193 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 708 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:42:32.635806Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact49
  • verified fuzzy40
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 3f3e3f5d-95c4-4dba-80a1-873ce12d6d39 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:01.002542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:87412531972f2ae5549824768aeceaf0eaadebec65a0c6d06f34b3b4a6ca5059

Observation 78d66ea7-8622-4c83-a5a6-a95c2d5e91be · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.997231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:8eab15acd9d19b55a4aa4279dace13a99295b8ee647a43a90f5e0b91d7607a15

Observation d45c2814-20a1-40ab-bbf5-4fccddb9e364 · outbound

This paper cites Minivla: A better vla with a smaller footprint.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Minivla: A better vla with a smaller footprint

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.451836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:add40b899b1641c6662bcefc0e694e6a99db7bdedc32cbd9ad242fb7f3e130c7

Observation 1222a5a0-d94e-4da3-9bf8-337b50151eb9 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization RT-H: Action Hierarchies Using Language

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.991878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:41434e957e8ff2a76ee9a92e69dbb0973bdcd9d4c6277842988b4b8ed184098a

Observation 89e3696b-734e-4753-bf03-c3f2e0044103 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.994437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:ee6c27d46240493eb79ea477072085e5ebc9e8fd1f5beca81a43189c68162dbe

Observation 328a572c-4a08-401a-a4f5-d2da11190dcc · outbound

This paper cites Roboagent: Generalization and efficiency in robot manip- ulation via semantic augmentations and action chunking.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Roboagent: Generalization and efficiency in robot manip- ulation via semantic augmentations and action chunking

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.460975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:3a72db80b0ab41867b132ad4e7da890a3704365249b2ecf7756f803ecef1d53a

Observation 10b90701-b7af-4b2b-9ebe-eac35d252f1e · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.980064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:251068cef085731a7cd59f89697f473818dda503e59fc798038a7700ccab8690

Observation 5e556176-8e55-4d4a-865a-26b45d0f82cf · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.986329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:9f0c501a07caba69e33e323858cba07f8152ad6893e9fac6dc98231efe52a995

Observation b1510de5-0c53-457a-8cbe-ebc43ff4fc69 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization RT-1: Robotics Transformer for Real-World Control at Scale

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.983104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:1f7c5b119aa9d9f472da3f203cfacafaebc17fe599f6edaaa218b6913aa03e91

Observation 4f1d3e48-3b07-4bb0-ba34-beeac5598e02 · outbound

This paper cites an unresolved cited work.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-22T18:05:01.459125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:86c5720f9f6a5b2dad0fc3084ee7307cb0ad6591f1acb8571c49a7f765359055

Observation 1b3ceba0-937f-4688-9dce-193331f2c144 · outbound

This paper cites Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.974652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:22f3d89ed5c6b9c3a987310f20fba7cc90285bdebe934edee8e94def821e39cd

Observation 75e1d32f-7ffe-4658-b2fc-ef8ea69bf78d · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.977492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:c93c7aa944b5c8b0fb6fdbda66e24fff444f2faabbb741002dead4239e3a8a85

Observation 64d3b1c8-7ccc-47a6-9849-95ad8e8a46b4 · outbound

This paper cites NaVILA: Legged Robot Vision-Language-Action Model for Navigation.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.989303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:f28e64f5045f43bb4848422d46b20f2cec999b83dd466707e9c5ba737a3c5769

Observation d5300036-5645-4478-ae0d-b91b5a3f4a5d · outbound

This paper cites Universal manipulation interface: In- the-wild robot teaching without in-the-wild robots.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Universal manipulation interface: In- the-wild robot teaching without in-the-wild robots

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.448514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:5ba046f2229a099c5ba00991e5eaea89d60f4dcd12cf89050e4113b30fd6f9b1

Observation d24f5f49-717c-4aa9-8545-1695fc521456 · outbound

This paper cites Racer: Rich language-guided failure recovery policies for imitation learning.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Racer: Rich language-guided failure recovery policies for imitation learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.457516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:8413e1e42c48d8fb4095016e3f4598adbdf77b9f0a8514fb39a83f32cb3a1330

Observation c6b63635-db63-4cd0-9250-900f99b42ee8 · outbound

This paper cites Robonet: Large-scale multi-robot learning.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Robonet: Large-scale multi-robot learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.462891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:e34e14f629f394512b865c6915cb652cc2089864994d9ca74a03e7f9d67668bd

Observation 018921d4-d02f-441f-862b-aefff6bb0a18 · outbound

This paper cites An unbiased look at datasets for visuo- motor pre-training.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization An unbiased look at datasets for visuo- motor pre-training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.464682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:a73b75df9d3ffef4781a0752c6e9c61a3126d20c2245291af5303748db3989b2

Observation 4903c79d-2e7a-4715-ba10-a683ddc771bf · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.971979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:abf20f990bad62f0ceb51e543831f5e6bcad91df1e2c0daae054f930d9266e0b

Observation d2b55c9e-f7d7-4036-804b-37bbcbe7668e · outbound

This paper cites Reviews-consumer technology.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Reviews-consumer technology

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.450122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:9e500a7eecc01fb4b1ff87cd4eb2b1e2beb5042f91fd65b4b97d48578cd02f2f

Observation fe721ed6-4894-4136-b725-f82d08392f76 · outbound

This paper cites Bert: Pre-training of deep bidirec- tional transformers for language understanding.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Bert: Pre-training of deep bidirec- tional transformers for language understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.455658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:e683527fb7838608f2618dcc7cf6d274309108d0738b81a0a66fa6b36f87a9ea

Observation 5bdf743e-2472-4524-8b28-1ed0ca327ba8 · outbound

This paper cites Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.453766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:b1eb32955a788431100be75cbe2cd74e31784c5ef6cc13d3f38edbc4b50ca358

Observation fc9d53e5-909f-454f-8b1a-a95d15c0e966 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization PaLM-E: An Embodied Multimodal Language Model

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.999789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:7ca6326196b6fd9cb2a9872ba7d46ffe95755780daa48281841b80b141cc9698

Observation cb2b7158-1d5a-45ee-b1c1-b61fcee8e625 · outbound

This paper cites Manipulate-Anything: Automating Real-World Robots using Vision-Language Models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.865624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:2f0dfa9eca11435fdd2b123d79bdb889dc7b7f3a19f2022af1e722e756bb2509

Observation 50ce92dc-96f8-4969-851c-c9724a6bcf1b · outbound

This paper cites Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.966765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:c924fb902d066e3dd5f3c37b7481f050259f2168264cf3dd2b571ea1c20957e8

Observation a416a26a-ea1d-41c0-a3d2-31ec3d1d38ab · outbound

This paper cites SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.923439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:b143f2eea6c55a1df980aa9106aa2376819110e65562e6da0d251550d62de73c

Observation caba3e9e-ae05-49b6-9d2f-545baa5da75c · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Scaling rectified flow transformers for high-resolution image synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.399541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:f34a193251a2c8260876308dd3bfcf48e30f977a1c2de9384b9b26720dde2136

Observation 60b016cb-36e7-43bc-b150-aef360bd979a · outbound

This paper cites Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.920809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:68f302febe8a9db3666433f36449634ca1256cf842f0e092ddd4fee90fdfc5cc

Observation 7c5746f0-1f9f-4cd6-b3f6-99c53ab6217c · outbound

This paper cites Anygrasp: Robust and efficient grasp perception in spatial and temporal domains.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Anygrasp: Robust and efficient grasp perception in spatial and temporal domains

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.397823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:dd40bd3e136fe990d34ec6a7d0132904b1e53fc41c30282421c09e0d0bb02427

Observation 41def0df-85e7-496b-bdca-b1d8718a43fe · outbound

This paper cites Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.428251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:65494bcd5c81558e7e6785399fe2218e753cda084afcfbe167744eea53a588ba

Observation 679b4fb0-974f-476b-ae0c-70c04fc373c0 · outbound

This paper cites Navigating to objects in the real world.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Navigating to objects in the real world

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.442969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:7fc725575833c747ccf11c77f08b36d592947446deb205b0615ef0d26b5b4784

Observation e9b44355-9293-4a62-a9a3-8778dca84f67 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.446733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:03d9c75703e80655631c371b59fb9abbc859cf210b76aeade1b925c7dd32c487

Observation 0fb3d65f-0631-40bd-ba3a-172e9030b655 · outbound

This paper cites Robot learning in homes: Improving generalization and reducing dataset bias.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Robot learning in homes: Improving generalization and reducing dataset bias

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.413044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:9fcadcbc2c98bfd2a1c47316b8aa86a946d5bd44a8a5a387dc42ed9081d2ac23

Observation 8db603b6-6d48-443b-818d-bdbbf9c95b76 · outbound

This paper cites Deep residual learning for image recognition.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Deep residual learning for image recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.405706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:d6e2c7cae6b6d8e86e594394c008c0c74f2c0953c0daa07eda57a3466c9b6efb

Observation ac1b73af-7c00-4fa6-9631-9c4996e1c3eb · outbound

This paper cites Masked autoencoders are scalable vision learners.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Masked autoencoders are scalable vision learners

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.395964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:3a0d0e7af7b94489a419496009b130e490ab7bbde27a992bd36087e31fc862b7

Observation 2f1cd2c7-f65e-4034-bc60-c64512065632 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.897341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:eb42b670c77a935e75769d22a4cc3ee7156a42b7b09a64c86af4fee4f5f3521e

Observation 8c61f841-1fe9-4e09-b874-027d4d476f79 · outbound

This paper cites Otter: A vision-language-action model with text-aware visual feature extraction.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Otter: A vision-language-action model with text-aware visual feature extraction

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.958640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:64449bd750640ec0b1b93667f964943a9aaee86a14215f7dbba82ee458dd8282

Observation 6a2ccd2c-0e9d-4f83-9356-a237e9600aeb · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.444807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:b68ffd51cfe7e093fdbf271bcce97f1ff35180dd0e9d97f89108850f6c29b73b

Observation 03668529-d3c6-4f04-9439-a58281f01c53 · outbound

This paper cites OpenAI o1 System Card.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization OpenAI o1 System Card

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.899688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:a80ecd6ac45fd6134c3bf4f0ea38904ec0ed6884f81d7881727ccd03aac1836f

Observation 2913e9ca-7bf9-4806-ba41-f2caa95b040e · outbound

This paper cites Robots at the tipping point: the road to irobot roomba.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Robots at the tipping point: the road to irobot roomba

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.424437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:4b9148acbd1740ce8dcf80f6531388afa2cd0b26ff52fd363f6eba59ee71bd02

Observation 45fad8e0-6958-4439-ad36-0978a75fc72b · outbound

This paper cites an unresolved cited work.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-22T18:05:01.416672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:0ec0bda7c056b15d774137ea6dd089c5b393f044bc5c77727349e479970654c7

Observation 24e088ea-b9ca-4b63-bfa1-28a8ce078cbd · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization OpenVLA: An Open-Source Vision-Language-Action Model

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.928639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:5956ed277faca7366847eae98498f9f136deb31c4accf826f7822233a8486829

Observation abe1abcd-4e08-4c5e-97c1-6b415b892c55 · outbound

This paper cites Segment Anything.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Segment Anything

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:05:00.902550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:e93ebf05fd33f99803913366724ce9101b484050a7baa0badfb0050d6cd98959

Observation 2c2695fc-1af4-4a13-9d6d-61c7a2daf9ec · outbound

This paper cites Interactive task planning with language models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Interactive task planning with language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.390008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:9a69862d5a3eaa6309ce45c3d7a336af9aa80e82fdd0a42ecaaad39e6015ea34

Observation 9dfaf2b2-2e11-45fb-91d2-6a8693942608 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.957536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:8ca02ce2528289b2a2e4199e1eebc5980eed03725f511769fc21b8b8bda42f30

Observation 729510c9-4d9a-4b2f-adcf-1e77007432be · outbound

This paper cites LLaRA: Supercharging Robot Learning Data for Vision-Language Policy.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization LLaRA: Supercharging Robot Learning Data for Vision-Language Policy

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.942249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:1e3216e21a0ca6f2d1dfa9ea33c7744d3ade5f3ac50bd7401062d4e52209b5ab

Observation 01f4aa74-3f39-4664-973e-0c22a3171904 · outbound

This paper cites HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.950134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:e9585842010d4cc94a65b966d76932b136f40a25ec24a56b7748a48d8301bfe4

Observation a4cf5d20-fc53-4f46-9d29-77fbbc64c6c6 · outbound

This paper cites Code as policies: Language model programs for em- bodied control.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Code as policies: Language model programs for em- bodied control

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.385804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:c04d23bb4b4d71083ea6013a00e782631d3522b717ea79c57ec0e8d32e73eea7

Observation c659049c-3b15-4da0-a082-7049d7d19bb9 · outbound

This paper cites Data Scaling Laws in Imitation Learning for Robotic Manipulation.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Data Scaling Laws in Imitation Learning for Robotic Manipulation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:00.368555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:0fcc7157f92744ba8a697c5e5aedb77cd2b01d69072d8a845b76e84575045f0d

Observation 70863e49-e320-4d4f-b285-dba65fec6b93 · outbound

This paper cites Flow Matching for Generative Modeling.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Flow Matching for Generative Modeling

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.929170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:2e2c988d08aeb1d594b78f7fa2a4115dfe332cf6aafdb59a3ad00284d45142f3

Observation 857188e5-2369-4e7e-90b5-0be4799080f7 · outbound

This paper cites Moka: Open-vocabulary robotic manipulation through mark-based visual prompting.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Moka: Open-vocabulary robotic manipulation through mark-based visual prompting

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.409419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:0a8617432beda2c8d865b0347af9d615753ce0bbe8cad6ad6ec0297d6d50332f

Observation f0ad5648-5195-416f-a61c-ac5f80899851 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.944757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:88eb13d6d46828619dfed1aee6acfcb61983505e50cdaeafa40dd96939ff2e3c

Observation 10654187-49b0-4ca9-87d4-c064cd22d9e1 · outbound

This paper cites OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.890907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:8649529829c05ee34f0d48536828afdbf816c4b910f9e7e69caac66f15c51249

Observation bb45dc09-c860-4ab2-80b8-b86a53be9b6e · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.931642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:cd0fa964a5b9d03c4be17109c365288eae40e0d41d6f13eca0a1e624b1155a70

Observation 4fffcbe7-bf8f-45d5-b2a9-eaf52fde9fa1 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.961170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:0aca03c95bb16119ac467745df977f64f8f3148b74c1872adde3eb23f348c966

Observation 33ff78da-1e73-4c88-9a77-25401e3fa3ad · outbound

This paper cites Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.944447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:3c18b440242afe42349f4d0c7cd8f0a2b5278d238b111b3b9280f18f1003383f

Observation d44ca5df-f078-4a0c-996e-dff2de4529d1 · outbound

This paper cites Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems, 36:655–677.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems, 36:655–677

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.422557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:87df491a67a9d44ea380b5ed455c5d6e6848ebd385d7e34175945ddef87cdfbc

Observation 91f068c7-a5ce-4092-832a-9a0b0e443802 · outbound

This paper cites R3m: A universal visual representation for robot manipulation.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization R3m: A universal visual representation for robot manipulation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.418980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:d71cb2466c0e5b787c1b0f8f6ce390159c3a8408aaf383a2acd016215384c185

Observation 753df86c-10a5-4030-87b2-d0e9b3d2b2d4 · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.962878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:15361e927dcd8567439743578772fd029fe3f7720e111d7260f152106f562f91

Observation ef0e990f-e717-404d-9658-4cb6820b5e1a · outbound

This paper cites Autonomously learn- ing to visually detect where manipulation will succeed.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Autonomously learn- ing to visually detect where manipulation will succeed

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.414846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:601e382b6db21ae34f7aa3ff54ff99d53f6b559bb516ddb33e2ef708d093ef0a

Observation c6264bba-281b-41fd-b8b8-bb7f49518a39 · outbound

This paper cites LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.919649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:0846e9c48db08ed2b2146e3cf31c16137edec8b6357efb6df2989a11435d6dd4

Observation 0e527388-67c9-4bff-ac4b-1bcc16ebcf55 · outbound

This paper cites Octo: An open-source generalist robot policy.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Octo: An open-source generalist robot policy

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.423304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:ea39f4720cf5c1ac1eebdc794a86217d33b0f9fd98c2678745643ba49df6d6e8

Observation 2ecd5116-f2e7-4cb9-afdb-72ec8e93560a · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.933566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:028837a4f40f4d49f39427f883c85c11286e14e2824e3d93d0d74e36f5762442

Observation b4860028-eeea-43b3-b91e-49406634ace1 · outbound

This paper cites FAST: Efficient action tok- enization for vision-language-action models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization FAST: Efficient action tok- enization for vision-language-action models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.419343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:5b5357714ffd864f7c5f7a129d060e0fdfadbb6762bd4a2429a5258a07ceecf6

Observation 7c1f1b22-cc76-4d44-ac4e-35010933035c · outbound

This paper cites Open-vocabulary Mobile Manipulation in Unseen Dynamic Environments with 3D Semantic Maps.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Open-vocabulary Mobile Manipulation in Unseen Dynamic Environments with 3D Semantic Maps

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.952247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:21b58d1691d848593732e4b2391362644165e8d98e2dfd62e419763ccdfadbef

Observation 41548a38-1c30-4e65-a077-217fd00ee66e · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Learning transferable visual models from natural lan- guage supervision

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.437420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:d5e47b3473e4bbfefedb01bdf16be0ac9471776113f18a2094b3e878772571e0

Observation 3a83061e-6c76-489f-8989-d1d738bef472 · outbound

This paper cites On Bringing Robots Home.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization On Bringing Robots Home

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.969501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:fd46cfb78bbe409a16ad859226252b39411991bfa9ba449eb9f7202023e14a34

Observation 3d6d025c-d723-4c03-ac61-baa72c34fb15 · outbound

This paper cites Gnm: A general navigation model to drive any robot.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Gnm: A general navigation model to drive any robot

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.430035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:1dc3c1ec97ec123d30f845fe4649d649f9cf6949a9b7f70e9cb9e0399a2588a3

Observation e8dedf86-1f6a-465f-977c-2584ea469877 · outbound

This paper cites ViNT: A Foundation Model for Visual Navigation.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization ViNT: A Foundation Model for Visual Navigation

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.947021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:db55104a1fb6babd47bf2af136127065f729c2fb49eefecca0a16358bb9783b4

Observation f97ab4cc-78e7-4a97-a4b3-034f09c98d90 · outbound

This paper cites BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.963893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:b6803a0879ee2373ea1a63f42221c32401733bd7f7dd091dde6cb527ce6803b0

Observation 4b94d6f4-8050-42ca-a421-12760bde3724 · outbound

This paper cites Yell At Your Robot: Improving On-the-Fly from Language Corrections.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Yell At Your Robot: Improving On-the-Fly from Language Corrections

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.887959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:e155c310c66cd91e5f28c17aa80f43adafa7269b6b6431179845f97eb13ec576

Observation 4ca952a4-2f8c-4c18-bd64-911acd1d9eb5 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.968391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:5d408c8f209c2acf2a2bf91a6e35e616c54904a649a6530c268dfcd4b19aa88e

Observation 1705a5b8-196e-46ab-a1ed-5d3bc823352c · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Progprompt: Generating situated robot task plans using large language models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.413431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:83c182df82daa2b67bc7a2c6e3f849374722dbbcf15852b50c552b29fdbe8162

Observation 8a97b39f-b308-47ce-a50e-cf8997c5a9b2 · outbound

This paper cites Open-World Object Manipulation using Pre-trained Vision-Language Models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.899383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:0060817d01e7686a7073744d7abaf9d7ca4240b7f47bf421d41fcee5ba16f3b9

Observation a427c21b-8a0f-4bcb-9ca1-c34cf5463855 · outbound

This paper cites From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.858804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:f54c3ce52bb85278ec64c33a1ae6ddfe77624086eee4bd59380f01fc9adbb90a

Observation ec8058ee-b505-4fbc-aa4c-de5c78b0bc54 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Gemini Robotics: Bringing AI into the Physical World

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.933864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:f27c016cddaeb0a10f56a3c00c23a6a775bb0f029f21760724b8db2e7e3ebefb

Observation 4690b45b-0cdf-4d00-b5c2-84c85f709926 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.441215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:17f43f9a37061be994f32eba5a60a415fbd605cd897fff8b175998cc726d3a14

Observation f1fb22f4-1a26-4e4b-a80d-ab3f1c1ff213 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization LLaMA: Open and Efficient Foundation Language Models

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.952613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:490b978404ea37d640b8ecce56c08c2491aaf4bcfa7c3846dadca9be8b8551e3

Observation 052b0a18-4771-4866-b28b-8f3c723cceca · outbound

This paper cites Attention is all you need.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Attention is all you need

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.387772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:e9e8cb81e2796c88cde06e4f845a800b752e92e5a4782d07b4532c802dc7e005

Observation e12937bf-8415-49a7-bbc9-cb24d6cb5fea · outbound

This paper cites BridgeData v2: A dataset for robot learning at scale.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization BridgeData v2: A dataset for robot learning at scale

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.409843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:412db8fa23aeec8f2fbc4c62320c896468a8b240742ea2097fa8b2521878373d

Observation f85ce96f-b335-442a-85ab-1699d27a5493 · outbound

This paper cites Llmˆ 3: Large language model-based task and motion planning with motion failure reasoning.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Llmˆ 3: Large language model-based task and motion planning with motion failure reasoning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.407943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:17dd883772c1cb4498adbbbac3749ca8ab31335502c55341156059486211887f

Observation a670d574-06e4-49f3-8ad0-99e646053eec · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Chain-of-thought prompting elicits reasoning in large language models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.406145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:81ffa8882e86e98bfa959dd0fe100e06b12594b11f53934b16e42ca7965d1e28

Observation c710c50e-2bcc-4709-b055-dcdbd03afaef · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.955724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:9e3e514a41044775c2c50a61dab3a132e3f8f62289c3a6d8629a5002a61947f5

Observation 18f4c3ef-2dc0-42c0-afa1-29edfaefb4a4 · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.976688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:892666b6de19dfaf56dcf06f832ea16dcd2e9a6969f62218d0558c4fa2d1387d

Observation cf3f0499-aec2-4f22-bca2-0f57d8c90e19 · outbound

This paper cites Masked Visual Pre-training for Motor Control.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Masked Visual Pre-training for Motor Control

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.939533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:2e374c93fc5e8c61b1f48dc2bb2916886713030edf8e090dd8ae49dd394ff922

Observation cc67c135-cb35-427a-b9e7-6bacdcf30490 · outbound

This paper cites Magma: A Foundation Model for Multimodal AI Agents.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Magma: A Foundation Model for Multimodal AI Agents

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.973902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:c7b9f9104690594aecff1a003299cd5576ddcfddbcb4385a68067fbb33f2cf10

Observation 72890800-e9f2-4b87-9548-89018afbe567 · outbound

This paper cites Capsfusion: Rethinking image-text data at scale.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Capsfusion: Rethinking image-text data at scale

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.439260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:9d41f209b48de95ca89bbf7735fdd983cbb0cf012e9b88fdc75e5f9d4c6fb060

Observation 828e06b8-607d-44b8-aecb-d3192bd68c6c · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Robotic control via embodied chain-of-thought reasoning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.431798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:4bac5842936299bd76fa832e3e9b447890a4149cd971273c919f2432a045632b

Observation 59a2ec82-c40d-4496-b57e-16411d91d8e9 · outbound

This paper cites Cot-vla: Visual chain- of-thought reasoning for vision-language-action models.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Cot-vla: Visual chain- of-thought reasoning for vision-language-action models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.435528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:610640942f6c62a740be9e43a530260dbeb0189e64216a50a4dcd5f25248a3bb

Observation 2dba29be-db06-4c0a-8c69-5dff826f3490 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.949393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:7d5d5c5367ec8d82216d0ef32b6a9cb8718b1dbe499ca0bb31ce13086150a510

Observation c380f0ec-d388-4497-84a5-7521cbedfed3 · outbound

This paper cites Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.884980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:92dca2b1a17a51af3eb604e5d24ccf5da70eb92c3f29f9853172fb38d1894c8e

Observation 4361a7fb-9b7d-44a6-b812-36d8e3f10e43 · outbound

This paper cites put the scissors in the drawer.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization put the scissors in the drawer

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.433563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:078a2718b8db9dc8899069456e459eef7af46b774cdbefad67749211ff3456b9

Observation 4cb6fad6-37d1-4747-a575-fcdc96a4e21a · outbound

This paper cites an unresolved cited work.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-05-22T18:05:01.426401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:d66c9c9341589133093234a3f97729a5e514563641d9e2a8c73504b646aced4f

Observation d4b8922d-a915-44b9-8e76-20c26a58ab8a · outbound

This paper cites action expert.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization action expert

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:05:01.420650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:0ca95311ec7a8b2bb7abc3dad6eac0a9886851fa8d2292189fbd0d0cb199d38e

Pith citing papers

Observation 1bff4d0c-4c89-49f2-abc0-fa8959b304d5 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.837969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:8d58b4287c4753d560aed62644023aa4e5ce180aaae8c63837ed9085b9d07b76

Observation db3bf548-5062-47a9-a496-ea1fd0b8bd69 · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:48:48.967737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:022fa0feed19f69813b2834d39b52d5d9449a2e914628790b4b40f1c0309e516

Observation c289f6b5-265d-4bab-9790-cadc7468ced2 · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.362088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:2286ad29f49314923cccea90f72635eb4d427c8ace5bc20b81e5064843e8fd9e

Observation b67af7ca-89d7-4017-8dcd-a9aed805bb63 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.406172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:70e4876e31c55a9f7322a8cbed434c00335a6cf51c6bae4b7bcb3fa631f79aa5

Observation 6ced7795-e819-4a0d-b6cb-46f522532940 · inbound

DreamGen: Unlocking Generalization in Robot Learning through Video World Models cites this paper.

DreamGen: Unlocking Generalization in Robot Learning through Video World Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:50:45.421014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T23:50:45.332466Z digest=sha256:a6cd935efc56c7d6319202892f292f3dca4457433bc92f62be801cde4df20954

Observation b4e42aa6-311a-4dc5-861e-cf83f6f7fa2d · inbound

DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving cites this paper.

DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:31:40.702988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T14:30:47.787654Z digest=sha256:18940f5b4520e26b9900f77fec8a99251391145a670dd59da472cfa92e8ee0e5

Observation 857dad4d-0e7d-4cc9-9770-0099d54f9e38 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:25:47.218063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:ac18ad895ed68e47239d27b61ebb0bc024a923ba4cdd1a44d8487bfd42b56242

Observation 45587987-2e58-4c9b-ba7f-001bd17d22eb · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:55:40.412498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:3e557077c58fce6790b8127385370e2aa9b1e1c1716592cd9fecf801366d22d9

Observation e5a96561-63e1-4239-81a0-1a43cb8bac19 · inbound

Real-Time Execution of Action Chunking Flow Policies cites this paper.

Real-Time Execution of Action Chunking Flow Policies $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:18:51.681308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T14:18:51.613045Z digest=sha256:5b9a59ffe71fddd689b6ee3f1c613ee3ce3d5f897fe5989bd75fbfd7de472bc9

Observation 72dc4682-8985-4280-81fd-d359e9130445 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.677015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:b69bb721d7b2ef7432b0f5503ca431625526d18bcb27018d20f5a3f3b07af5c3

Observation 101ca8ae-75dc-4e00-ab6f-89e1ace28cb3 · inbound

WorldVLA: Towards Autoregressive Action World Model cites this paper.

WorldVLA: Towards Autoregressive Action World Model $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:57:08.167210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T22:57:07.883617Z digest=sha256:ee91862d452767c38a336da0fb4ba2145c0d56b497eb5f36651094cdf5768d62

Observation 03d12432-816d-4351-be1d-ffe37b8cf870 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 127

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:08:35.406683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:69a459ecbf7581dbd783bacf3b05468a807e62c379dc961d617a226fb063a92a

Observation d617f648-b2df-43ba-8faa-50399d09bc84 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:42:41.574722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:ab1c7493474675a3c3cdc0d58babae05fc27100b762eff9014cdb395ae84f796

Observation 8e86f715-3493-4a84-9516-e8def77facdf · inbound

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation cites this paper.

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T04:32:56.770143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:32:56.397350Z digest=sha256:e136a82ec2e250386d9dd71ac706e5525fd779452c917e76b7dee2c0205cb053

Observation c250a78b-7295-4aad-a829-6045709cac23 · inbound

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos cites this paper.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:32:58.770684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:916453ded6a6c94ab7f6c8baf591a04ebcff45f30a7ec1b923770db04c39b349

Observation 2dea83ce-9844-4114-a00b-0f60d39b7f72 · inbound

Vidar: Embodied Video Diffusion Model for Generalist Manipulation cites this paper.

Vidar: Embodied Video Diffusion Model for Generalist Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:54:28.306948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T09:54:28.271928Z digest=sha256:73f30976852f25357fe5e41f32bb2f63cde0ec420d0111210e3fac4861af39a8

Observation 5d20a68c-29ed-4dcd-b978-f74f5fee51da · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T08:04:12.625571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:3c41c4096315ae3239c911e322b7aa305e82e8714fcf7be731724c7f4f8e8dec

Observation e151758a-4ebc-43a5-9c26-778274baeca7 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 204

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:28:16.256180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:b915ce3e382c591d3d72784302482ee1b5001856c257a8cb259c5de4cd3f6835

Observation 24cbae40-cbb2-42ce-af58-84229a872e20 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.635806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.635806Z digest=sha256:6c5912bab95b5e3ea992886564b1fadda901be6f4f3dfa238715aa10ad999500

Observation 6cb59a10-a635-46a0-9e1c-776520621a03 · inbound

Prompt-to-Product: Generative Assembly via Bimanual Manipulation cites this paper.

Prompt-to-Product: Generative Assembly via Bimanual Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:39:58.527031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:39:58.527031Z digest=sha256:451623c2cb0ab62e47fbdce6ceb7957181ee1a3fc5ea8f7ad4b83f9c1d6faa66

Observation 4a699971-f082-47c4-8c80-c75de34631a8 · inbound

Galaxea Open-World Dataset and G0 Dual-System VLA Model cites this paper.

Galaxea Open-World Dataset and G0 Dual-System VLA Model $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:31:09.947328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:31:09.947328Z digest=sha256:e1bf2c7aaa77c5150716dad8278f9f5ec282ffe04eaf874dd67c0868985b5be6

Observation e2b0527b-41a6-4c37-98c2-d8d7da7e52fe · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.277722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.277722Z digest=sha256:1755bc4e9e2d7a6cc19fb2f4a73573e6a8f0d9b3b50b3160cf2c05f80eb6559d

Observation eee92cf9-71f0-43df-9c11-9aaba1001183 · inbound

TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models cites this paper.

TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:29:06.880873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:29:06.880873Z digest=sha256:4ea1b600d9cc2de0893c72120c983f52264a4259b3d90992723f78dafa8bd350

Observation 41f9ae23-47df-4cfa-8699-e0e5f19bc2b6 · inbound

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation cites this paper.

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:10:13.390798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:10:13.390798Z digest=sha256:a9c75f17b95919e60277679a3fc5c511730959190f23f64d15b0aae8fb505641

Observation 8ccdd464-d875-4dd1-89a7-4b8d3fe74e44 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 224

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:24.499007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:ac67c2d816862f25c752ce563ad5af5b303b26cd1641188a80ea16d8b454d6b7

Observation f895124d-e9f9-4b5a-bd36-8e0cde952dff · inbound

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning cites this paper.

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:02:11.392205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T08:02:11.189795Z digest=sha256:3ef3846411ab821a688fd28af051c48b45bb69fa7c795c563eaa4b43e3c1b598

Observation 317d6c8c-a1f1-45cc-8389-fde0af5ed0f4 · inbound

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue cites this paper.

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T16:17:25.996881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:17:25.996881Z digest=sha256:60ff7fe5324989e256ed4527d623c575d21b4f38a9296b233473fe2acea20ba7

Observation 127896f9-22f6-4885-8a3f-34a7382a5108 · inbound

Training Agents Inside of Scalable World Models cites this paper.

Training Agents Inside of Scalable World Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:05:52.602287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:05:52.431747Z digest=sha256:868480934666c716685b7c3cded569ed1aafc9df31ccfab660491dfd729079c9

Observation 6fed43e7-171c-4653-83b6-24924e3dbcb8 · inbound

A Systematic Study of Large Language Models for Task and Motion Planning With PDDLStream cites this paper.

A Systematic Study of Large Language Models for Task and Motion Planning With PDDLStream $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T13:30:29.109251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:30:29.109251Z digest=sha256:06d01e88f41f527f261aa36ec89b041b2f8393e43b8b97db50b242c64f382595

Observation a4da73bb-5af4-4e59-8536-a7b8193d8678 · inbound

INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models cites this paper.

INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:59:13.209733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:59:13.209733Z digest=sha256:6522f95dc08dcf23287907baf04fa99a9592aa80ee6896cc1779932c13831100

Observation 6f88509f-7fdf-4324-b089-902bbd28b144 · inbound

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization cites this paper.

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T06:20:01.953718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:20:01.885711Z digest=sha256:b7cbe29e9a05d9ec1bdbb7b039c6b2bec7afc3f29a5a94687f92641d16c5283b

Observation f15583d8-99a2-4077-bb3b-05299ba444b6 · inbound

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization cites this paper.

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:53.815467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:38:53.815467Z digest=sha256:2e64356310719050c5dcb76318f25cb2710db3aa1a77ea7d3e37a4bc376be312

Observation 13847b29-9fb7-4c7a-8aea-0ed219205bcd · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:12.607785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:12.607785Z digest=sha256:37fd1e9994cbaafeb541898bffce8488f5629164fa14dad7745b5f2dcee0c17d

Observation c5c10fdb-c9c9-4880-9576-50d882d82c86 · inbound

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation cites this paper.

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:36.312378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:36.312378Z digest=sha256:eae849369ddef6282cd93b16732d4a3bb6185dce33f90971cce27267af0c0bfa

Observation c4e1984a-b5e8-4747-85e0-461ca5f17ead · inbound

Ctrl-World: A Controllable Generative World Model for Robot Manipulation cites this paper.

Ctrl-World: A Controllable Generative World Model for Robot Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T01:14:10.462298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T01:14:10.174044Z digest=sha256:7aa67ecf04c6f6ca9b94d01ccce7173c18f917a6ab78b1de1820ae50eb950f41

Observation 7e26644c-764d-4401-a13c-73b1c64c5931 · inbound

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving cites this paper.

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:48:01.060499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:48:00.943591Z digest=sha256:06c4226939c459ea887b81a9cb9b555d9c2d3aae5e01ff41f0fd0a69b397caeb

Observation 2c2d1d15-49a7-4538-98a7-2f3b8694c86e · inbound

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models cites this paper.

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:42.876795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:34:42.876795Z digest=sha256:42305fcbda704441eb396061a91846048b9d137f8b4723b8552abebfae1850d4

Observation 05942d54-c760-462d-aec7-289e9565d1b2 · inbound

What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics cites this paper.

What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T09:16:08.892396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:16:08.892396Z digest=sha256:dfcd58df5528787ad131e34b7957440fbcfa55bf1eb22df42fc8102f9b26a05c

Observation cd8a25d4-d20e-4821-b7ed-6d7c424de79e · inbound

RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation cites this paper.

RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T06:10:57.862148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T06:10:47.309028Z digest=sha256:54e544cb76d4259f5dd774e86df8897e0c82139a24a41b16802ee2569c113b00

Observation 70eb54f7-2acf-45d7-aa92-c5deb405cecb · inbound

A Compositional Paradigm for Foundation Models: Towards Smarter Robotic Agents cites this paper.

A Compositional Paradigm for Foundation Models: Towards Smarter Robotic Agents $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:10.367470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:10.367470Z digest=sha256:44fceac103097abe63750cb1d6f4ef1e00d7f512f6cb1b26d58c2ffb427df040

Observation 2d19ea9a-fc35-4144-92dd-8c6d4e76cb84 · inbound

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization cites this paper.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.868682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:c348e9a2279884230dc24d36f0d01cee8bd037ee8fe9b18839d2e40dbb8b22f5

Observation f22d87d0-674b-409d-8883-68ec379d1e4a · inbound

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail cites this paper.

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-18T02:35:13.194295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T02:35:13.126171Z digest=sha256:487f4f95ab5ada2c5876685e57c3e3ec6f6ec79ee83cb04f969e617b7fd0897b

Observation 3220ba0f-cba3-45c1-9046-468f63db3259 · inbound

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations cites this paper.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:05:33.754905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T01:05:01.553811Z digest=sha256:890dd6bd37078ab6115abf4ecb410230f4d50223c26d002277657498fea8dfc4

Observation 8734c05c-1182-44e2-9a35-0fa48370b3ec · inbound

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations cites this paper.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:04.706491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:04.706491Z digest=sha256:b87fadcbb5b4c848cde3b03672774af37bb867332d832e2bd09daf34e4facbd7

Observation 12199c74-cd1d-47a9-8c57-0bec1d2b0098 · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:51.090448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:51.090448Z digest=sha256:3a6f62eb3b877ae9dbe7b97b3e1fc67f9349e1aa975f5e415eb5de52d66153c3

Observation 1e5343eb-1f60-4702-a82a-269dbe94dcd3 · inbound

ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Model cites this paper.

ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Model $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T22:03:27.184247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:03:27.184247Z digest=sha256:57d1ebe32759b5ae8dc12f5ab4046d93d07292427c5a9d84b8d07848777149fb

Observation 9ba93a51-b85a-458b-88db-d0aa7b7143c1 · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:30:18.339633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:568cc16b9eab62aad74783cc7b06fa8971f294676f30df846a6119471838d917

Observation 5e9c034e-c047-46c3-a72b-9110abd22f73 · inbound

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models cites this paper.

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:10:48.920061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T03:09:09.713822Z digest=sha256:eac3b4745463852d7181e8153ec2bfa74a0e83c8d8696612e77b64fee5737989

Observation b1fd7e4e-a5f9-4452-8d6e-e485c52fed88 · inbound

Unify Robot Actions in Camera Frame cites this paper.

Unify Robot Actions in Camera Frame $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.191917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:351dc9ddab8da5f47921cd41cf89eca18ab4bd63fb01f8ba6366fec8f2cd2db3

Observation 2276f108-a3ec-49d7-bcaf-a04833d25917 · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:51.301625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:51.301625Z digest=sha256:f1d70d500ad989de6c9ffa3a91d75c4d8eea99c49493601e58c6b66eda93649b

Observation 11066cb0-96ae-43b5-81dc-de0f302b5d77 · inbound

VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference cites this paper.

VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:54.520626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:54.520626Z digest=sha256:91c9589134c5785eab235833a70f90027af20ac5d488784fd090d17a99b13a6e

Observation 959b51c7-0685-4d7b-b8ea-6f7a0fb70425 · inbound

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models cites this paper.

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:42.801522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:42.801522Z digest=sha256:7845bb0dca75e10246d65ad13f0ef39a14c91f80affa2fd8c9b17d026e09b046

Observation 2877a3d8-5e91-4fb3-b1a8-b40025c437f9 · inbound

IGen: Scalable Data Generation for Robot Learning from Open-World Images cites this paper.

IGen: Scalable Data Generation for Robot Learning from Open-World Images $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:58:55.029388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T02:58:36.214948Z digest=sha256:dd1b9f6824429d7205d96ae482c887e449f0107d9691558d423935cfab2120cf

Observation 95cf4eea-589e-48d0-b830-58bc45045b29 · inbound

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models cites this paper.

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:01:20.386325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:01:13.910539Z digest=sha256:1c44032a5a57051363e7e121e147eb1bc0176da49b80490cc7f5dbeadb5b69ae

Observation 02e4d1ea-30fb-4843-9d24-274fc9f9f59d · inbound

ImplicitRDP: An End-to-End Visual-Force Diffusion Policy with Structural Slow-Fast Learning cites this paper.

ImplicitRDP: An End-to-End Visual-Force Diffusion Policy with Structural Slow-Fast Learning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:05:07.647630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:05:07.647630Z digest=sha256:847a5a88c384edc724c4bf3f568617efff7eca6099ad50c915ea3c61bdc7e9a1

Observation 3856cc85-aa78-4f76-a532-ad5fcb631af5 · inbound

VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer cites this paper.

VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:38:20.732710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:38:20.732710Z digest=sha256:fd5122278b94f39ccf35e41c386b061017acb4e9500b8d6f12cb73068d8b73e0

Observation 6b47a868-7136-425e-9eb9-6fb24fb95377 · inbound

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models cites this paper.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:14.406664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:14.406664Z digest=sha256:aeb42b65f6de0f87bacdc03dd1018c55d1c71cbe91be2259d802121a25f11baa

Observation da3fe36e-dd70-4f7a-8681-45877e354712 · inbound

MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning cites this paper.

MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:23.426184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:23.426184Z digest=sha256:28f98a0f42955305f6b0924bea33519f5ffa3dfa457470ea958033c5fbe5a902

Observation d63030b4-20dc-40c1-b027-472928688250 · inbound

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs cites this paper.

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:41:00.323961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:41:00.142543Z digest=sha256:799b3e58c2a35ffe31599419dda1375a2f04fe607e34dad5434df63a4a0cc6a6

Observation 49176935-da59-4e4c-a172-d0113ba9cccb · inbound

Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation cites this paper.

Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T15:05:03.261026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:05:03.261026Z digest=sha256:d2f07f7e5d7b27379aeac01ccbf08cfc0d9c764954e3b2fb48fe6d82171a95d8

Observation ba07debc-8ef5-439a-8209-c1b2b6bdb2a3 · inbound

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision cites this paper.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.747597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.747597Z digest=sha256:a0ebab36e6f20bc90f6803c7c3528ea2a3b4ae2fd85cb2e40c6f2ed1579852e1

Observation 1bc8d88f-6cda-44f0-a3c2-8788ec401683 · inbound

Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding cites this paper.

Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:31:13.400040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:28:35.576661Z digest=sha256:f7815b035b8c0aca03e66742a57f26eb57fcfb32ac9029fb330f54bbfa5bd0d0

Observation f633d5bf-2293-42d2-bf5d-ae6e706e68b7 · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:41:10.923589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T18:39:59.449746Z digest=sha256:8de06d83ba41f8099b4f28afc6d6212e88ae718c9540f9868204ed5edd9d8f8c

Observation 32694325-e753-4d5d-9de2-3a90dc290465 · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T13:35:54.101247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:35:54.101247Z digest=sha256:829ae5fae914c8a2cf8a0b54dfc6b88a39f754f415ae0207a89c3cdf83d89218

Observation cb7779e4-cd41-46bb-84a2-f7208358222d · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:47.182992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:47.182992Z digest=sha256:7b06ee8a7cce533e7698416996ea4a588e090908c7750dae22b99981f0a0e53d

Observation c554426d-5ef4-47c7-937a-f1f8de4fc835 · inbound

Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot cites this paper.

Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:03:12.204186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T18:02:11.393346Z digest=sha256:aff9f89ed87938cfea25b0e55e4f47e21a48f08139135353bdaf193fa3b25258

Observation dfb145c8-acff-429b-8e15-964018781195 · inbound

Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot cites this paper.

Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T12:41:54.996220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:41:54.996220Z digest=sha256:50f4b3b989eb7470ab3763ef7796859303977ae599b87f7896f026abd76b1156

Observation 5d3eeb53-4de4-49ff-9163-060000a59307 · inbound

CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding cites this paper.

CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T12:37:24.605491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:37:24.605491Z digest=sha256:b00001ab2ecde2d7c1be5b881ac21f2b50868c2b5b5c8bd423b679c6acc01ac9

Observation 3c7bdf54-4323-4445-87c4-d80772fa3baf · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.806316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.806316Z digest=sha256:e26d1037f86577b27315ff12f457ebe7b69c838afd1a5b2e5eaef0fba94fab12

Observation 3befa379-7fdd-4311-a3ce-4bb9df6d84b2 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:08:01.749537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:c99eacc6ec28def0dd08edbb9ab0428abdcd5b5f4a764da6114943f56948df40

Observation 559cc542-074e-403f-bace-651f1620e661 · inbound

CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion cites this paper.

CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-21T15:54:14.486258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T15:53:44.767068Z digest=sha256:7217133d860fc53cb6150a9a3afefd5dcc7e3532652d1b132a9f0619519af15b

Observation ba848dd2-a97e-4270-b92f-4a85e068cd45 · inbound

SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models cites this paper.

SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T09:36:21.735961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:36:21.735961Z digest=sha256:80021c09c3d8c22cc9303d4b9532b7ad3a49332a3e0b2a16628b36bf2ce8d2ea

Observation a2574a6b-878b-45b6-8c80-5736e89514f3 · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T14:50:12.938758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:53769b7cf25f26a433a75e3d1a352e31581f6b0a02e05a8d22b2ccb677d09faf

Observation 6f4d1fdb-fa58-41f1-9dde-fb0e10ffac64 · inbound

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance cites this paper.

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:07:48.160072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:06:23.555610Z digest=sha256:1b01484c5ad2555ec07b6c054ce57c9316e8485fd76aa028fdba7de31dd01d8a

Observation ff30f187-bf5b-4e2a-b0ea-69750a3f8f9d · inbound

CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining cites this paper.

CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:27:36.858412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T08:24:44.943709Z digest=sha256:a90a9c680d111946e607b8e4d9ae4bb8b26725a65c838106a3b50a481f9d1621

Observation cba26dfb-47e4-4a94-8d5f-35dae4d9bcfe · inbound

PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation cites this paper.

PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T05:38:48.729216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:38:48.729216Z digest=sha256:7419fa0fd95d4fd095d116d92582f9533d9ad8c4c94c144a42e968e7e0a7b688

Observation ec841025-6bab-4d75-be70-b62660958626 · inbound

MobileManiBench: Simplifying Model Verification for Mobile Manipulation cites this paper.

MobileManiBench: Simplifying Model Verification for Mobile Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:23:41.983791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:23:41.983791Z digest=sha256:c816e9d150d83d113a304f2f95ac7926c59b07706d44f8623e701cb2eaef8d6e

Observation e8efcab4-8756-4f76-a201-a4a26be4ae40 · inbound

Action Hallucination in Generative Vision-Language-Action Models cites this paper.

Action Hallucination in Generative Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:30:44.245198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:29:48.843843Z digest=sha256:eafceada5116d5731504ce731b8f5a12ad58ed411e69ca276997bcafc7e4bbe8

Observation 822fa9cc-27c0-42aa-804c-b12f3cc2c5b0 · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.604747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.604747Z digest=sha256:57e0e1985366181c85e7904d603cc987250e5a3dda5632f8cf964b4be379ab5b

Observation f429bd4f-dda0-41d3-87bd-1f735a17ee52 · inbound

Consensus-based optimization (CBO): Towards Global Optimality in Robotics cites this paper.

Consensus-based optimization (CBO): Towards Global Optimality in Robotics $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:51:43.920039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:51:43.920039Z digest=sha256:8a63eba151db9bd3faf1e6180c6da3d8fc6bd4ae0ee818482f912ebad24939b9

Observation cbe83c48-fda6-4a53-9f4c-0bd5983d5603 · inbound

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation cites this paper.

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:00:26.056998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:00:01.741166Z digest=sha256:18468c63c7dbe035454454e51ea7bd4dbfbd6940edf0e7a7ec14668acc4776f3

Observation d84481d5-0f37-4a81-bff2-b68a4b21982d · inbound

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation cites this paper.

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:14:10.822640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T13:13:53.818915Z digest=sha256:85f0a08447966bf5946ca032914efa11dd3db82a23dca745ebd7dec2bbf44d32

Observation 453b2280-79e1-4fb6-a597-2683c9e4ee6b · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:37:13.930979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T03:36:09.272019Z digest=sha256:d4328efecd09d665da6c07a43f8dfca5d59952851d030d33cb385c6534d7f7e5

Observation 2e14a9a8-b0db-46fd-917c-51ac4d7493bd · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T02:53:16.234139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:53:16.234139Z digest=sha256:b0cea9fa78461936a9b8c07601b7251e5ab798925e87e89e443405fff5e83b18

Observation 8ce62710-b8d4-451e-a343-9dc33d815d60 · inbound

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning cites this paper.

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.306712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:07:10.387869Z digest=sha256:ca005962a7b3a6a4a0f9e70a983216100df3c7cf784b79894d56c63d67dd15ab

Observation c06f1fc0-9b34-480a-ad17-ba7a970d86c0 · inbound

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models cites this paper.

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:02:24.993241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:01:30.803128Z digest=sha256:e8ac93c298ae5cd072eed8bda707acdd30d295a9e6747d936804de75140abd28

Observation 62cccc27-cd54-4ada-aebf-f88eef954206 · inbound

RISE: Self-Improving Robot Policy with Compositional World Model cites this paper.

RISE: Self-Improving Robot Policy with Compositional World Model $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:30:32.154029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:28:37.997148Z digest=sha256:fc6c42ed816289c5d998fcfc43343ec811e0cbd80cb801724b838fa05cbff533

Observation 12fc5a08-103d-474b-a7c4-9b4e8fa421b1 · inbound

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning cites this paper.

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T03:12:11.669904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T03:11:52.645633Z digest=sha256:c6218ce789466c4e7b904334c06f115232fba779a057a1afe92e95b07e629f10

Observation ac8d9279-251b-419a-bea3-ce423503ecbe · inbound

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation cites this paper.

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:12.805125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:12.805125Z digest=sha256:5aa7fb03fb319e85c903c8421b19befa8b3670a75b57f42211bb21acf7519ad8

Observation 8fdfede4-577f-473e-b6bb-34c1cfcf2ae9 · inbound

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion cites this paper.

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T23:57:44.611376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:57:44.611376Z digest=sha256:ff52b6a3091009b530005aab1682f17e6749dc7b49cdaa0bedb1ac1ee1a8105f

Observation dfe368ce-6ede-4d57-a0e1-890224b89277 · inbound

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models cites this paper.

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T23:48:58.046107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:48:58.046107Z digest=sha256:1dd2a91bfae8537ee5aaad5bc97ad58eb219fd798a4409eea6e4e8919637f366

Observation 33274cb5-09d9-4684-af41-3ef5e931b8f8 · inbound

WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL cites this paper.

WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:51.672629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:51.672629Z digest=sha256:daafa1df383db6ce18a4834f0b921112052e0987f412709a64b736237cc215b9

Observation f3859444-12dd-4cc5-8c70-f4e5eab3d792 · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:18:15.507578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:8a64931d963938348cb58a01119551d6ebbd6a876b92f8a6aa237ee0ed66f927

Observation 0a3d9840-ce2f-4f5c-9780-6474f13409af · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:26.872226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:26.872226Z digest=sha256:6d8526a9fa9b56b62f3ab38d616c76238a0f9fe0327fa7f6f12e84fb89818752

Observation 92826e6a-2113-45d5-840c-7794690bd75e · inbound

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation cites this paper.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.161214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:2e0268dc6692f756f36c63730c6eb02c4af60d3b6a316029d44b98176403fe22

Observation 4c857a25-5d50-4912-b536-efb2ad166f3f · inbound

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models cites this paper.

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:20:17.357438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:20:10.435886Z digest=sha256:d055a0e60bd79071285181bf877027cca3ad728f0dce553da68fa653e7cb3127

Observation 9733372d-3292-4d28-92cd-2715eac8a56c · inbound

Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks cites this paper.

Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T21:12:11.016607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:12:11.016607Z digest=sha256:91b64abfba1e25f7ddd35e976f8a5d73aeab8045e24a0a678cdfed98b7ab8ccf

Observation 53349c06-8d18-40a0-9b86-0ec37d997b72 · inbound

CableRobotGraphSim: A Graph Neural Network for Modeling Partially Observable Cable-Driven Robot Dynamics cites this paper.

CableRobotGraphSim: A Graph Neural Network for Modeling Partially Observable Cable-Driven Robot Dynamics $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:09:32.221168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:09:32.221168Z digest=sha256:128c1faf794c1a22773aeb4a81148dde7afc70d12b029554ffe61e0306338cd7

Observation dd2cf00a-f745-4b5c-a5a1-479e0e8936f7 · inbound

Unleashing the Potential of Diffusion Models for End-to-End Autonomous Driving cites this paper.

Unleashing the Potential of Diffusion Models for End-to-End Autonomous Driving $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:54:10.261861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T12:53:29.732960Z digest=sha256:df4a97751b6819da324a13276844ed25573dab723de25ff1a18fafa134bc3475

Observation a842e040-2780-437b-9a7c-ac5160d82f15 · inbound

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation cites this paper.

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:10:16.047491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:07:46.733450Z digest=sha256:82f332db2e32d89ffdf36ffdfcca1eea05dcbe2135643ebed82ef2501a0f6ca8