Pith. sign in

Paper Citation Record · LEDGER

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

As of 6 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 100 inbound Pith citation observations for arXiv:2410.24164.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.24164 v4

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T12:38:24.425784Z

measured 160 of 160 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 100 of 1078 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:47:40.788104Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact35
  • verified fuzzy25
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 186b835d-ff67-4533-a818-63179ffb2719 · outbound

This paper cites GPT-4 Technical Report.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:38:24.569269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:6d609e59f707ebb8deeff7384f725c2bd20362ca8441dc08bab9c844bd80e494

Observation 66483837-3208-4533-b954-885452a870d2 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:24:06.862904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:7a10a6b8840bd8eaa29fb0717d87c74bf01d57443d8625a0323864a0333aab9c

Observation ef7df772-5ad9-49ca-aee2-1b4cb0812fbe · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.599723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:dafa1a32b1d67437ce2929aa962af2cf671ac50be5d41cec25360e848b2900f7

Observation cbcf5aaa-8a08-4d47-bb32-87102fe28d44 · outbound

This paper cites ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.572972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:3279abfbfe344775f0fd4ccb461e43e9192cc107c6b4d349f3bbc32f65ab307d

Observation 4c39feb4-ebb1-4228-9a90-81ca3f3fbecb · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.987351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:c348427571c118380742156db21eefb107444f2f8d98da9805e9090ec674efc6

Observation 7a4fdabc-a9db-4f26-876f-0ec0a1a75656 · outbound

This paper cites RoboAgent: Generalization and efficiency in robot ma- nipulation via semantic augmentations and action chunk- ing.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control RoboAgent: Generalization and efficiency in robot ma- nipulation via semantic augmentations and action chunk- ing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.607672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:4828b0be6e69b36fa243cd89f55baef63296b852059382d09c242f839a3af117

Observation 0ead8dff-ba03-423b-94af-1257f7b9ddca · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:36:04.729348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:1d04839eb601356f8905cfd4289a61b5c29659043ddbf28c310ccd9cc6e965bb

Observation 08636786-f950-4932-b653-6e6a3fb3aaa3 · outbound

This paper cites Scaling data-driven robotics with reward sketching and batch reinforcement learning.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Scaling data-driven robotics with reward sketching and batch reinforcement learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.584240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:e5bf2ff34cec857c79b70ce227f341eca2f39685993ff3fd6ba2f2610b5d5030

Observation 4e3f9108-50b9-486a-8e69-c038619a94da · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.615841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:fc9a734fd4705063ae186c87fbc1d7680c0737b268fc5ec09830a9a82d50e3c6

Observation fe746920-eb9e-419c-87ff-55614ada75bb · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:23:25.503833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:1003c40120fdb22b87a79aab5e4905238626eea4e3f4e90d6c75d7d4f72170e0

Observation b201e70e-56da-42ed-9ee4-7f5f261b8aea · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:86ec5d730759ca2d1a0fe5a325e9d7dd7b895d22994d91730e225205aca96b48

Observation e905828c-e40e-48eb-bb78-ec431a8f61b9 · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Glam: Efficient scaling of language models with mixture-of-experts

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.622903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:60ee0c340a1e0fdf25b4efa343af1e40f74cf88c0b9597a70a7fe8fdebf93259

Observation d230a8ac-84f8-4fcf-a997-c1ba5c5e9e14 · outbound

This paper cites Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:55:55.512010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:9e14aded26974cc56ac5719877a10b255191100fcce83f294d7a3ea7f1e99c4c

Observation e4ac397a-ddf5-45f2-98e0-a1deb0a55850 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Scaling rectified flow transformers for high-resolution image synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.627784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:ce681c77fd214707d1a4b58ac3741e842c2ee90e4685555c0e065f16099371b7

Observation a08f786d-d01b-4d72-b222-95095acd3c31 · outbound

This paper cites Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.597104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:0961c9d514536c8c0dbe6e3fee3c4b8104186d37355fa486c93d808d64f7153a

Observation 6b43814c-fe51-47c0-b32c-991389f29d7e · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learn- ing Research, 23(120):1–39.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learn- ing Research, 23(120):1–39

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.633090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:4b1c9242b1d0bc37e41bfbfc10f0ed473490b3eda796c16ee7efd2955fed9dfc

Observation 59d1a3e8-e72d-48e3-ad4a-cf9ce93cfa00 · outbound

This paper cites Zhao, and Chelsea Finn.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Zhao, and Chelsea Finn

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.635410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:a98d28b618023f046b32fc50833dba9315410b9074511175404cf8a005ba0331

Observation d8b245be-6e70-4664-a5da-da1bc2884d80 · outbound

This paper cites Robot learning in homes: Improving generalization and reducing dataset bias.Advances in neural information processing systems, 31.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Robot learning in homes: Improving generalization and reducing dataset bias.Advances in neural information processing systems, 31

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.637941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:113e1b578974f1dff35ebfcd5c04eadb67d26d508a3d5247c2413d28a8bb7643

Observation c71919a8-492b-4e79-9bb8-f368a347f265 · outbound

This paper cites MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.472611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:5ced9158a871c959346e2a78685b12d66c1798ea819076c9ab46834a0fa0ffe3

Observation ab2caac5-d90b-4c3a-9672-bb9039cf9700 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural infor- mation processing systems, 33:6840–6851.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Denoising diffusion probabilistic models.Advances in neural infor- mation processing systems, 33:6840–6851

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.642799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:7619499f0be05a2c167d3b42e8d0f34dc3347e835d00a9e5424bc0810a87959f

Observation 4f90ccf7-c52b-4875-bb01-e242188cd1a5 · outbound

This paper cites Highly accurate protein structure prediction with alphafold.Nature, 596(7873):583–589.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Highly accurate protein structure prediction with alphafold.Nature, 596(7873):583–589

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.645800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:701fec6befa56e464dd38cfa3bd7aa9a2c467155df9eed14216463291a8416d0

Observation 9ac88b53-d881-4a8e-9c0e-f27cac17364e · outbound

This paper cites Scalable deep reinforcement learning for vision- based robotic manipulation.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Scalable deep reinforcement learning for vision- based robotic manipulation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.648789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:d43873eddd018b0d040d375a4cf42796ba86e418156e652cc825a70c3e7616aa

Observation c821c1c2-7e56-4fb3-aaea-6e86a0350b64 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:51:19.504346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:6ed2f5285796cf350eaeac1633efb242d2f9f1b1873bbbae8871fce29bdd6ede

Observation 0ed1f587-1829-415d-87ee-b0d6b91a5e2c · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control OpenVLA: An Open-Source Vision-Language-Action Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:46:37.168344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:6c0de2a46e681dbd1bdb3312a492530851523202209729ad8b00b8abd638f55c

Observation 993eec13-d4b4-41de-8e9d-e7a161849f50 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:26:45.276836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:28bc59a2c7239a9e3e87db9aadc7f85ad839419dec148395d4aa30ef75b6cda7

Observation 5d27860a-43c1-483c-af8f-55899c72aa16 · outbound

This paper cites Learning hand-eye coor- dination for robotic grasping with deep learning and large-scale data collection.The International journal of robotics research, 37(4-5):421–436.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Learning hand-eye coor- dination for robotic grasping with deep learning and large-scale data collection.The International journal of robotics research, 37(4-5):421–436

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.661100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:d697161549e2785ad262f68243491f4f2d967dc8e39dd47482684b216acb46ae

Observation c3ea40f6-eebb-476b-977b-b9b1f53527f4 · outbound

This paper cites Mankowitz, Esme Sutherland Robson, Push- meet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Mankowitz, Esme Sutherland Robson, Push- meet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.664022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:0b6ebea1b8aa83256a326246a8b9fc014e1247758e9a1b92afe1eb0421071f13

Observation 1990ea1e-1107-4a4f-baf6-27a9a2b0c690 · outbound

This paper cites Flow Matching for Generative Modeling.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Flow Matching for Generative Modeling

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:38:24.489819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:598daa9d5cfe4f780c49cc51c3e9b0687b03aa6f216235365717b7adc4b1faf2

Observation 2977e68d-84b7-41bd-ada7-f78dfcb54bca · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.494101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:92ec61d71e517bce2631663f009e0d132a587d80ca78ae02a4e8d440cb8b9067

Observation e612d3ac-3a8d-4379-858a-e292a675148e · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Visual instruction tuning.Advances in neural information processing systems, 36

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.610349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:450abbed7ea8f7a678b3d8b03dbad16a516d363ff0d2af2ece379b855637bf35

Observation 6291f803-cfe0-4280-8f77-37ef6395435d · outbound

This paper cites OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.497934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:2d65c0ced7e1feb794869f1542f3b0c9a83b4a70014fe11b86cffe863416b7f2

Observation 36ba374e-170f-4016-8a3d-8ace4e6915e8 · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:03:14.510564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:f2c55f261750d7f27bb90a12e1239452e6e9ad14d2bc19615fce93ce8b0f13c1

Observation 1b0b3c6b-4fea-44e5-a5f9-2acfcc742e26 · outbound

This paper cites RoboTurk: A crowdsourcing platform for robotic skill learning through imitation.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control RoboTurk: A crowdsourcing platform for robotic skill learning through imitation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.620598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:39e6bcab1b38a6f3b230664f5d8ef610d3e5e1422683eacc8238017196b436f7

Observation 585fab9f-6e5a-4453-9c01-d963c5976fdc · outbound

This paper cites MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:47:55.567916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:096cfe04b7b797eddeca69f61fdaa60f118e3a217bb75ce36b0844132e5f2609

Observation 591cd36b-23c2-4c24-a64f-35b350cc45c4 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information pro- cessing systems, 35:27730–27744.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Training language models to follow instructions with human feedback.Advances in neural information pro- cessing systems, 35:27730–27744

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.630413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:72dcd0b210396a7b6a8a98dfffa96d545bfff91d403134f1d16283ca01f18dc3

Observation dd073e90-aa27-4f9a-b4a7-27fce03663d1 · outbound

This paper cites Scalable diffu- sion models with transformers.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Scalable diffu- sion models with transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.640505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:64ee6b13bdf5a4b3e542d4659e5c0b2dd492bdc7be7ff226a5c0dcf8d0655c37

Observation bc38e97b-1946-4688-85a2-cec54c0688e5 · outbound

This paper cites Supersizing self- supervision: Learning to grasp from 50k tries and 700 robot hours.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Supersizing self- supervision: Learning to grasp from 50k tries and 700 robot hours

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.651770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:aa80dea5ac6e36a92a29fa8097c5da28171381aa5eaebab118ce52c2f2a9bdaa

Observation c9813c43-2e1e-4bac-b39d-cbe66d9ae298 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Movie Gen: A Cast of Media Foundation Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.778393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:99d2b2ab2ddcdd4001dc13617b59e4cb2254ba679d417298600133b27336859e

Observation 194a8af6-2f32-4302-a671-22ed184741c8 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Learning transferable visual models from natural lan- guage supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.658072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:9fb5913671bba393fd1ed06b6dd579e21bfdd90e925129ddd04d74c84fae5e36

Observation ef130608-d600-46b5-805b-69e522cb750d · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control High-resolution image synthesis with latent diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.602234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:a3953b2956682ddf752edd4d660a7b13032f785a2178a05911ec0b983a2c4c75

Observation 2186145a-d3ab-43db-984f-6f6a60b33e97 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.604890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:12c8c79451e94f140dcda81460e82ecd416c05761d8b18c49528920ae697803b

Observation c3653234-5eae-4e51-9a46-de5bf658036f · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:02:34.528194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:2d629cc023abd32e872f65fc73808c2cc60ed113fb27d8b9238f63c7edc95c95

Observation f36a7891-6f8b-4805-bdea-83d2de6bafd8 · outbound

This paper cites On Bringing Robots Home.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control On Bringing Robots Home

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.520342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:42aff45b0eefa2b9c4568d046fd5f80e2b7dd702f20b8235e295378563e2e4ef

Observation 968ce4a6-e67e-4482-9899-d34d30bab563 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Fast Transformer Decoding: One Write-Head is All You Need

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:50:20.836519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:fe8ea91ff80dc26a5bc7fa0d9a72a06d59fa35cc9a91201ab188164c57d1a615

Observation d582be40-13d5-4787-9012-f91b99942e81 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:38:24.528600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:b0f94ea8ab1fd28d7f94156af8b1c1cef5a75bd5c260ce987469ecfcb10e572b

Observation d862559d-2aa1-4818-9fb6-8bb4ff588f21 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Deep unsupervised learning using nonequilibrium thermodynamics

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.612863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:5bac5456c211a404580570b5e533163aa741567d2a355f8908a371bf3f511055

Observation 0a682dcd-0704-4295-8db7-3ac96b14c373 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.532658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:6a8a17548e8b8365063b0103641c0c598bac6778abb542fbf0d2b9ab922b5300

Observation ed0ae4d4-95a6-4e9d-af1f-3e513cadb8e0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Gemini: A Family of Highly Capable Multimodal Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:38:24.536160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:dc59b5d6922375f9498e68d677eb5ce35631d185d1cd13544ecf5d9faece7d8e

Observation fad6ce8d-bf4c-4f20-9b99-f2551f9157fb · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Gemma: Open Models Based on Gemini Research and Technology

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:54:09.189864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:64e765a2b69185401f0a3a44b40596f57e8a545f4317a67d9681fb7b36e1675a

Observation 9d15a7cb-73ed-4287-ba12-1d7bf67cb273 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Octo: An Open-Source Generalist Robot Policy

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:26:16.174119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:5d94c14977bc25f2115a8eedbf5aaf57f09024b3ea1d705156c8c7b5f0e1bbf3

Observation 1a5ff08e-4d11-417d-ab61-52ab4e1dfe64 · outbound

This paper cites Attention is all you need.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Attention is all you need

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.625303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:bcab02806ab44c56c8f74da43b3eb5549ecbb7c8cad430cf104e882e532979ab

Observation 731de74b-d4ae-4f11-96b0-a6ec05cfd17a · outbound

This paper cites BridgeData v2: A dataset for robot learning at scale.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control BridgeData v2: A dataset for robot learning at scale

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.654745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:65e2df9af17ae0c42dcebb4539ef7c1677ae34d034cd59352ab681c134f48295

Observation c2530f8f-9bde-406f-a677-67c537ab8965 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Finetuned Language Models Are Zero-Shot Learners

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:14:13.786069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:358329e44e86d2758048143d3a02c27df0f5351e56cce12c605e951ef87c7150

Observation 15e242e8-bc3a-40c0-ab24-f0b2209e5ceb · outbound

This paper cites Emergent Abilities of Large Language Models.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Emergent Abilities of Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:38:38.557341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:43c8aa7302ffe9e3c7c748cce5146f20d3d75be673bd43f087637c7c9bb01692

Observation fea1c6ba-967d-4da8-b23a-dbf12b3f90b3 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:12:26.208161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:baab965ee78d28d464bb5e16e11a0457720994805ab8f5b281701e87ec43ad0c

Observation f60ea528-103c-41d6-8a6e-be2b0cf4ed97 · outbound

This paper cites More than a million ways to be pushed.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control More than a million ways to be pushed

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T12:38:24.618132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:0a810f5b475769a5710c5503d02fda730af3bac05768188b687c475598b8b445

Observation 7ec7b2e4-0148-4062-a61c-5fa50c4f4f64 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:36.384383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:a619c406a7c7911c5d90d3b933a0c4d61828e481697df725d70e50d7e9d4a5be

Observation 9ba0329b-450c-4bef-bc07-42e62e18d6d4 · outbound

This paper cites ALOHA Unleashed: A Simple Recipe for Robot Dexterity.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control ALOHA Unleashed: A Simple Recipe for Robot Dexterity

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.562457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:3ca064838ae325542db7b77ebf0f1a3d9e5efa71d3765f33688e5f5dc5035028

Observation f260777d-b39d-4545-bdb6-2db9fd5f0c7a · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:27.005672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:8a2df9428842f671ca36fabeeebe6bdf7a4c74c0b724030bf6c76d22de01c8cd

Observation fbfb63a1-1f00-4bd2-83ae-ad56fe9b29c9 · outbound

This paper cites Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.460447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:754c642f5781cb9072fdab37c2980772830291bdc305312d65e5a0466dd2d0df

Pith citing papers

Observation a0ea6fd0-44c2-40d9-9cd8-0ada0f20440c · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 120

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:25:54.432542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:afa8e76268dfbfb3712df6cd14de8e059830a679336adf0ed0294af3ae1cfe31

Observation bcb6b6ea-2692-4ab3-9e0e-7092a78ef483 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.683754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:c1b78f158647bd8e2c240ad12a61a92cd45bb788369a6ebbf1b873e1a0fb97d1

Observation 0600c0bc-a69d-4080-8a38-1d103ef5b0e3 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:31.885570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:2fb13f34a3c41c25d23d424b3e3cf3462ca55059364dc73642bc91dbd430eebb

Observation 2002e7c8-88ac-4260-b5bb-8c299c943d58 · inbound

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model cites this paper.

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:12:19.743290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T06:12:19.643111Z digest=sha256:41e161cc80089c841f1a02ee1843911c1b6af504e506589bc94d1c0833f53c82

Observation ec0dfc2a-d280-4e28-8223-2ac7edb35059 · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:48:48.938637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:ad6e2283f8d8b46ab22f98c1649f8ae59a66a6163733322dcf8277c1be10bcd4

Observation 2d4f3a82-d24f-407b-9751-e0f38e9fdf9e · inbound

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models cites this paper.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.185946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:2e4601022b87ff572044c810c5e15a44f2b7886e92126d1c58f60d67abb34f5f

Observation 22368c17-48cc-475f-b1a6-826d0c9a25f4 · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:32.232973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:9c40026a294e20a1a05d7e6a8acda3e489f90d47db1b04fb6b71c6074609b5c4

Observation 70a418a6-8288-4c5d-b055-d352eaaba0bd · inbound

Unified Video Action Model cites this paper.

Unified Video Action Model $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:50:29.708086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T17:50:29.675358Z digest=sha256:45349adb330f8560a06044a87085ccdcdf13c3cb2fda6f8155e842a55d170d8e

Observation 1bb1c304-9840-48dd-9b50-7271ac95850a · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:32:22.624008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:4f4de5ee099175450753d6d4bfd6138618800295f66e2eef44116fa4a4dcc12c

Observation 5aecb05a-2703-465b-8d19-b1d7e4e16bea · inbound

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems cites this paper.

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:09:24.576942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T15:09:24.367362Z digest=sha256:34fea0e15555461142955d264aeafbe6104f0c5e2f09ca752279cc7340def011

Observation 60c15d60-ad52-422a-86e4-8d335280e9a8 · inbound

AhaRobot: A Low-Cost Open-Source Bimanual Mobile Manipulator for Embodied AI cites this paper.

AhaRobot: A Low-Cost Open-Source Bimanual Mobile Manipulator for Embodied AI $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:02:17.521597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T00:00:53.307042Z digest=sha256:354fbdfa2bfc550bca818e8c9144b59f04a5eac7680b91e9f704f8bcbbd27122

Observation 61ae2709-9375-4bca-8b10-eb5b17f922b2 · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:00:48.811414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:b32603c0f133a7fecd849ac139fa9ec1c8803c67264a19042da20d2bd85d199b

Observation b8f7d2b5-4894-4ee9-b64b-3e88dbed82b0 · inbound

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots cites this paper.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-10T19:09:10.181938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:ac8ba06164ce7542062f1ff28983ed6ce79071da11a3c05bffa3de814b1ffc14

Observation 84d5953f-dc81-4d28-a135-e97261a8fa55 · inbound

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets cites this paper.

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:25:00.467741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T16:25:00.365534Z digest=sha256:3779cff6e7d054989b5b03b27c65d0271e9a9c1e46d58b0cab188c02297b53a7

Observation 2fccd485-6540-43ae-a635-d5c7e90a80f4 · inbound

On the Importance of Tactile Sensing for Imitation Learning: A Case Study on Robotic Match Lighting cites this paper.

On the Importance of Tactile Sensing for Imitation Learning: A Case Study on Robotic Match Lighting $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:16:58.624700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T19:16:17.765462Z digest=sha256:eba3400dc6e01629bb69ca85b8cb8a30f42a3913c3aa8a3ee94e97016ecddf39

Observation 5e556176-8e55-4d4a-865a-26b45d0f82cf · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.986329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:850ae470042060fa935ea0406adef98e24681b813c9aa74c49d7e245cc17b57b

Observation a73fecff-0909-4187-a945-73b7b40303be · inbound

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks cites this paper.

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:53:29.434032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T15:53:29.412890Z digest=sha256:1cd1946eed36f43b748539e9a60a21973f1672eabb75c103e8d3d4cbfb6e03e1

Observation 5ef22e29-3bce-495a-b431-e7415c487a2c · inbound

J-PARSE: Jacobian-based Projection Algorithm for Resolving Singularities Effectively in Inverse Kinematic Control of Serial Manipulators cites this paper.

J-PARSE: Jacobian-based Projection Algorithm for Resolving Singularities Effectively in Inverse Kinematic Control of Serial Manipulators $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:31:55.864774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T18:28:26.116236Z digest=sha256:27e23cc496e0788e95f7e4d252ad6cc929794668ca5de2fbbe2609c98e009160

Observation c73ba013-9396-40ba-a5d1-1bc598abcf31 · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.171711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:99de673cb6e0613a021e5a9659181de7bf9536e47266326df30d435ac5e5596f

Observation 2f4d5805-aa3f-4f2f-8659-48ebc2ccb1a3 · inbound

VLAs are Confined yet Capable of Generalizing to Novel Instructions cites this paper.

VLAs are Confined yet Capable of Generalizing to Novel Instructions $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T16:46:47.428158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T16:46:05.993833Z digest=sha256:b8a98c3ac489a2860e812209aa1fa82bcb429a67bdf95c914c21613bdf0a9558

Observation cc9c5646-0c0a-4933-9b77-32fe93602d9a · inbound

DreamGen: Unlocking Generalization in Robot Learning through Video World Models cites this paper.

DreamGen: Unlocking Generalization in Robot Learning through Video World Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:50:45.394919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T23:50:45.332466Z digest=sha256:b686f9740ae688e0fe798f981c40c2980fd9b58aefc5d062a92cbfcd3d22226b

Observation 79f4cdd0-7913-46d1-8983-d230889baa62 · inbound

Policy Contrastive Decoding for Robotic Foundation Models cites this paper.

Policy Contrastive Decoding for Robotic Foundation Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:11:38.508066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:09:48.762737Z digest=sha256:c44eb245a2c2c7dd8be2f0755beaabb457b58e654455244bc0fbfe351590144f

Observation eb9890e5-82ed-45bc-8a52-8d28c59fada0 · inbound

FLARE: Robot Learning with Implicit World Modeling cites this paper.

FLARE: Robot Learning with Implicit World Modeling $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:59:08.907271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:59:08.846629Z digest=sha256:5603ecb0d584fea329ac45ccd29c17258552f4a2dc6e779e6ff94d870b9210b8

Observation 9aec41ef-dc46-48f3-aae3-25b2891de423 · inbound

DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving cites this paper.

DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:31:40.583412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:30:47.787654Z digest=sha256:686883440509245fd7ba836c24752f4a91ac51887f8d9b16a589b64b81f8a492

Observation 58564182-be91-4a99-80f6-a761f9c71764 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:25:47.265869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:dcc72b6e32667ca8bc164e59ef88f2ad0fa546341a4da701203bb7c5870c2628

Observation 9c09501f-ec09-4c2e-98d7-99374eb8f6ef · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:55:40.480291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:5310394ed996b898ec83d51a3d234679c7dc583229ea058b85dc8bc94132139f

Observation 84b25f2e-3d98-49ea-8c63-0aafb05f01e9 · inbound

EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild cites this paper.

EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.709860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T12:56:30.634284Z digest=sha256:a94275f9bb6e112e8bb43957f1c34ea692eefd97ad0c0bbf96cab886726f3e81

Observation af0695d2-d1e2-4647-bb0e-e2c0b640429c · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:22:37.167640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:8db03dda4e1dbd59ee9eff8768124b544be606711e2cb41c3445b7aa9e7e319d

Observation a5806d79-07bb-4406-ba1e-7a1eba7b83fe · inbound

Real-Time Execution of Action Chunking Flow Policies cites this paper.

Real-Time Execution of Action Chunking Flow Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:18:51.674437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:18:51.613045Z digest=sha256:418b8090fd44ef5afb4e6718004a0af73b8ed7ede739d46610cb6a323d32561a

Observation cdb078b9-bc45-41fe-8e7c-458892ab1bd5 · inbound

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving cites this paper.

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:36:24.368787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:36:24.319361Z digest=sha256:07b5f3820e0e726ec623667eb33e1b4203002e668de78378d01678b7e9bab244

Observation 21c9af57-5d01-45fc-b924-7e0660213228 · inbound

Intention-Conditioned Flow Occupancy Models cites this paper.

Intention-Conditioned Flow Occupancy Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:27:14.531924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T10:24:52.160209Z digest=sha256:7755996fbfd60bdb66cd5878cdeb8bb5142fc98b53dde918af3e3b948a5d5e15

Observation 5b6f6463-1cb4-4704-8729-e5b45bc57d45 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.663856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:61af58baae77d7ee15e2cd34948c2642ea3659a430eb8d253392c141289e921b

Observation 34916f84-9160-4628-acb2-820dd22e26bd · inbound

GRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics cites this paper.

GRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:30:49.402120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T00:26:28.696429Z digest=sha256:c96bc7c9c03f2b02165516611444b24b1ba6e66625d0e5055947fdd0002cbdd7

Observation 584e080f-4f28-44a4-8cd7-f5f76fa11127 · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:55:46.398786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:490fdf8b1774d6f0d3c2ae3628657e622c0c3e9e1d00a746a0390c8e9387f568

Observation 11f12992-a548-4838-ba4a-fc3577db590c · inbound

ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation cites this paper.

ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:47:13.868635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T09:46:38.742771Z digest=sha256:3bbb9bcd682bd795ff83bd43dc3a09a87f8a4c143a4dbcd4c15c3fd99fb932fb

Observation 26903364-645d-4a93-a458-50ddb89eab29 · inbound

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation cites this paper.

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:40:27.315717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T06:40:27.206337Z digest=sha256:035132470a5079d07a85ed0d5b49bc6e6ff59fae1a870619c5314ecbeac6637a

Observation f0f2b177-6e27-4a48-9d8c-8a453ccfb118 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T14:08:35.173173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:a0b2bc21bbdbf1f11e4e631d2c455bd2f628f54fd27b94dd256c7fed1c940eb5

Observation 26e3b965-2d83-4fcf-83b9-f57dc77e32c6 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:42:41.465663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:e6ccc92956177b960e955022700b2179f195f1f54a7ff1bf15f73e0151e47e2f

Observation af8e91e8-8a61-4f2a-8a27-c598d6e0aa36 · inbound

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation cites this paper.

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:32:56.680934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T04:32:56.397350Z digest=sha256:f27defca456f8be5d1200126743de4e0fcc9dd9114a8ad45c40c9286a40f7009

Observation fafd0c86-d44d-478a-80d4-a1e4c6fa9d33 · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:05.299689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:911755fbabd5e162a24069809805d36427947982ce62dd8fec3b1c78e737ca65

Observation 20e9b048-92b0-4361-bb52-934b18440f96 · inbound

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation cites this paper.

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:52:57.541199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T03:52:18.984005Z digest=sha256:1cc54dab5b36c0d270568f8dcb86649733bfb6f59685431752459bcb85124dcf

Observation 4d5a8270-5228-48a4-85cc-23873b14b19f · inbound

Vidar: Embodied Video Diffusion Model for Generalist Manipulation cites this paper.

Vidar: Embodied Video Diffusion Model for Generalist Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:54:28.335954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T09:54:28.271928Z digest=sha256:941ca12eb177225f66f33e69191e8e99d84f5127ce41aa2f1836202c1ed58aed

Observation c0d62655-327a-46d1-aaca-c7d260a4f923 · inbound

Reinforcement Learning for Flow-Matching Policies cites this paper.

Reinforcement Learning for Flow-Matching Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:40.788104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:47:40.788104Z digest=sha256:ab11047fbb16cf5349ce6e0c9cd5aaa04cbb4665dd4656101084f95f4352c08e

Observation 5c4ec451-ad20-44f5-b1a5-974f783bb2dc · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T08:04:12.629932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:3eebde959d65fb79946985fc64e189a3e4c79f489e87abdae3713a8077a1cb14

Observation fd76e2bb-41aa-4466-9b3b-2a46499d2114 · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:33:38.444159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.444159Z digest=sha256:6fe7b8400ef20bed19003c47cc3eb3398e6b99defeed31eb2ca1eb36ca695dd1

Observation 08323709-7dd0-48ff-a5ae-b4252357730b · inbound

Towards Human-level Intelligence via Human-like Whole-Body Manipulation cites this paper.

Towards Human-level Intelligence via Human-like Whole-Body Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:48.271563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:48.271563Z digest=sha256:2cf08916c5d34f65b18b374fdb99575397b4097b58c7f9d0055123624a11095f

Observation 9bb34379-aaf8-4319-9ec4-5503ca708c52 · inbound

VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback cites this paper.

VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:34.064538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:34.064538Z digest=sha256:051844390c7be156b9b83ad7489c91896847da290db3e92f2a912e769343964a

Observation 490fde5d-cef2-4cf9-8512-2159be3cf686 · inbound

Test-time Offline Reinforcement Learning on Goal-related Experience cites this paper.

Test-time Offline Reinforcement Learning on Goal-related Experience $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:21:59.654946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T02:18:34.690441Z digest=sha256:5afd0718adfad49b556bf9bf1da89984647d9a54df26b49bae43ceb8fc5a3004

Observation 0ce2233f-4dad-46fa-b9aa-660cf115acc1 · inbound

Flow Matching Policy Gradients cites this paper.

Flow Matching Policy Gradients $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.618360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.618360Z digest=sha256:65be5626272071853149ecc58aeb2f7dc91d619d40a6bc485c816c11841d32f4

Observation 98fb2361-38df-4a16-ab48-2e47ff2f9c71 · inbound

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems cites this paper.

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-06T12:39:37.651115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:39:37.651115Z digest=sha256:a57fd2e2840240997a1cc24385f02987885982eedc66422cbc9143ec35764499

Observation d2bf74e7-b9a3-43b3-8576-2be2a04e15bd · inbound

Improving Generalization Ability of Robotic Imitation Learning by Resolving Causal Confusion in Observations cites this paper.

Improving Generalization Ability of Robotic Imitation Learning by Resolving Causal Confusion in Observations $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:51:02.091283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:51:02.091283Z digest=sha256:e1102859078048bf5d7d0e6846ef57125be4dd90607b1e52eda18413e2e166e9

Observation d9a7d4b8-7613-4fd1-9375-2c726b69c5fa · inbound

H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation cites this paper.

H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:48:18.867996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:48:18.867996Z digest=sha256:aa925e80b2e7123235d0c13088a09edcecfc38bbf319f5b98ce035a1d83d78c4

Observation 71a5747f-c752-4f73-bd4c-9bc07a1a5672 · inbound

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models cites this paper.

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T21:52:02.961866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T21:52:02.893886Z digest=sha256:61b9ee8af732b3e56362d2b5101c0255a65135f7d5d12448e71b601813d20992

Observation c9e8315a-733a-414c-a4c6-2f040681b95d · inbound

HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning cites this paper.

HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T10:11:47.387858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:11:47.387858Z digest=sha256:39488a15d2e33cd54bd226e80864d38d019ae302af9411f558507e99b6961fad

Observation d8964fb7-2688-4bd8-b183-841c56ef2688 · inbound

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation cites this paper.

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:28:41.964233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T21:28:41.904725Z digest=sha256:0180bab7f6429ad0eb4dcf1527fb12c57b98c2cb91363496b64fa3dab16e3ced

Observation 4ab12c15-ed62-4476-9a44-2dc02a708d4e · inbound

FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing cites this paper.

FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:29.431156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:29.431156Z digest=sha256:e5b7361ca735199950dbea89d243ee691522acc4a88190ccefc2cd2488ae009b

Observation e7e65a3d-ee2a-4f59-b585-f775af7c558a · inbound

Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution cites this paper.

Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:33.712603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:33.712603Z digest=sha256:3d31d84f1c7b53a9d5a2792ee78e4a1db6d8f6976f0cbc63f96bbc24d3757e1a

Observation 944057e8-bf5e-4fad-81a4-6ea10e59526a · inbound

Multi-Omics Analysis for Cancer Subtype Inference via Unrolling Graph Smoothness Priors cites this paper.

Multi-Omics Analysis for Cancer Subtype Inference via Unrolling Graph Smoothness Priors $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T22:51:58.490409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:51:58.490409Z digest=sha256:a3d0c0d3568be883da9c84d35e518fb9a08f0c4e129ee85e5831ee116d433b08

Observation 9ad5d33d-b8da-42b2-86d0-d8f04f8ae900 · inbound

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation cites this paper.

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:50:41.076925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:50:41.076925Z digest=sha256:9069a3d0b8106f311988d8460f0f32246481ee05c52723a4fdc0ee5efbbf15d5

Observation 1eb14267-c287-4869-b822-a7247872d9c9 · inbound

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation cites this paper.

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:46:39.871089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:46:39.871089Z digest=sha256:535edd0424fd52e66eab25202995ba229b085fe4fee9cae6a188a74b597123d8

Observation 1559dab7-3594-454f-bd9f-302cac670c54 · inbound

GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions cites this paper.

GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:59:00.765491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:59:00.765491Z digest=sha256:2ff45714f12db00924ae5a420d429ccf1ea9139f52dba488ee6aaf21d4865ff0

Observation f1a2d9de-0e4f-4892-a869-9e30b7767636 · inbound

DeepFleet: Multi-Agent Foundation Models for Mobile Robots cites this paper.

DeepFleet: Multi-Agent Foundation Models for Mobile Robots $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T00:16:55.429897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T00:13:27.533407Z digest=sha256:73d046e88e45cc9bef6ab3ced3da6b9e428177570e6d49ee4bf9142ca47e2f13

Observation e114be25-12c2-46b2-8898-f6ef1efcbe26 · inbound

Time delay as the origin of oscillations in anodic Si electrodissolution cites this paper.

Time delay as the origin of oscillations in anodic Si electrodissolution $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T20:50:28.780841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:50:28.780841Z digest=sha256:899d84d869a6ece5928f0cdfde24915b06970d5f2de61db9a5f30246bd71247a

Observation ba5700c1-d379-457d-9b22-d536e652a461 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:38.197244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:38.197244Z digest=sha256:dfc550a3125cd29e6886d5fa4551fb9c7c7781cdf1dfbcbf6f2f2fe6ffdeec52

Observation fa49e8e6-abff-4f81-a5b0-2f59041fffa2 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.972251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.972251Z digest=sha256:8a83ff2b91aa938c57daa8f336bb8f60db6dd873bc5db1eb801c3935f62dcb26

Observation 6b9ab49d-43d9-409d-b057-bbffb4c6db71 · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.119863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.119863Z digest=sha256:503ba4e8088c8967d29c73d20fc809f36d683d295f9789710b0568a0b21fc5e5

Observation d99a860e-408f-4259-b184-6a6754df008d · inbound

Context-Aware Risk Estimation in Home Environments: A Probabilistic Framework for Service Robots cites this paper.

Context-Aware Risk Estimation in Home Environments: A Probabilistic Framework for Service Robots $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:33:21.237689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:33:21.237689Z digest=sha256:73621d1678fadd8ee3caf0430816461016a644bdfd6824dc8cb44243d78b4637

Observation 0a2f361f-bb9e-42f5-bf36-5f7abb6d96a1 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.559657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.559657Z digest=sha256:c0b9af4902809939cbb5f5fca936d6d954672698b73d644652ec273ec651ef13

Observation 63e8f5d3-b3f2-4cfa-9638-0939ce5f95e6 · inbound

Prompt-to-Product: Generative Assembly via Bimanual Manipulation cites this paper.

Prompt-to-Product: Generative Assembly via Bimanual Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T14:39:58.474830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:39:58.474830Z digest=sha256:15dfc490b2ee1fa7300c58aaa0f979679dd08d39be26e0960e6e23459220dc70

Observation dedf9934-f427-451f-a357-b1117075386c · inbound

Mechanistic interpretability for steering vision-language-action models cites this paper.

Mechanistic interpretability for steering vision-language-action models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T13:50:17.340438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:50:17.340438Z digest=sha256:236ea665521a4ad136cbc52e03976eb825d25451c5738cc536101844b196f790

Observation 1fc43ed9-d9e4-4b61-be7c-5d1c4d81762b · inbound

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation cites this paper.

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:41.215161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:41.215161Z digest=sha256:1d93d87a4366bee61be096749894625badd25069fc7a5cd8d9335f904fabd408

Observation 6193c206-5a0f-451d-8efd-508c8acea7c8 · inbound

Galaxea Open-World Dataset and G0 Dual-System VLA Model cites this paper.

Galaxea Open-World Dataset and G0 Dual-System VLA Model $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:31:09.936405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:31:09.936405Z digest=sha256:76dad3040ef6c25420eeaf32e4e5b04d5d59ee6970512a1d4cc74a7f884b8c9a

Observation d7462c27-1579-4723-8cbf-57fbdbba4f08 · inbound

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training cites this paper.

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:15:01.261741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:15:01.261741Z digest=sha256:173bc5952bf47625a48bc9bcceeebe47ba916055c64e0ad5bf2370060fb5e8f9

Observation e74e5446-8423-4ded-ae86-5794975f9c70 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.253604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.253604Z digest=sha256:a0f0c3d6a410b623aa1127ee34cf8bb65c8ffdd2a6fb2dcdbb02886cf17f9662

Observation d79fd5b3-b3f2-44a5-bc6f-25749d7b6fcf · inbound

ANNIE: Be Careful of Your Robots cites this paper.

ANNIE: Be Careful of Your Robots $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:00:38.071920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:00:38.071920Z digest=sha256:f6bd38bfdc12cb94b77ef8e41efe4bb99c6ae9ef3b1c754f0590901e0422f878

Observation 49033faf-97a7-4f53-9d63-22125454ec02 · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.469543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.469543Z digest=sha256:8963285b8dd0b91f98e85e742e5193a55f7e42f1b4c72d3ae9482bb304fd71b2

Observation 425dfa97-a135-4d41-b75e-c7ccac9bfb33 · inbound

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies cites this paper.

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T05:48:47.041197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:48:47.041197Z digest=sha256:7b0d29dc8bccbe27053a31209e45b882f952aa435ca98ba7f4482c358689d439

Observation 5f54cd02-6b84-4070-bf28-5ec13872ec76 · inbound

Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen) cites this paper.

Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen) $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T00:02:31.054355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:02:31.054355Z digest=sha256:2107968bbb790e4af0f330e0a6f72c70ed7ae4cbb6686c9624d01b925476832a

Observation 50b658f0-973b-4902-9b3b-55aff3ec7ea7 · inbound

LLaDA-VLA: Vision Language Diffusion Action Models cites this paper.

LLaDA-VLA: Vision Language Diffusion Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.811254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.811254Z digest=sha256:d942050eae0f7df5ac8eb7c1d534f99dca29b811e3d3db5ef4d12445f2e23f59

Observation 85bfc1e7-8b0f-4d81-b632-90a41036c302 · inbound

RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction cites this paper.

RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:55.710564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:32:55.710564Z digest=sha256:3abdbaa0e2d610fa57016091be9fb0418b39013e6ae474b9b9f0b3b3f3bcdfcb

Observation f4b573ec-1b5c-4aa9-8bd7-c69c315de236 · inbound

TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models cites this paper.

TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:29:06.758383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:29:06.758383Z digest=sha256:f2430b53167af58e43a95d3e7e6eb44c2bfc67fc1f281d7d28e5c3cfa237f3d4

Observation 7b60584e-c3f0-43a7-a7a5-a7e828dfa9bf · inbound

Input-gated Bilateral Teleoperation: An Easy-to-implement Force Feedback Teleoperation Method for Low-cost Hardware cites this paper.

Input-gated Bilateral Teleoperation: An Easy-to-implement Force Feedback Teleoperation Method for Low-cost Hardware $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:05:26.681426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:05:26.681426Z digest=sha256:db0f506ad285699c283df824ba6ffbe632899d9a1c9ff169afefcb83833848f4

Observation 376dd412-306a-4f14-bfa4-4169b96b7374 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:24.981459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:12e3bb136d209abbc5ba2837ba5fba0d647b72aad176f7ad86c9892489a37fdb

Observation 5f3307b6-c5a2-47da-816b-9ff403a6365d · inbound

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models cites this paper.

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:47:08.802413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:47:08.802413Z digest=sha256:913844a66e44eb73accba641c19f208fe4ef14809e22943c441a1300ca6b921f

Observation 2cae877f-7bf8-4f35-ab0b-a02106c52495 · inbound

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning cites this paper.

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:02:11.305433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T08:02:11.189795Z digest=sha256:cb16dc30bc4a70e7546ded1a034bd24c8f3fb7a92bc0b02ac42ef4aad8743f77

Observation dcf27dbb-763d-4fc4-99d6-8988c21531b6 · inbound

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue cites this paper.

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T16:17:25.973193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:17:25.973193Z digest=sha256:ce20a13c06ab65afe624683b96f71b0470ea6dc079a2cb8e80e13edd17df5aa2

Observation 6e62abb8-f959-492c-b8de-133fde6de60c · inbound

N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout cites this paper.

N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T15:46:03.494365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:46:03.494365Z digest=sha256:17d4d94605bd36df3301a5c0d578ffcdfc546adee89a54651085b0e902733635

Observation a98ca1ea-ef4c-4ea7-867f-30f951b895c6 · inbound

VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation cites this paper.

VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:51:25.708500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T13:47:30.943307Z digest=sha256:8b14e7e69ab169dccff6a8b83a7bb6c82392c5819bbd44412165e16ac6d43f90

Observation 88a4699e-ac36-4b88-b175-bdf47f50ee3d · inbound

World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training cites this paper.

World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:51:23.555450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T12:48:32.123998Z digest=sha256:df15470ef79ab3ef2799f254cd92368ca8a4dd08bd4816bacf57b17079d47387

Observation caafefa5-b5e7-4b51-895c-510517a196cd · inbound

A Systematic Study of Large Language Models for Task and Motion Planning With PDDLStream cites this paper.

A Systematic Study of Large Language Models for Task and Motion Planning With PDDLStream $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T13:30:30.202233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:30:30.202233Z digest=sha256:e14fec4381e9866eacf6b7fd4af5d13cc6a0f813eb9a1197f16851ad6a1c02d9

Observation 34a3f0ec-986a-4507-9dc0-47abf2c7ae7f · inbound

RTFF: Random-to-Target Fabric Flattening Policy using Dual-Arm Manipulator cites this paper.

RTFF: Random-to-Target Fabric Flattening Policy using Dual-Arm Manipulator $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:24:05.276953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:24:05.276953Z digest=sha256:f7dcaf932660d411a94e65dff93c8ab5d7877d096d80746c013cde1abe30a92c

Observation ace31975-76eb-498a-a42d-dba1085d859c · inbound

INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models cites this paper.

INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:59:12.628530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:59:12.628530Z digest=sha256:e212393094ebe2e7de99bf45a18d5d592714725aee65eaf7f664bc3d46e476a9

Observation 1ec02467-3720-4f9c-a014-e23301308246 · inbound

AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation cites this paper.

AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T10:21:15.235184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:17:53.226866Z digest=sha256:fcfb18d4b08c18191ac8972e8401b520d8dd8df1c1d68dee0221f08764ff4e5e

Observation 23bc713a-ef68-41fe-ab00-40be527c7b94 · inbound

Flow with the Force Field: Learning 3D Compliant Flow Matching Policies from Force and Demonstration-Guided Simulation Data cites this paper.

Flow with the Force Field: Learning 3D Compliant Flow Matching Policies from Force and Demonstration-Guided Simulation Data $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:11:17.969000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T11:07:57.386640Z digest=sha256:d811d76bbf623e2ff9ae083385ba7cee68d7d6cb29367f8c3985e99cba4599be

Observation 0572aabe-83f7-4d86-8fca-57c064a3b0ae · inbound

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization cites this paper.

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:20:01.910970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T06:20:01.885711Z digest=sha256:f91e781380d436e358f70e82bff159f1579fae109eda71ef3daaea4eb80d496b

Observation 252522ae-b056-46e6-ae89-e10b48336282 · inbound

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization cites this paper.

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:52.995690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:38:52.995690Z digest=sha256:53e371a0853c9ce6c72fb163b4c43433687df8ebc01901530e6d62c3d39b655a

Observation af236d71-a32a-46f2-817b-9cfa4657666f · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-08-04T11:38:11.335522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:11.335522Z digest=sha256:ad653090444a59ad96b28c58af817b0eadd400586b19dbffb103f4a106ec0b96

Observation 6e8a0966-ba59-48b4-9bad-b7d286e6979b · inbound

USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots cites this paper.

USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:20:31.803123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T08:20:24.424886Z digest=sha256:a1368569d33a82f892a77ae13119bb1fcc6bde0e68628d04b58070e247017bdd

Observation 60e5f4df-3768-450d-9025-914d1e24d4e8 · inbound

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation cites this paper.

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:51:09.098383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T08:46:51.278024Z digest=sha256:42291d9b71f06b8e8f7c19871525b24f35abb42961e9127ca3789ff90b54ba2e

Observation aeea01ad-e29e-409d-96b2-737b5539f175 · inbound

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning cites this paper.

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:08.551806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:08.551806Z digest=sha256:2ba199463436b0b46ccf91c0fc56a18dc02b02e3beaf7cbbb285e2b7a2b245fa