Pith. sign in

Paper Citation Record · LEDGER

RT-1: Robotics Transformer for Real-World Control at Scale

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2212.06817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.06817 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 528 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:53:37.497994Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

39
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 38dc9863-dee4-4a8a-8ecf-fc1b67359e32 · inbound

Scaling Robot Learning with Semantically Imagined Experience cites this paper.

Scaling Robot Learning with Semantically Imagined Experience RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:59:10.529105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:59:10.352342Z digest=sha256:c80e23ccfdfdaf96d87b2cf89f6b368fb9aefff2f856ce1aeb548b27ba1e9c73

Observation ea8564a8-7678-4399-beb8-d57b5879c0f0 · inbound

PaLM-E: An Embodied Multimodal Language Model cites this paper.

PaLM-E: An Embodied Multimodal Language Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:41:14.065009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:9437acdafa2bddca0de60db0fb73a335ae5480da59876e394a47f602894bc0eb

Observation 14cca411-b4df-4221-9301-daeec0d9684a · inbound

Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware cites this paper.

Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware RT-1: Robotics Transformer for Real-World Control at Scale

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:15:35.971597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T04:15:35.550350Z digest=sha256:fce1691bd4af8c9710b453669ae022d50183c03c5c46f27c6a8061c248e0dcb4

Observation 9afb43a6-7151-494e-9421-911d3ec00edd · inbound

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models cites this paper.

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-13T08:57:22.474166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T08:57:22.299028Z digest=sha256:6bb7833dc6d36a56dbddb9f041970e34cba8496e773d77f2fc7bc25a0fb9159d

Observation 58f2b9d0-91df-4511-985c-bc33c579e840 · inbound

Cognitive Architectures for Language Agents cites this paper.

Cognitive Architectures for Language Agents RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:33:44.274096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:33:44.146134Z digest=sha256:99c9681b89780fc8e5f2a2554c1a6c400c87cb1e38bb6657ebc85271f78da23a

Observation f6d340fb-34ee-4eaa-8d37-6277f2d85d51 · inbound

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own cites this paper.

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own RT-1: Robotics Transformer for Real-World Control at Scale

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:44:02.671943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T06:40:00.328012Z digest=sha256:c52af8b8361e0321d56162c3f6472144c962df1539d40ae628989f6d95911365

Observation 44dc4df9-79be-448c-95f5-69737a72fd66 · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators RT-1: Robotics Transformer for Real-World Control at Scale

Reference 224

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T02:15:18.462173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:13ae1c67ddf2b871cc1b6858638f5767b77f644953d095ca88d3ba10cfbd6677

Observation 503c7138-8d10-4602-a699-9b5437278e5d · inbound

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models cites this paper.

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:54:59.156438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T05:54:58.940428Z digest=sha256:cdf7b0c455f1ca1f832376d778cb78b3d2a49950d37ada90ea69fa24cd236a9d

Observation 5608e651-9dff-49e4-a065-f9b1f7ae9973 · inbound

MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations cites this paper.

MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:47:55.105202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T09:47:54.977716Z digest=sha256:003b350d98a24ba1fb1cb955c995c399ef74ce19ce6aef9dafff59ffbd933fc4

Observation 3ea3525c-9e38-4116-acc9-45b49288e267 · inbound

An Embodied Generalist Agent in 3D World cites this paper.

An Embodied Generalist Agent in 3D World RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:22:18.646698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T14:22:18.606817Z digest=sha256:cba7e96502a794bd96785fba61479d1d1ab0a5f3e89001bb4809bc8ac4659300

Observation bf5e42b0-0596-4a18-a483-234fdfcb898a · inbound

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation cites this paper.

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:32:05.647023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T16:32:05.538507Z digest=sha256:a76aadaf233357303ee660e80cb130d4c762a2d86ee3c036c46df04fdec810ee

Observation c6736fd1-b98d-4059-bbef-fe156b650dfc · inbound

AppAgent: Multimodal Agents as Smartphone Users cites this paper.

AppAgent: Multimodal Agents as Smartphone Users RT-1: Robotics Transformer for Real-World Control at Scale

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T10:16:43.900058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T10:16:43.364787Z digest=sha256:48e4a0e80638526c2a2537a757c3b1a4f49ce7028e9cbb2afec4deaf7970ea69

Observation 613147d9-aa56-4435-b98e-e1e49f02bcd9 · inbound

Any-point Trajectory Modeling for Policy Learning cites this paper.

Any-point Trajectory Modeling for Policy Learning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:32:51.017858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:32:50.916085Z digest=sha256:099d8f46b0e8d089410136c9748c229845edd51f7238ed6cbebdafe24a176d94

Observation 2dc4e750-206e-4a51-9cbd-aeec0781f5f6 · inbound

Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation cites this paper.

Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:02:55.495579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T22:02:55.240949Z digest=sha256:39a1c4333188a22ac04de840371faf2b0e5c06aad932f31cd8f783655a46e9ad

Observation 2ba1bf3b-e7d1-4e04-85bc-9829b50a80e1 · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction RT-1: Robotics Transformer for Real-World Control at Scale

Reference 140

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T14:25:59.457133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:622498e7eba597bf210011c27c0817b936b5373850793ca02602531d3b4aa2a9

Observation ea57b152-7338-4dd6-b384-7835eb58b532 · inbound

3D Diffuser Actor: Policy Diffusion with 3D Scene Representations cites this paper.

3D Diffuser Actor: Policy Diffusion with 3D Scene Representations RT-1: Robotics Transformer for Real-World Control at Scale

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:00:05.004115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T22:00:04.812281Z digest=sha256:29996e4c8ebd4ac3d9faef0f0ffff01b72267fdbef5ace8d4599ca6aad0a2240

Observation f7220aef-c615-498a-80bf-d81877484e53 · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T19:22:35.366005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:9cee0153ba4922d8bf80309ea999959fc07ef17d1243cc357bb5ff2cdc831515

Observation 1d9ac655-1f0c-446b-81ed-130b7d8792e1 · inbound

RT-H: Action Hierarchies Using Language cites this paper.

RT-H: Action Hierarchies Using Language RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:53:27.785410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:53:27.642020Z digest=sha256:102ad3a13e41c475b3f52546b50e634a6223951f55e26f6a4a7c5217aff27d3c

Observation 26525cfe-1dc5-4f41-9bed-18f1fcf5a0e2 · inbound

3D-VLA: A 3D Vision-Language-Action Generative World Model cites this paper.

3D-VLA: A 3D Vision-Language-Action Generative World Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.262017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:22d86a325181185c5811addc7394e324e6847c0d0ab21f6b4490de1dfae2e6f8

Observation 9e43cc20-68bb-415a-909a-74a999b4ba84 · inbound

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset cites this paper.

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:51:18.956855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T05:51:18.508352Z digest=sha256:01e83c9fe085df071326d7960d5cdb83aa3054f5efbf88f5d5dcdc35639730c5

Observation f0b874fe-07af-41dd-96cf-1ece28c77511 · inbound

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments cites this paper.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.446816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:23ecab7db052dd3e58b51da6f56a31e69c7dd68e9cdd68b03cd3e01c8ac88380

Observation dafd1435-024d-47aa-b2ad-dded1a6f77de · inbound

RoboDreamer: Learning Compositional World Models for Robot Imagination cites this paper.

RoboDreamer: Learning Compositional World Models for Robot Imagination RT-1: Robotics Transformer for Real-World Control at Scale

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:47:30.273814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T20:47:30.225116Z digest=sha256:bb0e9e8d8332e6207bcaaba28c046beabd3d030a42c504d593ddf8644b58f72e

Observation 0feb4d6a-1dfc-41dc-88eb-13ae730911d7 · inbound

Octo: An Open-Source Generalist Robot Policy cites this paper.

Octo: An Open-Source Generalist Robot Policy RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:26:15.608064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:26:15.163358Z digest=sha256:9d010b5a049524bd17e542116aeccd432fe8dd6295a90f422a8bed3805473873

Observation c4e0899e-ac95-4a79-b21a-ce075e0adada · inbound

RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots cites this paper.

RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:46:30.233150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:46:30.198761Z digest=sha256:24a044b3cd7349c3a37714be4f4625c228a1f38133c2bac2aa01866518309747

Observation 17c70509-f1da-4839-8a68-8b8468f82283 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:41:14.065009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:5600d0491ec3a73ae1957633e51b5ba14861694faa4989d5dc6686b811000b06

Observation 56094f5c-d9cd-42fb-a593-1b42955fa160 · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.434788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:ccee06e3b89761753c09781ba9113d8d0ef580a6e109622f961c162429f1322a

Observation f937eda9-1f12-4337-9cad-9e0f9ab3135a · inbound

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation cites this paper.

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T16:12:26.085872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T16:12:25.980853Z digest=sha256:87bee45c5c4cd3025ee68d4016d8621b88049b173c44636df7112f3033c1d032

Observation f8e1e5f3-f24d-466c-8b82-4bec3c517e4e · inbound

Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models cites this paper.

Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:23:24.859969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T20:19:20.382009Z digest=sha256:e8ff50b562b0265fddf8e0aa7667440dd789d10dc622578ae44d1e41e024fd2c

Observation 4f97b871-d623-43ec-9b5a-a0bc11ef5c90 · inbound

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation cites this paper.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:17:01.334605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:f967daa3950c6f302b9bbb98a5f63be3255a3f702050a2ed353e3b123330235d

Observation 7ae26f0a-6d32-4375-a17a-a4d574c2675a · inbound

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation cites this paper.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.831021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:fe82334c011cc3bbd6716750bb1042c0601b4ed6108909187b4f9be42d5b06ce

Observation c12f8fb3-9635-4c99-bf9a-54dba4d0391e · inbound

Language Conditioned Multi-Finger Dexterous Manipulation Enabled by Physical Compliance and Switching of Controllers cites this paper.

Language Conditioned Multi-Finger Dexterous Manipulation Enabled by Physical Compliance and Switching of Controllers RT-1: Robotics Transformer for Real-World Control at Scale

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:08:21.014610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T19:06:58.600946Z digest=sha256:56b09f6a8c6371eddcfcb00dde0120777b1339e008d49d6a5fe565b3de67273a

Observation b49a8203-7348-4e1d-b605-566d6220eb0a · inbound

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning cites this paper.

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T16:06:09.751767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T16:06:09.448517Z digest=sha256:f3f8e6df689c39e4043ef1c304c043334298aa922eda0edde1b029caf468a8fb

Observation c8ecfdb4-8137-4178-babc-53f856d25a4e · inbound

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation cites this paper.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.417757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:faa3403cda77499c0a2dcf5db94949624b433ddd4b98a46e9bd450525f6ade90

Observation 21793fa3-57e3-474b-af4f-db9649f7fd31 · inbound

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies cites this paper.

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:22:44.317799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T08:20:05.898025Z digest=sha256:4625e4a5f45b4b8c59434351978ea8b6d3e6ad7187db8d83a54c20c7a6152ea7

Observation 09b49a95-e003-4e1e-848f-d79d9d334580 · inbound

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields cites this paper.

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields RT-1: Robotics Transformer for Real-World Control at Scale

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:52:43.929093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:49:37.878395Z digest=sha256:363599ac1de0a199bd0277aab3d02f60b85fbb1135bcde8d26d920f220342db1

Observation 1f3a0c97-3e64-4c3b-be50-08af785b1099 · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies RT-1: Robotics Transformer for Real-World Control at Scale

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:27:22.861089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:d55a2f474f31d7b400b1213e329bec4b5cc51a74b9d9c2e039638e43dabbaacd

Observation 0205a320-063b-4fa4-a7bf-8f828a9ac790 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.700111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:9ae75969e7372c1654267bdbd47605be978638714706aa4270a1a6a997da645f

Observation aca5e758-7e48-4e1b-a6fd-0aad8018bcef · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations RT-1: Robotics Transformer for Real-World Control at Scale

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:38:11.207969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:a789ffb5d9eb65efde5e45c37ed290a6f385b79f3d547f636b3569dde62c85a9

Observation 4ed60879-415d-45e0-b15b-cdbf7d349585 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:31.895348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:89db29ba99e82651ccb99fe59bde86075cc36f68c6ab1c74c509fbe07538beaa

Observation 9125a961-8243-44b2-8bb8-c6b91a6d4430 · inbound

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model cites this paper.

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:12:19.761397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T06:12:19.643111Z digest=sha256:06a12fc09f3c049d07fdcdf628179a9aacebb82d71bfc72284071beaf7da940a

Observation 0007f8ad-2d03-4cf3-9233-c02940787d2c · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control RT-1: Robotics Transformer for Real-World Control at Scale

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:48:49.047600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:da88122feab70ded383e346c36d7f97dfb239e70cf619eba0911dfd70f116c5b

Observation 94388843-6dc1-4022-8ce5-4f4fb8f9dd76 · inbound

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models cites this paper.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.192807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:9646339500a5a6d06234fc695ec537867ffb47ef39b6cc924e42e1f489c0e413

Observation 72085391-0f1e-48de-a08e-9e170d119666 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:32:22.612378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:261c716a1ac0bf056d98150d542dea2b9b4e9eb7bc8fe6a010acafe54a346a54

Observation 92d9c25b-8c46-4250-93ed-c5ebf8272a61 · inbound

Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction cites this paper.

Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction RT-1: Robotics Transformer for Real-World Control at Scale

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:27:21.474210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:26:42.319929Z digest=sha256:dd0ffb806ac06e19f6e8823a5a84e88df2cc4e578136287661bc63c197e42095

Observation 0efaa33a-616e-40bf-8936-1028b1fdf352 · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:00:48.926478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:4ac32744f685b47666836a9d3945521321699353a94df85d93a9aa105d803872

Observation e93f4a0f-58bb-45ae-88b0-defb21081e71 · inbound

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots cites this paper.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots RT-1: Robotics Transformer for Real-World Control at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:41:14.065009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:883d69197bd1adfe02994f4c687f9cd3adb702d3dfb5edc68896bc18c5f498d4

Observation fc5f0269-0b31-4d85-aa27-e85e1bba36a0 · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:21:45.228667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:a0b015889f1daa1a6ea204e9944d279c983f3221872067c15e96f496cb40d3e5

Observation b1510de5-0c53-457a-8cbe-ebc43ff4fc69 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization RT-1: Robotics Transformer for Real-World Control at Scale

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.983104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:1f7c5b119aa9d9f472da3f203cfacafaebc17fe599f6edaaa218b6913aa03e91

Observation bff1f616-1b70-4370-aac2-051c5d5253fa · inbound

Transformer-Empowered Actor-Critic Reinforcement Learning for Sequence-Aware Service Function Chain Partitioning cites this paper.

Transformer-Empowered Actor-Critic Reinforcement Learning for Sequence-Aware Service Function Chain Partitioning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:01:54.560213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T17:57:28.270468Z digest=sha256:6bf143dd5660ea0da2b9d344f3dd65987b9045a95c74f965c6c287d7e6664c18

Observation 68f18c5b-1efd-419f-b37f-c4c9c70bb4bf · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data RT-1: Robotics Transformer for Real-World Control at Scale

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.298380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:495e10d38d0532b746c00f62bf1fa1af71362ee36f20c4d745887f7fcb027c5a

Observation 0aa57541-553a-4478-b65f-32938c0a66a1 · inbound

Policy Contrastive Decoding for Robotic Foundation Models cites this paper.

Policy Contrastive Decoding for Robotic Foundation Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:11:38.494285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T14:09:48.762737Z digest=sha256:c41d8a966007edf906ac1d49f015a68a2e44919734d7b5740f5e627c8a8469e5

Observation 752185a5-1a6e-44d2-ac69-affe2c5b4d56 · inbound

FLARE: Robot Learning with Implicit World Modeling cites this paper.

FLARE: Robot Learning with Implicit World Modeling RT-1: Robotics Transformer for Real-World Control at Scale

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:59:08.983519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T15:59:08.846629Z digest=sha256:f8d81daa3703a53452961749fe91ab7e79ee20ac243719d0f61d0a29c4a3f333

Observation ed511c41-7dd1-47ac-8211-fe6676076bea · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:25:47.222464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:9f5bf7e2058bcc22e9920f37123de1956f2aaac6b32920f06844530fd1cf1289

Observation a4e64950-9ae9-4a71-9e69-3c85ea6ff1da · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:55:40.293360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:2ab7f78e510bb325f0236fe81028db68299ba19378dc4ab6f4f6f0f40425a03c

Observation 75c13768-3e92-44a0-9186-c6ef007c1c21 · inbound

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion cites this paper.

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion RT-1: Robotics Transformer for Real-World Control at Scale

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:52:17.897639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:50:17.902979Z digest=sha256:1e0eaef2421ab725d4d5bc71bfcb8add2aebff92f42ce93b0692b1ca02cdfae0

Observation 2cc2b862-3093-4b56-a6a8-0a488c781a51 · inbound

EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild cites this paper.

EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild RT-1: Robotics Transformer for Real-World Control at Scale

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.677233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:56:30.634284Z digest=sha256:e6cbb495af80a710c146fac91a932b628937d1141caed54282d76cef2f5bbfb8

Observation fd769117-8f8b-4c89-8f2e-13d42330c751 · inbound

Rodrigues Network for Learning Robot Actions cites this paper.

Rodrigues Network for Learning Robot Actions RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:27:15.959420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:24:47.221351Z digest=sha256:da112f7e4386192f489377d2819b55f307eaca280a32f4b38191b9e12adddfda

Observation 40907d13-51e1-4a68-b763-09913af2b1b8 · inbound

AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning cites this paper.

AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:46:44.231968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:46:43.955825Z digest=sha256:990b8ed0392c921b4abfc939f2b1fac7627d8160ed5b533809359472e27a2b56

Observation d5c9bba5-e3f9-4293-938a-4fcc5bad20af · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:55:46.331626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:b5ad5e6b11d9e14be8cc9e974586e248720da2a6fc94c4485482e33843adffbb

Observation e2fbd0d5-a7c2-467a-92ab-0cae3ca0170c · inbound

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation cites this paper.

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:40:27.359351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T06:40:27.206337Z digest=sha256:ed9ddfa4a6a040e591ec76bbb37ed11ee2708917cfefc69da5ab81d3822fe91a

Observation c846530d-9680-4f41-bfa5-d29ddc828b62 · inbound

WorldVLA: Towards Autoregressive Action World Model cites this paper.

WorldVLA: Towards Autoregressive Action World Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:57:08.007482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T22:57:07.883617Z digest=sha256:70bd5c20525207efc1fde6e15596c077fc70f6480cb3e08bb623595306d73b99

Observation 7edfa462-ebe9-43f5-803f-31616c344f08 · inbound

DexWrist: A Robotic Wrist for Constrained and Dynamic Manipulation cites this paper.

DexWrist: A Robotic Wrist for Constrained and Dynamic Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:32:07.864483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T06:27:27.164035Z digest=sha256:d850e5aac2dff3b74467ef41c4e9b4d216e022f67d73ef38a2e0b531aeb4880a

Observation d87fcb16-5b49-4a8b-8582-af294ccafdbb · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective RT-1: Robotics Transformer for Real-World Control at Scale

Reference 210

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T14:08:35.546539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:72a6aa8e7c80a329ae52ec0020efb746fbf8de0041759d056514395f3cabd8d3

Observation de944c56-498d-444e-a80e-313afc5e8cd8 · inbound

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation cites this paper.

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:32:56.667272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:32:56.397350Z digest=sha256:f00dcabd159cc3168095e97e9f93efd8a6d6fe3f9b578b36ff22392ddf981151

Observation ee730c38-d95a-433e-a641-b732bb5a3cf7 · inbound

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation cites this paper.

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:52:57.554667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:52:18.984005Z digest=sha256:a0015a0668620583c2979dfc68001861d13a3d8309c39c18b4b8c9d321074cdb

Observation 60716b35-40e9-4683-bda9-c7e95bcc3227 · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.634578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:597624efb4ddca8d40b068dda724aba8dec892493e98993baf5dfb5b11317764

Observation 80d8d40b-701f-4cc2-a289-d79af865b1a1 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:01.004579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:76b4a24df612598ee7334a1cf751555b4bc40b3e0d6efe62c165dda7d1e1a605

Observation 9f8a3c9e-ac97-45f0-9bd2-85dee69882ba · inbound

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models cites this paper.

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:52:02.948467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T21:52:02.893886Z digest=sha256:31151e66d96912fd7296682d1d185cd2ef571b5a347f9dc5ef7a6640662eccbd

Observation 23ea0b4e-5697-4c39-b686-95bcc92198fe · inbound

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions cites this paper.

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions RT-1: Robotics Transformer for Real-World Control at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:53:37.497994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:53:37.497994Z digest=sha256:edd72d4f6c5cc102b5e6e67ef6a3117ed314bb6ee20c81d96a08acb5fc7b8dee

Observation aa5d943f-cdd2-4cc3-b375-45ba2bee1a33 · inbound

GhostShell: Streaming LLM Function Calls for Concurrent Embodied Programming cites this paper.

GhostShell: Streaming LLM Function Calls for Concurrent Embodied Programming RT-1: Robotics Transformer for Real-World Control at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:30:47.750530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:30:47.750530Z digest=sha256:1f98b28758ed580eba2405dbdd1a36b4aca660b8508e84c141a30a1212af5118

Observation 2f5cbe90-e03d-45b7-b3a7-6b38b59a7cc1 · inbound

PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation cites this paper.

PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:07:31.071093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:07:31.071093Z digest=sha256:324cf8ee6846489d382adabde81c861973ac9aca9d252a0d558919d3381bd459

Observation ef6b4882-ee0c-4f74-90ea-d15c6d4ddb02 · inbound

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation cites this paper.

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:50:40.997985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:50:40.997985Z digest=sha256:8df8c72c4e6ec30cb34cdfc6bde55178902dbecf614d6fe93da738c752a63437

Observation 69d4df93-22b5-4fcc-8844-8bc81fdf9ecd · inbound

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation cites this paper.

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:46:39.823441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:46:39.823441Z digest=sha256:6dd2c22efef3eb31118544dbadf6a641f5bba4ac14d0c5ff58c9a3494b58e6e5

Observation c8deaee2-e282-4744-ac09-d11a2b5820e7 · inbound

AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning cites this paper.

AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:43.667661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:03:43.667661Z digest=sha256:507b42fea4f92aecf03c551229c43619727da22004a2b5ed7147e07b4eef150b

Observation 2ab685a4-4b7f-4a09-b1be-3b9db0277c27 · inbound

GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions cites this paper.

GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:59:00.915445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:59:00.915445Z digest=sha256:9839e25935d0c9f0ca90c78cd1a6cbe64188a732bd7ddbc996d82ea26e1fcd3d

Observation c419c722-960f-450c-abf8-f364c27ce629 · inbound

SwarmVLM: VLM-Guided Impedance Control for Autonomous Navigation of Heterogeneous Robots in Dynamic Warehousing cites this paper.

SwarmVLM: VLM-Guided Impedance Control for Autonomous Navigation of Heterogeneous Robots in Dynamic Warehousing RT-1: Robotics Transformer for Real-World Control at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:26.098762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:26.098762Z digest=sha256:2ba1f92e7243f30a654e9cd02e2df0d3a6661ed8a72ce6f2fd7b74b7e43aac4a

Observation bfb44e2c-d8e9-4a98-82f0-b1adadedab2e · inbound

DeepFleet: Multi-Agent Foundation Models for Mobile Robots cites this paper.

DeepFleet: Multi-Agent Foundation Models for Mobile Robots RT-1: Robotics Transformer for Real-World Control at Scale

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-19T00:16:55.399717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T00:13:27.533407Z digest=sha256:6af0efb135ef373853d47f3bfecafefdd45223612895e5ec1e66e117c9407e3d

Observation cae663c9-ceef-46cb-8ddc-516ed1e9dc8d · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:38.305098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:38.305098Z digest=sha256:7965e36440a727c9d9af156a37d972b84f9e279a11413a0614bf9a45b574ec3b

Observation 7e4f4873-108e-497f-a566-59a3242f80a5 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:34.191609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:34.191609Z digest=sha256:502f69a913621ccf6b3b1ff94221dddc843f673658de63988b30407312d113dc

Observation 2c495f89-e5a7-439f-8d27-ec80194f2e00 · inbound

Precise Action-to-Video Generation Through Visual Action Prompts cites this paper.

Precise Action-to-Video Generation Through Visual Action Prompts RT-1: Robotics Transformer for Real-World Control at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.378635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.378635Z digest=sha256:00e08aaced78fd8f9fef36c3ec5238d620be7abf4d82c31ac48eed04e072428d

Observation 39fee9e2-59d8-4708-b9a2-3fd4c2285b8b · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges RT-1: Robotics Transformer for Real-World Control at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.009894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.009894Z digest=sha256:5a4f4d0e081bfd8776f33fbf087b2da9a295b59cdfe5af7a77e4f2d05650b2ce

Observation d00f9ac8-b3e0-4ec2-8ddd-c3485f7ca83e · inbound

An LLM-powered Natural-to-Robotic Language Translation Framework with Correctness Guarantees cites this paper.

An LLM-powered Natural-to-Robotic Language Translation Framework with Correctness Guarantees RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:52.856652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:00:52.856652Z digest=sha256:4b52a48f6aa3391e6706ad2ae7ae5c0e938af77f431867289932f9990e4a8d12

Observation 62d756e0-feb3-44b8-b57b-9f3848d46a9e · inbound

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:43:24.444837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:43:24.417901Z digest=sha256:2eb994b188a2a9007999982b115b839e6407212f8fb92f5fe6a745c7b3c5d83d

Observation 59ac0281-9bee-4e49-9e43-07ca17fb57e2 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification RT-1: Robotics Transformer for Real-World Control at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.571191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.571191Z digest=sha256:dc33e7c02441e929c4123e172c1f75d4acfac122071af42b2b0f63c720d21857

Observation 1a4c6b88-bf8c-40a4-b512-c43c8d1d6b05 · inbound

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation cites this paper.

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:41.025653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:41.025653Z digest=sha256:8a262eaab60d43ebbaf0ffb651dd4479a92530fa8e5594987c3a3a0714b6384e

Observation 558aa886-49c3-49a1-b0f7-ee453ee5f2f6 · inbound

MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation cites this paper.

MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T12:23:06.522848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:23:06.522848Z digest=sha256:25b80d69051e11ec942630b1d223138110604048c008b0454ebc0de2fe861e02

Observation 8cbb395c-ed1d-48b8-9489-a93857a0a844 · inbound

Learning to Ask: Decision Transformers for Adaptive Quantitative Group Testing cites this paper.

Learning to Ask: Decision Transformers for Adaptive Quantitative Group Testing RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:30.170220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:30.170220Z digest=sha256:d4442dfdd3190a62a4de86cb34e8ce76d5c323ddf87e2fb27827b8829a31556a

Observation 192fc497-11f2-48b1-b603-b632616a995d · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.257110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.257110Z digest=sha256:93fe05e7b5d8f2b8dd1394bacc08d27431e729d87323e312b274cfce1ef23bd8

Observation 50d885f1-a294-4f5a-a4b5-d79854f048af · inbound

ANNIE: Be Careful of Your Robots cites this paper.

ANNIE: Be Careful of Your Robots RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:00:38.078835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:00:38.078835Z digest=sha256:77c08879c1edd2ba19ab7a70cf577533f4f678f405cceeef71305cd40cf7c39b

Observation d63c2495-1ddf-4d7f-88f7-592d5332d574 · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.488324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.488324Z digest=sha256:2e7efcc85981eb6027cd03a390eceee805c44be240f93ccc0c3a769f4d1c3143

Observation 850b3a3b-4f8a-4e47-a402-ece9fde14ab8 · inbound

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies cites this paper.

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T05:48:46.285406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:48:46.285406Z digest=sha256:a1dae88a6a97a5a8e989fa8b7202304e04d8721ed33c98b9a3058672658f7384

Observation d6dfc8b8-1200-484c-bdd3-d98d381fe61c · inbound

Plantbot: Integrating Plant and Robot through LLM Modular Agent Networks cites this paper.

Plantbot: Integrating Plant and Robot through LLM Modular Agent Networks RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:32:01.097740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:32:01.097740Z digest=sha256:cbf9804132f189a1ca0e4535013ecb151430d5d0c2cb05533becba1778e5f233

Observation 909c812e-2adb-4629-b094-ec1a00797727 · inbound

LLaDA-VLA: Vision Language Diffusion Action Models cites this paper.

LLaDA-VLA: Vision Language Diffusion Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.826083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.826083Z digest=sha256:b30e2a805c735b838c0f763846dd5dd04055beed6bb35eb6fc11c1e3f24f9e4e

Observation ac76b9b6-d31a-47d9-abd1-aea570868824 · inbound

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions cites this paper.

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:42:47.008808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:42:46.978148Z digest=sha256:630c921283a62eb8ffd8debb8ba8917751e868d6919744716bb376522563f780

Observation 4223f01f-21e2-455f-8151-4915ca379763 · inbound

Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions cites this paper.

Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions RT-1: Robotics Transformer for Real-World Control at Scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T22:16:15.379665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:16:15.379665Z digest=sha256:e29b433b95b26bdcbe183cea9ea4312a8c2629e907fb06d02b014063bd4ffac9

Observation 6244f361-bda0-4892-892a-bc93851a38f9 · inbound

RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction cites this paper.

RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:55.915566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:32:55.915566Z digest=sha256:33a28217197f0542cd45d66c8ac9b413c74e6b27a3f1990689543779e69fdf70

Observation 7866a3ef-ece8-4598-86d5-a7bdcefb6bfd · inbound

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation cites this paper.

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T20:10:13.484730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:10:13.484730Z digest=sha256:50eb30cb5b32533b3cc9bdaa362557ea7c9a973fdcc893b0c2ac42814da5083a

Observation 8ec99adc-64df-4064-a0c8-c71c62f2c9ec · inbound

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution cites this paper.

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution RT-1: Robotics Transformer for Real-World Control at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:12.581506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:12.581506Z digest=sha256:08b44b1f2c6e50f7222800e98f70f045874558fa104a05e6ec395fb49bd05837

Observation 368931d7-7fb9-48d5-a52f-578c8297282d · inbound

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations cites this paper.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.544334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.544334Z digest=sha256:02bd238d09bb00e60a27342c8ed0c1fafde8d08abbada60d3fb13fae82f944e0

Observation 0ea6386e-4b8d-4ac4-9fa0-1b34743e07ba · inbound

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models cites this paper.

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:50.033967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:54:50.033967Z digest=sha256:a13b7bbad40361d663a3cd98296cb295ff6ca1e85e356d6e3b8d5bc0d0c5d784