Pith. sign in

Paper Citation Record · LEDGER

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2307.15818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.15818 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 100 of 638 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:12:55.725471Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

270
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 96658184-8130-456b-a56a-c52d5da317bc · inbound

Cognitive Architectures for Language Agents cites this paper.

Cognitive Architectures for Language Agents RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:33:44.278357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T19:33:44.146134Z digest=sha256:784abdbda003c204cc1d503ab2171f1387de01c318ea684967f1c293032b3889

Observation ef194d76-9d83-44ad-b50e-1b4cf63ca8b8 · inbound

GPT-Driver: Learning to Drive with GPT cites this paper.

GPT-Driver: Learning to Drive with GPT RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:05:31.982543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T15:05:31.928650Z digest=sha256:1a341436f50dd8d49c75832b22e70ebe57d3295ea4915919c3d0092729f71307

Observation f50b1ee2-ebe6-4ea7-b375-ed4d48be7db3 · inbound

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own cites this paper.

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:44:02.621634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-24T06:40:00.328012Z digest=sha256:bbcc03f8622a99a78c1164d805894b2e3e181212c5d24d50d8c61fa8802ce96e

Observation 649fbcd2-8058-433b-a436-24504b3c9aae · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:15:18.658514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:8303c68fb7e277dbded9661c4f4aa5193c04700ec1758173b01da79961447cb8

Observation 23bba126-a8fd-4f01-b6e9-09e8094c89de · inbound

Open X-Embodiment: Robotic Learning Datasets and RT-X Models cites this paper.

Open X-Embodiment: Robotic Learning Datasets and RT-X Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:23:24.449479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T17:23:24.255829Z digest=sha256:8d78a8909fb5d5196da7143c9dcc5a69ac519c1e2b0d18b51a3bfd61179d6b1d

Observation 38c972c4-232c-4ccb-9a83-c2208992c99e · inbound

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models cites this paper.

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:54:59.000154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T05:54:58.940428Z digest=sha256:f94c8daa90d8007c6ee2ae6c659e73d1cbf6f0d2de567c0d6e4bef256e6b852f

Observation 3fd271d6-e777-4d79-844f-8ab2ab44ba53 · inbound

TD-MPC2: Scalable, Robust World Models for Continuous Control cites this paper.

TD-MPC2: Scalable, Robust World Models for Continuous Control RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 156

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T17:27:35.993143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-14T17:27:35.733800Z digest=sha256:27d2bb9da42468ed52e0bed8dc248c68ada7df94a1f4e8be7eeb5aa247996bdd

Observation 21e3ab03-a12d-4689-9be7-9794c12146e1 · inbound

Vision-Language Foundation Models as Effective Robot Imitators cites this paper.

Vision-Language Foundation Models as Effective Robot Imitators RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.605460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:2bf27097df0d5e3120930a8be3b879f9c0b28fb3a76f80a53b66e42ddddd5499

Observation e5f92c29-2975-4e70-9586-65035dd4c01b · inbound

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation cites this paper.

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:32:05.595167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T16:32:05.538507Z digest=sha256:ba8f14778162fb590e9a1c8c1e1d3a064ecfc723f038d9933f3f4552adb157aa

Observation 73ce41a8-5064-48c6-80d9-0e4fd8e93486 · inbound

AppAgent: Multimodal Agents as Smartphone Users cites this paper.

AppAgent: Multimodal Agents as Smartphone Users RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T10:16:43.896074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-17T10:16:43.364787Z digest=sha256:81d278af3dd553706d341544cc5b548ab44b931c25759742f98d1c90442bc251

Observation bef27680-7526-4a59-889c-cf316c3614fa · inbound

Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation cites this paper.

Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:02:55.532844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T22:02:55.240949Z digest=sha256:ea196ec68d41be6e4730376ae0f833fa14c492e8bd012ef02ee17e16bb6efc79

Observation 916b2d1a-9167-419b-aad0-b5a2b05d2e1a · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 120

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T14:25:59.347132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:c6097f6031d82737c890c90cbf916d852ec3b0221f88f111d8d824772333a3c0

Observation f022c4e6-8fd8-4677-aec9-aeda741ec060 · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-12T19:22:35.424710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:c76c0b5ccdbd547bbda61a845a1e3710a6232980b50a9a7cd9694ba9ea4b4a06

Observation b58cb528-c895-49c0-87a0-a7c36819cb45 · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:55:20.422893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:c93155a0d96b0a15c3b4a9add15d2a34e8c9801fd512581d6b6cf31004957687

Observation 30ab2bec-9c11-49af-ba4c-a82f90499660 · inbound

RT-H: Action Hierarchies Using Language cites this paper.

RT-H: Action Hierarchies Using Language RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T06:53:27.772992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T06:53:27.642020Z digest=sha256:0c701e1e69cfca7f810bdf35294f924313f23a6ef76b273613df44fd2921e849

Observation fb41528e-1b26-46b0-ae0e-888efea28365 · inbound

3D-VLA: A 3D Vision-Language-Action Generative World Model cites this paper.

3D-VLA: A 3D Vision-Language-Action Generative World Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.264601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:6d7c574944bfc78d0154c9723947217dfe913d452ccd2756924ae51776080353

Observation 37a39ef4-f617-4a6d-abd2-6c000a0eaf42 · inbound

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments cites this paper.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.450810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:adcb9ffaf0a8c5f8ffc99cd253efb029a696faabd22056eccd70eb774f4cb7dd

Observation 4882d3c3-7a9d-4410-9107-ade4035e05b6 · inbound

The Platonic Representation Hypothesis cites this paper.

The Platonic Representation Hypothesis RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 219

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:03:56.818142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T06:03:56.328012Z digest=sha256:e7bf72ae67833c548d35fd05d358c8650022fb61080f1178ca21c4f986ca857e

Observation cdb17367-46f1-4777-8115-4007acb99935 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:25:54.285080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:4bbc6a3a47807dfa51451a40bb396424cd63ac1f92b3cfa4bf0537b27c301292

Observation 69f71b86-68fc-4236-957a-6ae74cb5af71 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:36:04.729348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:5213f43efcb3c1b8f86fe7cc354e0603aeac25fe00bfc9bbdce668522ea2c086

Observation 54bf625c-a741-4d60-81c8-4c46b7298ba1 · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.440421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:7a7b63d7216938d88c3f7178198737bcd47ed37deb4b382947396a497d41a699

Observation 8158301a-499c-477b-ba24-38f401b22131 · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 107

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:25:17.975301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:354af298f30cb0a18506c9279229e779854876f1ea81780e6b4e4f5fc8046670

Observation 6bb2cd43-b995-43ca-9603-ddb731b32873 · inbound

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation cites this paper.

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T16:12:26.080139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T16:12:25.980853Z digest=sha256:00a374f060543f2477d6029b2023e113e2aad6f4e9c5351d9921bc2d293aaced

Observation a130f85b-dd09-4f7b-9b58-01a4f0dcb0df · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 150

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T12:04:10.698283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:47ae017ee7c114cf4054e79ab17e03b583ab5720007a17c13095a51e622e6dee

Observation 668a0c42-cfd3-4db6-ab7e-484a94654fc8 · inbound

Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models cites this paper.

Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:23:24.873272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T20:19:20.382009Z digest=sha256:9a316f6c570804d946517ed220fd10e076f6bbc13fdf99e5c9d6b619ae172429

Observation cb94fd1d-303c-4ee1-8144-d70b041977cc · inbound

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation cites this paper.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.896519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:03b0a723c55cb759da336908d31519d01e935bb7a98dae5a0905db09e384652f

Observation f9ac1071-0598-4bd7-bc05-8314ca9f4cb9 · inbound

Language Conditioned Multi-Finger Dexterous Manipulation Enabled by Physical Compliance and Switching of Controllers cites this paper.

Language Conditioned Multi-Finger Dexterous Manipulation Enabled by Physical Compliance and Switching of Controllers RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:08:21.025254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T19:06:58.600946Z digest=sha256:3ab72c51cdb8ae41bc378365d7da78cf80795b24386c0c41103cc0ca55bbc2af

Observation 9e164021-00cd-4c88-bc85-f7c0beeb88f3 · inbound

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents cites this paper.

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T09:29:27.362147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T09:29:27.173784Z digest=sha256:68d7fd57607af93baeb1c103e70827f292df0d62cd8aa898abf08030363c0b7b

Observation 0ead8dff-ba03-423b-94af-1257f7b9ddca · inbound

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control cites this paper.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:36:04.729348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:600317f464eedef221ab2424313768baf51ee0b184d99d43b0313210aa0f95c1

Observation ea61dd21-5e13-46dc-98ab-88fb0eb5a5a4 · inbound

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning cites this paper.

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T16:06:09.618905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-17T16:06:09.448517Z digest=sha256:34b6d11239a65f414acda4c59e6ac3750ff93d9e2adced86dc6c55c63171d8da

Observation 8ff72640-c87b-4848-8c52-9699682fb933 · inbound

DART-LLM: Dependency-Aware Multi-Robot Task Decomposition and Execution using Large Language Models cites this paper.

DART-LLM: Dependency-Aware Multi-Robot Task Decomposition and Execution using Large Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:12:55.725471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:12:55.725471Z digest=sha256:0d524dba8442b61465a4d456fea649601225f4b3a5b3bb5101634312ecc72a36

Observation 6caea9b3-b158-49f0-9e63-62b1e085c5b6 · inbound

ClevrSkills: Compositional Language and Visual Reasoning in Robotics cites this paper.

ClevrSkills: Compositional Language and Visual Reasoning in Robotics RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:10:46.109001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:10:46.109001Z digest=sha256:6477456f107746eeb835b05d2ab76cf6895c0188e65ce2aeb7f161c6bb391704

Observation a77a7c43-8faa-4108-ae38-69fb5e90ff21 · inbound

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation cites this paper.

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T21:04:21.793522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:04:21.793522Z digest=sha256:840bdde36924a3dbef30bdab2beaf90c2905aaf4978301051a1ad38671bc9507

Observation afa4cfbc-ac71-4591-86be-8a12bd8caafb · inbound

VeriGraph: Scene Graphs for Execution Verifiable Robot Planning cites this paper.

VeriGraph: Scene Graphs for Execution Verifiable Robot Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:13:13.782190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T17:12:31.645346Z digest=sha256:51095456854dc399dd2453634762f0081ab66ce9557fb36bf7e84f43f32f9a8a

Observation 48753607-abc0-47e3-afbd-49236c0fdff5 · inbound

Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms cites this paper.

Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T19:10:14.525121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:10:14.525121Z digest=sha256:4478def2f781412916ae878cdad5bac383369538491a159d85dfb7390e3f3a16

Observation b0b3482d-54be-4f3a-b3bc-955f090c82cb · inbound

Generative Timelines for Instructed Visual Assembly cites this paper.

Generative Timelines for Instructed Visual Assembly RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:47:39.697108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:47:39.697108Z digest=sha256:ef55f0cb1b712eec5b19aa75e8b1c45dc08895cbac85941ef2f3e4c587867bbb

Observation 4f109449-c97b-4e95-a2f2-ca3fef466c4b · inbound

I Can Tell What I am Doing: Toward Real-World Natural Language Grounding of Robot Experiences cites this paper.

I Can Tell What I am Doing: Toward Real-World Natural Language Grounding of Robot Experiences RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T17:06:34.128646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:06:34.128646Z digest=sha256:7d55964617080cc30d975f3b78af69f6ec553dbe496728d93b3d9d5c80444c8b

Observation 91d73b72-41e1-486c-8197-770a754d7160 · inbound

Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning cites this paper.

Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:24:14.226207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:24:14.226207Z digest=sha256:c6d98dbfc31e2f3158eeda371eb7366dd6d07f5849b009bf4b1f23b41dc8c637

Observation 245c755d-aa35-49be-a360-233464c34a6a · inbound

RoCoDA: Counterfactual Data Augmentation for Data-Efficient Robot Learning from Demonstrations cites this paper.

RoCoDA: Counterfactual Data Augmentation for Data-Efficient Robot Learning from Demonstrations RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:44:26.364806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:44:26.364806Z digest=sha256:65e24db899ecfaaf958fd2ba2c548a9d8b16d1b8934db0d752bc75c7d7142285

Observation a76e26a7-79d4-4e8a-83e5-5a2e6a124367 · inbound

ShowUI: One Vision-Language-Action Model for GUI Visual Agent cites this paper.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.319251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.319251Z digest=sha256:f9e62a2f30c379a6d6982bbc6e118a1e1938492e12488aa75a75e24fe1ab9f0f

Observation 41614331-65d7-47e0-8eeb-976ab87d2f51 · inbound

Prediction with Action: Visual Policy Learning via Joint Denoising Process cites this paper.

Prediction with Action: Visual Policy Learning via Joint Denoising Process RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:29:20.708898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:29:20.708898Z digest=sha256:05af6ced7787c71fdb7b5999f8140d2e956f9269573ea4fd45596d1c2db4e898

Observation fe84e8a9-6e0b-499a-9e62-6641a5e5dfe5 · inbound

Embodied Red Teaming for Auditing Robotic Foundation Models cites this paper.

Embodied Red Teaming for Auditing Robotic Foundation Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:05:51.525177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:05:51.525177Z digest=sha256:fdd9535239b7970bb527c736da1761fe723d134009176da5f6f68def77aff63b

Observation 0194a111-af39-4318-8b9b-20d84a1e7633 · inbound

GRAPE: Generalizing Robot Policy via Preference Alignment cites this paper.

GRAPE: Generalizing Robot Policy via Preference Alignment RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:23:15.557889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:23:15.557889Z digest=sha256:ca4750ebea3ac55b6b69e367122732aff7da5b7f572ac7721778bffcde206309

Observation 929fd749-adcc-45c6-a506-7b00056469fe · inbound

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation cites this paper.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.432800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:478d9abe5a144cdaca4499ba4126c886a0f84902dbc2eec60e2dc67b02409976

Observation 4561561e-a90b-4a51-905b-8febd9491325 · inbound

On Foundation Models for Dynamical Systems from Purely Synthetic Data cites this paper.

On Foundation Models for Dynamical Systems from Purely Synthetic Data RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T05:29:02.547059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:29:02.547059Z digest=sha256:0c5f96c34ef74bb3fdcd142cf52fbdf5d7a7af1b4aef87fd0b407b07c82d95dc

Observation 728e2d13-d0aa-4087-a869-d3d9acaae1ef · inbound

STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft cites this paper.

STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:54:54.400442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:54:54.400442Z digest=sha256:174de2edf3195a8e430b12fdcb4a908608ce7faf27a8faabd695944f10ed2b42

Observation 31b84b8e-6667-41da-a7ae-07a1a67e814a · inbound

Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control cites this paper.

Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:48:33.574671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:48:33.574671Z digest=sha256:e90bfd6d41e11e0c6acda177a1294e131f9ee5734fcc24e808d903a89451f514

Observation 7da82f87-4651-4a8b-8fc2-a7b3a5b42c37 · inbound

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies cites this paper.

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:22:44.360642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-23T08:20:05.898025Z digest=sha256:1fe510379e0a57ad971a001f33e28c045ad30aaa94896d74c88123301c9322a0

Observation abfa9331-448a-47a8-98cf-831b0e642247 · inbound

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields cites this paper.

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:52:43.886679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T07:49:37.878395Z digest=sha256:2eedf0ed34c01381efab637f5df37386aa733a69f8ec0832d9e8bc9d3c89130e

Observation d68b3fe0-05ac-4fab-8d7e-8ada58bf5b1b · inbound

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning cites this paper.

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:38:30.997068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:38:30.997068Z digest=sha256:831e43f9f357b72c8f7724a5df8bc415bb13c2049501dc00a73ccd44c2d1c769

Observation 27a61242-1ec9-4906-bc25-7f33746c3a3d · inbound

AIpparel: A Multimodal Foundation Model for Digital Garments cites this paper.

AIpparel: A Multimodal Foundation Model for Digital Garments RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:02:03.768214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:02:03.768214Z digest=sha256:d1f711ffacf63b87cd2d7a283c6ade9f9f9eff018205c68d01ce0113fcf3bcae

Observation b2b8ee52-f021-44c8-9733-42242a134d02 · inbound

Dissociating Artificial Intelligence from Artificial Consciousness cites this paper.

Dissociating Artificial Intelligence from Artificial Consciousness RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:55.478603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:55.478603Z digest=sha256:b9e98bfcb3b487b1ace57f25bcc0bb8b36425f8814f0f28e79e9d2e4cd1c24f8

Observation 801fa62b-b580-48a0-9527-024e954c2582 · inbound

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions cites this paper.

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:37:54.614751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:37:54.614751Z digest=sha256:ffc770c808e7ee2ac7c25c209f29d8eef24b47f972c2cc0b0dec9cef0a4c116a

Observation 09ae7cd0-ac43-4720-b14d-2d730d95ad82 · inbound

Can Large Language Models Help Developers with Robotic Finite State Machine Modification? cites this paper.

Can Large Language Models Help Developers with Robotic Finite State Machine Modification? RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:35:32.921030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:35:32.921030Z digest=sha256:ced224a4e0e496af20977c687043fb3a07b55099ee076001cf81bbbf1e215d86

Observation 5457b199-a646-4e3c-8f33-a7ca63a66dbb · inbound

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks cites this paper.

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:51:36.238946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T19:51:36.137985Z digest=sha256:d87ccb88aba22a9f5ffce114e009fa284a8950e199e692d7f63243edf56dc1be

Observation 8083658f-daf3-49e9-8445-d586424e8d39 · inbound

World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving cites this paper.

World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:52:15.134730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:52:15.134730Z digest=sha256:286854233fa96108d0f9b5e266540c07c2f86f210f5eea2451a7605f73769efc

Observation b519844a-7607-41f4-8b74-1cd4eb869f43 · inbound

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction cites this paper.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.172355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.172355Z digest=sha256:7ca38f30f36b37f20c662a1512d5ab380821ffd54f8618e9f518909291048212

Observation da14a212-209f-4beb-88aa-6134e084f8d5 · inbound

The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data cites this paper.

The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T19:23:12.074697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:23:12.074697Z digest=sha256:e15e9b4b60de0634c27bf4370aa74274d57c4cd1eed54f1e3094c680989291a3

Observation f51e188b-b378-4b15-b5b3-e540e52c8e9b · inbound

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons cites this paper.

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:55:13.202991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:55:13.202991Z digest=sha256:9a037a5ba22445cc60bab454c7a94dbc72b59a05675bb4f5b84af0d0b9865b58

Observation fb39e0f2-581a-4627-b27e-05764f095e0d · inbound

Learning Novel Skills from Language-Generated Demonstrations cites this paper.

Learning Novel Skills from Language-Generated Demonstrations RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:58.177565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:58.177565Z digest=sha256:f74f7e617d83e4398fa17dd53a750033c3dcd40f8b77cf5588aa18c70727a3a9

Observation 1b262064-4d81-44c2-b0f3-38a88d99d9f2 · inbound

Owl-1: Omni World Model for Consistent Long Video Generation cites this paper.

Owl-1: Omni World Model for Consistent Long Video Generation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:05.438400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:58:05.438400Z digest=sha256:98e8c4d3f01b23b699ba1a4b725f40ed36c9f503469bae2efd4a70a0941b23c4

Observation 14a40144-bb53-4fc7-8b60-7c81d6f01980 · inbound

Learning Camera Movement Control from Real-World Drone Videos cites this paper.

Learning Camera Movement Control from Real-World Drone Videos RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:54:39.688502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:54:39.688502Z digest=sha256:b11d15b016631076ddfaa777b50f2bc46e8c53d078ef865d44f47678c727803c

Observation c7e826cf-3740-4b27-a9a6-83344279e1bd · inbound

Doe-1: Closed-Loop Autonomous Driving with Large World Model cites this paper.

Doe-1: Closed-Loop Autonomous Driving with Large World Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:55:42.161217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:55:42.161217Z digest=sha256:37c60e0d918c7e228fc7efaf485dc641197a7ffd1e5eb4dbc6f94dd30e1c4fa9

Observation 4d00c172-07b8-4636-b490-b7b01f03a7f9 · inbound

RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning cites this paper.

RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:08.178444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:08.178444Z digest=sha256:90f9a5697b2979e011c9dc85961520a504c7d46249a0b99a7241395b8ecc0495

Observation f09baac6-b13e-4ac8-9ffd-647ba224fe8e · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:27:22.867493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:ecc89d911fca92dacf34edb9b7d81a387c9856a4dfa73eef6d41840ddb254883

Observation 1732c799-21f6-4b38-b370-1f9ef5864090 · inbound

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents cites this paper.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.013831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.013831Z digest=sha256:24d57875c8e3c5eafb425981b343744777d9fdf27657d6e0a64622d2a3b3464a

Observation ce788729-3152-4c03-b513-db31a68808d6 · inbound

Modality-Driven Design for Multi-Step Dexterous Manipulation: Insights from Neuroscience cites this paper.

Modality-Driven Design for Multi-Step Dexterous Manipulation: Insights from Neuroscience RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:05:13.175421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:05:13.175421Z digest=sha256:55ee81327613f5443268d74c17e18098fc4117e72493240bb41aedb5f04f7298

Observation 7cc30af0-d5ea-43f4-8c55-f0707043de11 · inbound

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning cites this paper.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.238803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.238803Z digest=sha256:06b50fdcb19a4ba501f39fbe07c7cb9a0f95fa8f980ef1e6055bf8b3f974a7b7

Observation 69986738-ce32-4f7d-a3b5-9c5142ebea0b · inbound

Meta-Controller: Few-Shot Imitation of Unseen Embodiments and Tasks in Continuous Control cites this paper.

Meta-Controller: Few-Shot Imitation of Unseen Embodiments and Tasks in Continuous Control RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:34:45.132484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:34:45.132484Z digest=sha256:35db2075a96acd80d5485d951c91a1259cedae12135fa015d5f9412ed2e15a08

Observation d68e3086-cdbd-443c-bf8d-f0aa26cbe441 · inbound

Task-Parameter Nexus: Task-Specific Parameter Learning for Model-Based Control cites this paper.

Task-Parameter Nexus: Task-Specific Parameter Learning for Model-Based Control RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:08:35.019456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:08:35.019456Z digest=sha256:eb8e52d04fee437aed2856fdd9bae0fab5a05bf381fb39da41db5d7940323d4b

Observation d2f43dcf-3eab-497d-80b0-bbe697b6ac30 · inbound

Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning cites this paper.

Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:41.081209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:41.081209Z digest=sha256:d38a6c1a19667d6b59eb8125a64734ff48830df338ee9daf02c999c25daacda6

Observation 635f13ad-a235-49ee-8cee-bfc845328e1a · inbound

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction cites this paper.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.337527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.337527Z digest=sha256:53a5585ef0afa976ed8c1279cf2543eeb82d533b30afe30e809397ff6dd51b43

Observation 8f70d6d8-e7a0-4264-823a-6d06bd4319d0 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.706165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:e3917521e94fc341e10e5350452334fefaa1ad471fbb3e84b140467d5c1b13ba

Observation efb69738-553e-4497-a1e3-e7834093d8a5 · inbound

The One RING: a Robotic Indoor Navigation Generalist cites this paper.

The One RING: a Robotic Indoor Navigation Generalist RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:20:28.032342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:20:28.032342Z digest=sha256:06529498b5d3918925a2011bffebc239548f108fde4ceb00db6d990e339835fa

Observation 870883cb-c452-47dd-b7f6-86fa504fe8fa · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:38:11.219125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:5dacf599abe89b186f7a4dcb838dd90d2484c380f33dc0c910ef5d59a77fad72

Observation 7e666c42-3e23-4fbd-8ecd-6c47303b384c · inbound

AutoLife: Automatic Life Journaling with Smartphones and LLMs cites this paper.

AutoLife: Automatic Life Journaling with Smartphones and LLMs RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:12:32.723483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:12:32.723483Z digest=sha256:5f778e42528270c6fb6f31a67071cc7075e1281ebc9aaabb6f11b8119a88738d

Observation 462edb99-56b0-4f08-b2a2-3ab34c1da5ce · inbound

System-2 Mathematical Reasoning via Enriched Instruction Tuning cites this paper.

System-2 Mathematical Reasoning via Enriched Instruction Tuning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:54.961904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:54.961904Z digest=sha256:375b572194dbd34df8a6b4351bdaa5862cb7719bbe8251eca29fc1997603b95c

Observation 1d4422b2-3183-4872-b068-6c6730d3f162 · inbound

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks cites this paper.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.537399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.537399Z digest=sha256:0ebe66d34bb07c7451554cc6270cafbce01be83f516cb5be48bf281bad1d93bf

Observation ddba8c1b-5791-42cf-9a91-2b6cc081af45 · inbound

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks cites this paper.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.620506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.620506Z digest=sha256:a38c115c2c76c94c4c3fc1ef1d8edea884a701a265d313d7427d750b63024d6c

Observation ce96daae-8f69-4ae3-b931-b06f74a29a5f · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.245779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.245779Z digest=sha256:82ef52f3bf75ea41f3c90652f71dd10723d35657edced0c8ada91377570d713f

Observation 6322b4e8-2a9f-46d4-ad72-c8609b6e963c · inbound

FaGeL: Fabric LLMs Agent empowered Embodied Intelligence Evolution with Autonomous Human-Machine Collaboration cites this paper.

FaGeL: Fabric LLMs Agent empowered Embodied Intelligence Evolution with Autonomous Human-Machine Collaboration RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:27:21.391885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:27:21.391885Z digest=sha256:24e82dbdc40cfc0d235d1e8ec86996b278be9a09b0ca0c14ece2abf1accdc233

Observation afa221d4-4a07-4cff-aeff-66825466c7bf · inbound

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance cites this paper.

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:37.339539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:37.339539Z digest=sha256:47e60eb9fa7dc832f002272a5be6fa5b55e3834034e54a291494fc7b088d999b

Observation 1c5ee29e-1ccd-49a0-9ad1-fff95276bad7 · inbound

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking cites this paper.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.991648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.991648Z digest=sha256:30b99f1974faa87f472caaed05e4e65083f208f1ed04a426acd4a4295b1388b8

Observation ec9c6166-5c3c-4088-9934-5f9b2747f732 · inbound

OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints cites this paper.

OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:50:41.032130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:50:41.032130Z digest=sha256:097607e702b5f30b54bd90745a6286347ef433bee4d355f0f67db3ff2cc12510

Observation b3bb3328-d326-4921-952d-9afe55280f7d · inbound

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives cites this paper.

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:44:55.692943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:44:55.692943Z digest=sha256:67a70d9cdc2bd069328eb48078047610bd1b2972bebeea687252ebaa57cabc3d

Observation db453c77-7842-449e-8aad-968d3c912f7a · inbound

Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation cites this paper.

Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:28.989604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:28.989604Z digest=sha256:b185d5d00f7f5e340c44c4ef67be78082801b450dd3425efae5b3948a6a4fd44

Observation 14aca93b-3052-4b89-83b0-feb16f436a09 · inbound

Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions cites this paper.

Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 273

Resolution
unresolved
no resolver link, observed 2026-08-10T21:37:05.443176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:37:05.443176Z digest=sha256:224a7a8168fd11b134bddbdf77e056e6753e717b9f484a88589ea94d8ea99e02

Observation b557a1db-8b23-4c04-b92a-d3b8486721ca · inbound

Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding cites this paper.

Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:20.698532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:20.698532Z digest=sha256:35678c6c2383afa5eededd910318a468e1c6eb0c6d06e276dc56184e5fcdf2be

Observation 12713fc5-4687-4d6c-baa2-32f687342492 · inbound

UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation cites this paper.

UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:29.333458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:29.333458Z digest=sha256:e7746d2bee04e8ced5643b5677cf878941564d454c8270edf9f2597d3a78d308

Observation 14c7e728-3a75-436b-9d65-cdf6aa305cc5 · inbound

Shake-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Manipulations and Liquid Mixing cites this paper.

Shake-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Manipulations and Liquid Mixing RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:25.875000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:25.875000Z digest=sha256:fea3338da6d698d28b02ed5e8bff70612d8fb8f2565657413a4342297ebdeade

Observation 7b1adefa-12e1-468b-903b-79afe464c0f0 · inbound

From Screens to Scenes: A Survey of Embodied AI in Healthcare cites this paper.

From Screens to Scenes: A Survey of Embodied AI in Healthcare RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T20:43:09.918001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:43:09.918001Z digest=sha256:5ede72ce346bc953208933f026b3b0c09818682415417ef2d76f3c04122a650e

Observation 289d7746-07bd-40cb-9d58-e247f5c2a9cd · inbound

LAMS: LLM-Driven Automatic Mode Switching for Assistive Teleoperation cites this paper.

LAMS: LLM-Driven Automatic Mode Switching for Assistive Teleoperation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:05.556413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:05.556413Z digest=sha256:bdf3c135e5ca357b6093e132c2282d914f2875ae88123d723708fc95036fde6c

Observation 447861ab-e23e-4abc-8f8a-4e8759b184d7 · inbound

Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving cites this paper.

Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:20:48.502809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:20:48.502809Z digest=sha256:28c771a6f3c3e09b0bf91aed508950029c893d9f13c21a610a46b32bfe41d66b

Observation d489a10d-73b8-4a93-985f-b51d0b174dad · inbound

Embodied Scene Understanding for Vision Language Models via MetaVQA cites this paper.

Embodied Scene Understanding for Vision Language Models via MetaVQA RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.519489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.519489Z digest=sha256:ce8c967e5a073ab2d28afd70e07280b3b0846f1ccf4b56081e9d154264b7253a

Observation e0b41df9-e0d6-4477-9494-b451fa7e1ab9 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:31.904673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:13dd0eb7c113f6b4469735c92eda5699e73fd44667dfa4db84d2df28ab8eaeb0

Observation 0f9a2674-08bf-4b9d-a39a-624b30438b75 · inbound

GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation cites this paper.

GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:47:36.403062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:47:36.403062Z digest=sha256:cc1ea0092e18c21c7ef4a28e282a774ff80b040023b6c6e111c5574c5e1b2029

Observation 065db545-3594-4e8d-a719-9ef9786643fc · inbound

Universal Actions for Enhanced Embodied Foundation Models cites this paper.

Universal Actions for Enhanced Embodied Foundation Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T19:29:39.705636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:29:39.705636Z digest=sha256:e820bd2507a51c06802573da79900c6c1d336eae5b0f0073c0261e5436722e6d

Observation 5768a73a-6f9f-4afe-a8fe-95db8df26036 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:33.831351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:33.831351Z digest=sha256:51f2d42aefca20bebb9863b81bc086c5251a1269887386c851bc43a7530014a3

Observation e5535967-32e9-4051-8d77-8d402f89e2af · inbound

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models cites this paper.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.607231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.607231Z digest=sha256:bd7d8a11869ccc331224b33e57b9c4f0aa9a7d9515ba76906f8331133e069aec

Observation bf8d3fb4-5f69-4029-a1e2-c59cca2d2d0c · inbound

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model cites this paper.

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:12:19.863742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T06:12:19.643111Z digest=sha256:4cd581e301e3d63456935aa17e1ec168acf45c4b8ec26ad8ce8a5447042866e2