Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T21:44:27.562453Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 92 inbound Pith citation observations for arXiv:2311.01378.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T21:44:27.562453Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:22:28.270135Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-05T11:41:02.629883Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1cbe03ef-32ab-43bd-960c-956b0c11ab89 · outbound
Vision-Language Foundation Models as Effective Robot Imitators Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ae782965-892a-4f8c-8340-0a8814ea5859 · outbound
Vision-Language Foundation Models as Effective Robot Imitators OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b76ed509-3bc4-4db6-8bce-c1d6b97074d4 · outbound
Vision-Language Foundation Models as Effective Robot Imitators GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 21e3ab03-a12d-4689-9be7-9794c12146e1 · outbound
Vision-Language Foundation Models as Effective Robot Imitators RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 16fb9d8a-ec60-4dba-9f3f-d33c6e9f8f4e · outbound
Vision-Language Foundation Models as Effective Robot Imitators Language models are few-shot learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cc775273-3999-4bae-a26f-35f8c8eb9903 · outbound
Vision-Language Foundation Models as Effective Robot Imitators Universal Sentence Encoder
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 013ea9a7-7e86-44fc-b6ae-ba413f602e8e · outbound
Vision-Language Foundation Models as Effective Robot Imitators PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c4b870d8-c583-4d48-959f-e5c60d8f7618 · outbound
Vision-Language Foundation Models as Effective Robot Imitators PaLM: Scaling Language Modeling with Pathways
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0fde5446-ce46-40d3-8080-6026e75ffaf5 · outbound
Vision-Language Foundation Models as Effective Robot Imitators PaLM-E: An Embodied Multimodal Language Model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a2ed258a-4c73-4e87-a83d-d994d9cfca3e · outbound
Vision-Language Foundation Models as Effective Robot Imitators Language-Driven Representation Learning for Robotics
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 149dc577-de54-4a50-a79e-d9dd3eb604de · outbound
Vision-Language Foundation Models as Effective Robot Imitators M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1f06a7a9-cea8-49fb-ad6c-4f29bc989509 · outbound
Vision-Language Foundation Models as Effective Robot Imitators Robotic indoor scene captioning from streaming video
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e5fa4bbc-eb12-4e2d-9344-5bebd5c1a737 · outbound
Vision-Language Foundation Models as Effective Robot Imitators Energy-Based Imitation Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9f768c3-f6c6-4e19-8ced-8946163901b7 · outbound
Vision-Language Foundation Models as Effective Robot Imitators Goal-Conditioned Reinforcement Learning: Problems and Solutions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7323d1d8-b846-4324-822a-759850c23c31 · outbound
Vision-Language Foundation Models as Effective Robot Imitators What matters in language conditioned robotic imitation learning over unstructured data
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 68418dd1-9afa-4db1-a4de-4e00b7713817 · outbound
Vision-Language Foundation Models as Effective Robot Imitators R3M: A Universal Visual Representation for Robot Manipulation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation df1f8824-7038-47c7-af60-9fcdbb989026 · outbound
Vision-Language Foundation Models as Effective Robot Imitators Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 612f1bec-e0ef-481d-b834-23a4ac1ca143 · outbound
Vision-Language Foundation Models as Effective Robot Imitators Instruction Tuning with GPT-4
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f3b40001-ab12-4ff5-b376-dde687899770 · outbound
Vision-Language Foundation Models as Effective Robot Imitators Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f5d8f62-1dab-4419-8d7b-df71be54d9e7 · outbound
Vision-Language Foundation Models as Effective Robot Imitators LLaMA: Open and Efficient Foundation Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7ba2036a-28ce-4fb5-9b14-2f9c48d94575 · outbound
Vision-Language Foundation Models as Effective Robot Imitators EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 78aefb08-78d4-476d-a35f-ae965cd083a0 · outbound
Vision-Language Foundation Models as Effective Robot Imitators X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 78db5617-2162-4a2f-a754-1bfd92eb2af7 · outbound
Vision-Language Foundation Models as Effective Robot Imitators Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 73934fbc-1691-4f82-9cbc-9d7fc1677cc1 · outbound
Vision-Language Foundation Models as Effective Robot Imitators We loaded the pre-train 14 Preprint Table 5: Comparison of co-trained models and fine-tuned models on the CALVIN benchmark
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2861e5a5-e962-4179-9624-2a09f20465a4 · outbound
Vision-Language Foundation Models as Effective Robot Imitators As for V oltron, we also include a version that fine- tunes the representation layers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cfc1417d-b307-4bef-8db1-9221044fe4ca · inbound
Agent AI: Surveying the Horizons of Multimodal Interaction Vision-Language Foundation Models as Effective Robot Imitators
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1050991e-bc6c-437c-8c5a-6284f60016c6 · inbound
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations Vision-Language Foundation Models as Effective Robot Imitators
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b23c365e-ba9c-43f8-a8d5-e0e29b49fb2d · inbound
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation Vision-Language Foundation Models as Effective Robot Imitators
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8f2515bb-e095-434d-b654-653a3a14d2a5 · inbound
A Survey on Vision-Language-Action Models for Embodied AI Vision-Language Foundation Models as Effective Robot Imitators
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8b8992ab-39f2-4bd9-9008-25defdb0c618 · inbound
OpenVLA: An Open-Source Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 826d4c58-8d89-4f86-a768-72cedeec22af · inbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0350bb56-cf30-4593-a6f5-ce0ff128479c · inbound
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7486c90b-204f-4f2b-8b9e-e26d128809bb · inbound
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Vision-Language Foundation Models as Effective Robot Imitators
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ecc731-396f-4840-86c5-745e17fe6df1 · inbound
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons Vision-Language Foundation Models as Effective Robot Imitators
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a307198-f89e-4159-b195-7ec5d128335b · inbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Vision-Language Foundation Models as Effective Robot Imitators
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 97f2b47b-c758-4e7a-8a9f-0696a40473f4 · inbound
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Vision-Language Foundation Models as Effective Robot Imitators
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b620cb81-1c65-4201-ac9c-6e1e5feead9c · inbound
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning Vision-Language Foundation Models as Effective Robot Imitators
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a64c9f9a-58c5-41cf-967c-ebc2d33f5a63 · inbound
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance Vision-Language Foundation Models as Effective Robot Imitators
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a1560bc-e52e-4e33-bd79-dae0622a990d · inbound
OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints Vision-Language Foundation Models as Effective Robot Imitators
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bba8958-450f-46b2-93ec-4f1f13e1f4df · inbound
Visual Language Models as Operator Agents in the Space Domain Vision-Language Foundation Models as Effective Robot Imitators
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 134239e5-fb44-436c-a382-b566955b0ca3 · inbound
RoboReflect: A Robotic Reflective Reasoning Framework for Grasping Ambiguous-Condition Objects Vision-Language Foundation Models as Effective Robot Imitators
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5e07be3-73be-4491-8298-d82fb1aca33c · inbound
GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac41728e-41a9-47c1-bddc-354ccd085d3e · inbound
Improving Vision-Language-Action Model with Online Reinforcement Learning Vision-Language Foundation Models as Effective Robot Imitators
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ed24c6-9829-42df-af65-7ad72d392227 · inbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Vision-Language Foundation Models as Effective Robot Imitators
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5a5dec7-0536-402e-b65b-24890623dc09 · inbound
RoboBERT: An End-to-end Multimodal Robotic Manipulation Model Vision-Language Foundation Models as Effective Robot Imitators
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f51fffb1-90e2-4b88-aa30-54da0b47c14d · inbound
Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dac971e-7165-498f-bab5-2a201487cda7 · inbound
3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Vision-Language Foundation Models as Effective Robot Imitators
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70c4fcec-1f1e-4baa-af95-4b606682321e · inbound
GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffd48dfc-9335-4f6a-9b3b-2c2d7b535ee2 · inbound
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success Vision-Language Foundation Models as Effective Robot Imitators
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 80991980-ebd8-45ff-9f24-8ebd1953dba0 · inbound
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fd9bba19-1692-4008-bfa8-67a6719c60d6 · inbound
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Vision-Language Foundation Models as Effective Robot Imitators
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 00183152-6ce8-4b2f-b379-b4eff41b1f29 · inbound
VLAs are Confined yet Capable of Generalizing to Novel Instructions Vision-Language Foundation Models as Effective Robot Imitators
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c52d2050-4d24-444c-86b3-ef4f2f636ada · inbound
FLARE: Robot Learning with Implicit World Modeling Vision-Language Foundation Models as Effective Robot Imitators
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f249808f-6129-4452-aa95-f196ffe61df3 · inbound
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b16272f-66ac-4e50-a783-e906c8945fd4 · inbound
BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization Vision-Language Foundation Models as Effective Robot Imitators
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a517b294-5f40-42d8-90fe-a0e1c789012c · inbound
Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b8e407a-0420-49b2-a79f-b764eee9cfff · inbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Vision-Language Foundation Models as Effective Robot Imitators
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c774c3c4-96de-4b34-b385-85d1e4b2264a · inbound
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning Vision-Language Foundation Models as Effective Robot Imitators
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a71858-8c5c-4c83-9ddb-8808717a9a11 · inbound
SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ca80984-43d2-4b6f-9dd9-e4f89fb2afb8 · inbound
CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6f21b8b-0c64-4036-950b-5921283d297f · inbound
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc8f8a7-6d31-4b44-b8b3-bbad4f0e1b1b · inbound
GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d93450-37ea-40c6-8ed9-3283ec429749 · inbound
Dynamic Double Space Tower Vision-Language Foundation Models as Effective Robot Imitators
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b4db716-9f2c-4606-9d99-6282f0170e54 · inbound
ROSA: Harnessing Robot States for Vision-Language and Action Alignment Vision-Language Foundation Models as Effective Robot Imitators
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0019b4a4-5b4e-4be1-a351-23cdb4da1fd4 · inbound
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16e31991-7784-4d6d-aad9-0f1592599885 · inbound
WorldVLA: Towards Autoregressive Action World Model Vision-Language Foundation Models as Effective Robot Imitators
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ccb96fcb-72d1-4808-86e8-e17c38e1c6c1 · inbound
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a21d8c-9bb6-4cc0-b235-b45d901d2b1d · inbound
GR-3 Technical Report Vision-Language Foundation Models as Effective Robot Imitators
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bb2bda5c-3247-4b2b-a1a6-d06443aac02e · inbound
Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning Vision-Language Foundation Models as Effective Robot Imitators
Reference 111
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35f40323-ed57-40d4-b829-b4e32bb89201 · inbound
Two-flow Feedback Multi-scale Progressive Generative Adversarial Network Vision-Language Foundation Models as Effective Robot Imitators
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e40e6f-bbe4-4281-b24c-86264bed0b64 · inbound
ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training Vision-Language Foundation Models as Effective Robot Imitators
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dce0a682-6ed5-4854-914e-3d1a0ef91947 · inbound
FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies Vision-Language Foundation Models as Effective Robot Imitators
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5b98ee-3931-467c-b4b6-6d209502f48d · inbound
LLaDA-VLA: Vision Language Diffusion Action Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b022643b-7b10-4474-b2e0-7bb1e4370340 · inbound
TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b2874ea-6c8b-49ce-8efa-7104d79a790e · inbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2672b566-8f58-4b39-a5f8-0b221330dcda · inbound
Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception Vision-Language Foundation Models as Effective Robot Imitators
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7816679a-32e7-4762-ba51-ee9b28fac194 · inbound
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference Vision-Language Foundation Models as Effective Robot Imitators
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b49db4d9-ec83-487b-8134-59b8054e7b2f · inbound
Large Video Planner Enables Generalizable Robot Control Vision-Language Foundation Models as Effective Robot Imitators
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3cd072e4-9009-4ed1-9522-9bc9dcc93d22 · inbound
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109a6f98-431f-40d0-be5d-ec143c757832 · inbound
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2fca8ba7-75e2-492c-90b4-926308054eb2 · inbound
Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Vision-Language Foundation Models as Effective Robot Imitators
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0156e3b7-b707-4466-aee0-8f412012e693 · inbound
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs Vision-Language Foundation Models as Effective Robot Imitators
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3248ea7c-a93b-453a-b155-b54935d2e0af · inbound
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b5d37b1-f4c6-48f6-beab-1444b650778f · inbound
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation acf4a9f8-ca4f-4cc7-a719-92cd552dba1c · inbound
Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks Vision-Language Foundation Models as Effective Robot Imitators
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9976df8-3e13-467b-b346-626101002b82 · inbound
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment Vision-Language Foundation Models as Effective Robot Imitators
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f223ed7b-05f6-4852-b60d-e27e6b6246b4 · inbound
R3D: Revisiting 3D Policy Learning Vision-Language Foundation Models as Effective Robot Imitators
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2457dd32-3bd9-4ea0-b750-6b37fd5912d1 · inbound
Steadily moving semi-infinite fracture in plane poroelasticity Vision-Language Foundation Models as Effective Robot Imitators
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d989c838-42b4-4de4-9e6c-bce3fb67cb03 · inbound
XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments Vision-Language Foundation Models as Effective Robot Imitators
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3cae46c5-e933-4d17-ade1-e02b383aab9c · inbound
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed14748f-ba53-4b04-8253-407f2af5c6fd · inbound
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy Vision-Language Foundation Models as Effective Robot Imitators
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eda8b2af-9ba1-4373-9ccc-ee5dc093d9a6 · inbound
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy Vision-Language Foundation Models as Effective Robot Imitators
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 99a02836-19fc-4e0d-8fdf-06d50d8168b6 · inbound
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy Vision-Language Foundation Models as Effective Robot Imitators
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 63d57f8c-ef74-40da-b75e-15824cdc024f · inbound
Nautilus: From One Prompt to Plug-and-Play Robot Learning Vision-Language Foundation Models as Effective Robot Imitators
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c4c0af91-7269-4c83-b653-1933350d8d1e · inbound
Nautilus: From One Prompt to Plug-and-Play Robot Learning Vision-Language Foundation Models as Effective Robot Imitators
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f162522-895d-4f9f-9e24-27ed585a84e7 · inbound
World Action Models: The Next Frontier in Embodied AI Vision-Language Foundation Models as Effective Robot Imitators
Reference 266
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7c181494-535f-451a-9f0a-6c9783ba3979 · inbound
Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation Vision-Language Foundation Models as Effective Robot Imitators
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 89061257-a158-4bfc-b6c6-86b99145f0db · inbound
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Vision-Language Foundation Models as Effective Robot Imitators
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 514e7fa4-60ff-4164-b8a9-058e32ba8694 · inbound
DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation Vision-Language Foundation Models as Effective Robot Imitators
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 976c9f63-34ef-4bcd-bdcd-c1bb6d6532a9 · inbound
ProgVLA: Progress-Aware Robot Manipulation Skill Learning Vision-Language Foundation Models as Effective Robot Imitators
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 32c9a440-5ccd-4b66-8b71-731a5e88279d · inbound
GEM: Generative Supervision Helps Embodied Intelligence Vision-Language Foundation Models as Effective Robot Imitators
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 23cdc0b3-ffa8-4502-b64e-1d0fdd5b567b · inbound
PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 69a65c01-7b88-4177-8bbb-e1f733ce53ad · inbound
PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41f5cdd1-553a-4629-8013-e5488a2c66e4 · inbound
Wall-OSS-0.5 Technical Report Vision-Language Foundation Models as Effective Robot Imitators
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b3cd827f-3772-4f8b-b639-290db411c575 · inbound
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling Vision-Language Foundation Models as Effective Robot Imitators
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b5bfa74-4664-4ff1-beb6-78e8eea08836 · inbound
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Vision-Language Foundation Models as Effective Robot Imitators
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f594631-394e-41eb-9553-f68c7f431a11 · inbound
Remember what you did?: Learning Behavioral Memories for Partially Observable Object Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b83e0735-4bf9-4e9c-923d-8f30147b996d · inbound
KEMO: Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies Vision-Language Foundation Models as Effective Robot Imitators
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b32f87a6-b19b-4c5c-9c5f-99324bec936d · inbound
VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon Vision-Language Foundation Models as Effective Robot Imitators
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation da9a171c-ec69-422b-b275-2573ef587e77 · inbound
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Vision-Language Foundation Models as Effective Robot Imitators
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 178ff601-9fc6-4c5b-8cfa-38572c73e082 · inbound
EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54fea40b-07c2-4328-bae1-71a7d0e0786e · inbound
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Vision-Language Foundation Models as Effective Robot Imitators
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e991754-2032-4230-9537-a18f423502a0 · inbound
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 207
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5fb9ad4-a1da-4803-aea9-f43f45530565 · inbound
A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference Vision-Language Foundation Models as Effective Robot Imitators
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8da74966-0b46-4aee-a303-caf09a14c868 · inbound
CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d692ef-3573-4fbe-9135-d9fe257aecc2 · inbound
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Vision-Language Foundation Models as Effective Robot Imitators
Reference 142
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901f4227-b934-4ab4-95f0-eac1b048c079 · inbound
Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.