Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:00:46.861455Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 36 inbound Pith citation observations for arXiv:2412.18194.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:00:46.861455Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:01:33.819892Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T06:15:00.866473Z
72 of 72 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 47224534-008f-4427-bd1e-9f223a33f2df · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd34ac40-22af-4f78-9715-7eb736ab3190 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 772b3218-2ec8-40e1-93b8-b2928e469a7c · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks RT-1: Robotics Transformer for Real-World Control at Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddba8c1b-5791-42cf-9a91-2b6cc081af45 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b53a24d2-03b9-4335-b103-b4f6cab9c1eb · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73403817-69a1-4947-bd7b-df39ab032607 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f2d464e9-0e00-4eb9-902b-c872c8c75760 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf6de0df-a01d-4f6f-943a-63ca9bd089e5 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Diffusion policy: Visuomotor policy learning via action dif- fusion
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cb116d2a-e860-4665-8733-544f8661ffec · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Training Verifiers to Solve Math Word Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51663c34-eceb-461b-94ee-90b4919c4b08 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Opencompass: A universal evaluation platform for foundation models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 59ef7340-ef16-48a7-967c-a109562377f1 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Objaverse: A universe of annotated 3d objects
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 65650ae0-68eb-4b4c-8f83-dc1cafe6f6a5 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 363bf8cf-c257-494c-b525-0ac0eda573ab · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks PaLM-E: An Embodied Multimodal Language Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81637d6e-375c-4332-8858-198ee1abbbf2 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13570e40-50a2-4043-b2ff-3ff3e83b7333 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Graspnet-1billion: A large-scale benchmark for general ob- ject grasping
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 65990cba-eea7-44fb-9227-5b80a78d2da1 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Robust grasping across diverse sensor qualities: The graspnet-1billion dataset
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cb0b3943-ed5a-438c-8adc-38cf5635768a · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18e65911-0a32-4784-b332-1af383b54c1a · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Arnold: A benchmark for language-grounded task learning with con- tinuous states in realistic 3d scenes
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b199088e-947c-4b27-b0e4-b531aca79050 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Maniskill2: A unified benchmark for generalizable manipulation skills
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3b9f3ad6-d36b-415d-acc8-87dbe9233181 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 354ab407-a8e2-4678-876c-57ae76989fcb · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efaab614-9124-41a8-9311-62274aebe441 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks CoPa: General Robotic Manipulation through Spatial Constraints of Parts with Foundation Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b571f8c-2af3-4261-942b-3f1a187304c1 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d32fa0-b2f8-40a6-97a3-3d292d95bb83 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Rlbench: The robot learning benchmark & learning environment
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 58e85c39-8861-4cef-b840-513f9695ea16 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Sampling-based algo- rithms for optimal motion planning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 06925268-58b7-46a8-87d3-5dd03fd4784d · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Anytime motion planning using the rrt
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 70ca534b-5685-46a6-a734-9e31da0be195 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks OpenVLA: An Open-Source Vision-Language-Action Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8447ef19-f16c-4134-bb7a-f455cf5a956c · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Segment any- thing
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f0847ff-31ad-4255-a792-6c3c6c46cd06 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Behavior-1k: A benchmark for embodied ai with 1,000 ev- eryday activities and realistic simulation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 29b8c085-48fb-458a-b72c-4031b7a46324 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3585a43-410f-441e-b75c-bf60465da25c · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e5cae09-34e7-4c7e-ab31-af6fd2849741 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Evaluating Real-World Robot Manipulation Policies in Simulation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1258bb16-501a-4d33-b7f2-a4ca00792633 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Code as policies: Language model programs for embodied control
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83c8cc33-ca3c-4a93-a98b-4ff0f0e8015b · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Data Scaling Laws in Imitation Learning for Robotic Manipulation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a50d88f-e442-4535-980d-eff6d518f483 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Libero: Benchmarking knowl- edge transfer for lifelong robot learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c68f9783-1fba-4123-bfec-8f3fa25461c7 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Moka: Open-vocabulary robotic manipulation through mark-based visual prompting
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee19f1a-ab1b-49da-b27f-7ce2d96b8967 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Visual instruction tuning, 2023
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5118c81-408c-4eca-8b97-e85f7d822354 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b83b6f-7da3-4c33-af54-ae081f624c49 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Visual instruction tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bbfb59a-abd2-4fd3-9b19-deda30e6b289 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 344bc8ac-e910-4439-9106-a3484cb9b891 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d9181730-91d3-4001-9edb-38328b5d888d · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0479a13d-90c6-450b-ae5e-577be2ca03c2 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Robocasa: Large-scale simulation of every- day tasks for generalist robots
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0f391d6a-6be4-47c2-b7f1-ee6c54998ea8 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks A survey on domain-specific languages in robotics
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 91b7d84b-6db7-4c21-ab4e-0250096fe101 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c89441b9-e34f-471c-a313-e99ddf01f7a7 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Imitating Human Behaviour with Diffusion Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fe79577-1ad2-4723-90d4-da8068785d13 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Pre-trained models for natural language processing: A survey
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 16692531-6535-4aab-9f05-013f4850cfb8 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 253a08cd-c33d-4672-a28b-abb7d94b614b · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Moss: An open conversational large language model
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d7720e3-a74a-4eef-8d82-b2589f701cec · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Robolang: a sim- ple domain specific language to script robot interactions
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fa3da2a8-2901-43b4-94ec-205545356345 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Habitat 2.0: Training home assistants to rearrange their habitat
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ddf1e40e-26b9-4664-a412-0be7e6a8cec8 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bdf7a57-6332-4348-ad6d-d858f18fd838 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Octo: An Open-Source Generalist Robot Policy
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2f58db1-d105-4dc8-a18e-8f5e796efd46 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Mujoco: A physics engine for model-based control
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1920ee13-d62f-4957-8820-5822f6720ca3 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks dm control: Software and tasks for continuous control
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e5da3e45-261e-4ebd-891c-e7e0a8ece3be · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Bridgedata v2: A dataset for robot learning at scale
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 91e01401-ebd0-4b26-9845-bd7e8b334096 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Chi, Tat- sunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 578157f5-bcbf-4c3d-aa6b-9bea1da1fc2a · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b7f1037-5fb4-443c-ac2d-7051a902aa84 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks HomeRobot: Open-Vocabulary Mobile Manipulation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee6cf526-502f-4589-8acf-6806d6a5e2f2 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks obj2mjcf: Cli for processing composite wave- front obj files for use in mujoco
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4a3fc966-ab64-4ba3-aa4c-6b26063527eb · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Clip2: Contrastive language- image-point pretraining from real-world point cloud data
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aac0719-92fb-4c4c-afae-62fc72185af7 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks RoboCAS: A Benchmark for Robotic Manipulation in Complex Object Arrangement Scenarios
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb6f1346-3c36-4927-8711-8f40243e840e · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Task Descriptions All Tasks
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ad4088fc-db38-4f0a-bd34-ebad1c6a57cd · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Hammer nail, 6) Press button, 7) Insert,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d5559c5-a4fa-4f8b-93de-4381ca00f975 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks placing an ap- ple on a plate
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 46a1de3f-cdb1-4242-b320-9885c371bd4e · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3e9e9eb1-4a68-411d-9cf1-0de93b0d7473 · outbound
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 323bc9c9-c8be-42dd-8806-5e18280f8c3d · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks VLA Setting To assess the generalization ability of various VLAs, we primarily fine-tune OpenVLA, Octo, and RDT-1B using our dataset
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 567905c6-188e-449d-80a6-6c2789567296 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ce7a7a4c-b0b5-4b76-a28a-f7a0f15cd34c · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Once the relevant information has been collected, the success of the task and the accuracy of target identification are assessed
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 47696a2e-728b-4574-8358-38193af4dd01 · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks After evaluating individual tasks several times, the scores for each time task are ag- gregated to yield the final score for the model under each configuration
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 943084df-e1fa-4219-babe-447db74d26bb · outbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks put the strawberry into the basket
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eec0d6ab-a241-4887-885e-500239218926 · inbound
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ce3870ec-e29e-4367-b05b-e6080f75c688 · inbound
RobotxR1: Enabling Embodied Robotic Intelligence on Large Language Models through Closed-Loop Reinforcement Learning VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b89ef46c-9358-43f9-9fb1-9c91ead27897 · inbound
Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23195fa-c2df-4fea-8328-7b4a0421e611 · inbound
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04cad5a2-d5bb-4cd1-aa82-961911dc6ef4 · inbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e21363c4-de58-43e7-addd-118d46879e4d · inbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ad70840-7b9f-43fa-b15c-2e79d06dc35d · inbound
RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1deb9fcd-83d5-4411-a38a-1555e06bbfb3 · inbound
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 276
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f6521287-567f-486e-a43c-569efabdd807 · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2d85b8e6-4ffb-4fe8-823f-1a8e01e82eb8 · inbound
4D Visual Pre-training for Robot Learning VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc7d9943-095a-469d-b367-3aaa6d1ef8b4 · inbound
Galaxea Open-World Dataset and G0 Dual-System VLA Model VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e85c76da-2b42-4b5d-be22-58099d2e4007 · inbound
RoboBenchMart: Benchmarking Robots in Retail Environment VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 961ea27c-946d-4443-9f19-e2a797e9bb96 · inbound
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd4cb20-7d0a-4d33-833a-a0527cd597af · inbound
Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f09b8d10-664f-4342-8997-3abb78eb47be · inbound
vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c7685e57-c745-4ce7-bce3-845f5c60ca7b · inbound
Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 97a3e32d-23e3-47a3-9110-8386a18aba9e · inbound
VISOR: A Vision-Language Model-based Test Oracle for Testing Robots VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c39407f3-b62f-4472-ad39-ef5a1164ba67 · inbound
VISOR: A Vision-Language Model-based Test Oracle for Testing Robots VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9816f0e6-a24b-4a4b-a4df-317a72ce36f7 · inbound
World Action Models: The Next Frontier in Embodied AI VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 240
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 88885839-87d0-4bda-81c5-fa7e8efecb9c · inbound
DexHoldem: Playing Texas Hold'em with Dexterous Embodied System VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ed0293e3-433d-4eaf-91e9-64e563944e18 · inbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 92f94aae-21c1-4a8e-b7aa-8531b06582f0 · inbound
RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 713e4e26-5890-49bd-aad6-ce5ce4ff1326 · inbound
Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dab1c090-3dd6-499e-bb53-5f2e630a387a · inbound
VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 26ed2b1d-1f73-49cf-9e10-c90681737d50 · inbound
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation da963171-54b4-4abf-8d30-2b04a5f3235c · inbound
A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d2911b14-c2e2-4ee9-85ee-99d4b08fa3d1 · inbound
UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a59bb101-0cca-4d05-b79d-351c3ec5b00f · inbound
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e772ff54-ed4f-4dbd-83ce-6f4a1296f6cb · inbound
World Action Models: A Survey VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 197
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 95799076-95f7-4d13-8d90-ac53db763ee6 · inbound
Bridge-WA: Predicting Where and How the World Changes for Robotic Action VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f9ddcd77-ad4c-4ab3-b7de-2247332469e1 · inbound
ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 43b3b4f2-9de9-4796-8040-d5b0b3f21b88 · inbound
ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 481223e8-41e8-4ee7-9ff0-c15a83959a0f · inbound
Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb710edc-4f47-4d25-8e93-ae7a0192a90c · inbound
$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d2ef7b-364e-4d66-b04f-60f0462db2ae · inbound
GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd8f21c-ebc7-4c37-8d1b-6f4fb0fc269b · inbound
Self-Evolving Embodied Agents via Skill-Harness Evolution VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.