Pith. sign in

Paper Citation Record · LEDGER

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

As of 24 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 36 inbound Pith citation observations for arXiv:2412.18194.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18194 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:00:46.861455Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:01:33.819892Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved38
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47224534-008f-4427-bd1e-9f223a33f2df · outbound

This paper cites GPT-4 Technical Report.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.605686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.605686Z digest=sha256:ec14f831207f2b03a9dd2ddc9f8152b83f0198369312658a62581520adabd646

Observation bd34ac40-22af-4f78-9715-7eb736ab3190 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.610999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.610999Z digest=sha256:61ed40dabc8d2c4bc7c4355ffab3da1d759ff9f9aaf1764ecf99853d9c5b5340

Observation 772b3218-2ec8-40e1-93b8-b2928e469a7c · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.616427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.616427Z digest=sha256:38e6a6b4ae65bce5ae4cbd15bcf9b5778c2c55177e89156e19cec4d16fa3df5c

Observation ddba8c1b-5791-42cf-9a91-2b6cc081af45 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.620506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.620506Z digest=sha256:a38c115c2c76c94c4c3fc1ef1d8edea884a701a265d313d7427d750b63024d6c

Observation b53a24d2-03b9-4335-b103-b4f6cab9c1eb · outbound

This paper cites Language Models are Few-Shot Learners.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.624916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.624916Z digest=sha256:7d3053393c9ecd5bf1f60c9bf3e7e16d06a89e9eac3f061f6689c4ac65c9009a

Observation 73403817-69a1-4947-bd7b-df39ab032607 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.701307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.628756Z digest=sha256:97186297c4329f8253e4cfd3a166195b20ef6285f2e26e73e176c9434ce8d440

Observation f2d464e9-0e00-4eb9-902b-c872c8c75760 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.632220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.632220Z digest=sha256:088fc5a88cc75318756abb16281c669977d88b6ee67276b354cc8e672f76740e

Observation cf6de0df-a01d-4f6f-943a-63ca9bd089e5 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action dif- fusion.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Diffusion policy: Visuomotor policy learning via action dif- fusion

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.688889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.636182Z digest=sha256:50faf00d4eb957f6dae107e34c38559fbfd2ae6e1363f94ed2f1c097d136bbc6

Observation cb116d2a-e860-4665-8733-544f8661ffec · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.639509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.639509Z digest=sha256:f6035035d45f06ab54c0835f6e95edac5141de569a0fb0bb2a0bb284a3899865

Observation 51663c34-eceb-461b-94ee-90b4919c4b08 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Opencompass: A universal evaluation platform for foundation models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.673876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.643066Z digest=sha256:9db247387798c9aaf45dce4db8b3ee297d07ec9431c4092ae0ff72b2d238f85b

Observation 59ef7340-ef16-48a7-967c-a109562377f1 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Objaverse: A universe of annotated 3d objects

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.661742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.646783Z digest=sha256:08f8432ee167ef49e168497a0445f934cba36130c8857fed65495e01c2a0d1f7

Observation 65650ae0-68eb-4b4c-8f83-dc1cafe6f6a5 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.650056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.650056Z digest=sha256:196a830520c472c733adaacd276bb8bcefef7db26254e8096cd534f4133787d1

Observation 363bf8cf-c257-494c-b525-0ac0eda573ab · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks PaLM-E: An Embodied Multimodal Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.653222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.653222Z digest=sha256:027059ec696860db507f9ce92fe35636db5b0a1292753d76f3fd3d8deecd98a7

Observation 81637d6e-375c-4332-8858-198ee1abbbf2 · outbound

This paper cites The Llama 3 Herd of Models.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.656759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.656759Z digest=sha256:dbaaf98e581bb57be48f2292e39d3f53f7205e3fab98a06b8bbbb573ca38c1b3

Observation 13570e40-50a2-4043-b2ff-3ff3e83b7333 · outbound

This paper cites Graspnet-1billion: A large-scale benchmark for general ob- ject grasping.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Graspnet-1billion: A large-scale benchmark for general ob- ject grasping

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.649750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.660116Z digest=sha256:037858dc514db65c1da28067d3e4cd11bfaaba15665f151c42dc26b5ee9b0288

Observation 65990cba-eea7-44fb-9227-5b80a78d2da1 · outbound

This paper cites Robust grasping across diverse sensor qualities: The graspnet-1billion dataset.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Robust grasping across diverse sensor qualities: The graspnet-1billion dataset

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.637082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.663276Z digest=sha256:4d7e534e992f6878f2ee80cb910a348c98ff206f4afcbc0c2e5c238820c27bea

Observation cb0b3943-ed5a-438c-8adc-38cf5635768a · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.666514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.666514Z digest=sha256:3825a6b4ef7be302bf63d0890a96f2de872410e5c939a118465c1d07c015abed

Observation 18e65911-0a32-4784-b332-1af383b54c1a · outbound

This paper cites Arnold: A benchmark for language-grounded task learning with con- tinuous states in realistic 3d scenes.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Arnold: A benchmark for language-grounded task learning with con- tinuous states in realistic 3d scenes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.624955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.669817Z digest=sha256:645686539824d67b9b3310d783c7c6e0dfa54757ade82cb6668f3088518a47d9

Observation b199088e-947c-4b27-b0e4-b531aca79050 · outbound

This paper cites Maniskill2: A unified benchmark for generalizable manipulation skills.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Maniskill2: A unified benchmark for generalizable manipulation skills

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.612570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.673051Z digest=sha256:7d5f46646acf7b27a1bc260352adf009d983730affbfd7243479b7ec812fb8a1

Observation 3b9f3ad6-d36b-415d-acc8-87dbe9233181 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.675975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.675975Z digest=sha256:56db833017978822b30a61091e9f7e6fcc1d5f3b2ab3ec83d630cbbffb1e3bd5

Observation 354ab407-a8e2-4678-876c-57ae76989fcb · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.679083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.679083Z digest=sha256:77d851ecd9ea94d67fff3add93f25e6760aff842c2f9c899669ce427481d8aac

Observation efaab614-9124-41a8-9311-62274aebe441 · outbound

This paper cites CoPa: General Robotic Manipulation through Spatial Constraints of Parts with Foundation Models.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks CoPa: General Robotic Manipulation through Spatial Constraints of Parts with Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.682449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.682449Z digest=sha256:68ffaff4c3efa4cde07a44452b92b4f78af50903f1cad17dfb7877425ed42007

Observation 8b571f8c-2af3-4261-942b-3f1a187304c1 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.686565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.686565Z digest=sha256:372d129b9cd9b19206faff5865b7e7872e9015b03a44b72b9cec31957c3d8195

Observation 35d32fa0-b2f8-40a6-97a3-3d292d95bb83 · outbound

This paper cites Rlbench: The robot learning benchmark & learning environment.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Rlbench: The robot learning benchmark & learning environment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.600606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.690769Z digest=sha256:1bf85e810bc981c01d8d696d5d53cf9aa83d9cd89b033d6cf5d80ba3d22e605d

Observation 58e85c39-8861-4cef-b840-513f9695ea16 · outbound

This paper cites Sampling-based algo- rithms for optimal motion planning.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Sampling-based algo- rithms for optimal motion planning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.589682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.694846Z digest=sha256:12e5bd7dab09af064901da6db687cf65800abd88ffef88cfa41985f7fa564fec

Observation 06925268-58b7-46a8-87d3-5dd03fd4784d · outbound

This paper cites Anytime motion planning using the rrt.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Anytime motion planning using the rrt

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.578414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.698544Z digest=sha256:90404c87faa89f1f8d68bad6010ad9b16d97d6ef2b49e08d864c9d3d04f89aa0

Observation 70ca534b-5685-46a6-a734-9e31da0be195 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks OpenVLA: An Open-Source Vision-Language-Action Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.702618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.702618Z digest=sha256:3ed776dfe3505404e9561aaf52091d657803c2da8d2ed0c6b1c4df9865a8945e

Observation 8447ef19-f16c-4134-bb7a-f455cf5a956c · outbound

This paper cites Segment any- thing.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Segment any- thing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.706604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.706604Z digest=sha256:fe13d4fcd9a62f55d7f58feca711503c9a19eb0040c94f7054a3687fa777a450

Observation 9f0847ff-31ad-4255-a792-6c3c6c46cd06 · outbound

This paper cites Behavior-1k: A benchmark for embodied ai with 1,000 ev- eryday activities and realistic simulation.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Behavior-1k: A benchmark for embodied ai with 1,000 ev- eryday activities and realistic simulation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.560815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.710322Z digest=sha256:6d072c7d435b806b55e001bc96c7bcb59261903a42d908b971d9298a35637b86

Observation 29b8c085-48fb-458a-b72c-4031b7a46324 · outbound

This paper cites Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.714101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.714101Z digest=sha256:f6f6ed11c263090865153249a683e97317d86354b0cf311a51db024825d129c8

Observation e3585a43-410f-441e-b75c-bf60465da25c · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.718287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.718287Z digest=sha256:0812baa85bd790f5fe0918d419b43028ee2f505eb8ff30b709973d1b61b87ccd

Observation 0e5cae09-34e7-4c7e-ab31-af6fd2849741 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.721956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.721956Z digest=sha256:b14f23e49cc5218f017d0e2a7821484591c1b6e9f3e48141819b55b208f59390

Observation 1258bb16-501a-4d33-b7f2-a4ca00792633 · outbound

This paper cites Code as policies: Language model programs for embodied control.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Code as policies: Language model programs for embodied control

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.725120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.725120Z digest=sha256:4020a2b4c6945db5dc9844e58593cbaf9c3eaf13ca294c79d97e8c1292912f73

Observation 83c8cc33-ca3c-4a93-a98b-4ff0f0e8015b · outbound

This paper cites Data Scaling Laws in Imitation Learning for Robotic Manipulation.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Data Scaling Laws in Imitation Learning for Robotic Manipulation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.727996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.727996Z digest=sha256:cbd7e6f89cba011bebda14c60d9bab42bab48e9d9973e261888b14623766bdac

Observation 1a50d88f-e442-4535-980d-eff6d518f483 · outbound

This paper cites Libero: Benchmarking knowl- edge transfer for lifelong robot learning.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Libero: Benchmarking knowl- edge transfer for lifelong robot learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.542248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.731453Z digest=sha256:e75effc7ba17e9439c750eec4ddf58c33a546bdea5268126677d15d29a3e8ceb

Observation c68f9783-1fba-4123-bfec-8f3fa25461c7 · outbound

This paper cites Moka: Open-vocabulary robotic manipulation through mark-based visual prompting.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Moka: Open-vocabulary robotic manipulation through mark-based visual prompting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.734513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.734513Z digest=sha256:2533e58f8c85fe58eb52d3b2b19ac4ff039767eceab899cbb112d7cbb83d039d

Observation 0ee19f1a-ab1b-49da-b27f-7ce2d96b8967 · outbound

This paper cites Visual instruction tuning, 2023.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Visual instruction tuning, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.737815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.737815Z digest=sha256:ec9f56ff33074c4c9900369e9e29b72b097e2f668c9c4ecb004dadfa75dfbd8d

Observation a5118c81-408c-4eca-8b97-e85f7d822354 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.741209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.741209Z digest=sha256:9aa951179f376203d8bc7035c5eabab67fd7e4d5fe65bbbae92f486c0fb435a0

Observation 86b83b6f-7da3-4c33-af54-ae081f624c49 · outbound

This paper cites Visual instruction tuning.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Visual instruction tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.744756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.744756Z digest=sha256:5619c314b5912a65cc6877e0920b24e22331fb53d4b86344c029f5ee62b5c28e

Observation 6bbfb59a-abd2-4fd3-9b19-deda30e6b289 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.748166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.748166Z digest=sha256:f093b8cb9971b1274efdc23101671fca77f10f9c369e63d9ab6e06eba4eb5524

Observation 344bc8ac-e910-4439-9106-a3484cb9b891 · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.503250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.752024Z digest=sha256:7e3b50f35e9ffaf65bb6a80d9cebadf987b46844bb286a789e6727b2e9fe49d0

Observation d9181730-91d3-4001-9edb-38328b5d888d · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.755419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.755419Z digest=sha256:8fa90b15afe247add8c02017180e1f1db74328ff266543ae5e5a23be53369c94

Observation 0479a13d-90c6-450b-ae5e-577be2ca03c2 · outbound

This paper cites Robocasa: Large-scale simulation of every- day tasks for generalist robots.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Robocasa: Large-scale simulation of every- day tasks for generalist robots

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.492960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.759340Z digest=sha256:aae99c141f44cfa436f8e5501464bdb92a41156dcff2860b4558d0a740682512

Observation 0f391d6a-6be4-47c2-b7f1-ee6c54998ea8 · outbound

This paper cites A survey on domain-specific languages in robotics.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks A survey on domain-specific languages in robotics

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.482583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.762718Z digest=sha256:087892b86c377eb935f34321e17407f64be57ed4117c70062f31d387bed50ff9

Observation 91b7d84b-6db7-4c21-ab4e-0250096fe101 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.769815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.769815Z digest=sha256:4652aecc37c3b610791007dcb4f17f08ed92173c64eb0097a6f434f4b4507560

Observation c89441b9-e34f-471c-a313-e99ddf01f7a7 · outbound

This paper cites Imitating Human Behaviour with Diffusion Models.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Imitating Human Behaviour with Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.773345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.773345Z digest=sha256:2a0b176d924c3b9537f62c9b7fc62a11d6d0762930bdb29697e4fcb18bd1b9fe

Observation 0fe79577-1ad2-4723-90d4-da8068785d13 · outbound

This paper cites Pre-trained models for natural language processing: A survey.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Pre-trained models for natural language processing: A survey

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.471289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.776967Z digest=sha256:c7e07232312e2bdf11938dd71c55be0f92772769b4726e37ef11f1764312b614

Observation 16692531-6535-4aab-9f05-013f4850cfb8 · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.458956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.780227Z digest=sha256:868958bc95815f84896a8e32c228bc83ce071ffd3c9c1fcc69fe778a486ef26e

Observation 253a08cd-c33d-4672-a28b-abb7d94b614b · outbound

This paper cites Moss: An open conversational large language model.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Moss: An open conversational large language model

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.447051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.783358Z digest=sha256:4d70ad90c76f117158bf765c11c138c76df260a243ecb8cb0ed1db535e3f3f60

Observation 3d7720e3-a74a-4eef-8d82-b2589f701cec · outbound

This paper cites Robolang: a sim- ple domain specific language to script robot interactions.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Robolang: a sim- ple domain specific language to script robot interactions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.435564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.786771Z digest=sha256:6273e03aa413dd93f95191f5fbbbedc24b06517befc9c2710a3d9895932334a5

Observation fa3da2a8-2901-43b4-94ec-205545356345 · outbound

This paper cites Habitat 2.0: Training home assistants to rearrange their habitat.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Habitat 2.0: Training home assistants to rearrange their habitat

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.424746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.789984Z digest=sha256:36daf3584d1d21488c016df37f794ccdc4681119ba7a7bdcef4f1a98f86edff7

Observation ddf1e40e-26b9-4664-a412-0be7e6a8cec8 · outbound

This paper cites ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.792973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.792973Z digest=sha256:a1ddaf2dec0ba276c33b19d67e64686a70f81a550affcfbf21b99c8fddf2af14

Observation 8bdf7a57-6332-4348-ad6d-d858f18fd838 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Octo: An Open-Source Generalist Robot Policy

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.797028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.797028Z digest=sha256:383215518554c1dac05d603d0d0a9f8170bbde746462d12ebca5bc694eca8f7f

Observation f2f58db1-d105-4dc8-a18e-8f5e796efd46 · outbound

This paper cites Mujoco: A physics engine for model-based control.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Mujoco: A physics engine for model-based control

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.414091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.801220Z digest=sha256:ba9adc34b44e39ba77727ec13051c391d1c29ed809479830c9e5cb80a224ad58

Observation 1920ee13-d62f-4957-8820-5822f6720ca3 · outbound

This paper cites dm control: Software and tasks for continuous control.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks dm control: Software and tasks for continuous control

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.403970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.804992Z digest=sha256:03d14089dd09bc9e66ee921e7e009ecb249de4289df19182fc306ce8695968ec

Observation e5da3e45-261e-4ebd-891c-e7e0a8ece3be · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Bridgedata v2: A dataset for robot learning at scale

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.392728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.808812Z digest=sha256:5fb2b3916a1df9cbb28aa2839e104fe24dc4cf1dfdf374a488e5b9d022567e59

Observation 91e01401-ebd0-4b26-9845-bd7e8b334096 · outbound

This paper cites Chi, Tat- sunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Chi, Tat- sunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.382148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.812648Z digest=sha256:2b3e98389f7f84bec352425831912538bc047d195c557e6fc059f05f73ac4611

Observation 578157f5-bcbf-4c3d-aa6b-9bea1da1fc2a · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.816625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.816625Z digest=sha256:76efa31293339f4226bf5f021e5cdb0c6330b3cb0d99419f1ece601f22e4ac18

Observation 3b7f1037-5fb4-443c-ac2d-7051a902aa84 · outbound

This paper cites HomeRobot: Open-Vocabulary Mobile Manipulation.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks HomeRobot: Open-Vocabulary Mobile Manipulation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.820650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.820650Z digest=sha256:d2121fa340c2e337c94e0190053199b78b82112d8af6156587a99e8c9e8a1c3b

Observation ee6cf526-502f-4589-8acf-6806d6a5e2f2 · outbound

This paper cites obj2mjcf: Cli for processing composite wave- front obj files for use in mujoco.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks obj2mjcf: Cli for processing composite wave- front obj files for use in mujoco

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.369445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.824368Z digest=sha256:27e7be7077d8d09f28fe8ccd7ee3aafaf0ac32dd2ffff0d9d44b868259a3f074

Observation 4a3fc966-ab64-4ba3-aa4c-6b26063527eb · outbound

This paper cites Clip2: Contrastive language- image-point pretraining from real-world point cloud data.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Clip2: Contrastive language- image-point pretraining from real-world point cloud data

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.827976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.827976Z digest=sha256:ce583187083fbc8e38588e817e067116302e08c806d070860af2caea53f5f4e0

Observation 5aac0719-92fb-4c4c-afae-62fc72185af7 · outbound

This paper cites RoboCAS: A Benchmark for Robotic Manipulation in Complex Object Arrangement Scenarios.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks RoboCAS: A Benchmark for Robotic Manipulation in Complex Object Arrangement Scenarios

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.830929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.830929Z digest=sha256:1966ebfd4499c61157ad94279f06698cb3229ed0d42dab23ff5c0ce71d34ed00

Observation cb6f1346-3c36-4927-8711-8f40243e840e · outbound

This paper cites Task Descriptions All Tasks.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Task Descriptions All Tasks

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.349385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.834264Z digest=sha256:357b3f02d04ed8a216c47365a0bb70d5817d3160b4652f354ea84566136d50e2

Observation ad4088fc-db38-4f0a-bd34-ebad1c6a57cd · outbound

This paper cites Hammer nail, 6) Press button, 7) Insert,.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Hammer nail, 6) Press button, 7) Insert,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.338238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.837382Z digest=sha256:b7c7193378408bb097f32f2aefcad657a6647db87e4775902c3e80b89b840f94

Observation 3d5559c5-a4fa-4f8b-93de-4381ca00f975 · outbound

This paper cites placing an ap- ple on a plate.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks placing an ap- ple on a plate

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.327717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.840201Z digest=sha256:9650a499f89c1b2819af221870350b86c0a0bdc1b00ed3c2c5d8ebbca71c1af7

Observation 46a1de3f-cdb1-4242-b320-9885c371bd4e · outbound

This paper cites an unresolved cited work.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:00:47.316542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.843417Z digest=sha256:cb1f83ead3f7d0f697b4b0662f913587279e68497ddb072ccffba462e9bda3b2

Observation 3e9e9eb1-4a68-411d-9cf1-0de93b0d7473 · outbound

This paper cites Apple", {.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Apple", {

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.305271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.846050Z digest=sha256:218e19526d4bb4068123aabd2d7aa84f29e6b0106eedcaf1ff2413a712e53484

Observation 323bc9c9-c8be-42dd-8806-5e18280f8c3d · outbound

This paper cites VLA Setting To assess the generalization ability of various VLAs, we primarily fine-tune OpenVLA, Octo, and RDT-1B using our dataset.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks VLA Setting To assess the generalization ability of various VLAs, we primarily fine-tune OpenVLA, Octo, and RDT-1B using our dataset

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.293449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.849239Z digest=sha256:4ae6773d77f7a1b0651c46632a3624bbc5bb3ff5e51062ad6af1a4428003fb2f

Observation 567905c6-188e-449d-80a6-6c2789567296 · outbound

This paper cites an unresolved cited work.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:00:47.279718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.852357Z digest=sha256:cdc8f1d49bf55ac3682521fe859e45ed4736eba1befae24a15408e33ac90e4ff

Observation ce7a7a4c-b0b5-4b76-a28a-f7a0f15cd34c · outbound

This paper cites Once the relevant information has been collected, the success of the task and the accuracy of target identification are assessed.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks Once the relevant information has been collected, the success of the task and the accuracy of target identification are assessed

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:00:47.266436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.855549Z digest=sha256:7e47e60d5a7e9897020e4238da07570d52ccbc23bd7b75ad9d3338435254f5c4

Observation 47696a2e-728b-4574-8358-38193af4dd01 · outbound

This paper cites After evaluating individual tasks several times, the scores for each time task are ag- gregated to yield the final score for the model under each configuration.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks After evaluating individual tasks several times, the scores for each time task are ag- gregated to yield the final score for the model under each configuration

Reference 72

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T05:00:47.252671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.858521Z digest=sha256:e634c9753ba02b9bac1e968a7f02bfcb4604db4598cde461dc6749139f21af4c

Observation 943084df-e1fa-4219-babe-447db74d26bb · outbound

This paper cites put the strawberry into the basket.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks put the strawberry into the basket

Reference 73

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T05:00:47.238391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:00:46.861455Z digest=sha256:5b690e32efadb7750506383e9b7edac4141bbecae132133aa67e886c761f33a5

Pith citing papers

Observation eec0d6ab-a241-4887-885e-500239218926 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:32:22.473964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:db5965b3b5db8f314d4d2e193cd4f71e118080bd2d7ab17f000e93496f11ebcf

Observation ce3870ec-e29e-4367-b05b-e6080f75c688 · inbound

RobotxR1: Enabling Embodied Robotic Intelligence on Large Language Models through Closed-Loop Reinforcement Learning cites this paper.

RobotxR1: Enabling Embodied Robotic Intelligence on Large Language Models through Closed-Loop Reinforcement Learning VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:01:33.819892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:01:33.819892Z digest=sha256:1b97d3bb66462ce7003c8e6d7e84c4350724604e37b08ce6d57622824556f980

Observation b89ef46c-9358-43f9-9fb1-9c91ead27897 · inbound

Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments cites this paper.

Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:10:03.406175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:10:03.406175Z digest=sha256:81d92f0e778884504044cfdcac74bee7ffb023af83d6f2b75c80eecdc222f477

Observation b23195fa-c2df-4fea-8328-7b4a0421e611 · inbound

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation cites this paper.

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:50.256841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:50.256841Z digest=sha256:52865cc8f46e600eefacfe8eb3044e3825236ef3197200fc17685722a6936e1a

Observation 04cad5a2-d5bb-4cd1-aa82-961911dc6ef4 · inbound

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks cites this paper.

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:53.465809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:53.465809Z digest=sha256:177ca554549fcfa50d8df7c1526a8cd05d55ecdac118b787529411637612431b

Observation e21363c4-de58-43e7-addd-118d46879e4d · inbound

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models cites this paper.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.627121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.627121Z digest=sha256:68a8df74636e3302f83fd46cb341835f8fd023f38ef8abaf0e66498473b3abe2

Observation 5ad70840-7b9f-43fa-b15c-2e79d06dc35d · inbound

RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation cites this paper.

RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:15.602361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:15.602361Z digest=sha256:4d9a9bc200806cc8073b481b7c6e3bb31a1d02de1e2d85902d6f1907d3f01b1b

Observation 1deb9fcd-83d5-4411-a38a-1555e06bbfb3 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:08:35.362092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:5a2a240fcae5a836a314209643c0a509206cd19d0c30593780ddf7bdef12ead0

Observation f6521287-567f-486e-a43c-569efabdd807 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 119

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.824554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:6d6c6b6fd1f2072d83e18d73296858fe54c9fe8fef18db935cbfa4733c73368c

Observation 2d85b8e6-4ffb-4fe8-823f-1a8e01e82eb8 · inbound

4D Visual Pre-training for Robot Learning cites this paper.

4D Visual Pre-training for Robot Learning VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T17:03:21.446755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:03:21.446755Z digest=sha256:002a58da495c15966fbd780f137e7c43d56ad47ce1e7b841ba6eb9c30a21357c

Observation dc7d9943-095a-469d-b367-3aaa6d1ef8b4 · inbound

Galaxea Open-World Dataset and G0 Dual-System VLA Model cites this paper.

Galaxea Open-World Dataset and G0 Dual-System VLA Model VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T13:31:09.974877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:31:09.974877Z digest=sha256:018e5c2066611ab66a8433838f0564715b5283a8bb7faa61f529a9c5eed18f90

Observation e85c76da-2b42-4b5d-be22-58099d2e4007 · inbound

RoboBenchMart: Benchmarking Robots in Retail Environment cites this paper.

RoboBenchMart: Benchmarking Robots in Retail Environment VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T22:29:59.056501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:29:59.056501Z digest=sha256:d26cd1e3a149cb0b8b3e7d8b12e71fe97c03794d412866aa790d86ddbb7acd4e

Observation 961ea27c-946d-4443-9f19-e2a797e9bb96 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:27.342980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:27.342980Z digest=sha256:de8a38ceac31bb3d8c783a5885ad7543535e51ef374210a717a437562fd8aaef

Observation 4cd4cb20-7d0a-4d33-833a-a0527cd597af · inbound

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models cites this paper.

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T23:49:03.695895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:49:03.695895Z digest=sha256:566460ddc45dee29107e4eab74f6044c0c776e65605eebb463a5fa3a5235c88c

Observation f09b8d10-664f-4342-8997-3abb78eb47be · inbound

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models cites this paper.

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:25:30.899585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T11:24:45.578472Z digest=sha256:ce6bf1805d6b84dbf1c8a4f3e5337b73c44d17d52d5656d9ac187eced30cfa24

Observation c7685e57-c745-4ce7-bce3-845f5c60ca7b · inbound

Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation cites this paper.

Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:10.905387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T09:09:13.191350Z digest=sha256:d14663a8090890669619f2f216b45027864cc69427000267fa6eee1ac65e3c8b

Observation 97a3e32d-23e3-47a3-9110-8386a18aba9e · inbound

VISOR: A Vision-Language Model-based Test Oracle for Testing Robots cites this paper.

VISOR: A Vision-Language Model-based Test Oracle for Testing Robots VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:21:23.375027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T05:17:03.803350Z digest=sha256:f3846168e4142d13fc4866f7d22ca4d783d2a136d9ee0b364f70bc0dced5e4f9

Observation c39407f3-b62f-4472-ad39-ef5a1164ba67 · inbound

VISOR: A Vision-Language Model-based Test Oracle for Testing Robots cites this paper.

VISOR: A Vision-Language Model-based Test Oracle for Testing Robots VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:29:08.992580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T22:24:25.871817Z digest=sha256:faaf269df1a015257bbe8edca638d0240b84827bcfb8465788999e1d74dd6a46

Observation 9816f0e6-a24b-4a4b-a4df-317a72ce36f7 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 240

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:18.033580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:ec22a02d41cfe94e25009fba9c4b92c72b4846190b386dc5609b4ad5cd59574b

Observation 88885839-87d0-4bda-81c5-fa7e8efecb9c · inbound

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System cites this paper.

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:43:11.069321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T09:38:43.801252Z digest=sha256:2b52b09058b8cf49b7ed7f616295a49e7b40bdbc4bd64f9ecd0b22fee35770f1

Observation ed0293e3-433d-4eaf-91e9-64e563944e18 · inbound

Colosseum V2: Benchmarking Generalization for Vision Language Action Models cites this paper.

Colosseum V2: Benchmarking Generalization for Vision Language Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:33:38.665724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T16:32:32.885345Z digest=sha256:852008cf16fd22dc6ad27bd8f73cadf4deef8c0d68a676536cfb6f85c768a3f1

Observation 92f94aae-21c1-4a8e-b7aa-8531b06582f0 · inbound

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models cites this paper.

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:20.691103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T14:39:27.174418Z digest=sha256:c2b1d61bdfcbb5f2f5250adb56a471e55276f0db35d837a4d5ba624653bc0b39

Observation 713e4e26-5890-49bd-aad6-ce5ce4ff1326 · inbound

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation cites this paper.

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:46:33.226420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T09:40:04.685274Z digest=sha256:88cf602177bbe380709a755bcda63ac049b1a63d9cdd8411c33e504c9992e629

Observation dab1c090-3dd6-499e-bb53-5f2e630a387a · inbound

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation cites this paper.

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:18.169340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T21:40:00.330510Z digest=sha256:f87ee0af73323d5b4ed84599698d2d64a23c0435930b431fcee24bed0cec1137

Observation 26ed2b1d-1f73-49cf-9e10-c90681737d50 · inbound

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks cites this paper.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.622739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:14bd11e02137c8d09a959daca5e3cc68bf1a9aef00d157c5d4ce5eb3e320271f

Observation da963171-54b4-4abf-8d30-2b04a5f3235c · inbound

A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation cites this paper.

A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.125593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T13:06:35.282837Z digest=sha256:ccdfa580f491570781870b93fbfef399e311dfc994488238a3fb583f9cd7d6eb

Observation d2911b14-c2e2-4ee9-85ee-99d4b08fa3d1 · inbound

UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data cites this paper.

UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.785689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T12:59:26.898710Z digest=sha256:e49bef88b0e4c61c372a2fa705d2ef8566381347a66b51b06dafeb48cba8e0c1

Observation a59bb101-0cca-4d05-b79d-351c3ec5b00f · inbound

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models cites this paper.

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.015844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T20:55:29.230477Z digest=sha256:ce0ce54bc9f24c1d74a9a3a1619ea04d1218a15d1644df2bea3813912fc654bf

Observation e772ff54-ed4f-4dbd-83ce-6f4a1296f6cb · inbound

World Action Models: A Survey cites this paper.

World Action Models: A Survey VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:35.430139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T17:11:12.686936Z digest=sha256:b5ea2b4b94b6402e80bc004cd6c06a958dc0510349f680bd0e9a80d2ec015bf1

Observation 95799076-95f7-4d13-8d90-ac53db763ee6 · inbound

Bridge-WA: Predicting Where and How the World Changes for Robotic Action cites this paper.

Bridge-WA: Predicting Where and How the World Changes for Robotic Action VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:38:04.579410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T11:28:50.286894Z digest=sha256:c93fedc252d1f76050142b5187fd55fde14b949fd243a337bfe1cf4f4fba494d

Observation f9ddcd77-ad4c-4ab3-b7de-2247332469e1 · inbound

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement cites this paper.

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-08T08:24:47.094418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T08:19:27.960645Z digest=sha256:8851684f8848f9c71edd9cabb390b9468047db42015824bf548c8a3799a88ff3

Observation 43b3b4f2-9de9-4796-8040-d5b0b3f21b88 · inbound

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement cites this paper.

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:42.200463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:31:42.200463Z digest=sha256:2cdc1e61f1726336f5b81e5fb7c1d94cc90d059e06a354ccf8d05d5801a99abc

Observation 481223e8-41e8-4ee7-9ff0-c15a83959a0f · inbound

Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution cites this paper.

Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T02:25:54.279975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:25:54.279975Z digest=sha256:97aff09c229a91cf15530ac91d901722bb45c30e4756d6bf17a8584586acdad1

Observation eb710edc-4f47-4d25-8e93-ae7a0192a90c · inbound

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation cites this paper.

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-30T12:42:18.656193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:42:18.656193Z digest=sha256:4b1d3c59a3e4113b7d48c31a36e353740aa128192ee30837df1618b414f329e0

Observation d0d2ef7b-364e-4d66-b04f-60f0462db2ae · inbound

GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning cites this paper.

GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:47:48.475113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:47:48.475113Z digest=sha256:5e6d6d417af61f375cd0edcc6980dec288dcb15c22497f3233cc0284e9c1b067

Observation 3fd8f21c-ebc7-4c37-8d1b-6f4fb0fc269b · inbound

Self-Evolving Embodied Agents via Skill-Harness Evolution cites this paper.

Self-Evolving Embodied Agents via Skill-Harness Evolution VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:49.772799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:49.772799Z digest=sha256:eed8d9ddde33cf6e58763a940aaefe877726e8ff60fe9cf899b52807d06d6193