Pith. sign in

Paper Citation Record · LEDGER

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2412.18194.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18194 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:03:21.446755Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eec0d6ab-a241-4887-885e-500239218926 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:32:22.473964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:9e2951601f8e2aff8efa75bc13af6c1346d392eb76371edefcba68a084d4f664

Observation 1deb9fcd-83d5-4411-a38a-1555e06bbfb3 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:08:35.362092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:47964bd75606dff7e1d7a7095d8183e8ba127e9059e304e492550a28f9e1ada7

Observation f6521287-567f-486e-a43c-569efabdd807 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 119

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.824554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:4cf1dfa93c58caeb9789de0abb35a722c0d0466adbbd6f24f8d405c47ba3f8e0

Observation 2d85b8e6-4ffb-4fe8-823f-1a8e01e82eb8 · inbound

4D Visual Pre-training for Robot Learning cites this paper.

4D Visual Pre-training for Robot Learning VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T17:03:21.446755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:03:21.446755Z digest=sha256:08a86cb706040f174c5d017598f9f76354bfc1567c5ffaa5301167e5370a1d80

Observation dc7d9943-095a-469d-b367-3aaa6d1ef8b4 · inbound

Galaxea Open-World Dataset and G0 Dual-System VLA Model cites this paper.

Galaxea Open-World Dataset and G0 Dual-System VLA Model VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T13:31:09.974877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:31:09.974877Z digest=sha256:a602861c8e921f5681e9f4f9fcd1532476f8180abdd1ab20dcb52331f9c65331

Observation e85c76da-2b42-4b5d-be22-58099d2e4007 · inbound

RoboBenchMart: Benchmarking Robots in Retail Environment cites this paper.

RoboBenchMart: Benchmarking Robots in Retail Environment VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T22:29:59.056501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:29:59.056501Z digest=sha256:66ec08e7914a649e9d992fc514a728264e0adc086e65ca4a4b0e54d3bd017b6f

Observation 961ea27c-946d-4443-9f19-e2a797e9bb96 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:27.342980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:27.342980Z digest=sha256:64d2b7f9891879e5362d9d7675a046f7c61bf920f55c6c95591140de7adb81a8

Observation 4cd4cb20-7d0a-4d33-833a-a0527cd597af · inbound

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models cites this paper.

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T23:49:03.695895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:49:03.695895Z digest=sha256:a315f0a74a6ef0e9c96bf0c91d560c62f46f5e11fc3e5d2ee318daaa814409a9

Observation f09b8d10-664f-4342-8997-3abb78eb47be · inbound

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models cites this paper.

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:25:30.899585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T11:24:45.578472Z digest=sha256:a5f6f7ffd8e0cb8951e679244494dda0b26e785e6248c6edc2f70ce2c69982a6

Observation c7685e57-c745-4ce7-bce3-845f5c60ca7b · inbound

Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation cites this paper.

Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:10.905387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T09:09:13.191350Z digest=sha256:0686ed60ed21fc0fc0b638330fd907c51a57b75c1a82d2d8a47f329976145902

Observation 97a3e32d-23e3-47a3-9110-8386a18aba9e · inbound

VISOR: A Vision-Language Model-based Test Oracle for Testing Robots cites this paper.

VISOR: A Vision-Language Model-based Test Oracle for Testing Robots VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:21:23.375027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T05:17:03.803350Z digest=sha256:985fe945e9236babc771692ae0a3fec90f1a580afa09f38f7415679b0223c477

Observation c39407f3-b62f-4472-ad39-ef5a1164ba67 · inbound

VISOR: A Vision-Language Model-based Test Oracle for Testing Robots cites this paper.

VISOR: A Vision-Language Model-based Test Oracle for Testing Robots VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:29:08.992580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T22:24:25.871817Z digest=sha256:28787b192063bbf53eccdd51e28b3bf81c38a0168aacbcdf79ec0f1a364e9d46

Observation 9816f0e6-a24b-4a4b-a4df-317a72ce36f7 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 240

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:18.033580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:2693f36fc188a46435aa6c16394e511f7b024a6b54917d4d1f64038827d484de

Observation 88885839-87d0-4bda-81c5-fa7e8efecb9c · inbound

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System cites this paper.

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:43:11.069321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T09:38:43.801252Z digest=sha256:98dde22bb26aa886bdc70aa014bcd64895c4ced834d29c402f3b28b6067320ca

Observation ed0293e3-433d-4eaf-91e9-64e563944e18 · inbound

Colosseum V2: Benchmarking Generalization for Vision Language Action Models cites this paper.

Colosseum V2: Benchmarking Generalization for Vision Language Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:33:38.665724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T16:32:32.885345Z digest=sha256:d39c0158aa9c091e13a2afcf490c7891bd04db4e18137ad4bddc96158d354733

Observation 92f94aae-21c1-4a8e-b7aa-8531b06582f0 · inbound

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models cites this paper.

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:20.691103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:39:27.174418Z digest=sha256:517e18eaee29b52f05852d1630945c42f00bbaf98ebff016e6fd33a1cec289ba

Observation 713e4e26-5890-49bd-aad6-ce5ce4ff1326 · inbound

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation cites this paper.

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:46:33.226420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T09:40:04.685274Z digest=sha256:63816f80bcf411a4a7c500c1ebee0e9ce40466f2564072fd54cf280e45a862a2

Observation dab1c090-3dd6-499e-bb53-5f2e630a387a · inbound

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation cites this paper.

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:18.169340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T21:40:00.330510Z digest=sha256:07c286994bc2349473ace89a0972656b6dc3908042d2e151e776fb4341f9bcf2

Observation 26ed2b1d-1f73-49cf-9e10-c90681737d50 · inbound

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks cites this paper.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.622739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:7b61e959c155040340659a96845010c5364f12ee5723036ba1ed7046e367a576

Observation da963171-54b4-4abf-8d30-2b04a5f3235c · inbound

A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation cites this paper.

A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.125593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:06:35.282837Z digest=sha256:78030eaa6f7b3ce1a346333d5cf6fe700b0fd1ac62f14619cb69dc0b7b86eb5b

Observation d2911b14-c2e2-4ee9-85ee-99d4b08fa3d1 · inbound

UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data cites this paper.

UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.785689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T12:59:26.898710Z digest=sha256:bcbddfe55065d297ea8efa63e9f01586848e02798deff8d45ec9f649da3a9b3a

Observation a59bb101-0cca-4d05-b79d-351c3ec5b00f · inbound

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models cites this paper.

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.015844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T20:55:29.230477Z digest=sha256:7512c74ecb3eb77091c37d8af8920fa3026d4e6dd5ac1ce7466709017b6e61d6

Observation e772ff54-ed4f-4dbd-83ce-6f4a1296f6cb · inbound

World Action Models: A Survey cites this paper.

World Action Models: A Survey VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:35.430139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:11:12.686936Z digest=sha256:93bb0deec7e677a92b6964d673288fb159ebbe2549b7cd6bf65fb08878d3054b

Observation 95799076-95f7-4d13-8d90-ac53db763ee6 · inbound

Bridge-WA: Predicting Where and How the World Changes for Robotic Action cites this paper.

Bridge-WA: Predicting Where and How the World Changes for Robotic Action VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:38:04.579410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T11:28:50.286894Z digest=sha256:f12a68d90a091f4e94d182e43a5dc89751d2988a6bdb17fa09008e8ad845f986

Observation f9ddcd77-ad4c-4ab3-b7de-2247332469e1 · inbound

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement cites this paper.

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-08T08:24:47.094418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T08:19:27.960645Z digest=sha256:20bfe198bec878e79f4789a1c6f3fca873894d9431bbff61c23f0df349cceece

Observation 43b3b4f2-9de9-4796-8040-d5b0b3f21b88 · inbound

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement cites this paper.

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:42.200463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:31:42.200463Z digest=sha256:f72bd5d3e575002d0d702973d2a31329534420bfc81f14ccc57bc46d24c5dd47

Observation 481223e8-41e8-4ee7-9ff0-c15a83959a0f · inbound

Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution cites this paper.

Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T02:25:54.279975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:25:54.279975Z digest=sha256:5c5bdf3637e8fb6be67d2423932efd8589273dfb1005acb3eb914e6c9681a4f9

Observation eb710edc-4f47-4d25-8e93-ae7a0192a90c · inbound

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation cites this paper.

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-30T12:42:18.656193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:42:18.656193Z digest=sha256:64877342b375d1e4266de25dd4369474fbca65c0c70c708659e1538bfa74763d