Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T16:32:32.885345Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2605.27759.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T16:32:32.885345Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:10:33.754125Z
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cbe8b940-f904-47cc-a9f7-093dc2d361d9 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models ChatGPT: Optimizing language models for dialogue,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ef84ecb-c270-465e-ad73-4d5150abf36e · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models SAM 2: Segment Anything in Images and Videos
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 871f02dd-90ea-4e65-8f78-18aaf8dfe943 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Rlbench: The robot learning benchmark and learning environment,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 007ad77f-bb6e-4cb2-a361-d0c19f0da85f · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Pyrep: Bringing v- rep to deep robot learning,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee08941-b285-4b45-94ee-9ec22ae7e100 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Coppeliasim robot simulator,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96ded4e1-af25-442d-b617-452f495d1a1b · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models The colosseum: A benchmark for evaluating generalization for robotic manipulation,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 305a473a-f0a0-469a-b9a9-ae0c7d4bcce6 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Libero: Benchmarking knowledge transfer for lifelong robot learning,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7b5cba4-ef93-4676-bf50-d2d78382f044 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Libero-para: A diagnostic benchmark and metrics for paraphrase robustness in vla models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 97244e5d-e855-4cd6-9e75-f0cfefbbe4e2 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Roboverse: Towards a unified platform for robotic manipulation,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fadeaf97-ce4f-4859-83fa-3559ad2d939d · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Roboarena: Distributed real-world evaluation of generalist robot policies,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a810712b-3d35-4a88-9818-311295004413 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Robotwin: A platform for scalable robot learning,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c008b7f-7035-4f78-8166-1ca8716c9fb4 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Bimanual manipulation benchmark,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2fae33f-3283-4603-937b-b5b101451653 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed0293e3-433d-4eaf-91e9-64e563944e18 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9018405c-5ecd-4b40-a2ca-d5eb26d4301e · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dc307761-1811-4e52-954e-bd759c6ee3f8 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b6cf6b50-5057-467f-a215-4c54461ed9ed · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models R3m: A universal visual representation for robot manipulation,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24c839a1-4e12-4702-b24e-314a64b00d40 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Mvp: Multi-view pretraining for vision-language robotics,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d886665-2f5b-4bcd-af67-2ca6bfc621d6 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ddeea52c-adce-41ad-bd81-3064d602b1b4 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Cliport: What and where pathways for robotic manipulation,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7edb617b-9c80-4cb8-b47a-e3e9de1a8977 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models V oxposer: Composable 3d value maps for robotic manipulation with language models,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d008cbd-7351-4fe6-949d-36c226da8d61 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models C2farm: Coarse-to-fine imitation learning for manipu- lation,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46af9365-c207-41fe-bcba-f9fa6c0ca0d7 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Kite: Keyframe imitation for task execution,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc441fa5-ed95-4bb0-b60c-8f8a177c9bc2 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Learning fine-grained bimanual manipulation with act,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2519e6b-d901-4b7b-971c-c5c0b1c9c898 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Peract: Perceiver-actor for 6-dof manipulation,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7214ea25-e6eb-44db-899d-0178760cdb31 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Rvt: Robotic vision transformer for manipulation,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed9aae14-7382-4577-b9f5-01be80caea64 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Rvt-2: Scaling vision transformers for robot manipulation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16acdf44-8564-4078-80e8-7a29ff69cfc8 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Act3d: 3d feature fields for manipulation policies,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e66fd3f-cc8a-41ff-a4a6-07fc83cb5798 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models PaLM-E: An Embodied Multimodal Language Model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eda6ce01-dac8-4d17-8859-a9f0d804c420 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 04dd3932-b846-4a2e-8a09-cbe562287af1 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models OpenVLA: An Open-Source Vision-Language-Action Model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6939e931-bf43-420f-9f91-22b8b7d5a64f · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Provable Preconditioned Plug-and-Play Approach for Compressed Sensing MRI Reconstruction
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b9efcf2f-556f-4d20-afa6-7d120d273d23 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models AgentNet: A scalable framework for multi-step agent trajectory generation.arXiv preprint arXiv:2501.00000
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8935a05d-cb8a-408a-ab8c-58eae96e8d62 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models π0.5: Vision-language-action models for open-world robotics,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db2112e6-6e00-4320-881c-d048ee87d685 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Open x-embodiment: Robotic learning datasets and rt-x models,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edd948c0-0bc0-432a-8d09-e6cd64e37bf6 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Lerobot: State-of-the-art machine learning for real-world robotics in pytorch,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52a26050-24c8-49f5-9ef1-9193e4f087f0 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Maniskill3: Gpu parallelized robotics simulation and rendering for generalizable embodied ai,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c5cdd1-3ccc-4880-ae85-c69912764229 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Learning Transferable Visual Models From Natural Language Supervision
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 05b17789-7c24-4604-abf9-714219389994 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Sigmoid Loss for Language Image Pre-Training
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e9b8fb3b-eeaf-4144-861f-f26d0d872b66 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Deep residual learning for image recognition,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9cb44e3-7f76-4668-8ac2-4b1608b3ed26 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models MolmoAct: Action Reasoning Models that can Reason in Space
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 835a08be-64b7-4fc1-9347-eb9d519d2e45 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models VIMA: General Robot Manipulation with Multimodal Prompts
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a5df6c8f-ac0b-452e-a79c-4c60d8b2a9cc · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Learning an actionable discrete diffusion policy via large-scale actionless video pre- training,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a001da23-370f-4b61-b88d-9d130fae1590 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Unified Video Action Model
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71a90fad-2969-478f-b485-6e1d027a3a2d · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fd864f6-be9f-47d9-8fe7-65fd4d9e7605 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ba119abb-c646-4ed6-90bf-fe89c6db2a7f · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51944ed5-f013-4f5a-a327-7100e862b164 · outbound
Colosseum V2: Benchmarking Generalization for Vision Language Action Models Contrast Sets for Evaluating Language-Guided Robot Policies
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9fc87436-d7ab-4b77-bd4e-777affaf6790 · inbound
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models Colosseum V2: Benchmarking Generalization for Vision Language Action Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 926ecd5d-215c-4b2e-bbaf-d13d5fa8c494 · inbound
Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models Colosseum V2: Benchmarking Generalization for Vision Language Action Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.