Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:21:15.562904Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 21 inbound Pith citation observations for arXiv:2505.15517.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:21:15.562904Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:56:37.191428Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:59:46.871127Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2810f855-b6ec-49ef-bff6-84e1f1ed0461 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Learning transferable visual models from natural language supervision
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b306a3-2f10-4c8b-aeca-419db50ceaf2 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Qwen2.5: A party of foundation models, September 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f7afe0-db68-498b-9577-5941013cbed1 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b283527a-94e2-4907-8815-859bfbf2d01c · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42254d0-0d1d-4b9c-8e6f-866eea83b1ce · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Claude 3.5 Sonnet
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9087d7c-72a7-4cd8-a609-f6258e7b940c · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets GPT-4o System Card
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 88ad1e3c-4542-4184-9ad1-b9195e891e99 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Gemini 2.5: Our most intelligent AI model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fc579134-c768-4c91-bc3b-cbe21cff9066 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef00d9c1-6dce-42c5-9560-73ebf9d79a41 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c840a04-665a-46c8-86d0-de2aaf072892 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets OpenVLA: An Open-Source Vision-Language-Action Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ae5ddd-834d-4fde-8be4-6d2b444708e5 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Gemini robotics: Bringing ai into the physical world, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e53a4628-f2d8-499f-a70b-27291e9247b3 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Hi robot: Open-ended instruction following with hierarchical vision-language-action models, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49756c7c-1796-49b5-ac41-6576e0eeaaac · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets EQA-MX: Embodied question answering using multimodal expression
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5486305-0b77-41c1-abe9-ad1f99cd4ad6 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19ba6cb-a6b2-49fa-9a97-67f7ed454784 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Embodied agent interface: Benchmarking LLMs for embodied decision making
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e4dabe6-a413-4a42-a05a-e30dad133706 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets ALFRED: A benchmark for interpreting grounded instructions for household robots
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a5b6fda-a852-4c0e-87a3-444df8ed09ab · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Habitat 2.0: Training home assistants to rearrange their habitat
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 89fac9a5-be2f-424e-96fa-2ecab99ce51e · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets AI2-THOR: An Interactive 3D Environment for Visual AI
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28ffec7f-042f-4d23-a4be-d9ff5dd7e60b · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets RoboVQA: Multimodal Long-Horizon Reasoning for Robotics
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbbefaac-f9b0-4556-84d0-13326ffaa9a0 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Robobrain: A unified brain model for robotic manipulation from abstract to concrete
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8b530ecd-9b84-45c7-abe9-90bc9fa4d57f · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets End-to-end training of deep visuomotor policies
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8bb77a2-732e-4716-a7e7-3ec37fe1b30c · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Octo: An open-source generalist robot policy
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb2f140-4df4-4fd6-8fd0-da4d723b5616 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Diffusion policy: Visuomotor policy learning via action diffusion
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fadc7e27-88de-4086-9c27-191a49deb0b7 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70eef0d6-9a22-4c5e-9eeb-66ec350dfdf6 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Doremi: Optimizing data mixtures speeds up language model pretraining
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d8c58291-f7e2-482e-bbb5-1a69728bb5fa · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Remix: Optimizing data mixtures for large scale imitation learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 570b1f98-6c35-4d47-b7b7-fefbbd8c99e2 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Waveform Manipulation Against DNN-based Modulation Classification Attacks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 677fab60-abfd-4446-93b7-ff132fdf478e · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Rt-1: Robotics transformer for real-world control at scale
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2536aae8-b720-4680-bf63-10254c34c0d8 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0380f50-e9a3-4b7f-9da8-cb92844a76ed · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets https://physicalintelligence.company/blog/pi0, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7139a414-069e-42f3-816c-a11d6f3d5eb5 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Droid: A large-scale in-the-wild robot manipulation dataset
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0516040-c53b-4e48-83ed-da47be961464 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Latent plans for task agnostic offline reinforcement learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4707bcc1-3f63-47b7-b916-875eb7225b2d · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Grounding language with visual affordances over unstructured data
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 723865a2-2ef1-487b-850f-bfe2f1c56e39 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 721a2928-b8d1-4200-824a-7be393c81ebd · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Berkeley UR5 demonstration dataset
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e9d37e-1a9e-40d1-b3bd-f87294bb7b31 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Robot learning on the job: Human-in-the-loop autonomy and learning during deployment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a37763b-cb4e-483e-bcb7-c240ec7642cd · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Real-world robot learning with masked visual pre-training
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 401421ae-7219-4c66-a579-481f77a0dbed · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets The surprising effectiveness of representation learning for visual imitation, 2021
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6e486d93-47b0-470d-92b1-8d69a030be3b · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f58f7461-6395-421a-b962-82b99b8a7214 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Modularity through attention: Efficient training and transfer of language-conditioned policies for robot manipulation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation caa322b1-315a-47a2-8d3b-be8fcd050494 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Learning modular language-conditioned robot policies through attention
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d21542a-2063-4be4-b112-479ab68a7519 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Viola: Imitation learning for vision-based manipulation with object proposal priors
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fada99ea-02d1-4d8e-afe3-83c6258a0d68 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c6f6d8-9800-4059-a84c-c4873317280e · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Watch and match: Supercharging imitation with regularized optimal transport
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2c736d5f-51b0-4d83-b430-8bb752612624 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Vqa: Visual question answering
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d6a418f-6093-407c-94c2-7a97097d269f · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a429751-fb69-4ff3-b70e-16c03ab4fd54 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets What is Right for Me is Not Yet Right for You: A Dataset for Grounding Relative Directions via Multi-Task Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa95f25d-8413-4f1c-bdc6-dff3c0d17766 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Embodied question answering
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9b0369e-a073-456f-96b9-c14cb3d243d3 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2320ef47-b91b-4d3a-a440-3a5a62ef015f · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Christensen and Gregory D
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c0dbc78-7d6e-4cc4-8b85-fdebb3c0ba27 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 323bc312-f722-403f-a991-2854cc6ed622 · outbound
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Configuration D
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3958f27a-cd4c-4d43-a2c9-5896bb4b176d · inbound
Demystifying the Visual Quality Paradox in Multimodal Large Language Models Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71e0cde9-4a6e-48b7-bf08-c46222e1b4e5 · inbound
HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5371277-8e79-4797-ab8d-c1eabf6250ec · inbound
Foundation Model Driven Robotics: A Comprehensive Review Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84347f74-b78d-436e-b668-2d701380c58b · inbound
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab07615c-fe48-4386-9ebc-e6962892d043 · inbound
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4d926e-04c7-4446-82ee-35e3b02bbd32 · inbound
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d10b362-674a-4162-8a8b-2ffdfcf7c6aa · inbound
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6597ac9a-428c-4f63-81ed-4b8b7612e738 · inbound
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13d6f012-16b6-44f1-a40d-981388de4813 · inbound
Rethinking VLM Representation for VLA Initialization Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 69813e01-6f70-4049-8b54-9152fed034f7 · inbound
Extending Embodied Question Answering from Perception to Decision Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 78ecdf25-e94e-4225-b0d2-a12583a15c8c · inbound
GEM: Generative Supervision Helps Embodied Intelligence Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6657bccb-16d3-4b2f-b549-c6a75779ff99 · inbound
Wall-OSS-0.5 Technical Report Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6d129ccf-4039-44f3-a1d0-427ab9ba868a · inbound
Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ea05806d-e451-4a73-87e4-ebb4d88fd28c · inbound
RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93a450d5-0507-48c7-8596-c619c4fb2a52 · inbound
RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9149a18-0ed3-4570-b74c-381e2a561269 · inbound
Vesta: A Generalist Embodied Reasoning Model Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 974cc6a3-9b1b-4118-884a-d95e485da3ce · inbound
Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a99f9c32-9729-45bf-81f1-806a6e7e4970 · inbound
EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c34079-fe85-4677-91b1-0d98d1ed7f7e · inbound
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ac0bafc-1318-471d-8f61-4b03a1718018 · inbound
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21826f6d-47ea-4d4c-be0f-410311188457 · inbound
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.