Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T18:53:07.734871Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2606.00054.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T18:53:07.734871Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T07:04:08.886475Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-30T07:04:20.747493Z
82 of 82 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 18ac73da-f463-4ba2-a28b-c89268783562 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Cosmos World Foundation Model Platform for Physical AI
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9c938cc1-a1f6-4d36-a5e9-f11b48349e71 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Agibot world colosseo: A large-scale manipulation platform for scalable and intelli- gent embodied systems.IROS, pages 3549–3556, 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 20361032-6f3a-4fff-ae6d-c9dacb06e552 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Egocentric-100k, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0bbf8ae7-e72c-46b3-8cf3-7e13617476d2 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Affordances from human videos as a versatile representation for robotics
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3df0448e-7df1-4993-b5db-1d90f923dae2 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Hot3d: Hand and object tracking in 3d from egocentric multi-view videos
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7e79e161-99d4-404f-b8f2-91e9bace544c · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Gen2act: Human video generation in novel scenarios en- ables generalizable robot manipulation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1bd363c-1639-456f-9ae9-a333752b12a8 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Motus: A Unified Latent Action World Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bb8bf0e5-db77-432b-9b84-dd01ba1ba51a · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 38512e7b-48d3-4151-b281-913c2531b82d · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eedcf644-ffc9-421d-be6c-e7b478a662dd · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f4c6c7a4-7e1d-4452-8b76-239ee9134886 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Scaling robot policy learning via zero-shot labeling with foundation models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f104c49-2d16-41d7-8e5c-ba8ee5f9b914 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee4dedce-2e0d-4727-88ba-fb7c5f20506a · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Affordance learn- ing from play for sample-efficient policy learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 95238e06-8f83-4b8b-aa24-589727f1bb20 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data RT-2: Vision- language-action models transfer web knowledge to robotic control
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86fe84fc-bc9e-4c51-8f57-b77d256c6f0a · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data UniVLA: Learning to act anywhere with task-centric latent actions
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 14b53126-07ae-4fe0-bd9c-04e0f117560c · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data In- n-on: Scaling egocentric manipulation with in-the-wild and on-task data
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ec35ef4-52a8-49df-bd10-24eb53b1b237 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data A Short Note on the Kinetics-700 Human Action Dataset
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a03b5a39-a11c-40b5-b4e8-40b8d7f9dda8 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 172dd323-3d20-42be-9cab-eb70a0bdf684 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation feec6cc3-8c0f-4f6f-b7dc-cd8173b86ce7 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data VidBot: Learning gen- eralizable 3D actions from in-the-wild 2D human videos
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7b2be9a-46c4-4d73-99b2-06c5560ada06 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90dbf41a-e9e8-4317-af53-269b5f9365da · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 32eeaeb9-73b2-4fb4-a8b0-47eff7ebc859 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Moto: Latent motion token as the bridging language for learning robot manipulation from videos
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 62cf13e8-28a2-4065-b37f-c431084ea8df · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data The EPIC- KITCHENS dataset: Collection, challenges and baselines
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b4b0f713-14dc-4854-a954-937cfce6f683 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Tam- ing transformers for high-resolution image synthesis
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83a00802-ea26-40a3-8241-196af6cdf826 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Arctic: A dataset for dexterous bimanual hand-object manipulation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8ea4912a-e097-43de-a47d-94ac4085f9f4 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 87c53b22-5c50-4df3-9ea4-164fc839f0cf · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Learning la- tent action world models in the wild, 2026
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76b32393-54fb-4001-80b6-d48d4d3fbdba · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data The ”something something” video database for learning and evaluating visual common sense
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d988bbc4-8cf9-4c95-b5bc-8aa9c4b2336b · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Ego4D: Around the world in 3,000 hours of egocentric video
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5367fe3e-02a7-4d90-932d-fd08659b0f57 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Ego-Exo4D: Understanding skilled human activity from first-and third- person perspectives
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e31e212-597f-4f5a-86a1-d13710801c66 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Lelan: Learn- ing a language-conditioned navigation policy from in-the- wild videos
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 144c5b1b-61f8-4315-a9b9-0348280ee8aa · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5fdffa2-0e6c-4b55-be1e-d01eca5edba0 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Video prediction pol- icy: A generalist robot policy with predictive visual repre- sentations
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 329018f2-72cd-492c-bd47-e69180bbaa95 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 153d1cf6-1429-4da8-96bc-f836a05f04d8 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Emergence of human to robot transfer in vision-language-action models.arXiv preprint arXiv:2512.22414
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 18fe7038-584f-4da5-b998-49cb2650963a · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Droid: A large-scale in-the-wild robot manipulation dataset
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43627b06-b2db-4ab3-8ca3-ead90c48b194 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data OpenVLA: An open- source vision-language-action model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65c30495-6782-4bdd-8129-f91b8ab546be · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Masquerade: Learn- ing from in-the-wild human videos using data-editing
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d92a5420-96c3-4473-a06b-3bbdacc18c42 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data In the eye of the beholder: Gaze and actions in first person video.IEEE transactions on pattern analysis and machine intelligence, 45(6):6731–6747, 2021
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5f64c9d5-a79b-4fe2-b256-fcf315daf888 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 71edcfae-2d9c-46a5-be2a-735e8fae973f · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Evaluating real-world robot manipulation policies in simulation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 606c1874-ca19-4553-bc0f-a625873fe92c · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Scalable vision-language-action model pretraining for robotic manipulation with real-life human activity videos
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d4411059-2c4c-422e-a189-26b0855789ed · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data HOI4D: A 4d egocentric dataset for category-level human-object interaction
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f39635a-6bca-46dd-89bf-4475e76d9042 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Libero: Benchmarking knowledge transfer for lifelong robot learning.NIPS, 36:44776–44791, 2023
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f569bf4-4816-4589-8f13-bf0b7e26407a · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Taco: Bench- marking generalizable bimanual tool-action-object under- standing
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7ab91d45-1471-4c74-b056-ea02fcfc11e3 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b81a492-ca1b-4c7d-ae4d-a3ea150ec2fa · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data VIP: Towards universal visual reward and representation via value-implicit pre-training
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a7e82cc-b8b7-4d15-97e7-03f55b538b76 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fb5629d5-c291-436e-aa8b-2da65eabe43d · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Grounding lan- guage with visual affordances over unstructured data
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 15e68163-16c9-4640-bd48-26bee433dff8 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation efd18abb-8f9c-4d42-addd-9b48a6fb8cf0 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data R3M: A univer- sal visual representation for robot manipulation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd7696d7-53d8-4403-b85e-56e7a7dda7e5 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data DINOv2: Learning Robust Visual Features without Supervision
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9cdec0d1-4af6-41f3-a151-b0ad7b47cae4 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Open X- Embodiment: Robotic learning datasets and RT-X models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5f4d637b-0e20-42ec-8173-c24cfd2c2295 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 953c51a5-78dc-40b5-90ac-62856ac1e7e7 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Dexmv: Imitation learning for dexterous manipulation from human videos
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 389845df-7d91-47f9-b581-86ad9aeaaf7b · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Embodied hands: Modeling and capturing hands and bodies together
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 576821bc-7506-499f-9199-90fb03a22849 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Vipra: Video prediction for robot actions
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c399ca70-e946-41b2-9cb4-8b5424075a9d · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Assem- bly101: A large-scale multi-view video dataset for under- standing procedural activities
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f68176b-c9b9-4b73-8094-0f2688a036bd · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Wave humanoid robot
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05dd8d5a-264b-4ebf-a384-e86bce466c59 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b32ba99c-c08b-4f86-a544-776805df94ba · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Gemini Robotics: Bringing AI into the Physical World
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3046b3f0-d022-4e41-b18a-0d4a94a9b3db · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Tesla ai day 2022
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 11f20381-9bfb-4d21-b363-8f8c9d18d0c5 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Neural dis- crete representation learning.NIPS, 30, 2017
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc218ccf-2419-4487-b341-e9f17f755bed · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Bridgedata v2: A dataset for robot learning at scale
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 18604bc6-b6c5-45bd-9c41-54b9459c7c76 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Holoas- sist: an egocentric human interaction dataset for interac- tive ai assistants in the real world
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2857ae1a-cf8c-47b7-a62a-84c5c670841e · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Gensim: Generating robotic simulation tasks via large language models, 2024
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f391e4a6-3a55-476c-a21e-0515ee4d95db · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Any-point trajectory modeling for policy learning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3118baa4-8fbf-4b22-bda3-f1967bc205ca · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Unleashing large-scale video generative pre-training for visual robot manipula- tion
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27d7170c-9200-4b90-8d4d-b8a25be6e0ae · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Masked Visual Pre-training for Motor Control
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 269bca05-ac26-4422-9c6b-98dbf958e131 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data A0: An affordance- aware hierarchical model for general robotic manipulation
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a707b6e-a375-4e66-b430-47f35023bdd6 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Magma: A foundation model for multimodal AI agents
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b95a5978-599b-4eb8-a88c-ceba7b7e504b · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation db5b6368-63f2-49b6-b3d7-e41a8df41f3f · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Latent action pretrain- ing from videos
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9562be4b-65d1-4414-9562-a570989a887e · outbound
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bfa07efb-88b5-44d9-8e6d-e3e2f65905a8 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6fdb062d-9680-4156-89ef-993a3362db72 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Motiontrans: Human vr data enable motion-level learning for robotic manipulation policies
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fbcd2d81-cb21-4130-8d32-5b86d58cd5c9 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Hermes: Human- to-robot embodied learning from multi-source motion data for mobile dexterous manipulation
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6caa56c7-1f6d-4819-999b-0ee45f8db512 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Oakink2: A dataset of bimanual hands-object manipulation in complex task com- pletion
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 08e0b236-24c0-4550-9dbd-52fbc22692dd · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Clap: Contrastive la- tent action pretraining for learning vision-language-action models from human videos, 2026
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 82709cb9-2929-4f90-9aa9-52c33c84bc01 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data FLARE: Robot learning with implicit world modeling
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4cfc41cf-317e-435f-996b-5b2f28bcbf23 · outbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Zhu, Pranav Kuppili, et al
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f856e00c-ec8c-4ffe-a690-6a5d236da996 · inbound
CORE: Common Outcome Regularities from Action-Free Visual Demonstrations for Robot Manipulation From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.