Pith. sign in

Paper Citation Record · LEDGER

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

As of 6 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2606.00054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.00054 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:53:07.734871Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:04:08.886475Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T07:04:20.747493Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact26
  • verified fuzzy54
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 18ac73da-f463-4ba2-a28b-c89268783562 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.198844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:29b778ec5188e89d2eeb9396faa856a91872105bcacc0b7caef7fa1c3210e00e

Observation 9c938cc1-a1f6-4d36-a5e9-f11b48349e71 · outbound

This paper cites Agibot world colosseo: A large-scale manipulation platform for scalable and intelli- gent embodied systems.IROS, pages 3549–3556, 2025.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Agibot world colosseo: A large-scale manipulation platform for scalable and intelli- gent embodied systems.IROS, pages 3549–3556, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.673530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:2735587b3d1def07850abd05efa1a8b2555beb380e0a3438da627786036a813a

Observation 20361032-6f3a-4fff-ae6d-c9dacb06e552 · outbound

This paper cites Egocentric-100k, 2025.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Egocentric-100k, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.671801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:6a17b4de0f9c7ddfcca32ac3a9336bdd6926231b2420ad1bc12cb370198e72b7

Observation 0bbf8ae7-e72c-46b3-8cf3-7e13617476d2 · outbound

This paper cites Affordances from human videos as a versatile representation for robotics.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Affordances from human videos as a versatile representation for robotics

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.664883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:aae5ec5b847e00d109c524b044adc7fe0dda1d0f67447b9d662f09a3dcfe7b08

Observation 3df0448e-7df1-4993-b5db-1d90f923dae2 · outbound

This paper cites Hot3d: Hand and object tracking in 3d from egocentric multi-view videos.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Hot3d: Hand and object tracking in 3d from egocentric multi-view videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.661120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:b985ce4a77c316a0848c5131a396e21d4248717dee9d210eb50a1a2baefacc79

Observation 7e79e161-99d4-404f-b8f2-91e9bace544c · outbound

This paper cites Gen2act: Human video generation in novel scenarios en- ables generalizable robot manipulation.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Gen2act: Human video generation in novel scenarios en- ables generalizable robot manipulation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.687973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:b74d79eb2e3e862b119cfa477495df7cd9c1844f17cbe57639196e3aaa12495f

Observation e1bd363c-1639-456f-9ae9-a333752b12a8 · outbound

This paper cites Motus: A Unified Latent Action World Model.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Motus: A Unified Latent Action World Model

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.177500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:cc62f6bed0b087cc932ed4f133408eb73b9d4eedcf82c0c67f20a1eb3af49b8b

Observation bb8bf0e5-db77-432b-9b84-dd01ba1ba51a · outbound

This paper cites H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.204443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:d91362ad882e14318388af47541d0a9f5137579887197e573f3b57ec78175283

Observation 38512e7b-48d3-4151-b281-913c2531b82d · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.207724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:9be4082a5edc50f4f559154b647d2af220bea18c701da45793c34e4b866e16d5

Observation eedcf644-ffc9-421d-be6c-e7b478a662dd · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:55:00.191452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:7ab20620061160d85f1ad76ee98921158a29ab25e03486bfd58ed8c999806398

Observation f4c6c7a4-7e1d-4452-8b76-239ee9134886 · outbound

This paper cites Scaling robot policy learning via zero-shot labeling with foundation models.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Scaling robot policy learning via zero-shot labeling with foundation models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.682644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:2bc9e67f31709a7030ae71bd04cfdc380e9742d8fe6d570d3d36868bf44a83e7

Observation 1f104c49-2d16-41d7-8e5c-ba8ee5f9b914 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.174869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:143058b2dfe6ebd239011c5b999967a53c225bf036feb34be03d0e3c8e57e377

Observation ee4dedce-2e0d-4727-88ba-fb7c5f20506a · outbound

This paper cites Affordance learn- ing from play for sample-efficient policy learning.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Affordance learn- ing from play for sample-efficient policy learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.675254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:fde64eaab47d659e11037e2df284b3da6c4b1a6cd39ff9e4d1c4794106f93b19

Observation 95238e06-8f83-4b8b-aa24-589727f1bb20 · outbound

This paper cites RT-2: Vision- language-action models transfer web knowledge to robotic control.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data RT-2: Vision- language-action models transfer web knowledge to robotic control

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.659313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:9ec55c2a9c58e495c499e58da44e1e58218cd6eb8fcafa045c161d6f5d2c3d39

Observation 86fe84fc-bc9e-4c51-8f57-b77d256c6f0a · outbound

This paper cites UniVLA: Learning to act anywhere with task-centric latent actions.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data UniVLA: Learning to act anywhere with task-centric latent actions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.693239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:6c2812823e056d4a4025b94f14bfb390544fddfd026919f108e8321fcf8d5af2

Observation 14b53126-07ae-4fe0-bd9c-04e0f117560c · outbound

This paper cites In- n-on: Scaling egocentric manipulation with in-the-wild and on-task data.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data In- n-on: Scaling egocentric manipulation with in-the-wild and on-task data

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.197710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:d81fde887f169dfbb5db7db371aa9bb2c30e8f589dd301902cac57dce932b41e

Observation 4ec35ef4-52a8-49df-bd10-24eb53b1b237 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data A Short Note on the Kinetics-700 Human Action Dataset

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.183297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:d28e26c44cfe7ab9a2abaf241a7430592baad650b55b259a885fe79c9e119fec

Observation a03b5a39-a11c-40b5-b4e8-40b8d7f9dda8 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.192350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:6062680879ee7f139b1ffeeecae78b02836e05bc1919235afed428891a5b47b8

Observation 172dd323-3d20-42be-9cab-eb70a0bdf684 · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.167016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:32405f6ff76b751641e4229525c4d78a4b74e7acb7cfb3173a1d57b3be2408ee

Observation feec6cc3-8c0f-4f6f-b7dc-cd8173b86ce7 · outbound

This paper cites VidBot: Learning gen- eralizable 3D actions from in-the-wild 2D human videos.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data VidBot: Learning gen- eralizable 3D actions from in-the-wild 2D human videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.686094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:f4313f125aee835f1624c1fc44b33c937da4be943d7bdd526ce73f902aec1105

Observation d7b2be9a-46c4-4d73-99b2-06c5560ada06 · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.201414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:324072dd8077e0ff2d03f97833cb064c1410e6cfab0fe843d0d4e20f00725619

Observation 90dbf41a-e9e8-4317-af53-269b5f9365da · outbound

This paper cites villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.193835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:1cc415f4d92865e418ca27afe8abc97e96b02ec710b528e6846c65fcb4e337d6

Observation 32eeaeb9-73b2-4fb4-a8b0-47eff7ebc859 · outbound

This paper cites Moto: Latent motion token as the bridging language for learning robot manipulation from videos.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Moto: Latent motion token as the bridging language for learning robot manipulation from videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.684391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:16fd6050a321a0af0cf5076a7e8d080364b5e84b9ae3a3dff16de0c41754775e

Observation 62cf13e8-28a2-4065-b37f-c431084ea8df · outbound

This paper cites The EPIC- KITCHENS dataset: Collection, challenges and baselines.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data The EPIC- KITCHENS dataset: Collection, challenges and baselines

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.679139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:7d80c5c67513d9b80a42b40c4b405f5d278ee65fe4fa61a5067eb49ea62f17c7

Observation b4b0f713-14dc-4854-a954-937cfce6f683 · outbound

This paper cites Tam- ing transformers for high-resolution image synthesis.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Tam- ing transformers for high-resolution image synthesis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.691478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:033ade7f17fe0f0b8ced9b74e584e8f84bc74643427fc043a11f2a1ac640db4f

Observation 83a00802-ea26-40a3-8241-196af6cdf826 · outbound

This paper cites Arctic: A dataset for dexterous bimanual hand-object manipulation.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Arctic: A dataset for dexterous bimanual hand-object manipulation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.680903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:d69e89616948192ed1967f6a2926df7e08a92b5e8a1863f56418b3b1a611e068

Observation 8ea4912a-e097-43de-a47d-94ac4085f9f4 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.172475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:eb0fb94c04d76cefc6f06e370496102e786ddd6097df336aa9f442191a39fda2

Observation 87c53b22-5c50-4df3-9ea4-164fc839f0cf · outbound

This paper cites Learning la- tent action world models in the wild, 2026.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Learning la- tent action world models in the wild, 2026

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.670126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:fd0c37dc065820418e234ab5427f4dcef493ac4c0b41bf19c9b5b8b823162ca7

Observation 76b32393-54fb-4001-80b6-d48d4d3fbdba · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data The ”something something” video database for learning and evaluating visual common sense

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.663964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:a6a2c9d66a0956c2e16ee93446abbbb389a9adb4aaf4bc556d256b1edcbd9947

Observation d988bbc4-8cf9-4c95-b5bc-8aa9c4b2336b · outbound

This paper cites Ego4D: Around the world in 3,000 hours of egocentric video.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Ego4D: Around the world in 3,000 hours of egocentric video

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.666666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:d60f7b4fa36b7c44b6e2a93687af54e9b3fdfc8301aaa1367183f37d062ed471

Observation 5367fe3e-02a7-4d90-932d-fd08659b0f57 · outbound

This paper cites Ego-Exo4D: Understanding skilled human activity from first-and third- person perspectives.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Ego-Exo4D: Understanding skilled human activity from first-and third- person perspectives

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.689662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:82d2bbf1ef63d6adb26d29d76491b8939ec308df4fd08c82780fcdf3a874cce7

Observation 0e31e212-597f-4f5a-86a1-d13710801c66 · outbound

This paper cites Lelan: Learn- ing a language-conditioned navigation policy from in-the- wild videos.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Lelan: Learn- ing a language-conditioned navigation policy from in-the- wild videos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.677342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:0cd07d89c738576c95d99ef4b002b5859f06c095d61e1825e12d75a30cc421e7

Observation 144c5b1b-61f8-4315-a9b9-0348280ee8aa · outbound

This paper cites EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.169609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:1006867825f404105fde5f9fb53d7b348a35c753578befebd87d5798d8fe4516

Observation a5fdffa2-0e6c-4b55-be1e-d01eca5edba0 · outbound

This paper cites Video prediction pol- icy: A generalist robot policy with predictive visual repre- sentations.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Video prediction pol- icy: A generalist robot policy with predictive visual repre- sentations

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.668385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:9abfaaf0bbe50a2fac29eea8cdf54da42059a10aa4bdf77a613a179baa3e2f3e

Observation 329018f2-72cd-492c-bd47-e69180bbaa95 · outbound

This paper cites GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.196583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:91891f94f866323573a22878e6be7393aa0537154beafb0184de50d553b98cae

Observation 153d1cf6-1429-4da8-96bc-f836a05f04d8 · outbound

This paper cites Emergence of human to robot transfer in vision-language-action models.arXiv preprint arXiv:2512.22414.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Emergence of human to robot transfer in vision-language-action models.arXiv preprint arXiv:2512.22414

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.188954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:24c9dd162033da018681eacbe23ee688655f02a2b962ed5b71151db3f39ace77

Observation 18fe7038-584f-4da5-b998-49cb2650963a · outbound

This paper cites Droid: A large-scale in-the-wild robot manipulation dataset.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Droid: A large-scale in-the-wild robot manipulation dataset

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.662834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:bb694fda8587532fbe6e25beeb1b89e4e8706aa31dd40801edb414edbde5ab6c

Observation 43627b06-b2db-4ab3-8ca3-ead90c48b194 · outbound

This paper cites OpenVLA: An open- source vision-language-action model.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data OpenVLA: An open- source vision-language-action model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.642727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:5a462dab4b86ad0eb2b06f0cb641b711006cca1a303dd4c32618510f50374b6b

Observation 65c30495-6782-4bdd-8129-f91b8ab546be · outbound

This paper cites Masquerade: Learn- ing from in-the-wild human videos using data-editing.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Masquerade: Learn- ing from in-the-wild human videos using data-editing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.655220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:f8f311351a8b468562ccedf1bbc9be1a1706947d7148aefffd344ece4a01807a

Observation d92a5420-96c3-4473-a06b-3bbdacc18c42 · outbound

This paper cites In the eye of the beholder: Gaze and actions in first person video.IEEE transactions on pattern analysis and machine intelligence, 45(6):6731–6747, 2021.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data In the eye of the beholder: Gaze and actions in first person video.IEEE transactions on pattern analysis and machine intelligence, 45(6):6731–6747, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.651685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:6ad584e2b6a8ce9364389fc9f8abc976329340e8cfbbfb76989b475fbbd5f890

Observation 5f64c9d5-a79b-4fe2-b256-fcf315daf888 · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.212811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:b813451dbf2372b7dbb62739a6f7e1555caae48e0a2c9135877b8be9f94ddb06

Observation 71edcfae-2d9c-46a5-be2a-735e8fae973f · outbound

This paper cites Evaluating real-world robot manipulation policies in simulation.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Evaluating real-world robot manipulation policies in simulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.620159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:b2c9f0fcde05613318e8571006ccd786a5e3029f4c0c47961a08de0c5bef0c4a

Observation 606c1874-ca19-4553-bc0f-a625873fe92c · outbound

This paper cites Scalable vision-language-action model pretraining for robotic manipulation with real-life human activity videos.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Scalable vision-language-action model pretraining for robotic manipulation with real-life human activity videos

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.218035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:c2e712bc033f1ba0935292cb052163eacda7be300d34fc00bb508dee57a4e236

Observation d4411059-2c4c-422e-a189-26b0855789ed · outbound

This paper cites HOI4D: A 4d egocentric dataset for category-level human-object interaction.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data HOI4D: A 4d egocentric dataset for category-level human-object interaction

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.616732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:2af0724c9d892bf46060e432850f6af2e483d0e66452123f64f95d807077c21b

Observation 9f39635a-6bca-46dd-89bf-4475e76d9042 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.NIPS, 36:44776–44791, 2023.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Libero: Benchmarking knowledge transfer for lifelong robot learning.NIPS, 36:44776–44791, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.646214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:503e3b959b772f549f29c14d1433a50fca787b5f2710b04a4d2613c35cc4fe78

Observation 6f569bf4-4816-4589-8f13-bf0b7e26407a · outbound

This paper cites Taco: Bench- marking generalizable bimanual tool-action-object under- standing.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Taco: Bench- marking generalizable bimanual tool-action-object under- standing

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.644426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:edc3d3c42c2fa4d0d653c3971522c35588d3be9199d926149af583fbf76af8c1

Observation 7ab91d45-1471-4c74-b056-ea02fcfc11e3 · outbound

This paper cites Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.210434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:fbfd456f68dc0ba00117ac09b7f0e74c943c817f2703f7803eb1cfb74642a601

Observation 2b81a492-ca1b-4c7d-ae4d-a3ea150ec2fa · outbound

This paper cites VIP: Towards universal visual reward and representation via value-implicit pre-training.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data VIP: Towards universal visual reward and representation via value-implicit pre-training

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.637109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:4b689a5a61634d52a7f78a24136d4389aae69c08d39c080dad57e224015a0e91

Observation 8a7e82cc-b8b7-4d15-97e7-03f55b538b76 · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.618333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:8cedc3e9e59823d8c198ce91585aae1edfd632ffd0529bd4e45cdf9836ae16f0

Observation fb5629d5-c291-436e-aa8b-2da65eabe43d · outbound

This paper cites Grounding lan- guage with visual affordances over unstructured data.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Grounding lan- guage with visual affordances over unstructured data

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.603084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:f5495b919fc7cdf9382b28c40ed6bb683fc2b84a3f33afd8ba6537eccad9e273

Observation 15e68163-16c9-4640-bd48-26bee433dff8 · outbound

This paper cites HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.633187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:56d244f0fa7d0b6151db3cc1cc7eb4b62c14ae460e2bca1ddc587e9caa9f6ad3

Observation efd18abb-8f9c-4d42-addd-9b48a6fb8cf0 · outbound

This paper cites R3M: A univer- sal visual representation for robot manipulation.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data R3M: A univer- sal visual representation for robot manipulation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.599782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:b1e8f28191ee16597eef81437e009181e2b895cc60eca3459527904eff6f3038

Observation bd7696d7-53d8-4403-b85e-56e7a7dda7e5 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data DINOv2: Learning Robust Visual Features without Supervision

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.234166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:31200cb59284909c546c8a556447560b7dc6a24f9c4c7915424487a4b518dfd8

Observation 9cdec0d1-4af6-41f3-a151-b0ad7b47cae4 · outbound

This paper cites Open X- Embodiment: Robotic learning datasets and RT-X models.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Open X- Embodiment: Robotic learning datasets and RT-X models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.603472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:0d0a974ac054e3ca32e48c4b7dcd4f2fbed14841fa05e12f17042574a06ba7c6

Observation 5f4d637b-0e20-42ec-8173-c24cfd2c2295 · outbound

This paper cites mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.239698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:a12251fe3dd70902b3ce043f8115720d749bcfa0bdf4a17c6c8018118515e3c6

Observation 953c51a5-78dc-40b5-90ac-62856ac1e7e7 · outbound

This paper cites Dexmv: Imitation learning for dexterous manipulation from human videos.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Dexmv: Imitation learning for dexterous manipulation from human videos

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.631459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:e3ef0cb860e9168e072ea5ef3a8d7c65bb563ac76c4e406b3fb66ecc84fcca6c

Observation 389845df-7d91-47f9-b581-86ad9aeaaf7b · outbound

This paper cites Embodied hands: Modeling and capturing hands and bodies together.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Embodied hands: Modeling and capturing hands and bodies together

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.635013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:cec9e40f8ce600eacbe6137bfbe4916df3f8cafd4a0016980e8fb9ed3de7dc7a

Observation 576821bc-7506-499f-9199-90fb03a22849 · outbound

This paper cites Vipra: Video prediction for robot actions.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Vipra: Video prediction for robot actions

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.215340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:39b8ccf6528d5aba88eac0d02dc03c9c64fec6206005f9a34334473b559ca49c

Observation c399ca70-e946-41b2-9cb4-8b5424075a9d · outbound

This paper cites Assem- bly101: A large-scale multi-view video dataset for under- standing procedural activities.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Assem- bly101: A large-scale multi-view video dataset for under- standing procedural activities

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.638947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:c325dd061d9f976898e3dc0f556f4ae4ca482ee85c9a6661506480d670c53e77

Observation 0f68176b-c9b9-4b73-8094-0f2688a036bd · outbound

This paper cites Wave humanoid robot.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Wave humanoid robot

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.605085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:d6901f6edaa534da88ed58f7aaa483ed610bf2758c3d453ea3e4a3efb4b41aa7

Observation 05dd8d5a-264b-4ebf-a384-e86bce466c59 · outbound

This paper cites Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.226075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:f6ac47364227eccf4196f27a54733c975f2d868f241cadfeafc9119935957ecc

Observation b32ba99c-c08b-4f86-a544-776805df94ba · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Gemini Robotics: Bringing AI into the Physical World

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.223143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:6a6ec53cb9c818b8646560a369d61c17cafdfff0c20f211fba357bf0054404f8

Observation 3046b3f0-d022-4e41-b18a-0d4a94a9b3db · outbound

This paper cites Tesla ai day 2022.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Tesla ai day 2022

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.619919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:4011be60fd6a72c11cc83be76218090b0501a4e1678cafb3c200f0d21eec3462

Observation 11f20381-9bfb-4d21-b363-8f8c9d18d0c5 · outbound

This paper cites Neural dis- crete representation learning.NIPS, 30, 2017.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Neural dis- crete representation learning.NIPS, 30, 2017

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.640829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:8281f36d4071d194ce21a680a983b3a848cef2cad17f2b7d37d7d4ac98a77388

Observation fc218ccf-2419-4487-b341-e9f17f755bed · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Bridgedata v2: A dataset for robot learning at scale

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.647937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:aed9992f9e98665c75ff6fa7e054aecb193d460d6671de8e713fc544b2fffc01

Observation 18604bc6-b6c5-45bd-9c41-54b9459c7c76 · outbound

This paper cites Holoas- sist: an egocentric human interaction dataset for interac- tive ai assistants in the real world.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Holoas- sist: an egocentric human interaction dataset for interac- tive ai assistants in the real world

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.653436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:c348774d9ba783c0093d766d0f7418a6dd776a55f28d9fd0aa3ca79958cf25b8

Observation 2857ae1a-cf8c-47b7-a62a-84c5c670841e · outbound

This paper cites Gensim: Generating robotic simulation tasks via large language models, 2024.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Gensim: Generating robotic simulation tasks via large language models, 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.629643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:0d18bb1d1f67ed90909f6a00a5fabdf6f31070bdae33bc23b1ea9d65b08c095f

Observation f391e4a6-3a55-476c-a21e-0515ee4d95db · outbound

This paper cites Any-point trajectory modeling for policy learning.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Any-point trajectory modeling for policy learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.618477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:8089d0f9fac6950df9a51be6e52db231ea61733cc6f6afd1be6d9bfd45000bcc

Observation 3118baa4-8fbf-4b22-bda3-f1967bc205ca · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipula- tion.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Unleashing large-scale video generative pre-training for visual robot manipula- tion

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.625882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:0e9e7f6d8e82b4dc7f4e0c91050bf67f1e80b4951285d90db8764f0fdc4b2345

Observation 27d7170c-9200-4b90-8d4d-b8a25be6e0ae · outbound

This paper cites Masked Visual Pre-training for Motor Control.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Masked Visual Pre-training for Motor Control

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.220581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:eda8a4fee23533e1e2584f9ab81a1484db932ddbb85cd390b515d109b8f8bc41

Observation 269bca05-ac26-4422-9c6b-98dbf958e131 · outbound

This paper cites A0: An affordance- aware hierarchical model for general robotic manipulation.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data A0: An affordance- aware hierarchical model for general robotic manipulation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.623556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:e1962b5de300036768e95f7dbd776d65603d2a4f9ab11e302190cf2d35aaf4dc

Observation 1a707b6e-a375-4e66-b430-47f35023bdd6 · outbound

This paper cites Magma: A foundation model for multimodal AI agents.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Magma: A foundation model for multimodal AI agents

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.627765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:0342f233bf2b420405f0fc7f58540ff13148c9684df2b4324258473daf77922d

Observation b95a5978-599b-4eb8-a88c-ceba7b7e504b · outbound

This paper cites EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.228649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:73f74a0b6dc0db66c0ebf124c3736c25f421a9402bf59b7af1e1452e28b4b82c

Observation db5b6368-63f2-49b6-b3d7-e41a8df41f3f · outbound

This paper cites Latent action pretrain- ing from videos.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Latent action pretrain- ing from videos

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.629913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:a274e43b6e06183984f9921238017e32c16903b17137459b271c3e32c6593a36

Observation 9562be4b-65d1-4414-9562-a570989a887e · outbound

This paper cites Yoshida, S.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Yoshida, S

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:55:00.237030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:edbc10093c9d67b121dddaa0833f2edcd643a2841aef82d94daac2885689e8ef

Observation bfa07efb-88b5-44d9-8e6d-e3e2f65905a8 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.231722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:0ad35dfed44cc0895535a1e6c196d94ba82f51ddffa5793d132f5b83cbbf31fd

Observation 6fdb062d-9680-4156-89ef-993a3362db72 · outbound

This paper cites Motiontrans: Human vr data enable motion-level learning for robotic manipulation policies.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Motiontrans: Human vr data enable motion-level learning for robotic manipulation policies

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.614821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:d4b169bb00b60902ccd2ab8263dcdbf1f73dd993b465d26b609c59b70966ff56

Observation fbcd2d81-cb21-4130-8d32-5b86d58cd5c9 · outbound

This paper cites Hermes: Human- to-robot embodied learning from multi-source motion data for mobile dexterous manipulation.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Hermes: Human- to-robot embodied learning from multi-source motion data for mobile dexterous manipulation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.649852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:159b7a35af01488756cfcec7622e9ba1913860d582f349be27ad66b2b276a40f

Observation 6caa56c7-1f6d-4819-999b-0ee45f8db512 · outbound

This paper cites Oakink2: A dataset of bimanual hands-object manipulation in complex task com- pletion.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Oakink2: A dataset of bimanual hands-object manipulation in complex task com- pletion

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.613032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:efb4a0c6417cd19d251a7a4add07aed708d58d82f40d41fc833dde1d5ccc2c1e

Observation 08e0b236-24c0-4550-9dbd-52fbc22692dd · outbound

This paper cites Clap: Contrastive la- tent action pretraining for learning vision-language-action models from human videos, 2026.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Clap: Contrastive la- tent action pretraining for learning vision-language-action models from human videos, 2026

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.621908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:477d9fde34d6bf2e025d75ad2f52f5249c5f899580afe0ca27d2ae69901ad5a3

Observation 82709cb9-2929-4f90-9aa9-52c33c84bc01 · outbound

This paper cites FLARE: Robot learning with implicit world modeling.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data FLARE: Robot learning with implicit world modeling

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.623717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:a87770d17785bed1c3a9898aecc0d936d38cdae21ca35afa3f12cd6acfaa1939

Observation 4cfc41cf-317e-435f-996b-5b2f28bcbf23 · outbound

This paper cites Zhu, Pranav Kuppili, et al.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Zhu, Pranav Kuppili, et al

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:34:27.631645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:459c96d934acc6125df032f4da0863c852e141ec7357d5f2f9747625c8b512cf

Pith citing papers

Observation f856e00c-ec8c-4ffe-a690-6a5d236da996 · inbound

CORE: Common Outcome Regularities from Action-Free Visual Demonstrations for Robot Manipulation cites this paper.

CORE: Common Outcome Regularities from Action-Free Visual Demonstrations for Robot Manipulation From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:04:20.748949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T07:04:08.886475Z digest=sha256:f3266058f6e9cbc3422a3b8e65890f79d03a361f01806d3f915df9ce9f7bcee5