Pith. sign in

Paper Citation Record · LEDGER

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

As of 7 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 72 inbound Pith citation observations for arXiv:2507.12440.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12440 v3

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T04:32:58.733165Z

measured 155 of 155 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 72 of 72 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:20:20.579410Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

83 of 83 outbound references displayed

  • verified exact18
  • verified fuzzy41
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 0ed6bc88-415e-4e99-aafd-20845899eef2 · outbound

This paper cites Vuong, S.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Vuong, S

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.930156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:9ca0c0398ad8e8e390d4ba1be04f9f83b2489d6e666108764b13906797e5e679

Observation 9c44f7f5-a543-4b4c-a5a6-bddec513fd05 · outbound

This paper cites Khazatsky, K.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Khazatsky, K

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.933745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:fe9301004985440b433c653ff6f112f3222804e990b438d3f30380fa123febd0

Observation c93141f6-9c67-4065-9b40-fefb354c28be · outbound

This paper cites OPEN TEACH: A Versatile Teleoperation System for Robotic Manipulation.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos OPEN TEACH: A Versatile Teleoperation System for Robotic Manipulation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.791730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:053b480ac2993f73b4fa1d21afcc4ab82955942d36deb1c5f8c77cba6ebde0a6

Observation 8b923843-f2b7-45a8-87c7-0eba72127999 · outbound

This paper cites TeleMoMa: A Modular and Versatile Teleoperation System for Mobile Manipulation.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos TeleMoMa: A Modular and Versatile Teleoperation System for Mobile Manipulation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.788177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:b996ffefaabc4989ca929b0fc338860be82b40adb22cd61a2aaed2eaa09702e4

Observation a640a85a-966f-49b3-b0a0-79ee35a61d48 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.935393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:027470c78a5e1f4e5e233170c577e2f72ee0ceef04018c48db79ed055984db29

Observation aab5372b-9d8d-49ba-a048-16f4e16b30bf · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.936997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:0af98965c6e6c1516825ff14a6cf5ec2cf285426386e6ba3a9d741aca9a9d6f5

Observation 7b4ddec7-c528-4651-8ff7-608598c44a8a · outbound

This paper cites Fang, H.-S.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Fang, H.-S

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.938826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:ce9a2d38f0d9e432d4cad86812b125a29263a0b132bd02dc4f6df42bad3d1212

Observation ada99f2f-06e6-4937-9026-67a493d30ded · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.940601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:755b288e71bd5dd6ba7dbff29b463ebb3df5082f89b2bff08a7f4b6550518c0e

Observation 1bae9bd2-4dd4-49de-997e-52e812636c92 · outbound

This paper cites Naceri, D.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Naceri, D

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.942466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:658054716f3471322b9a7d564ad9af19e180b69776b3a4813513ca2a00c06d69

Observation 9032e7d9-70bc-48af-be0e-d67cbcff8383 · outbound

This paper cites Open-TeleVision: Teleoperation with Immersive Active Visual Feedback.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Open-TeleVision: Teleoperation with Immersive Active Visual Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.795129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:6ea9e94200a010edd80f6c9a3b2af4628d2325c9dc724e871a1382d7c1c4e835

Observation 23390a69-66f6-41a6-931b-fa9e32351cba · outbound

This paper cites Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.816737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:16881c9cc2a583592a54bd9a85a122ea27e788dcab44beba5d6a98ff025a9f1c

Observation d32eb02a-43ed-4527-b766-9bf74f0ff7f4 · outbound

This paper cites Gaze-guided hand-object interaction synthesis: Benchmark and method.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Gaze-guided hand-object interaction synthesis: Benchmark and method

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.820398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:5ef42f2c9f5bcf757bf15b78013d2514889da29a26ced1b2afe3a37bacfa8298

Observation 0ee0858d-8750-4c2c-85e8-21a43172222c · outbound

This paper cites Ghosh, H.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Ghosh, H

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.944070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:6db410ff24f84aa6dc4cc77d0241e7d0560765975793d0dd88d82dd1e678c2db

Observation d76022e1-97c2-4756-b9e5-95ff04cc5f06 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.945844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:2871311470052fa0e0370a198ae59e712f53ab4f86a985f69bbc3892f9fdd257

Observation e5db4faa-709d-4796-8112-becec2c57d6a · outbound

This paper cites Black, N.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Black, N

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.947603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:28e87c45781f529578c3c39acfbacd708a7bbc26345bbfbab65d877a7c9f6e76

Observation 1573cea1-701d-47a8-ac5d-8790141f15b4 · outbound

This paper cites Brohan, N.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Brohan, N

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.949241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:8099f03212240c718b500a2901497e34f04b5403d5a23c6f9ea29bd40d1ee25e

Observation 3b146c72-5b74-48c8-bed7-50e10250e9fd · outbound

This paper cites Isaac sim: Advanced simulation for robotics development.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Isaac sim: Advanced simulation for robotics development

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.951081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:92fd2e5a234a5e926c904c9cc1709d1d770604209506642f14ae626fcfba41b9

Observation bdcdb99b-d1c1-497a-86e9-fd906deb400b · outbound

This paper cites Romero, D.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Romero, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.952952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:c0b5f4f0afcffcf500a1807e77a7d83617b51d021a9d0bcf1fe8cd07ef8a309f

Observation 9539cc4c-6c4c-4019-a970-da5e8b7ff12b · outbound

This paper cites Rodriguez, M.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Rodriguez, M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.954859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:c4972d29af0feff6757ed15cafeb99119afe73269afcad40bf724df64ea98c5c

Observation 7dfe254e-eabd-4e75-a992-e1a46506b1f6 · outbound

This paper cites Rosales, R.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Rosales, R

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.956752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:fb53be2af1c545a3932b3121a359fc1eca015904abc8545026c2b7f229e8c3c0

Observation f871a194-e135-4782-8743-cc76b36d39f5 · outbound

This paper cites Prattichizzo, M.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Prattichizzo, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.958403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:8b396cd86d4ece7ea62745d4a95db98b9eb28f8fe711495d99809c2cdc82f62f

Observation 2b979ad0-6171-4519-abce-ddcb2843364b · outbound

This paper cites Ponce, S.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Ponce, S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.960051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:c4de0fe805d6754d582c843997806f441f04d7c391f1162cf922d0f33dcb7145

Observation a79c1865-d90a-44e4-ac4a-9e723cbd5e2c · outbound

This paper cites Ponce, S.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Ponce, S

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.961753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:409b8182be4a60e036ac45300b83bfd4f9cd74d3d82a07140b6e86e5569710f7

Observation 811a25d3-444d-4074-8465-009786708e4f · outbound

This paper cites Zheng and C.-M.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Zheng and C.-M

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.833025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:f4ec14f282c6b3dbc9fad8a25de8270d16df4feeab7f9277a9c781c4951a7097

Observation e7ad1e08-c8bf-49f2-8c5e-ae013f93e608 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.835302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:e598c9ed43f60332f28ef18b24c8fbd5abb6cfe1e6c95926b839e9b7a5e7f2b3

Observation 038af182-6a54-4ba4-a5e9-e0af5dbf2429 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.838469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:c0f600395353f09483d1e166a2ed1c8a83e7241a99c7fb035e8a90f54d5c612f

Observation 1238de21-5c5e-4e1e-86db-2dd1c7005122 · outbound

This paper cites Nagabandi, K.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Nagabandi, K

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.840439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:f489fe021620886898558060651039f184ae73edaaa28d9ef9a5bf2a818d302e

Observation 8a6a568e-ff5e-48ec-8e50-196e0d7136ca · outbound

This paper cites Jiang, S.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Jiang, S

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.842310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:1a9170eaed5de59b4de43157b0cc92ae8de7fa90c9d24368dc7541f44435b447

Observation 36b87ba5-ec61-4175-a493-59555a067100 · outbound

This paper cites Corona, A.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Corona, A

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.844244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:2196b7e64f27b7167e16adb3f336d0522fe30df985d25a22617c4303d39f624c

Observation 1e57fd03-f052-41d9-a27e-db9a4e676531 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.846223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:7e8351f8352f03b52a3b87b9f936596d3296f5e85a2dc0a2d01766159608956a

Observation 613eea68-9074-4920-a4be-7060e7de6a3c · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.848122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:cb261cf7d8637c9660753285ab4aeab1e5d7d3984301d01a06d8f90c4ccc3de1

Observation 8f0b2e07-4e1b-43aa-bdd2-04947160a258 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.850149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:7553f29be77dbca1e3505dc263f4e306e0ef38ff0fd4b017fe1ae6b5f043954b

Observation 98748107-6853-49e5-b002-d5e8b398b198 · outbound

This paper cites Brahmbhatt, A.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Brahmbhatt, A

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.851971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:76ac108d8294cc68ccb387f509c12961ab44cc8264f333e4385c8e35817961b4

Observation 9e2bc1f8-2562-49bf-bf0f-53cc8645c3e2 · outbound

This paper cites Turpin, L.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Turpin, L

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.853991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:072038eaa93b5c8972d77624685583bf4045bf6b5657682b44af7d22a6239b36

Observation 35db358e-6c79-48ff-9af3-c40788dbeff2 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.856005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:7a1da9ef6f8ae9840bfd941552d76b352d7023f4448aa5efc0f7e8192debc723

Observation 0baa12ed-e81e-4785-9fb6-4e04cce48a93 · outbound

This paper cites Yoon, Ryan Hoque, Lars Paulsen, Ge Yang, Jian Zhang, Sha Yi, Guanya Shi, and Xiaolong Wang.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Yoon, Ryan Hoque, Lars Paulsen, Ge Yang, Jian Zhang, Sha Yi, Guanya Shi, and Xiaolong Wang

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.799201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:ed611a5922d7c039a0ebda9bc5cc4438261c44faae3e4f281e3f4828dd97a273

Observation 7f6ecd0f-6c98-428b-a03e-ef603ccb88c4 · outbound

This paper cites EgoMimic: Scaling Imitation Learning via Egocentric Video.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos EgoMimic: Scaling Imitation Learning via Egocentric Video

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.802658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:792a3aaced15a0899127beed49dfd8739640f54c1b7475a8b2f27cb2c048f48e

Observation b31a7f52-4eb5-4ff9-bab5-1545dae98eff · outbound

This paper cites Achiam, S.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Achiam, S

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.857851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:8331cc84626ba4aa923342cfdfed2c5f300964d28f615279a32a251c0a9efe34

Observation 865a6579-9aae-404e-961d-abe09e76db61 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.862966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:ba5b1354d2eb114f4ec7d61ac7d48f74ebda20c8d89a2729a47503e16c3a9f33

Observation 16ae8019-65d2-430c-9454-ff145993fb34 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.865027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:8b94dafe89ba547839988f015289b3dc5af92e9f9eb189421f75fa47485d6749

Observation 9b81d31d-b7ad-44c8-abe3-376da3b75bc4 · outbound

This paper cites Pratt, I.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Pratt, I

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.867350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:e51e1f8d7a9e06b41a52f31c36ef24e1b299899f9af6a95ec395b13b5a9cc67c

Observation e4eaae4f-6425-4526-9d66-077d003cf108 · outbound

This paper cites Alaluf, E.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Alaluf, E

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.869525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:2c2c278d3834ed3b58612edce6d6ec3e97211e50cb9783e7a4e3bdd97ab53031

Observation eab2daca-4438-4827-afb3-4ff0a3abaca6 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.871297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:dd6078c05dd3a5d8c59569cbbc0044bfa0e4792b13573693950674c193cfdfe1

Observation b3bab134-7c4d-4926-918e-e615bab6e6e1 · outbound

This paper cites Huang, S.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Huang, S

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.875025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:06e2529f26302324a3ca587b2a1abc4cecfacfa72467e58bc27c0792705f283b

Observation 7e40dc2c-45d6-4952-be4a-661bde1a0830 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.876850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:d8faf729d06de713e258a7c001d348926e64a18b4584d8f27c7a6aa5daa135b0

Observation 25767cf3-3856-497e-9de7-3690fe020f14 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.878581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:3ad30fd6e59c0520cec13bf27b35ef7f0addcdfbc119453d4a0c5ba5697b7c16

Observation c250a78b-7295-4aad-a829-6045709cac23 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:32:58.770684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:a18cfe8609b452bf3627d09a227c0d7ccd83d6743a994d50f84adb4d43c73eb0

Observation 7dba06be-5bb9-4ac7-8204-ee09a4721e96 · outbound

This paper cites Mandlekar, Y.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Mandlekar, Y

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.880503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:de8ee7abae884177406bb90a114562d1eab604df3588cef151055d6971ad05d5

Observation d125605d-8cfc-48de-a49d-ff9bff0b9ac2 · outbound

This paper cites Mandlekar, S.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Mandlekar, S

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.882214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:63e6c6dd13de75d582984eb6b8bbd08c44b003b7c6c2755f8a34ab7c329409d0

Observation 00f86eea-1d94-48df-bc8a-7c31ac50b801 · outbound

This paper cites Dasari, F.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Dasari, F

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.883934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:fc610ff5195a29ff6e50af59b1c20aa233f0e286da8b4b2261161c83f5e72132

Observation fb0f664c-479b-41ba-a2e6-643f98d96d53 · outbound

This paper cites Kalashnikov, A.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Kalashnikov, A

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.885838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:c14225056db2d86ac37e597fccd0d688f35f1c96da5e9b68c2f993f9558d749f

Observation c9a26a40-af76-4340-a70a-fc35b26658f6 · outbound

This paper cites Damen, H.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Damen, H

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.887833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:619d9bfcdf5be5fccc5392864dbe396566f600b4c237393a81b82291eca720d4

Observation 1ef81e9b-4bf1-4f77-9598-24af79387c88 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.889544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:315cb983ee9dd50b8ac08a64d961934089a595009cec34c086ad7e2a22dca5d7

Observation b0bba765-2663-4f9a-8912-f94e8ad932a4 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.891383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:722e0391f58a1216140e7e160968d6d59e9e241b2dc9e8d71b63e1f56625708b

Observation c2894ad7-535c-48ec-8c1d-ada1371d6031 · outbound

This paper cites Grauman, A.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Grauman, A

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.893655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:03b1b5530df4bca77bef126a2c7df375ed93f0c8c7ec084a49563b2e710b0d2f

Observation 276e07fc-29d1-405e-b6fe-7d69ceadd9fe · outbound

This paper cites Grauman, A.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Grauman, A

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.895517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:c930bdcbad4e180f95154de7dadf61c192068bb1f9d2bba9743484847248da24

Observation 4eaf3de1-56dd-4111-82f6-250aef88e7a4 · outbound

This paper cites Mahdisoltani, G.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Mahdisoltani, G

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.897184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:068cb7c54cfffa52e34a818cf604d67beaf2aef2dc2d7c718a8e780b87c82550

Observation 3e26c793-cd4e-4c95-9fa6-29d38e2a2e54 · outbound

This paper cites Damen, H.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Damen, H

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.898988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:82d5b519381ecb912a6659175ed25a1e106d1f07094565cc1b71f76bfd628b41

Observation bde0bcb1-1ce9-4257-99b0-554f30c9b210 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.902881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:dfd497c3e61c2eca386006347af33d664cd1ef9f4ca3b41e7c2f0811fb976ec8

Observation 709e3014-4166-4e07-9d2d-b58736559dc5 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.904780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:1b924ff0b6e5aa226eefc1b589ee2a5a7f1b80cdbf1bbe7e44ad37b230917ff3

Observation f91308d9-9a25-4f82-9dfd-7b17254ed7ee · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos R3M: A Universal Visual Representation for Robot Manipulation

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:32:58.805832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:4c777ac7db73f075379636c0bbed0ad84843217392cabe00701746b4b7db7ec8

Observation f2633556-5944-414a-97cd-90aa1509c3bd · outbound

This paper cites Majumdar, K.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Majumdar, K

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.906632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:beb3ad1cadf82ab3c2a76d07a9f3e76a20b9552f4fcc2b497d7058d3eb17d88d

Observation 17691082-9687-4709-a809-a9494f6760e0 · outbound

This paper cites Karamcheti, S.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Karamcheti, S

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.908375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:9916dad2ce01e38589926c682c8cd377fde4239333ad215b75b7ecfa5f576adc

Observation 6b35b50a-abcc-4d3d-8c11-6daeeeac83e6 · outbound

This paper cites Spatiotemporal Predictive Pre-training for Robotic Motor Control.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Spatiotemporal Predictive Pre-training for Robotic Motor Control

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.823660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:17838295cc338b1d681ae948cf93c3115675d09f0457a7cd8247eb4214595bfa

Observation dd9b74d3-11c6-42be-86e6-da0f5e495af5 · outbound

This paper cites Learning Manipulation by Predicting Interaction.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Learning Manipulation by Predicting Interaction

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.830323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:3424683ab3c567f51dd91372e958d6c1acb8b30e8dfc484100c8703e82163a2d

Observation 162c26cd-a0a7-4313-83cb-fec5c051812f · outbound

This paper cites Latent Action Pretraining from Videos.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Latent Action Pretraining from Videos

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:32:58.826620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:ffcf4ba3ffcdd4d33a3aac3c2712e5618246b368ca40ef1f8ab1ed1af6f86485

Observation a3e33c1f-b5b8-42cd-b82c-98ab774ed1f7 · outbound

This paper cites Lirui, C.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Lirui, C

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.910710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:c8626a13c6a9b9ad4b850eb1f5116f7a67e2f0e716ec9bbc3a0688b941619654

Observation e3b33f33-822a-4f75-b31c-5aaa08be4c8f · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos NVILA: Efficient Frontier Visual Language Models

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:32:58.776597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:5bb56a60a6e072465838ff5fc76a52e682e716cf73391c8d866b0bdd1f6f11e8

Observation 9ffaaa22-2da9-45f5-a5df-aee42c07c9b0 · outbound

This paper cites On the Continuity of Rotation Representations in Neural Networks.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos On the Continuity of Rotation Representations in Neural Networks

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.780219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:bd0c267c4b2acd3f9d669263c7c12dfb5853544c04ee63c7804977be22b79b1d

Observation 7ac3faed-88bc-4171-87c8-d9726ec89d29 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:32:58.784465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:55637ee7be5102e9ab756739ffecee90637e2ceef025b111c0b08e51d1f878a7

Observation 3e4ffae3-17a8-4163-8553-271383b92846 · outbound

This paper cites Orbit: A unified simulation framework for interactive robot learning environments.IEEE Robotics and Au- tomation Letters, 8(6):3740–3747, June 2023.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Orbit: A unified simulation framework for interactive robot learning environments.IEEE Robotics and Au- tomation Letters, 8(6):3740–3747, June 2023

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.765652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:0002a550c4324655138f8b5b888159a5836b1ae2f28f4b2ce79ab34bdd261ec9

Observation 435674a3-24b8-4773-843f-ad7925994291 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.912751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:db2328336bfa9d9606916de8276f31fff7e30c25aded436b20bc979b9f88fc00

Observation e6f1e820-eb4b-4566-bfdf-602ac649fd5a · outbound

This paper cites Robotics.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Robotics

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.914497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:b7ad1d4fb7e28b825e7a461a38a4c115c8bd590e498bf6ba9e0f828714462e38

Observation 7e58ed22-46f9-4e21-b05f-e05154f06b29 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.916409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:2e36d2548eb41ab888707885ecf2b22860627d05ecfdc12c8fd461f656097ecf

Observation 658dc08a-a6d9-4897-b482-b56c658884d7 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.918483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:4a06cc9728cb69bf267c9262f62265805c90a5cd93c7d62401ddbb0d78ee7b7e

Observation 86371c3f-7f10-4e18-ba62-39d56061c571 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.920523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:6218564ab05d602dcedb68e6601ba0903d276171eb5650aba05e727ae303b970

Observation 3f450265-fcd6-4a08-b940-2dd66dbf4a27 · outbound

This paper cites an unresolved cited work.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:32:58.922162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:ff6a635b673694d8663af551bfeb2580f55091c441c72edcaf119ed3824d668c

Observation 6bd75f0f-7626-47c2-880c-6b59314dc797 · outbound

This paper cites TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.809560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:48d5016991c76051644147752e59a448de90c89677650d3d4c7fd87f56580c1c

Observation e0d8c426-7f7d-4c62-8b9b-a8c896f144c2 · outbound

This paper cites Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.812986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:66e3db60a1558e52550af6375a972b91c14e939d619cc3997513e3f4bdcbfdaa

Observation 30c213a6-8018-4c72-ae15-738af907cb88 · outbound

This paper cites Oquab, T.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Oquab, T

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.924023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:02d0339f12b844e078e94112fa16a8156e3184cdba2a4af1fb45bf322423a3ae

Observation d737640d-a9a3-4478-8500-6b90a9fd273f · outbound

This paper cites Darcet, M.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Darcet, M

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.925776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:99754bc29145bfde108db4a4cd5a490e44958f1d47798a68a5b9a376fe785fd7

Observation d2fc7b23-a0d4-44d6-b283-0acab4b130f8 · outbound

This paper cites Nvidia omniverse: A platform for virtual collaboration and real-time simulation.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos Nvidia omniverse: A platform for virtual collaboration and real-time simulation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.928341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:36a259ec64ce79cd3a1672faef7e0a003940c24369b425c649140b1f6130089f

Observation 2e1729d2-b841-4302-83cc-4fd095e8512b · outbound

This paper cites 22 14 Appendix 1 Dataset Details Language Label: The combined dataset includes ego-centric RGB visual observations, wrist poses, hand poses, and camera poses.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos 22 14 Appendix 1 Dataset Details Language Label: The combined dataset includes ego-centric RGB visual observations, wrist poses, hand poses, and camera poses

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:32:58.932077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:7b782c3c61d90d308e45deeeab867674a08a89f071e1ae0c691d64f3f2d634a6

Pith citing papers

Observation c8eb7e53-f544-443a-bdd0-da4c063eaf84 · inbound

HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation cites this paper.

HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T15:20:20.579410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:20:20.579410Z digest=sha256:4130cb441a25fc9364b6b8372c0f0956e8366bf1ed6bd3041bda5d31b0ec4893

Observation 78290916-5698-45fb-8855-3e5e8279fec2 · inbound

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration cites this paper.

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T18:50:16.771640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:50:16.771640Z digest=sha256:fb7db22efefc74d7632b6dc7bf0c435dee712a0eeb0708f20f034552038f6b19

Observation f85aafc8-57c6-4c01-9bd7-9893a48d9ece · inbound

Scaling Cross-Embodiment World Models for Dexterous Manipulation cites this paper.

Scaling Cross-Embodiment World Models for Dexterous Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T00:28:28.407653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:28:28.407653Z digest=sha256:6137f1cc54ff545e523a3f2410ce6b318715858750cfc794674185fdd3d87943

Observation ad14006c-7330-461a-b4c3-6483b931bc30 · inbound

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos cites this paper.

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T17:02:33.997887Z digest=sha256:5ce4e3650f04ab9794df3edf56ab2b7119d45069025e96f15d5885f66661fb51

Observation e6e77eb6-0f66-40f4-8c65-6b9edf9ecf2a · inbound

EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration cites this paper.

EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T01:19:59.046127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:19:59.046127Z digest=sha256:0e0f254b43d6f050becb87270c1388342c1a91fcb17d6f83636257ca6fdc3251

Observation d3442e1a-da2f-43de-8da1-61d77661401e · inbound

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion cites this paper.

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T23:57:44.698790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:57:44.698790Z digest=sha256:52efdb882c0b98418987701cf4f8dd5713d14572dc5cf82124fc8b4c5c0420bd

Observation 932f7c20-2c5a-4502-a56b-5c850330bb7d · inbound

NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning cites this paper.

NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T21:12:28.816574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:12:28.816574Z digest=sha256:15eeece12bc56ab796539041b1538699c07b0d597b43c41fd95584667beb003a

Observation 26322bfc-c4db-4a3b-833b-27c7c5b48b45 · inbound

HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations cites this paper.

HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:34:09.117855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T11:30:47.668896Z digest=sha256:c68299e84ba07c97777ac74e1a05475c41e9e98346f9404c34b05e07e44a6b2b

Observation f1c6ebd1-fa47-467f-b286-b9573dd6dcfa · inbound

AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation cites this paper.

AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-13T18:05:58.357155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:05:58.357155Z digest=sha256:9e2fbce163c6867e189ed1aa34665128b730eaec8c6f3fda75b334c79327690e

Observation aff72e32-ae8a-4f7c-b80e-ed568a4dca57 · inbound

Grasp as You Dream: Imitating Functional Grasping from Generated Human Demonstrations cites this paper.

Grasp as You Dream: Imitating Functional Grasping from Generated Human Demonstrations EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:21:30.169431Z digest=sha256:96d5014dc10e3a5eaf38c02dac3b2f974c9aca71e5d0ef5d0b468d1fe9088ffc

Observation 65dbfa02-8a97-4b70-bb3e-fe86bc17275d · inbound

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World cites this paper.

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:07:41.489995Z digest=sha256:16405fb6d566f04694cde5cf7e2d4058039d7783bc8f34800be80e187ab174b7

Observation 565988b3-96fb-4146-9501-9d899833eeea · inbound

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World cites this paper.

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T08:25:22.011013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:25:22.011013Z digest=sha256:e6c8f88bb88a5ef9fda8611c741444bafd575966ff1fd507b84cc73c2aacf878

Observation c74a357c-428d-45f6-8605-4ff72320bcfb · inbound

HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation cites this paper.

HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:44:32.243968Z digest=sha256:8cbb9493c2481cd1888ba602e3986aff17b2d865cfba7ee38ee4d05c3935c1a1

Observation 0f6b5d9a-0edc-478b-8ebd-8942f31965e1 · inbound

HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation cites this paper.

HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:49:57.294769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T09:49:34.983940Z digest=sha256:bf21298312f6d3abb5e3f6853ebbcf537bd18e33d238b5d93531315e33857c11

Observation 1e6a4e8d-c484-4181-a36a-80d8c798b6a9 · inbound

ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human Demonstration cites this paper.

ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human Demonstration EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:24:54.053447Z digest=sha256:62f33ead5266d0477b1cf56d64ae767178026c819146751333937f8e8cfb44c4

Observation 2f849ef2-1c7a-4974-be45-c6cf61e388d2 · inbound

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment cites this paper.

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:36:23.197843Z digest=sha256:262f901d8c6ac1350ee093bea4bc2fe16dadf5d50acd5e14aaade923ff65db90

Observation 5b3a441a-bf1a-4da0-905a-6e077ac80f7e · inbound

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities cites this paper.

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:42:34.409651Z digest=sha256:f4fe692f1c81834bf3d5364bc4e97f268c612bdf833170c6fd1bc721422d5c6d

Observation ab2236a3-384b-4c1e-a969-5621dad563a7 · inbound

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling cites this paper.

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:04:25.672157Z digest=sha256:abc02b64bda6b1dedd6aeb3703ea1bdbf0ee3574b84f06a805db172f2e2b494c

Observation 89107eb2-2940-4a2e-b0a6-e3085a931b34 · inbound

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors cites this paper.

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:07:24.208555Z digest=sha256:929fabf30cb14976327d226791047996295f71356364376e36ca092d21cf186a

Observation a9991ad8-1c8d-4bfb-9443-9dfe0f7e1350 · inbound

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors cites this paper.

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T18:39:49.173828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:39:49.173828Z digest=sha256:0b80704ccf7ea8fe80e9aeaac3f6663e54e95c809b4982acae41efa9116d45ca

Observation 56074592-06cd-44b0-be07-313183bcd6e6 · inbound

GazeVLA: Learning Human Intention for Robotic Manipulation cites this paper.

GazeVLA: Learning Human Intention for Robotic Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T11:37:30.784513Z digest=sha256:4053c454d493a7f7098de2b2a967bc8c5d167cbee480c110335b287e3be15d07

Observation 97742f1d-b1dc-4bc0-a5b0-282eab0d5f3c · inbound

EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks cites this paper.

EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:19:34.743194Z digest=sha256:06f97800f58ffda3cfc18b556bb5d6ec23d296f838b7ad33585d9662be75d796

Observation 33c0d231-2c86-4a9f-be89-b633977734a6 · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T02:51:27.662262Z digest=sha256:ea80bb8e1e7c51d8939e85efa390380863ff9bd7b3d0f71369027bbdd5debaec

Observation 142db0f4-95a6-4790-bbda-d1ab0ad3fb2c · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:21:29.219837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:16:58.104663Z digest=sha256:7823770f460f56d1d0b17a6408b0a28adc7cdf622689fd2393a76dbc006806c1

Observation cd341fd6-ad13-4b0b-aa5d-476c670cb09c · inbound

Being-H0.7: A Latent World-Action Model from Egocentric Videos cites this paper.

Being-H0.7: A Latent World-Action Model from Egocentric Videos EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:48:01.461993Z digest=sha256:899c70bb727334b3f745ff42c0634ae8a56e86c762129930eabfbf1509f42174

Observation 40e74a6d-e5a1-44f1-940d-46776c42825a · inbound

OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation cites this paper.

OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T06:54:49.169238Z digest=sha256:b6abe3a6175c5308d60dc1e0db9cfe211167336b60aa739d612344055cf26b89

Observation c48085dd-ca5d-4c85-b219-fdb75ce68d75 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 204

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:6917cd0a9078d09895c029f0df25c2d990d42c13ac0e7ef5be20c6defb91a538

Observation 8d727f0a-fc92-4900-b9dd-1d7d4b590b40 · inbound

SCAR: Self-Supervised Continuous Action Representation Learning cites this paper.

SCAR: Self-Supervised Continuous Action Representation Learning EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T20:38:44.286516Z digest=sha256:5f1677660ae939cc0bfb4aefc6d2d58ce4d4430b4f962e9780c783dea57e3925

Observation fa2b2816-8762-4c3c-9096-c4e43be05bc6 · inbound

EgoKit: Towards Unified Low-Cost Egocentric Data Collection with Heterogeneous Devices cites this paper.

EgoKit: Towards Unified Low-Cost Egocentric Data Collection with Heterogeneous Devices EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T21:20:56.830588Z digest=sha256:1b37727ad0bb1e75675e11da43a90722072d6c0749a2c952db326b47044bf265

Observation 0f7901b9-33cd-4cae-b2a0-1a5fa9e5c865 · inbound

StableHand: Quality-Aware Flow Matching for World-Space Dual-Hand Motion Estimation from Egocentric Video cites this paper.

StableHand: Quality-Aware Flow Matching for World-Space Dual-Hand Motion Estimation from Egocentric Video EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:11:21.610700Z digest=sha256:a91e9072caa13410d5f0a479766f4368f29d2a837c597ccddc3e74b543168062

Observation fabb0792-6596-47fa-b9dc-4b6e2b2624f8 · inbound

Dexora: Open-source VLA for High-DoF Bimanual Dexterity cites this paper.

Dexora: Open-source VLA for High-DoF Bimanual Dexterity EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T09:43:53.032153Z digest=sha256:2bd841be89972de718653a8ecf77575b43276e1799457b599db766af2c728528

Observation bb346e45-957f-43f9-ac86-1dfd7103aeb5 · inbound

Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum cites this paper.

Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T04:32:58.962571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:16:15.510851Z digest=sha256:8a75cbe19794c062e05a11b40b3e803299972407e9dbfdb677b6dbd3cb858017

Observation f10432fc-2773-45f3-943a-838afbd94419 · inbound

HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos cites this paper.

HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:45:48.692760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T01:02:58.035784Z digest=sha256:69f0a074a81d568ff615baf009dc2ef7646ba2a0e6651bb73e23cf300c4ef647

Observation e6b7e8f7-4376-472e-abb6-ceb3f62b59b9 · inbound

Grounded 3D-Aware Spatial Vision-Language Modeling cites this paper.

Grounded 3D-Aware Spatial Vision-Language Modeling EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:13:15.386870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T08:08:36.012761Z digest=sha256:7b4d68f6dfdfaaed79fba8965bde6cc3186aa05d059644dbced6e37e57472e04

Observation b95a5978-599b-4eb8-a88c-ceba7b7e504b · inbound

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data cites this paper.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.228649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:73f74a0b6dc0db66c0ebf124c3736c25f421a9402bf59b7af1e1452e28b4b82c

Observation f79ce0d9-38b5-4985-ac81-51426ad7fb61 · inbound

ActiveMimic: Egocentric Video Pretraining with Active Perception cites this paper.

ActiveMimic: Egocentric Video Pretraining with Active Perception EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:26:58.952670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T01:19:47.785250Z digest=sha256:e2a3f8c3b9a90166b4bf9f228a00233e36ef784c0e8b20fd2679abdced86ffb2

Observation a8d1bbf8-b4ed-49f9-8a3d-2034fd8e6545 · inbound

What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos? cites this paper.

What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos? EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:58.976851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T01:04:26.608953Z digest=sha256:6a3edf10656a2b26c2f5ff1b869208cd1ce3aeb53eb25b248fa4c498669583da

Observation eb25e211-2b4f-40c8-9972-461e2ad3a8ee · inbound

SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation cites this paper.

SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:47:28.059830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T19:29:15.489648Z digest=sha256:c8e9a3f2fdef535f87d6a0f78497a7c460424dde02c8dc2d87b9355c57c0b7ae

Observation d3656c35-33f6-4663-8b7d-c25e52ed22ef · inbound

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control cites this paper.

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.330508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:52.734998Z digest=sha256:99c2421d0f607e1d9ba246bf11536c780e42cfa991215d8762dac8fe96f82c7a

Observation 1a6ae559-f93f-4a7e-a9fe-577820a3819d · inbound

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation cites this paper.

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.719272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T16:26:25.136391Z digest=sha256:afaa71ad784d7bc691021376e472249ef37e358f57dd51c5ac3679dadb314959

Observation ccc259b1-1094-49f7-9d6e-f222ebc55d83 · inbound

LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition cites this paper.

LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.197162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:02:54.183839Z digest=sha256:c9676a4fd567d4cae3470b0959337bbe801c0d3f19606042781b5dfca8392f15

Observation df504181-4c65-40e4-9a47-4e05af4bb914 · inbound

$\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models cites this paper.

$\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.707779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:50:55.866910Z digest=sha256:e1d9a7b4a105a8d182d8a6fc52ca3d6dd9573dbb5a0bd39599bf19707d27a503

Observation f1340118-19e1-4b7d-8edd-c71c044bbf4c · inbound

EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations cites this paper.

EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:38:05.121651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:27:43.453005Z digest=sha256:d57932fde89557b2fbb7c1b87b2b47b0dfc49addb976df8bc866ed439a9c632d

Observation 7922ea34-3b38-457b-8fbb-cdafda1e772f · inbound

EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment cites this paper.

EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:33.076842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T06:36:40.317160Z digest=sha256:9673a9e055b5efd78b86abe0279a74c14c72baf28f03e0fa28caf7ea4a3ad347

Observation 3f93b938-4c8d-4bdf-860d-f854ff00101c · inbound

FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation cites this paper.

FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:28:31.721392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T07:02:17.427032Z digest=sha256:a78dc99c1d1c56ccc400470a9cbf0ad40b8011f19962d4334edc6cccce171859

Observation b0a28147-1a24-4c47-aa17-1de436a5d9a1 · inbound

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack cites this paper.

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T11:31:33.814742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:31:33.814742Z digest=sha256:cc0bb9891425d10fe915250d38e99bb383448942ed2abd3a129d67926a89ed68

Observation a21ea5b0-d478-4cd8-bc39-fef362b3d092 · inbound

T-Rex: Tactile-Reactive Dexterous Manipulation cites this paper.

T-Rex: Tactile-Reactive Dexterous Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:48:45.836541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T03:48:22.394092Z digest=sha256:247bc9c4e619e0441725c0c0be036e82717282de82e64d9e6019e33e960581a4

Observation c8fa37e1-ecbb-49f1-ad94-e889a471146d · inbound

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining cites this paper.

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:58:47.605592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T03:25:39.450667Z digest=sha256:c0792bdbe4d5c4f5dfdbf465dd55366e1db324776eae3a7e3dbf5420db900c8f

Observation b2c97fd8-6c75-4771-9235-58d325786214 · inbound

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining cites this paper.

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-27T03:30:26.569142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T03:25:39.450667Z digest=sha256:4cb5d37fd041b0faaa7c37529ee73f0efa1087933ff2e5cf67a4438d51c8bf48

Observation 3f728eff-8671-4dce-b4ad-0d0782c52100 · inbound

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction cites this paper.

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:09:14.740739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T21:27:47.702578Z digest=sha256:43ed7a728dc59de966e6640a5cd6fddab3b3e2b89ccd5f4ac56ea5d5f49ce0c1

Observation fcb465fd-1fb9-4042-a1a4-aaf2326d7ce3 · inbound

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos cites this paper.

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:19.370892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T20:51:21.209882Z digest=sha256:8a1dcde94a8d40fdbfb2bdf4440c13b6b93c27162413d17871893c2ea85f8117

Observation 3c0ed274-eed9-4717-a7ff-a1651544ec89 · inbound

Robot Self-Improvement via Human-Video Dynamics Models cites this paper.

Robot Self-Improvement via Human-Video Dynamics Models EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:19:37.582669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T14:40:13.855741Z digest=sha256:e5a2aaa11dfd0f7a2c0d2d97a5ca815175a4d99d35f02b3644250d95446bf545

Observation 0e10fb32-2be9-4c54-a351-15c785ad5163 · inbound

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data cites this paper.

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:19:44.691062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:47:54.951618Z digest=sha256:e07c4a209c58b53b29d9e0ce7c5ffeb6aa72f215f108f1893f512a2bd5d8a3cb

Observation 3f022312-5323-40e6-839e-a5bdd55acc8f · inbound

OpenHLM: An Empirical Recipe for Whole-Body Humanoid Loco-Manipulation cites this paper.

OpenHLM: An Empirical Recipe for Whole-Body Humanoid Loco-Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:29:41.808247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:38:43.339469Z digest=sha256:9b49add6132b542e7f4e7d5fd45f2182c053484a666f4b9bc0c122d1f0fb32a3

Observation a93b0869-f00c-4480-988d-94f1e06b26b1 · inbound

LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation cites this paper.

LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-04T11:29:50.824933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T07:58:17.225491Z digest=sha256:9352125d4cb9efa10ff282d70af6dd62eed53c7374f230bd34f69ff199a6c9d2

Observation 363bcb00-4a44-42cd-923b-dd79a16e2d23 · inbound

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought cites this paper.

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 108

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:19:57.788956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T00:46:17.094339Z digest=sha256:3ff2658e307bd31a8a2a04cecefab7d37b89c37b6223e428638822f60e60238b

Observation 7431e20f-57f7-4d3e-accf-5bc403869779 · inbound

Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding cites this paper.

Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:20:00.473782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T23:47:12.032082Z digest=sha256:fe918c41880a0b5b882b7fc40ca26b6369c1b951819a3037de95badf434fb96c

Observation b3f5f51c-379b-4bcb-82aa-967b6acabd28 · inbound

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly? cites this paper.

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly? EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.320457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T01:13:08.086026Z digest=sha256:2b4f65594dc72afd64d6d81714c999a926ae6b24d8b68f96d6978f626175f0df

Observation d83d8a1e-aa7f-475f-804c-21e5f767c959 · inbound

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly? cites this paper.

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly? EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-12T12:00:50.387540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:00:50.387540Z digest=sha256:85f1e31e74e356764ce9157c5d6dcc4097205ca1b8954d09c1ca90015e4f20df

Observation 69eeb5ac-c36e-4839-b093-e724b23c8af2 · inbound

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots cites this paper.

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-01T16:55:51.451925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T04:23:04.622902Z digest=sha256:252fb2675bbbfca13a4276e46652fc50e9fe65f143112b75cd4acdd176896e30

Observation 5fde026c-793c-4a22-9fe3-73798e4de340 · inbound

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments cites this paper.

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:55:41.355626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T05:02:17.092213Z digest=sha256:d6a0c19229079cf7f42ceb9e62b82d5cc19ff61dc98e50eb1dd723a14a7d3607

Observation 12b12969-e47f-4eba-906d-ee8abef1d384 · inbound

EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes cites this paper.

EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T15:07:03.962903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-02T15:04:09.575186Z digest=sha256:8f447a7ce4261ab0eadb11f5524030e17f63f9cd93711477d792be492752913f

Observation 5214ee80-de4c-4955-bb8e-b7b497330267 · inbound

Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation cites this paper.

Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:26:53.965852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T11:18:32.259684Z digest=sha256:7b812b97bdc096971222da5a03e0c4a52d92ea5d6bdaa2f5a4e68d8682c6935d

Observation 18ce32f8-ccd3-4bbb-83fe-7fbd7bb17eda · inbound

WSA$_1$: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control cites this paper.

WSA$_1$: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T22:51:22.538166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:51:22.538166Z digest=sha256:d8566979cee58a44cb7eef183a9bdbc476e14e26995a008ba6f2a65ce8d6c2ac

Observation 5a02135a-b814-4c23-a705-10d9051c8fde · inbound

EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data cites this paper.

EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:47:31.663947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T18:39:34.356696Z digest=sha256:46349ca031ffe5140b09587b52fbc5bf547edc00200376eac9c3bcf35738abbe

Observation f1521096-62c9-424b-8ad1-8040816255d5 · inbound

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos cites this paper.

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T17:30:54.988498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:30:54.988498Z digest=sha256:8cf5916a3f4236ff3d25c00fa89b4c86388e258ab1344884ed7ab51741bf55d3

Observation a58bd9c8-4c2a-4b09-9131-e0a0307058df · inbound

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning cites this paper.

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T03:25:28.118760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:25:28.118760Z digest=sha256:355c0d18a482c647a2916312179c932d76d5d8037a29068d9c3dc4ad410c4d01

Observation c4c67866-0373-46da-8f5a-9db043203143 · inbound

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning cites this paper.

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T02:42:49.111022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:42:49.111022Z digest=sha256:036deb4b5f3a7feceb7ef60539ea6993e4012a0f46bb350edb91952bdfd963ba

Observation 45e4ddb7-eb0d-4b11-adcf-24fcbb078b2f · inbound

Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting cites this paper.

Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T22:04:55.957529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:04:55.957529Z digest=sha256:27dbfe6aa07d2466f7814bc9ec3faddf89669af4f479d9e16f2f4459518e3b8d

Observation 19195c38-fbae-483b-915e-b10e7145e626 · inbound

EgoRecovery: Acquiring Failure Recovery Ability Through Human Recovery Demonstration cites this paper.

EgoRecovery: Acquiring Failure Recovery Ability Through Human Recovery Demonstration EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:20.233779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:20.233779Z digest=sha256:017d0532585f99674d123d6820d6c7b561e312c3a938f814549e259b43f49212

Observation 48498a43-221c-4383-9281-9d9f02c6db98 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T08:51:08.859285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:51:08.859285Z digest=sha256:d3c2f49bae54edc3f326d4ecd5b70e6695d4bcf7324f677776d0eb7aa6b784d3

Observation 0726349b-97cf-43ef-854b-8890bac421c7 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T01:22:54.356838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:22:54.356838Z digest=sha256:311532f089f19cf8926143355c4a8015773535e58fd9635077ae70e0f13ceb01