Pith. sign in

Paper Citation Record · LEDGER

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

As of 14 August 2026, this Paper Citation Record lists 100 of 161 outbound references and 63 inbound Pith citation observations for arXiv:2507.15597.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15597 v1

Coverage vector

measured 100 of 161 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:33:46.663820Z

measured 163 of 163 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 63 of 63 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:10:14.892153Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T04:16:48.668605Z

Reference resolution

100 of 161 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved96
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f1216af-6b17-4952-b576-49cb6f98da3a · outbound

This paper cites Review on human-like robot manipula- tion using dexterous hands.Cogn.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Review on human-like robot manipula- tion using dexterous hands.Cogn

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.030744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.030744Z digest=sha256:128773889acf542b34d45e2bd34a542f3077954e4da47152cde79ca1cd859e5b

Observation 213d718d-a32c-4af2-9165-83202181cc38 · outbound

This paper cites Human- like dexterous manipulation for anthropomorphic five-fingered hands: A review.Biomimetic Intelligence and Robotics, page 100212, 2025.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Human- like dexterous manipulation for anthropomorphic five-fingered hands: A review.Biomimetic Intelligence and Robotics, page 100212, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.117746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.117746Z digest=sha256:f46b54fd1fcc117e902f56ce101cd133e260f531ae243c53b09bb4356b2ee7e0

Observation 7fba0bd0-220a-4636-88c9-994dbfdec581 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.186546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.186546Z digest=sha256:bb0e5a2d8b82ea68db369efa07ed475c19d3d5356341c7e0362228898b85376e

Observation 6a33fae2-b40e-46a0-9aa4-30f4344c74ec · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.258306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.258306Z digest=sha256:b74f8b6cf3701a6e50caa1c4bd7ad6be0e44df8ebea9a67280f31781b0cacd0f

Observation c761ec84-fc8d-4f83-90a1-a10fa9f9c1b3 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos OpenVLA: An Open-Source Vision-Language-Action Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.333345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.333345Z digest=sha256:d891cec22091ee55276e17a2a44148d890f9ec5dbd481abb8d92591f5bb9e9b9

Observation fd76e2bb-41aa-4466-9b3b-2a46499d2114 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:33:38.444159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.444159Z digest=sha256:6fe7b8400ef20bed19003c47cc3eb3398e6b99defeed31eb2ca1eb36ca695dd1

Observation c3755938-f091-4b94-bf74-424a6f6857db · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos A Survey on Vision-Language-Action Models for Embodied AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.504555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.504555Z digest=sha256:26225dc8b8e206bca47284ca3e5cc25f98f4c987fd84a41d9eb98d17a0bda2d5

Observation 083d9651-4bab-4508-bd8c-d5c12ae2c2ce · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.565294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.565294Z digest=sha256:cd8b72acc643a4fbc4a1c6a7bec98a42f23bda9330001adf2fb2f536f56b0a1a

Observation 0167b96c-8cbb-478e-8f81-72518a88ed38 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.669232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.669232Z digest=sha256:60508dd698cf4ad7f568185a335abfc1c4a498244355ccc8f0a06dd13dbf5672

Observation 15bf4cc7-8ccd-4bf6-85d5-88b7d152e3ca · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.782713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.782713Z digest=sha256:5437f29cbd060aca324175800f95f2e0f501297c39d98cfe14e2f3e18d64aa9c

Observation d16fe479-8fab-4ff1-bafc-196a245d4018 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Octo: An Open-Source Generalist Robot Policy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.850217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.850217Z digest=sha256:f475678804b8fb1e1b052263360cabd9d03d0f84035bca48cb1a4f986aa43cf0

Observation f6b63b9f-f0e5-4395-9dc3-3f997dcb7cb8 · outbound

This paper cites Benchmarking Reinforcement Learning Methods for Dexterous Robotic Manipulation with a Three-Fingered Gripper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Benchmarking Reinforcement Learning Methods for Dexterous Robotic Manipulation with a Three-Fingered Gripper

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:33:52.228371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:33:38.935193Z digest=sha256:d7ce9329d13272e2e26dd574c19d818def99bd5aade0986b1e8bb23d9fba9720

Observation 85fbd045-d31a-4aeb-8290-96b34abab5cf · outbound

This paper cites Dexterous manipulation through imitation learning: A survey.arXiv preprintarXiv:2504.03515, 2025.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Dexterous manipulation through imitation learning: A survey.arXiv preprintarXiv:2504.03515, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.037596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.037596Z digest=sha256:c0ff2ae896641102fa1cd0bb8121acd382795990295b44fa54c0ca1caac4b656

Observation 9ca9382e-bac0-4847-b679-5c588f9d4f30 · outbound

This paper cites Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.082418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.082418Z digest=sha256:01e052ba431087cccf3a058b6254b71b038e1a3bee499cf4e7db7bf0971ac5cf

Observation 491785a3-b285-43a1-84bd-44dd8f44e829 · outbound

This paper cites Scaffolding dexterous manipulation with vision-language models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Scaffolding dexterous manipulation with vision-language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.147164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.147164Z digest=sha256:518c250322a1703c6d5d6dea614defb16082f6a35254aad7d8bbd88978251c1d

Observation 184e660e-cae0-459c-a0b8-4082e56b2734 · outbound

This paper cites Unidexgrasp: Universal robotic dexterous grasping via learning diverse proposal generation and goal-conditioned policy.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Unidexgrasp: Universal robotic dexterous grasping via learning diverse proposal generation and goal-conditioned policy

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.204910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.204910Z digest=sha256:90d5d3bb36fec09077c970a9212ad86be15e891be9aea2296a26e845674460c4

Observation 6029deac-11f4-4de2-8f49-475f5190838a · outbound

This paper cites Unidexgrasp++: Improving dexterous grasping policy learning via geometry-aware curriculum and iterative generalist-specialist learning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Unidexgrasp++: Improving dexterous grasping policy learning via geometry-aware curriculum and iterative generalist-specialist learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.260785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.260785Z digest=sha256:4535555c19a9a75fcd173bd2d7964bb1f11868191f30de3b44302c1c27158ef1

Observation 9012a185-501a-4aa7-b078-aa123dc4d97a · outbound

This paper cites Dexgraspvla: A vision-language-action framework towards general dexterous grasping.arXiv preprint arXiv:2502.20900, 2025.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Dexgraspvla: A vision-language-action framework towards general dexterous grasping.arXiv preprint arXiv:2502.20900, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.321889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.321889Z digest=sha256:192e5f52debe5d7a4cad19142567c95286b83ccc6c1d8b0b1775f20de38eb7dd

Observation 9c96aaab-977f-4a81-9243-af342b15fcf0 · outbound

This paper cites DexVLG: Dexterous Vision-Language-Grasp Model at Scale.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos DexVLG: Dexterous Vision-Language-Grasp Model at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.445246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.445246Z digest=sha256:b0283c098780b761e3f672348f954fd3eb1c2f7b6ca03e8013bd394b12d94244

Observation ddc36849-c43d-428b-b6ad-a01e4bd224c4 · outbound

This paper cites Efficient residual learning with mixture-of-experts for universal dexterous grasping.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Efficient residual learning with mixture-of-experts for universal dexterous grasping

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.511611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.511611Z digest=sha256:ed07f3bb143278b85253bb39430666dfe63e72ccbab44c8d8ddee18144040caa

Observation 13d52d96-73ba-4d36-b6ac-1a52ae2cc874 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos R3M: A Universal Visual Representation for Robot Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.565406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.565406Z digest=sha256:90a0f42344d497fc6ec0a29551450cdd1d05ce2f3bafd0165912783f2ebcbb26

Observation 1c462fb1-22bf-4e43-905f-15c05c20b67a · outbound

This paper cites Real-world robot learning with masked visual pre-training.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Real-world robot learning with masked visual pre-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.683285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.683285Z digest=sha256:c0a63a02cc7f94e01d8568b9bcf14ed8507f7cdc660ff3d773e8fcc99ae861ac

Observation 135da890-1f3d-4ffa-a1a0-76593fa951cf · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.742191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.742191Z digest=sha256:ba5fc491995c8bed3fd81097d10179ec4fd217769fa948dc62268d93ca491a1d

Observation d704d37f-98f8-4199-a36d-79442114c205 · outbound

This paper cites Improved baselines with visual instruction tuning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Improved baselines with visual instruction tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.817847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.817847Z digest=sha256:17ea626d21f8f39d7be0652cd1b012076150414e7ddc3a863ad0aa561fa1bafb

Observation 2097f02f-9b4a-4ef1-aa62-ffc549579545 · outbound

This paper cites Humanoid policy˜ human policy.arXiv preprint arXiv:2503.13441, 2025.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Humanoid policy˜ human policy.arXiv preprint arXiv:2503.13441, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.886667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.886667Z digest=sha256:108a3b5204b6e5e3f3558036090a47e17739cfdc59708347b18804c2b03caba7

Observation 741b91a3-801d-4279-96aa-da1d43cc7a09 · outbound

This paper cites Integrated linkage-driven dexterous anthropomorphic robotic hand.Nature communications, 12(1):7177, 2021.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Integrated linkage-driven dexterous anthropomorphic robotic hand.Nature communications, 12(1):7177, 2021

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.964864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.964864Z digest=sha256:f51655fa3b5a29b35a195b35e49f683d33b5544e8577a8529ff62f512c07fe89

Observation e8b02f06-166e-421e-ad27-2187dbe4f505 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.066734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.066734Z digest=sha256:c33621f7eb0f8514f19dbe1e3959450485cca3f0ef35d7f7e3630d789e28161a

Observation 19f3022c-e89f-4ee9-bf75-12b6dc9d4e00 · outbound

This paper cites Autoregressive image generation using residual quantization.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Autoregressive image generation using residual quantization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.128808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.128808Z digest=sha256:a44ad413f4143eed15a2c361da427e4fd87ee56be336115edf7458cb5648e61f

Observation acbb01f9-ab9e-405a-9aaf-dc55f47d8d26 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.225699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.225699Z digest=sha256:a0928ec6ee19fce53b7f972ea11b782876aa946d48aefa107dd49aa8886096e2

Observation 2ea0e8ec-bf17-403b-accf-af76bacde7dc · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.309722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.309722Z digest=sha256:fe4e7d9327cdef5033ce221200d8f125165d7f55cb16f994dc4733b66d0ffca3

Observation 1988342e-245c-4e46-b539-17e5a62fccee · outbound

This paper cites Improving language understanding by generative pre-training.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Improving language understanding by generative pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.390951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.390951Z digest=sha256:f638d9d8d2000042aa3636519f1e4c1f7d7105c3c29551f16d92a3ed8f8efe37

Observation 5ed2d5c7-ac6d-4597-bfd7-852ab5fb44e2 · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.524143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.524143Z digest=sha256:d2865ed4383dc40d1a0624038877975e9cf3e0acd8cf3f8d5ca83dafb65b3e27

Observation 97e4cffd-b982-4108-a1c4-743565f197b9 · outbound

This paper cites Language models are few-shot learners.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Language models are few-shot learners

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.627678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.627678Z digest=sha256:530f1d87a504658f58a14981d2e9a35d3993964fd0390fc0e208c35af076946a

Observation 60782969-dfa1-4753-9d7f-d8d02d600f00 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.742365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.742365Z digest=sha256:2aa1ffc2407e213154a6bf1f73056f7b8bcd0f51105cf1b18876e6187057e9a2

Observation 4374b8c0-d9c7-49ff-b2cb-24044a597685 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.877665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.877665Z digest=sha256:8d96702bb69f4098f089e17789f0663591a5263ced58ccba09f04aebfb5614f7

Observation 8c0f0ddd-94f6-4668-a86d-3723dcc5bfa3 · outbound

This paper cites UniCode: Learning a Unified Codebook for Multimodal Large Language Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos UniCode: Learning a Unified Codebook for Multimodal Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:33:51.856771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:33:40.966946Z digest=sha256:7abee3d922a9def9c5e2a49f7c30a4ab068e9091e717eb920b8ed86ef77ffeda

Observation deb167aa-01ee-46b0-9fbd-898b7204f23c · outbound

This paper cites From pixels to tokens: Byte-pair encoding on quantized visual modalities.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos From pixels to tokens: Byte-pair encoding on quantized visual modalities

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.028548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.028548Z digest=sha256:8c55c861055a167734e0954643418dd75f33611488b4a237b80837541d11f192

Observation adf0c352-4347-4667-bd28-45f795d9c910 · outbound

This paper cites Unified multimodal understanding via byte-pair visual encoding.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Unified multimodal understanding via byte-pair visual encoding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.136954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.136954Z digest=sha256:79fdf894b79b193d1c577142ce4868ba4fbc7f75c052f0a923bc10fc709b72bf

Observation 066c0331-e95b-4295-b5a3-562d86f09264 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.212733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.212733Z digest=sha256:210c3b2d9540abbd465f849264a58bc39caf9a7b73242afd5038a3f46fe4521c

Observation 0342fd5e-004f-42f7-ac47-d2370206bec7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.297194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.297194Z digest=sha256:656d1b711793de151fddd41f4b923f07cb59f80dbf3768fc256f1637b50527af

Observation ad47a498-55cc-41c1-8ab6-3ddd7213b88d · outbound

This paper cites Qwen Technical Report.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Qwen Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.390792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.390792Z digest=sha256:431d27d310d5651fceb4e95f7d70668f7b51ed4adfda3c99733f58c182ecd2be

Observation 16384af9-1a5c-4ebf-8af8-82e3d0a36ef1 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Learning transferable visual models from natural language supervision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.519049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.519049Z digest=sha256:c43b1c5af374af91cf4f4a371a5610f8fb0919bcf2bcf439f74132ac111ebca3

Observation 4aa01e2e-e516-487d-b31b-24f3b03e132d · outbound

This paper cites Sigmoid loss for language image pre- training.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Sigmoid loss for language image pre- training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.612755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.612755Z digest=sha256:fef3271f9bf4ad44728e03a9537aa2bc3b3b82d7efdd25aee6cf2cf933966593

Observation 795e4fc5-e425-4936-b631-b248f5c0f6af · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Flamingo: a visual language model for few-shot learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.653621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.653621Z digest=sha256:79d74e56b9276af222399b33ac84549389ba6e3165bf4735e8d97a1ef313f135

Observation 73eea413-02bd-4805-9e29-ebe6079d4b4b · outbound

This paper cites An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.744204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.744204Z digest=sha256:7c5f2a7610ee9448bfe689c58ebbbb940329b1ca580c882754b2a8089b5453e4

Observation 4e4c7e95-305b-4de5-a648-8eb8351a6023 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.848263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.848263Z digest=sha256:9d3c0b4c3c2e245fa41918761a262cbe81fdf2a8a5425482ab44d97573d3ad20

Observation 2c79e090-2cd3-4c00-8792-c604cdfd504e · outbound

This paper cites Otter: A multi-modal model with in-context instruction tuning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Otter: A multi-modal model with in-context instruction tuning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.991890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.991890Z digest=sha256:d1061714b91c85f14ce7f9b03ec38d3b3f2d94a1b3e2c277e04afda435f09e46

Observation ae659049-9bdc-4c7c-a9a9-ae018ab13db5 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.059251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.059251Z digest=sha256:af9d0d88dfbc5142f9431b9a11191dc2bab81c7d4f03233566b1da4b2a3c3c9b

Observation 27a8bc03-cdfb-4104-b4c7-46b0ec370c46 · outbound

This paper cites Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.174500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.174500Z digest=sha256:44ada4430eb052936b68c8cda977f8ad255aa552924fd8f6c1d9801ce62bd7f2

Observation dc1b8b98-c779-4bba-ba14-eacaaf7e8583 · outbound

This paper cites VideoOrion: Tokenizing Object Dynamics in Videos.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos VideoOrion: Tokenizing Object Dynamics in Videos

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:33:51.694013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:33:42.245773Z digest=sha256:aec7c783580032ebf9ddbe2849a1d7e097ed87c0b7847fd7b4c37bc83ee773a7

Observation 38d25635-d466-4879-9c91-7c7a075b8501 · outbound

This paper cites Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.364931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.364931Z digest=sha256:fba47c5eadb04b2f614bd1730b63189518dea1f92f0f94ba3fd6a882cc300d68

Observation f2b40deb-04fd-4eac-9224-ba01623ceca6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Gemini: A Family of Highly Capable Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.485380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.485380Z digest=sha256:ad4301173b346fab7ffd7abb6e03cce5f1fd82a11f7c1ad471777f9ea4df8e30

Observation 4144850f-7008-4400-b3f4-71b67d53e5f1 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.589633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.589633Z digest=sha256:701850fa2113e8bf2631e0c2a01aba536dbec3b7c3011c9f91ad8d7703e32494

Observation 7303ebdd-6754-4c31-abb3-8adc9def0a10 · outbound

This paper cites GPT-4 Technical Report.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos GPT-4 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.732915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.732915Z digest=sha256:4fb3383b8d7f4a41215f792415cf8e36a8c32ff246010564bb1ae0f8140a2977

Observation 0d05e1ec-5676-4ae6-8442-5ff22a4731f8 · outbound

This paper cites Qwen2.5-VL Technical Report.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Qwen2.5-VL Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.826045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.826045Z digest=sha256:5a4fad933cf78c956cc4fd7ddd9ca9df04f6a8a1073a476e8737afe7a9ffaa51

Observation 2359026e-9846-4af3-baa0-3e9031579db4 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.934724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.934724Z digest=sha256:3d59b7f1ad250d84e96cecada52b1436db1812d55e69658d2919f4483bcd1eb6

Observation bdbd68bd-fb7c-474d-b25b-d70e32d48fb4 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.026591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.026591Z digest=sha256:d6168efbfe901b252f582cdf0aee65f4d91d99aab0457aaaecab273d50b5e94e

Observation d9fe530f-52cd-4c1a-99ab-7acedadf48b7 · outbound

This paper cites Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.146808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.146808Z digest=sha256:28cd646928620d20093063b48c67e90903dfabe7403f00e18cf2f55b67be5f76

Observation edecf294-eb4f-4955-bde3-9c9d01281c77 · outbound

This paper cites The kit motion-language dataset.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos The kit motion-language dataset

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.239350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.239350Z digest=sha256:e7523e46ad98244e3c7d0329c2f5756908535e9ea3ce1babf9114709d24cca50

Observation 6faaf8af-74e8-4ad0-a2a7-ccd77fdca5b5 · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Amass: Archive of motion capture as surface shapes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.309159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.309159Z digest=sha256:dbd3335aa401d9405f394b7e0fd81b40b35cb22d7f02020272f3b8e966337c19

Observation d53026f1-1d85-4640-81e7-8aba6ff3d6ff · outbound

This paper cites Generating diverse and natural 3d human motions from text.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Generating diverse and natural 3d human motions from text

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.383085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.383085Z digest=sha256:f013f5f180300434b862ef2e68643300e61a4f2a52329bd4ae8103f76e60414d

Observation 60971193-2d7b-4864-a387-f741b9783ea8 · outbound

This paper cites Babel: Bodies, action and behavior with english labels.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Babel: Bodies, action and behavior with english labels

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.482132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.482132Z digest=sha256:9d8432f41c6da85f2e2a83aae93a1437d784d7205c24074e3295a6814e2bd4db

Observation 5ee93827-7900-4bec-87d3-aa21d56eef54 · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advances in Neural Information Processing Systems, 36:25268–25280, 2023.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advances in Neural Information Processing Systems, 36:25268–25280, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.571970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.571970Z digest=sha256:cd6641eedd1748c4e430d707e91a9c8dd05af16e7c19edc0e2245740c301497c

Observation 1e470422-15ee-460f-a820-a5e001100406 · outbound

This paper cites Egobody: Human body shape and motion of interacting people from head-mounted devices.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Egobody: Human body shape and motion of interacting people from head-mounted devices

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.663928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.663928Z digest=sha256:4c35d0cd7bd6d068dfb55b3b9ad7abc0a36392d29c336194d3b96fe3c6278dcb

Observation 18c5d703-6a66-4c4d-aba1-7872741a2f59 · outbound

This paper cites Nymeria: A massive collection of multimodal egocentric daily motion in the wild.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Nymeria: A massive collection of multimodal egocentric daily motion in the wild

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.756773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.756773Z digest=sha256:346ed966aae749b5f6b239e608bdabf1cb9b68314c966d1c13438946c6c3a8af

Observation 5a028043-3caf-45a2-948c-f1761def8109 · outbound

This paper cites Scaling large motion models with million-level human motions.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Scaling large motion models with million-level human motions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.859592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.859592Z digest=sha256:eb7d30d82a64f44bc0fd284c6844c50d2c247eaa88636e7249366e46a06c44dd

Observation 9e5e0376-e134-4548-85f5-bed4f8b76961 · outbound

This paper cites Smpl: A skinned multi-person linear model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Smpl: A skinned multi-person linear model

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.946231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.946231Z digest=sha256:972ba2497deb89a34cae02456b69abdafa905d29d908f5d6b5de453bf98e85b1

Observation 8bfadaba-b22d-4fbc-9e10-e9213c25b9d1 · outbound

This paper cites Expressive body capture: 3d hands, face, and body from a single image.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Expressive body capture: 3d hands, face, and body from a single image

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.086686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.086686Z digest=sha256:17f786fbb32c4b8d3f4bcc9fff6b09eeab8c2edc89c09a3958f1bbda26f948ef

Observation 1d7f0419-75a2-4b7a-8313-4339200e13ea · outbound

This paper cites Human Motion Diffusion Model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Human Motion Diffusion Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.187290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.187290Z digest=sha256:e831308eae66f8d5bcb12642a678471a87660ca3bc37f19154e565b5cc304fbd

Observation f64da1b0-6537-49c6-ad33-c267df7b40fb · outbound

This paper cites Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.341154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.341154Z digest=sha256:f4289b88b67d1d336543444d115fc1e4c728cf016bbdbe8d119f0c596b35c6be

Observation b1483f9c-a7ff-4d5e-9aaf-a978a0a6bfe8 · outbound

This paper cites Executing your commands via motion diffusion in latent space.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Executing your commands via motion diffusion in latent space

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.449291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.449291Z digest=sha256:0f6acfcccdbe16325cd9e55cd69b2dfbd3115d9d121c6621cc70ea755dd47b36

Observation 4e06db4b-fddb-477f-9953-fcfc76856df7 · outbound

This paper cites Physdiff: Physics-guided human motion diffusion model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Physdiff: Physics-guided human motion diffusion model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.595272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.595272Z digest=sha256:017ab78ac10ac5c25e6263044e47824c9e8bcc6c82ee61b0f158fb8613858d9b

Observation 0acc39f5-fa8f-453b-8df1-d67450ec27dd · outbound

This paper cites Remodiffuse: Retrieval-augmented motion diffusion model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Remodiffuse: Retrieval-augmented motion diffusion model

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.691840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.691840Z digest=sha256:fe57e8578313c6c523b2cdaea0cc6a90b23259ddeacc72163b35e65148514e63

Observation 59819418-f3cb-4b2b-a203-e1ee2f99d6bc · outbound

This paper cites DiverseMotion: Towards Diverse Human Motion Generation via Discrete Diffusion.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos DiverseMotion: Towards Diverse Human Motion Generation via Discrete Diffusion

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.817249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.817249Z digest=sha256:13a852952bd6dc4c59c6b1a36a1177174fe249d8dd41ae4f0bc30f5cb72df3d4

Observation 682795f0-2e56-4979-9ad7-04e47722f7b2 · outbound

This paper cites Large motion model for unified multi-modal motion generation.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Large motion model for unified multi-modal motion generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.965632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.965632Z digest=sha256:c688c502cac2b0b5a5ac92207c19292b80aa3ba72f9143cc8c00b540cc04c826

Observation 3781b643-41af-4e61-9e88-dfb34367abfe · outbound

This paper cites Neural discrete representation learning.Advancesinneuralinformation processing systems, 30, 2017.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Neural discrete representation learning.Advancesinneuralinformation processing systems, 30, 2017

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.087540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.087540Z digest=sha256:660057fc6387d491c75bd165948b0059df5ce4dd0f56c14f4ba01e516c48f5c5

Observation f2ca6c3c-e873-4be6-8e5a-76ab3fc3d47e · outbound

This paper cites Generating human motion from textual descriptions with discrete representations.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Generating human motion from textual descriptions with discrete representations

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.136041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.136041Z digest=sha256:a346e241580976c97ca39a391367417e7c14c395ad87ee3e70add4b4620c9afc

Observation faaf87c1-c4a7-4048-ae4f-ba422cd2cc71 · outbound

This paper cites Momask: Generative masked modeling of 3d human motions.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Momask: Generative masked modeling of 3d human motions

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.196778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.196778Z digest=sha256:f07ea5d2aa96a3e88172f926b4d1315e255269cfe497226249a9f69827478914

Observation a6c50ce6-706c-492e-82a5-fb1296c08074 · outbound

This paper cites Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.323096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.323096Z digest=sha256:741ea611f59a3889cdc05f59c106db86905bd7514b227526ca848aa14c4d09f9

Observation cabb6102-fbb4-441b-a064-d4dd6e18247c · outbound

This paper cites HumanTOMATO: Text-aligned Whole-body Motion Generation.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos HumanTOMATO: Text-aligned Whole-body Motion Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.392908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.392908Z digest=sha256:44272171e70215580f6b859a467208ed3e4151c6890390fa70f34006c531629a

Observation 9fde13bd-7364-46ce-a8b4-bf498e37734b · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Finite Scalar Quantization: VQ-VAE Made Simple

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.489989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.489989Z digest=sha256:1368c5100b0f107698ebc8336239166effc564e8121101e79bdaa7aab233067c

Observation 42a3761b-fb1e-4cd7-9020-57caeba12fd1 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.572085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.572085Z digest=sha256:dec2bf1f8faf2c6d0018890750daa9482adeda412c98434dcab4b2948d1598d0

Observation c9e1c9c6-8b7c-45e9-8f6f-66736c3c6272 · outbound

This paper cites Motiongpt: Human motion as a foreign language.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motiongpt: Human motion as a foreign language

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.640936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.640936Z digest=sha256:4c08e8625b8e12154ac90ba68d63edd24e232c4367674ae34c1e995a76d14fdf

Observation df7acbec-d0c8-4121-9b14-ad743b74c024 · outbound

This paper cites MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.729880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.729880Z digest=sha256:bb1cd1e932c7fdb345a513905a6942e07dd4a1413a7f805d9e3d505ca8b28ee8

Observation f7cd73b6-a9aa-41d4-97c9-144cc1f960c5 · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.839617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.839617Z digest=sha256:8c2d109bce4befba809813a8c7dd6fbd6bfe4b9b65a74355929858ed08e2826f

Observation bfed5262-2915-4001-b16e-cc2be9cb474e · outbound

This paper cites Avatargpt: All-in-one framework for motion understanding planning generation and beyond.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Avatargpt: All-in-one framework for motion understanding planning generation and beyond

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.967329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.967329Z digest=sha256:4e77020b9d64c8c70aa5b0f948822f730c60ac5361779d55c51b6b22a6be91f2

Observation ba2656d6-4ed0-42c8-92a0-fb815962cbee · outbound

This paper cites Motionchain: Conversational motion controllers via multimodal prompts.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motionchain: Conversational motion controllers via multimodal prompts

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.024326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.024326Z digest=sha256:17f59dfcb092f505ba7ba4e4291a7f4406561743d17b3987f6bdf56f7b7e7927

Observation e9b8a8fc-126e-426a-913b-64b2385a169d · outbound

This paper cites LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.059708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.059708Z digest=sha256:8953f91b6f08016bf011e84ce2bde02e6e448266470330bf350c6af4fdfe307f

Observation 5d245ff8-9b14-45cf-bd77-9f159ac67b35 · outbound

This paper cites Human motion instruction tuning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Human motion instruction tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.130864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.130864Z digest=sha256:25b0f3655adee91bd49984f8e524d086bddd27d30631d222fb8e769a5de0a613

Observation 8a7212e1-8afc-45b2-a49d-cba95fb4f3e0 · outbound

This paper cites RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.196891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.196891Z digest=sha256:2bc013cb6727b5fe1c720f391dae6b481667e653a666fe6c4bdc3fbe0166dc0b

Observation 4a9fdfc3-20bd-4aa5-a976-0eb87aef1e2c · outbound

This paper cites Perpetual humanoid control for real-time simulated avatars.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Perpetual humanoid control for real-time simulated avatars

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.255945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.255945Z digest=sha256:2a115762cc2057cf90dc2daf9fc3e015a3ac6f2f6ce122abfa4d9cb7638d7fa5

Observation 2b6f6754-fb9c-4da2-ac3d-db188e302e23 · outbound

This paper cites ExBody2: Advanced Expressive Humanoid Whole-Body Control.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos ExBody2: Advanced Expressive Humanoid Whole-Body Control

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.336690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.336690Z digest=sha256:d63bb9d45ba5e7afafc8c3aca2655a0d1ae7889d0f9a5a92e0331ac637f20b7b

Observation 6dad6a80-a450-43de-b8e7-b6141c91e102 · outbound

This paper cites Reindiffuse: Craft- ing physically plausible motions with reinforced diffusion model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Reindiffuse: Craft- ing physically plausible motions with reinforced diffusion model

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.386963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.386963Z digest=sha256:cbccd0ff612ffb66790fd6c5eb8835d65d0fa64229aa073d92a97a79a0ef116f

Observation ae8d7db9-e877-4a54-8368-084b19507e00 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.451798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.451798Z digest=sha256:9af2d11fef3ba2230359add6468cf49e552926fc3c4816462d8cc155f07976c9

Observation 900fbd39-23f3-426f-9049-b2f95871ed97 · outbound

This paper cites Hand-object contact consistency reasoning for human grasps generation.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Hand-object contact consistency reasoning for human grasps generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.508374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.508374Z digest=sha256:ad800f5c04c3989a4069cce58a9b9307df74f135ac371f6b33e0b4a62e4683d9

Observation 91a3701f-955f-4d51-b7f4-43fe3e148d98 · outbound

This paper cites Joint hand motion and interaction hotspots prediction from egocentric videos.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Joint hand motion and interaction hotspots prediction from egocentric videos

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.541539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.541539Z digest=sha256:5c5ce57bf5e54c3324369f8d696056898f70b910e721f75eb336d15107f1fc47

Observation 079824de-be87-4b1b-97ba-a373f732fc17 · outbound

This paper cites Hot3d: Hand and object tracking in 3d from egocentric multi-view videos.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Hot3d: Hand and object tracking in 3d from egocentric multi-view videos

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.564839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.564839Z digest=sha256:415a315276a395efde059ede696b76c0c8fefe55ca0f5dfaf919cacb72456a13

Observation b19ece4c-6d3d-4fda-867d-6e11e371444c · outbound

This paper cites Hoi4d: A 4d egocentric dataset for category-level human-object interaction.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Hoi4d: A 4d egocentric dataset for category-level human-object interaction

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.641803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.641803Z digest=sha256:74ec1fc0dbc5ae8093d4b9a0bde50d25ce3b649f31dcc7fd84848ab603ef3a4a

Observation e9bcd271-54de-4248-a716-a2d7b3e2b4a0 · outbound

This paper cites Oakink2: A dataset of bimanual hands-object manipulation in complex task completion.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Oakink2: A dataset of bimanual hands-object manipulation in complex task completion

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.644330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.644330Z digest=sha256:1b7f908328e53b3a736856a8848f5f5dcdfc1b61b33bc55a988f9d0f25016f08

Observation 21d286b6-3268-443e-8562-e447361b5c5f · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.663820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.663820Z digest=sha256:3f5fe17ade2482677bdfdf9b9dc93cb4dd6dbe69c4797933f2f243fa91153947

Pith citing papers

Observation 8b59d164-430f-4446-9e55-2a5eac4ef086 · inbound

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration cites this paper.

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T18:50:16.766190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:50:16.766190Z digest=sha256:984c304c4c3f536f519a68dadae65d77906adb1330d2fb12ae2bd109c7da2b10

Observation 8ea14750-e58d-4048-b56d-b8b74d8d3eb6 · inbound

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model cites this paper.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.531852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.531852Z digest=sha256:475f71ac01fedc8b00c43ddc27e1f1ff75aed10ecec1fc48216a45238aa1aa0b

Observation a8f85282-66ee-412a-84da-64e4d1894b45 · inbound

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting cites this paper.

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:54:18.265883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T17:53:19.002603Z digest=sha256:e821637fcbc9cfea3863bafda4e6909e8ee20a02a5bdb3796f1c8bd6560ab1e2

Observation 228ccdd7-920f-40ef-bcab-b572bd8fc53a · inbound

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models cites this paper.

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:42.870351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:42.870351Z digest=sha256:ca81173be76fe6847ba321699c03187052b8c83096d369a1fafc76bd04ad0520

Observation 14292983-4c79-4716-84d6-caa3591e71cf · inbound

DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter cites this paper.

DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:18:55.692685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:18:55.692685Z digest=sha256:6e4488ec4705d5c537a30737668a37febd0b75372bc8dbd7ea6639b8f64275c3

Observation ce3dc394-c737-4508-8cc6-ddc2e5d46e0c · inbound

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos cites this paper.

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:02:34.356382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T17:02:33.997887Z digest=sha256:4561c6949bd1cba25ea02fdfd9b3c5db754cceed3a30f52cca032cb630026a39

Observation df45d407-7e35-4e6c-ae92-b93786349036 · inbound

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models cites this paper.

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:02:24.979717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T06:01:30.803128Z digest=sha256:cb4186a4dd5c94973d5e90da056b088d4a49728bb7323f8c6973a8948028e045

Observation 4bb248b3-fd46-4cde-a901-5771d349e93c · inbound

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion cites this paper.

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T23:57:44.645995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:57:44.645995Z digest=sha256:87bcff4f1b34673216264a1c8d1c8989f39172fdbdd002b178f8239743a99bc5

Observation 66e903d4-96b3-4591-9df8-7feafe535f33 · inbound

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation cites this paper.

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:30:20.893877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T21:22:41.935691Z digest=sha256:a289ab8d181a70a9a4859c1fd1c492cbeff0c3048c0d0c83517c3365898d8431

Observation 995409de-51e6-4c60-8313-a0d448e7afc8 · inbound

Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves cites this paper.

Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T21:11:31.022938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:11:31.022938Z digest=sha256:f51136d23a45c41a245ec9cf837da7f511d5667c75c82c19444a24d3f5559565

Observation 3f1506ab-6795-42ac-a247-a675d0621889 · inbound

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World cites this paper.

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:35:57.002389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:07:41.489995Z digest=sha256:3d2b4bc04956c47f32c77fc8392bb352cf6f95d585316ba3d9f1726320251837

Observation 0d1f7b55-f0f4-4ffb-a777-628617c0a3dd · inbound

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World cites this paper.

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T08:25:22.011013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:25:22.011013Z digest=sha256:fe9d9e30becee9565b33b7e3b71ee6e981e055ed158665ffbb9ce151496399da

Observation 792b1e76-3bab-4a37-bf0e-e6655d513d82 · inbound

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment cites this paper.

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:03.890345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:36:23.197843Z digest=sha256:8ea8694a70bf034e09880dd501351e960083011f8f5d38befe49ba66a5402c12

Observation 64153462-9782-4f60-8a31-79fdad4c8103 · inbound

A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies cites this paper.

A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:30:22.689485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T12:29:38.306670Z digest=sha256:0823422f69a0dbd9370095bddde479139f6c63ccb713a0f9d992617ad777b512

Observation 2a2c6c44-858e-4db3-8980-fbea3cac73d0 · inbound

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models cites this paper.

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:29.409174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T04:27:18.284698Z digest=sha256:b6f00db4f25712f6796b087c46d806777f8d1b31b02e267febd5bf5fb668590e

Observation fdef3828-756b-4c33-8992-1d37bb0ebb4b · inbound

EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks cites this paper.

EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:16:11.740209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T06:19:34.743194Z digest=sha256:af633ff40742cef462a89f05a30dc8f3f408b004f86c98ea3eb62c8f17bcc4ff

Observation 8fceb112-ea8a-4399-9ffb-052c51f8bd42 · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:26:11.725302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T02:51:27.662262Z digest=sha256:529650889bd765dcdbc8c92b6d80c48f1007a46961a5d8bcb0afb6265e7ecfcb

Observation 17e17457-c4f2-43ba-8fd9-251ff72234fa · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:29.156747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T11:16:58.104663Z digest=sha256:018114045ec3577aef866a8065f51668141bfa7cb6bac92a4e39a3f9f8d9d139

Observation ced91d16-378a-404b-8909-be2b52314578 · inbound

Being-H0.7: A Latent World-Action Model from Egocentric Videos cites this paper.

Being-H0.7: A Latent World-Action Model from Egocentric Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:56:08.488051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T20:48:01.461993Z digest=sha256:dd491cdae1372434338c694206437e52d31a04a64fe9b45399c703a03ea9d36f

Observation 13379a2f-4b85-4ade-9d08-19553221a7e8 · inbound

World Model for Robot Learning: A Comprehensive Survey cites this paper.

World Model for Robot Learning: A Comprehensive Survey Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:11:04.931247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T20:38:12.709629Z digest=sha256:92b2db80b3e1a8011b027f2b0344e3bc971d53c629e6b20f2be9765e2b21f217

Observation fe5defcd-459a-4c6d-9907-f07a38569958 · inbound

HumanNet: Scaling Human-centric Video Learning to One Million Hours cites this paper.

HumanNet: Scaling Human-centric Video Learning to One Million Hours Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.148923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T00:51:08.414394Z digest=sha256:df35884aa670768f4bcd77fd2e4cad809c6a97be3ac22da78b27d1cb03aee517

Observation a47c22ac-14c5-4628-abbb-cb6baa225d67 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.846258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:31d109c27043fe003772534de5bad8bd763a298b66bad0b304792ccc4fc00097

Observation f15d5d9b-626c-4120-bb49-7f6ebdfa7de5 · inbound

Towards Robotic Dexterous Hand Intelligence: A Survey cites this paper.

Towards Robotic Dexterous Hand Intelligence: A Survey Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:06.046963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T05:44:23.945064Z digest=sha256:210d8674af34b3ab26745bd0a04afc0dddc1911051b8109903cd8ad2e53462a8

Observation e6258c08-aa48-4067-98fd-59a673b365a8 · inbound

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention cites this paper.

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:09:42.122084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T03:09:26.858170Z digest=sha256:42884a25102bf231d5c602862f91a4da9772abe807a968e035192f35849bbc2f

Observation 51ee9ed2-a371-49ed-b5cc-a6fd44d4c33b · inbound

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention cites this paper.

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:34:05.565221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T08:31:57.077346Z digest=sha256:e54158cc9e8a762e8fb333698af51bc2f906f5ad3f62d6f24b1e8dd146d52a51

Observation 7b6ac8ad-7cde-43e6-9872-23eaf06c23d0 · inbound

Dexora: Open-source VLA for High-DoF Bimanual Dexterity cites this paper.

Dexora: Open-source VLA for High-DoF Bimanual Dexterity Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:48:11.778910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T09:43:53.032153Z digest=sha256:eaecd902d69c5b0d807f3fe5cbd601a10e3fe2bf82a918ae11f9310cde34bef0

Observation b88443ab-14a4-4068-9482-5c141b4c9267 · inbound

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models cites this paper.

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:24:04.188694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T00:18:33.150920Z digest=sha256:3641f8755b1295036da9478cd91206f8bb7cec41b023937ecd7151a4b921a21e

Observation 3b4d04cf-4c1d-4503-9490-b72a87ba08e3 · inbound

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models cites this paper.

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.768238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T07:18:59.266263Z digest=sha256:6682ad7bbfa6a254f664301aa6cf3228a5edd9c48d784db6c3db0efa233a89f7

Observation 7ab91d45-1471-4c74-b056-ea02fcfc11e3 · inbound

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data cites this paper.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.210434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:5efd2bd422b625a7d538cd0fbf5b53972cf80f1fe78684976dc573162dc0b5db

Observation fd4dd992-0dc2-4880-be98-213785842ab2 · inbound

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling cites this paper.

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 220

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:27.773712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T13:33:03.368006Z digest=sha256:677008cc742885f8404023fea4b1461930c4302b3ff14ea9da28db87fe0ea31a

Observation e1fb98e0-3d75-4e42-bbdf-70f7cd08c5c7 · inbound

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation cites this paper.

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.363432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T10:28:22.430294Z digest=sha256:5068d86a7df0c8b6d08b11b66d7c5293a23a479dd73589cc8468637659cf1a20

Observation 9bb1aec4-8dd7-4f1e-8274-da739a390b14 · inbound

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis cites this paper.

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.909752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T01:05:26.382826Z digest=sha256:693d21ec357d9250775415ebfa866beeeda4cba19f4f75dbb33584308bbc3905

Observation 389ecd41-aead-427a-bbcc-05e33a2970c7 · inbound

RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning cites this paper.

RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:57.100854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T01:44:32.716179Z digest=sha256:087cdc7eeacca412d650b68862d53b5efb7c8329ba1128d6ccc320f2af742e33

Observation 79d0f620-2bb8-42da-992b-afc11cad27bc · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:47:09.505386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T22:25:17.522240Z digest=sha256:4d6d841de5e35583f4f7601dc836a355c55bd64cdcb85e5455eca8b9b831da21

Observation ca0da9aa-21c0-490f-bec8-32def5ec9703 · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:35:34.493805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-01T07:17:42.045939Z digest=sha256:e6ea1f45ed141ac160132c96451b84a2a18e4a7a72899748a99fd39e32a589ac

Observation 462e6594-bd69-49ee-bb7c-2f578b6c608b · inbound

$\omega$-EVA: Envision, Verify, and Act with Latent Interactive World Models cites this paper.

$\omega$-EVA: Envision, Verify, and Act with Latent Interactive World Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.742396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T16:10:02.176202Z digest=sha256:ca390a6001604099010ebd532287ddb66224c49920ddde4445442943d6713cb1

Observation 18ae0df7-1766-4f47-9266-4fe04473699e · inbound

Next Forcing: Causal World Modeling with Multi-Chunk Prediction cites this paper.

Next Forcing: Causal World Modeling with Multi-Chunk Prediction Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:57:38.513373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T13:31:53.905704Z digest=sha256:343863a32157a00f8fdb23db9ff66d50329a3c5fca6e048c3d01ffc0caded078

Observation bfee454c-5b31-4cba-9671-a3779d34a43c · inbound

LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition cites this paper.

LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.220598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T10:02:54.183839Z digest=sha256:a5bb5f5f858279e4f3b02432471308082313df588aa4c8127015de012bee85a1

Observation 67082a9d-78e1-4134-b360-31a571b052fa · inbound

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models cites this paper.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.855989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:2c7bd05c4ed769e52b29d5fdcfe89def8ad79f3d4c4bf45b448b21db38b12373

Observation b1983b03-9216-4236-8dcf-73e1abbe4232 · inbound

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos cites this paper.

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:19.497076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T20:51:21.209882Z digest=sha256:0aa3f95702ffdcf79ea48c98f7b10ed37b354a39d05a42e1d8782651b0118a6b

Observation 25489053-9cea-4a60-8674-87b32178ba29 · inbound

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning cites this paper.

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:20.363664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T20:49:01.269513Z digest=sha256:e28f81b3baf8d57eb55e9b693bf53591fda7dfb45e45b5eb4a2244207ba56e95

Observation 6ed30d62-4531-400c-b6a4-28aeb16a039e · inbound

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining cites this paper.

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:29.683310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T17:53:24.431287Z digest=sha256:ed1aef83e31c7d8fa5e548af1b9ccd319cd8eb6ebbdea545151a35eba7b6547e

Observation 73810ba6-6a3f-488b-922b-b89b9011c5fd · inbound

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data cites this paper.

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:44.736280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T11:47:54.951618Z digest=sha256:6758eb8a5f69b5249a00c53b9e7f3e5c1733b61dbc5e8a44d2e108e1a58009ce

Observation 8e24bd8d-ec6c-4acb-8ff3-dd8b1f91a3b6 · inbound

Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos cites this paper.

Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:58.480120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T23:57:48.395172Z digest=sha256:1693c4daa839e0c7f10b9de9ce202aea3a5a62d88728fa02164a50276c6ab7e5

Observation d836be21-89c4-4eb5-b6bd-f59877630a17 · inbound

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots cites this paper.

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:55:51.368357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T04:23:04.622902Z digest=sha256:959a5c467b4353d34f310b0e1898c2df723a013a1456d27b0f7cdb84495e8e79

Observation a6366124-29b5-4006-a58d-fff533a68f5b · inbound

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments cites this paper.

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:55:41.344117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T05:02:17.092213Z digest=sha256:ab0b9f5f7fab70a3305340a614c8d453bedff809ae2871ba39b2c2453839c77e

Observation f5181d8a-4432-4e9d-b577-41cd6ea9a2ff · inbound

Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation cites this paper.

Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:26:54.003775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T11:18:32.259684Z digest=sha256:82f8717aaf2e59eb99a1ab1acef8b68b9e05b87b4306b629feb17112b60e7392

Observation c898bc55-29e6-4ff2-b04c-63542be34257 · inbound

From Foundation to Application: Improving VLA Models in Practice cites this paper.

From Foundation to Application: Improving VLA Models in Practice Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.445289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:2cdfce18c6e45396a3ec14ac73b09e194fa1db2a81819d586adcd2d2c0e1e948

Observation caedafdc-980f-48df-93e8-8e0073c6f14a · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:16:48.670270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T04:12:29.153764Z digest=sha256:b817f769bd35aadfd511f067aed34d265cdb0a0011b5a87686451db0baab8980

Observation f742ffaf-db65-4417-84e8-dc83c1b193ce · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:38.754872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:38.754872Z digest=sha256:1d385370116c23c4e5d2bc6f0bce34946196adb70ea68120930f0cec33172b60

Observation 2766f5a6-6e0c-4c25-8cc5-2eea594df53e · inbound

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos cites this paper.

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T17:30:54.988498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:30:54.988498Z digest=sha256:518c13e7f0654b7d17ae17c5d4b24d6d5d29baac8248ae7e02e62524ba15c8f3

Observation 2138be29-39f8-4103-9629-8bd6580c280d · inbound

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning cites this paper.

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T03:25:28.115858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:25:28.115858Z digest=sha256:52ed540a688cb486f84f94e54c4522deb6b527a3d6cc83e0fcfbc7a9bb330445

Observation 5efc2714-0f33-4327-81ef-d8033792d4ec · inbound

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories cites this paper.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.119524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.119524Z digest=sha256:ec8967f96e7a98a8b240bda7f505be4f46b374f85a6dcc60bcd8b73d3419882c

Observation 58695d96-07e2-4795-85c2-02cc499f5d1d · inbound

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis cites this paper.

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-01T19:06:43.487945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:06:43.487945Z digest=sha256:591d6a219cb8f9de538ab56567f1806b509b6f93bdeab39e500bb3b9cf22ad57

Observation baf6ea43-c392-463e-9ea5-59d2c684839b · inbound

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments cites this paper.

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T23:29:10.712814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:29:10.712814Z digest=sha256:70b0b37f65a5be6f4d904275cbbaece0c27ecbffc4535ab41e5da53f7da03dfa

Observation fcd9e4a2-c9cb-44ad-a3cb-036e28beb184 · inbound

Data Pyramid for Embodied Manipulation: A Survey cites this paper.

Data Pyramid for Embodied Manipulation: A Survey Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 237

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.854691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.854691Z digest=sha256:c78677f6e5bc4265b1c5e877af3ec0825d3ee7624175db1c584c2bc8f0e2d10a

Observation 012798cc-eb2f-46ec-a04b-3a3a4f9b30b8 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 227

Resolution
unresolved
no resolver link, observed 2026-07-31T08:51:29.293605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:51:29.293605Z digest=sha256:311266fc38e2791ba00ee16c9f7330a1b9670b58a90f54851ac27ec3cd73781e

Observation 719960a9-0211-4c33-8590-b2815129ed21 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-04T01:23:10.898271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:23:10.898271Z digest=sha256:47bb00ffb17dba662c884b69893de38eda291d7d38bea5ab0398107dcbdd2b98

Observation e0a809f7-8b0d-4924-99dd-b9e9b18e3015 · inbound

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph cites this paper.

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T18:35:26.658756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:35:26.658756Z digest=sha256:f5c04e99cfde208d869549aff70e12e3a7c6942909241ff4fe64111eb72b0a99

Observation 6434a5c5-d393-40ae-b61a-1c6cd850d6d0 · inbound

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph cites this paper.

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:10:14.892153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:10:14.892153Z digest=sha256:35cee42df307271b6aa522131b20eacfdebfafacf9d31c5c8a7cf9fabebd9b21

Observation 67c89dc8-ab09-47eb-9697-ebc544dd42f5 · inbound

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data cites this paper.

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T04:25:52.097084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:25:52.097084Z digest=sha256:f550195ed2dc718a64896876bd23a0a42ccdb131cc29e3b1c67422825ad56849

Observation db8ecbb9-fb0a-4255-93ec-bd6f50059c1b · inbound

DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units cites this paper.

DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T00:58:35.689675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:58:35.689675Z digest=sha256:0f7db8377719ffa67cee6c3797f96127267dbf6db63e6e0b2856cac2871f4931

Observation 36d4981d-d08b-4f6a-85bd-92b8aed13ce4 · inbound

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation cites this paper.

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:32.515534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:32.515534Z digest=sha256:f3343eb1a360fc1cfbe5520c3ec4eafede0edef66418c1e50847ec0f930c4110