Pith. sign in

Paper Citation Record · LEDGER

Learning from Massive Human Videos for Universal Humanoid Pose Control

As of 21 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 12 inbound Pith citation observations for arXiv:2412.14172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14172 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:29:46.938716Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:11:12.977121Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T18:15:21.360726Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2f564c5-9444-417b-a4b5-0ed7c5570e54 · outbound

This paper cites Human- to-robot imitation in the wild.

Learning from Massive Human Videos for Universal Humanoid Pose Control Human- to-robot imitation in the wild

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.262409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.536471Z digest=sha256:a82ec70fd0e2dfc02fff681a71b60adbf999ebb570cb9e36efdb4b314ef07965

Observation 5a7f4b1a-5685-4bb5-b74b-1b22ade4404a · outbound

This paper cites Affordances from human videos as a versa- tile representation for robotics.

Learning from Massive Human Videos for Universal Humanoid Pose Control Affordances from human videos as a versa- tile representation for robotics

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.246202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.542027Z digest=sha256:469387b74ac2ce9df4d9f15659b34d5652db4f6ec6e1a21ab90279e9fd40faaa

Observation 308e47e4-a51b-4d4d-ad9b-fa30e79de5d4 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Learning from Massive Human Videos for Universal Humanoid Pose Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.547914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.547914Z digest=sha256:4f2cb086532bbc5947dd2fa32b2e7a4645a91570fe80924899995621d55f1217

Observation 1bb15ec1-a7db-4d7e-b6d8-f1ff694044a6 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Learning from Massive Human Videos for Universal Humanoid Pose Control Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.230438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.553333Z digest=sha256:d114b47f653727e11048bb02c6f3f737fc14ed97c865d47081076c6159aa3405

Observation 4d620fe6-aca3-49fb-8375-cddc6d5ba51a · outbound

This paper cites Rt-1: Robotics transformer for real-world control at scale.

Learning from Massive Human Videos for Universal Humanoid Pose Control Rt-1: Robotics transformer for real-world control at scale

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.214657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.558290Z digest=sha256:82cfcf19218bc0a2195f778163145151a615140b2ea4240cb80141b95e4d81ba

Observation 60a506c3-a396-40f4-8831-0320133b3d99 · outbound

This paper cites Humman: Multi-modal 4d human dataset for ver- satile sensing and modeling.

Learning from Massive Human Videos for Universal Humanoid Pose Control Humman: Multi-modal 4d human dataset for ver- satile sensing and modeling

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.198665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.564057Z digest=sha256:b60456debc597f4b9ff9e7766bc20aa89eaf0301859960e76f7a8a6d7c1407b0

Observation 2b2fa1f7-a3f5-4c9b-903d-acf8cad4804e · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Learning from Massive Human Videos for Universal Humanoid Pose Control A Short Note on the Kinetics-700 Human Action Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.570427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.570427Z digest=sha256:5ce93188e13ad04c6994cf7935c1ee1ec111e0b071f6b4db7667727667728d70

Observation 4b26f8c8-5f8c-477a-a893-9f635a6b3a40 · outbound

This paper cites Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.576024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.576024Z digest=sha256:11a1c8cdfe188b1bf4a52a714a63b98f9da4eb120643390f71503f7be4fd4385

Observation 57989dae-2787-40b3-9e70-b74f1bcb509c · outbound

This paper cites Expressive Whole-Body Control for Humanoid Robots.

Learning from Massive Human Videos for Universal Humanoid Pose Control Expressive Whole-Body Control for Humanoid Robots

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.581588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.581588Z digest=sha256:0f1e8e03f051db8e8784d56b47450def5dd0084ec482ac14f76f99520b232f66

Observation 3bffc8b4-0ebc-48cc-ac6c-591ad609275b · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Learning from Massive Human Videos for Universal Humanoid Pose Control VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.587377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.587377Z digest=sha256:cd72e95967a3cf5c19ca72ff5293b67b03068654662fe10953d54b65e0cc7372

Observation 278ff0fc-6389-48b3-a6f8-36114dbc7bc7 · outbound

This paper cites Haa500: Human-centric atomic action dataset with curated videos.

Learning from Massive Human Videos for Universal Humanoid Pose Control Haa500: Human-centric atomic action dataset with curated videos

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.182463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.592997Z digest=sha256:05a5196b3ed5dd813790aa067c4bde9126883b8dc18ba766aa0bc9b0cef5868d

Observation eb06418c-c5e0-478a-8ba4-05098bb59e37 · outbound

This paper cites Video language plan- ning.

Learning from Massive Human Videos for Universal Humanoid Pose Control Video language plan- ning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.166929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.598049Z digest=sha256:4ba2bdcb8c342179dc1c562c0b3709f4df19e8699f9ca724f9432ae867e79d05

Observation b922c7b1-30cd-45c7-af6e-d7455ed72025 · outbound

This paper cites HumanPlus: Humanoid Shadowing and Imitation from Humans.

Learning from Massive Human Videos for Universal Humanoid Pose Control HumanPlus: Humanoid Shadowing and Imitation from Humans

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.602989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.602989Z digest=sha256:50d7970b92396b4ca3306398630ce151df5e26543ded49c34bb2cec7853f680c

Observation e01ab8e4-c80e-4fc0-bb86-92f286ee1c10 · outbound

This paper cites Humans in 4d: Re- constructing and tracking humans with transformers.

Learning from Massive Human Videos for Universal Humanoid Pose Control Humans in 4d: Re- constructing and tracking humans with transformers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.150060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.608188Z digest=sha256:a6e4673dc8c1a0db2b94b65718704da43b1c55c11ed6656caf417ea75054e686

Observation 37186e09-e50f-4a07-ba78-b5e2a85bf85e · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

Learning from Massive Human Videos for Universal Humanoid Pose Control Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.119189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.613033Z digest=sha256:ff46d51cbff869ffad1e2387e26f2efaaae73a15c9a1a97e2c4ab798f88cf9ba

Observation 5587e8bb-166b-4092-a9f3-0626e88f4cee · outbound

This paper cites Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer.

Learning from Massive Human Videos for Universal Humanoid Pose Control Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.617363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.617363Z digest=sha256:b895fcede7a8e4e1554293887682c7a08ebb99b2b46b489225280e1ed7dcc980

Observation fe179203-9085-403c-a846-227519eeb38e · outbound

This paper cites Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.622418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.622418Z digest=sha256:4f82ef617818e9684d766d2523f917180be6bcef73743d917692a64d13c12206

Observation cac6edd0-aae1-4390-a800-6454f9c51af6 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

Learning from Massive Human Videos for Universal Humanoid Pose Control Generating diverse and natural 3d human motions from text

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.091669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.627544Z digest=sha256:9cf048de44a02d55c1143d2154fcbe41f8266dbfe2dbd9c5597f54d3182b1309

Observation 3c4be4ac-4728-433e-b851-35550912fb76 · outbound

This paper cites OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.633592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.633592Z digest=sha256:93ba26732e5b2de8ff8c31a6f0251fc94622ebfa1468332cc1130d01ba6055dc

Observation 4f5a56bb-33a0-4a9e-8214-bd25fdcb3a02 · outbound

This paper cites Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.639956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.639956Z digest=sha256:515b063fba9cde2a831a198b35658f0d3e018beef6f03febb2bcd6df7668538c

Observation 44e1aba8-3b7c-4a90-b374-45faeffde534 · outbound

This paper cites HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots.

Learning from Massive Human Videos for Universal Humanoid Pose Control HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.645736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.645736Z digest=sha256:350a342d7c65f53a4790e26532e8caefeeed8e5f6778040915dadb4d32439782

Observation 956b06a8-b3a9-4e95-b742-f724b9ffae7b · outbound

This paper cites Motiongpt: Human motion as a foreign language.

Learning from Massive Human Videos for Universal Humanoid Pose Control Motiongpt: Human motion as a foreign language

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.076822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.651226Z digest=sha256:aaa7419f9acce45ac1bba7923785a778017b80ce2a0703eb5b3e7109ea416df0

Observation 32bab377-2f38-4602-9fb6-ca890e9b9468 · outbound

This paper cites Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions.

Learning from Massive Human Videos for Universal Humanoid Pose Control Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.656230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.656230Z digest=sha256:6ed3fc628955254c5d29ca350386c6e94cec4e07fb731875f632c6b8adf922d6

Observation 908142aa-3004-4d2a-9cfe-ee310cdb2244 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Learning from Massive Human Videos for Universal Humanoid Pose Control OpenVLA: An Open-Source Vision-Language-Action Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.662385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.662385Z digest=sha256:6c9676dd0f299bc16781dc83dab1d12ea80f6f0e1efa22a16cab78a8b209a79b

Observation c0fa429e-0150-42d4-86ea-4eb0665ed30b · outbound

This paper cites Adam: A method for stochastic opti- mization.

Learning from Massive Human Videos for Universal Humanoid Pose Control Adam: A method for stochastic opti- mization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.061524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.668329Z digest=sha256:f096d2e189e54e2c915856f45142532bebd3b21b693dde5bedc94f03b85fe035

Observation c40744b4-17e1-412a-94e0-1c061b45e7d1 · outbound

This paper cites Segment any- thing.

Learning from Massive Human Videos for Universal Humanoid Pose Control Segment any- thing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.045121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.673251Z digest=sha256:39c287f8719e8c76fbca0106d3fb4128791a071b40d99088d591d091fb1ce8eb

Observation 217d7b6c-05cf-4931-a882-03cf9945ddcb · outbound

This paper cites Vibe: Video inference for human body pose and shape estimation.

Learning from Massive Human Videos for Universal Humanoid Pose Control Vibe: Video inference for human body pose and shape estimation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.030750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.678183Z digest=sha256:9bb11a5f5ff2bdb442bd70a75a45ac89c129a7ffea67eb881c7c5d0c1b2c66cc

Observation ca51d2b7-a4ec-4c68-bf3b-15e1f43cf470 · outbound

This paper cites RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation.

Learning from Massive Human Videos for Universal Humanoid Pose Control RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.682966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.682966Z digest=sha256:c108c0dfdf46927ddb5a4245c38c7b3604e46ea9f6080babe79218d7aa18ba4b

Observation 8a85e45e-db21-4b27-8c02-3ada4aa16fff · outbound

This paper cites OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation.

Learning from Massive Human Videos for Universal Humanoid Pose Control OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.687823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.687823Z digest=sha256:222878acc0f708c0ecea6810794758ff2ba0306987d338e24ccebffb53ca8dad

Observation f274362c-2130-42e9-8c26-430eff8810e6 · outbound

This paper cites Robust and versatile bipedal jumping control through reinforcement learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control Robust and versatile bipedal jumping control through reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.015089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.692355Z digest=sha256:907d66d9850157f24cffcab9a199a7694fe7162c1d38dfa441800affbb09d54c

Observation 3603b37d-41b9-4728-a233-10d76baf0975 · outbound

This paper cites Intergen: Diffusion-based multi-human motion genera- tion under complex interactions.

Learning from Massive Human Videos for Universal Humanoid Pose Control Intergen: Diffusion-based multi-human motion genera- tion under complex interactions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.998739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.696590Z digest=sha256:80aa0efe4285f08ea6aefcc3e3cddbb2a3f30cc422b13ca40638f50206e4f927

Observation f596ef4d-21e1-4a82-9deb-12573753a192 · outbound

This paper cites Motion-x: A large- scale 3d expressive whole-body human motion dataset.

Learning from Massive Human Videos for Universal Humanoid Pose Control Motion-x: A large- scale 3d expressive whole-body human motion dataset

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.982370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.702254Z digest=sha256:e69547462d5880caaf909245931ff5c7e576cc718df2e71b5129a6dde580cbc3

Observation 96108273-4c8e-4a15-8d60-c719cb406700 · outbound

This paper cites Smpl: a skinned multi- person linear model.

Learning from Massive Human Videos for Universal Humanoid Pose Control Smpl: a skinned multi- person linear model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.966149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.706861Z digest=sha256:f308daf1cca27faef5f96be14cc526fd185e727e36cb0a5e5f5e54c7e63af251

Observation 517272c3-5557-4c01-86ad-23a4afe65238 · outbound

This paper cites Perpetual humanoid control for real-time simulated avatars.

Learning from Massive Human Videos for Universal Humanoid Pose Control Perpetual humanoid control for real-time simulated avatars

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.950302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.712113Z digest=sha256:292706d3c4555fd5970f9bad6202442a041d66a547e73bed6574c6ec66ac20e3

Observation 798ca62d-351f-4ecf-9523-5480b283fbc6 · outbound

This paper cites Universal hu- manoid motion representations for physics-based control.

Learning from Massive Human Videos for Universal Humanoid Pose Control Universal hu- manoid motion representations for physics-based control

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.934415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.716847Z digest=sha256:d08daf2f0af9b5bf1bad56e06785852153050c47a010cac9255edcbc36305be0

Observation ea053ca2-a14e-4fa4-be02-ea855ee10b59 · outbound

This paper cites Vip: Towards universal visual reward and representation via value-implicit pre-training.

Learning from Massive Human Videos for Universal Humanoid Pose Control Vip: Towards universal visual reward and representation via value-implicit pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.918913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.721637Z digest=sha256:31e29408454c6dfb37d0743ee538253fe0c8bb3627bb3459cb0a0ce54859ae7e

Observation a95dbc38-23fa-4a82-b71c-d2600a61a412 · outbound

This paper cites Troje, Ger- ard Pons-Moll, and Michael J.

Learning from Massive Human Videos for Universal Humanoid Pose Control Troje, Ger- ard Pons-Moll, and Michael J

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.903796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.726818Z digest=sha256:6f33b8bcf0fa6a5501ef06ab6f595a3bd125f8318a4949d363692dc9d311d9f8

Observation 65f9f69c-0650-4628-8869-f0e6ec985c9c · outbound

This paper cites Struc- tured world models from human videos.

Learning from Massive Human Videos for Universal Humanoid Pose Control Struc- tured world models from human videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.888403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.731534Z digest=sha256:ee6f8774dc67e4a08cc5c5b25372b31a36a07d1f5ccec6ccb49e09bb9ede94b5

Observation 19274c76-e8b9-4aa1-b6bf-a7628c4a8286 · outbound

This paper cites R3m: A universal visual repre- sentation for robot manipulation.

Learning from Massive Human Videos for Universal Humanoid Pose Control R3m: A universal visual repre- sentation for robot manipulation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.871488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.736234Z digest=sha256:fa7e189289aecda6dd22d71acace00e5d784c5061286130ed5e8e26c233f2f65

Observation 99fe3a29-021e-4cb5-b20f-55cdef8871fb · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

Learning from Massive Human Videos for Universal Humanoid Pose Control Open x-embodiment: Robotic learning datasets and rt-x models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.853314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.740954Z digest=sha256:220e15af9b824997aa2a408fb8feaebee792a89def90bea54539786cfe59451c

Observation b364eb6a-cc80-40e1-b49c-5947c2333a09 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Learning from Massive Human Videos for Universal Humanoid Pose Control DINOv2: Learning Robust Visual Features without Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.745609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.745609Z digest=sha256:dc0f98667b29b726fe7cb7e8c4c370f951d463f566f4502c776c5b9dd5836e13

Observation fe1b216b-3fe6-43bd-827d-d850f6677914 · outbound

This paper cites Amp: Adversarial motion priors for styl- ized physics-based character control.

Learning from Massive Human Videos for Universal Humanoid Pose Control Amp: Adversarial motion priors for styl- ized physics-based character control

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.835861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.750403Z digest=sha256:c7fc040f7d2bbf2c7bc6839710b45d686c83aa192e7575e8ff7c885b68dcd988

Observation 10f5b3cd-20c4-41db-a2c9-55a494f34368 · outbound

This paper cites Ase: Large-scale reusable adversarial skill embeddings for physically simulated characters.

Learning from Massive Human Videos for Universal Humanoid Pose Control Ase: Large-scale reusable adversarial skill embeddings for physically simulated characters

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.819228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.755389Z digest=sha256:eef6015d4e243ba0c2129a58ded1f7d1785de287ee59d936102fe66669cc43c4

Observation 78ac66b6-88a2-4d11-86fc-a6dd473ebcc9 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learn- ing transferable visual models from natural language super- vision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.801583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.760711Z digest=sha256:e20db8a44847918f012cb5c62d5006db09c85d1249ec2b9a9d6a943e484491af

Observation ee266928-3e85-49ad-8d69-b8bb4bc1b5b6 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learn- ing transferable visual models from natural language super- vision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.781821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.765413Z digest=sha256:28dee630fbcb9d2ff969a7e987d86cb9c016905fc6495e30acd5ff0b87be08a8

Observation 0daf2c91-ad0c-4166-88df-476db7ec7403 · outbound

This paper cites Robot learning with sen- sorimotor pre-training.

Learning from Massive Human Videos for Universal Humanoid Pose Control Robot learning with sen- sorimotor pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.763624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.770997Z digest=sha256:67e25bf3475cdc02f27ad9c8913879b2fc9da8f547f364be2b962e999002d8d4

Observation 10a7882d-d20e-4a0f-97c9-dcdca3b827cf · outbound

This paper cites Learning Humanoid Locomotion over Challenging Terrain.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learning Humanoid Locomotion over Challenging Terrain

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.775903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.775903Z digest=sha256:0fe2caadeffed56b2942fabc3aff736bfa50a6e7eb38f1c8ae4652b9414f6839

Observation 2c3abc2c-05fe-4465-b3de-c461f20efa04 · outbound

This paper cites Real-world humanoid locomotion with reinforcement learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control Real-world humanoid locomotion with reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.747592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.780721Z digest=sha256:f64a7045c7ca4de5f7a10926106a17e2482e67adc928fe1e4a98f8b5d91737fe

Observation 5072d50b-85ac-4216-9185-58c8aa090ba8 · outbound

This paper cites Humanoid Locomotion as Next Token Prediction.

Learning from Massive Human Videos for Universal Humanoid Pose Control Humanoid Locomotion as Next Token Prediction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.785407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.785407Z digest=sha256:7b0b51971b53e5206ac9e592184d9e6650d4f475c2c988a5d30364aed0c3ecfc

Observation 5d758e81-3be2-4853-881d-1c3cf994a5f3 · outbound

This paper cites Real-Time Flying Object Detection with YOLOv8.

Learning from Massive Human Videos for Universal Humanoid Pose Control Real-Time Flying Object Detection with YOLOv8

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.790270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.790270Z digest=sha256:a3038638935061c22cf4c2bef13108e6ec41ef9b243fbd9fb8a4420b0e838c1f

Observation e672f7f7-57b8-428c-8d85-e373661896e2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Learning from Massive Human Videos for Universal Humanoid Pose Control High-resolution image synthesis with latent diffusion models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.730944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.795145Z digest=sha256:add4520e204e04e9fa80323c468f843423db04b7a0d6a0bf5faa630021ee4f52

Observation 254612ce-8c2f-4b76-beb8-e991700c0aaf · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning from Massive Human Videos for Universal Humanoid Pose Control Proximal Policy Optimization Algorithms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.800392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.800392Z digest=sha256:8e81149fce8169dfc8c628da54015e9296c01943e86f7711ac91bbf540cc0145

Observation 502b82d9-b4f0-43a8-96ae-960214be5daa · outbound

This paper cites Deep imita- tion learning for humanoid loco-manipulation through hu- man teleoperation.

Learning from Massive Human Videos for Universal Humanoid Pose Control Deep imita- tion learning for humanoid loco-manipulation through hu- man teleoperation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.716151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.805497Z digest=sha256:5ca400f94eca24cc5684463c9156d74ac20b7cd34f7b77d4841a619d9e4e2647

Observation 33deea5b-0ebc-4203-9b5d-2474c9124a42 · outbound

This paper cites Human motion diffusion as a generative prior.

Learning from Massive Human Videos for Universal Humanoid Pose Control Human motion diffusion as a generative prior

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.810547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.810547Z digest=sha256:1f09848ee2469ecbb84b176f347bea462cfeaed4fc14646c4b23970915f7e1de

Observation 8c3ae30a-4b76-4b83-ba13-ae8d1bb6a003 · outbound

This paper cites Sigurdsson, G ¨ul Varol, Xiaolong Wang, Ivan Laptev, Ali Farhadi, and Abhinav Gupta.

Learning from Massive Human Videos for Universal Humanoid Pose Control Sigurdsson, G ¨ul Varol, Xiaolong Wang, Ivan Laptev, Ali Farhadi, and Abhinav Gupta

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.690173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.815541Z digest=sha256:a7663bedb58cba1b2c4d51eff240fd726eff90dcb876ae961cd459c483a8ce5d

Observation 7914f6a8-3442-4fa8-a1c5-39be0e96e3fb · outbound

This paper cites Grab: A dataset of whole-body human grasp- ing of objects.

Learning from Massive Human Videos for Universal Humanoid Pose Control Grab: A dataset of whole-body human grasp- ing of objects

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.674122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.820635Z digest=sha256:adb9736a1730d0131f523626f7254d260fcf6f276fd0163eeb6d3f126f2c2b99

Observation 94c5396d-f752-4489-bb9c-a6f390277232 · outbound

This paper cites Humanmimic: Learning natural locomo- tion and transitions for humanoid robot via wasserstein ad- versarial imitation.

Learning from Massive Human Videos for Universal Humanoid Pose Control Humanmimic: Learning natural locomo- tion and transitions for humanoid robot via wasserstein ad- versarial imitation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.656426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.825451Z digest=sha256:da927d50a564599cd28751b6b01c580ed9f083a8c9817423e106e4c13d45baad

Observation ca47395a-ea0b-42e3-97ae-66ec84079f37 · outbound

This paper cites Calm: Conditional adversar- ial latent models for directable virtual characters.

Learning from Massive Human Videos for Universal Humanoid Pose Control Calm: Conditional adversar- ial latent models for directable virtual characters

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.639383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.830885Z digest=sha256:d627c55d50a63002075c7a61c1c6521b0eb2b3c45bb7c2807ba443ca86547d3e

Observation 3e5067d9-fb3e-4365-9bde-ef2346c2fe07 · outbound

This paper cites Human motion diffu- sion model.

Learning from Massive Human Videos for Universal Humanoid Pose Control Human motion diffu- sion model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.623011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.836307Z digest=sha256:762081ce7d119312ab13ac711f334461e84ac94138ae519b641ad3b07fe7a2db

Observation bd2e05f3-8ec0-4b6d-beee-a0476f0be197 · outbound

This paper cites an unresolved cited work.

Learning from Massive Human Videos for Universal Humanoid Pose Control Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:29:47.606132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.841442Z digest=sha256:756c2b06db5632f01594352ae4f359ec2e91531657022c24e35e4ba2517ac06f

Observation 0a0d4209-fb94-4419-996c-43c326de9651 · outbound

This paper cites Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance informa- tion processing.

Learning from Massive Human Videos for Universal Humanoid Pose Control Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance informa- tion processing

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.589903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.846487Z digest=sha256:421b37e948748a8050296387951d8a4105e3ce82e54023401e9baf94c75f9944

Observation c2937247-faa9-4cc8-a806-b2180b0c35a5 · outbound

This paper cites Neural discrete representation learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control Neural discrete representation learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.573317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.851796Z digest=sha256:6fe263c3144487408962e4596af211071ce2575c1882e03435911bdfb2e298f0

Observation fceb11b7-b548-4501-9ce9-34e27cc0f58b · outbound

This paper cites Attention is all you need.

Learning from Massive Human Videos for Universal Humanoid Pose Control Attention is all you need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.857444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.857444Z digest=sha256:e6944559d6fa895003f303b13f23f5f59cfaf325e88002b5145373d59470aaf5

Observation d09ae25b-e332-47ce-b25e-74e1bf9ee26e · outbound

This paper cites A scalable approach to control diverse behaviors for physically simulated characters.

Learning from Massive Human Videos for Universal Humanoid Pose Control A scalable approach to control diverse behaviors for physically simulated characters

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.546454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.862317Z digest=sha256:b62eeb90e51e8d3b32163584aa57a3656c7ca379ad2a31a5fe8d428bacef415f

Observation 8430144f-b435-446c-8bf5-dc82c031e98d · outbound

This paper cites Masked Visual Pre-training for Motor Control.

Learning from Massive Human Videos for Universal Humanoid Pose Control Masked Visual Pre-training for Motor Control

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.867522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.867522Z digest=sha256:6564e4fbd9f4c6c393fd34dcd6172aa0df4a915d7f8e12b03d9d87cdead56642

Observation 1f6e4142-c8f6-42ed-a677-7c129db4e9a2 · outbound

This paper cites Omnicontrol: Control any joint at any time for human motion generation.

Learning from Massive Human Videos for Universal Humanoid Pose Control Omnicontrol: Control any joint at any time for human motion generation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.528594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.872612Z digest=sha256:feb074c9cd4b737ec95afe78d3b2e0da6e2ee516c3c9e23d5281a2cf7d87b0bf

Observation d28dbbe8-b7a0-4635-9529-1842d532c4a0 · outbound

This paper cites Flow as the Cross-Domain Manipulation Interface.

Learning from Massive Human Videos for Universal Humanoid Pose Control Flow as the Cross-Domain Manipulation Interface

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.877246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.877246Z digest=sha256:b80d3843786931e86e01f9574043e849597a54710af0526a29cdb03aece500aa

Observation 46a9c989-981e-4610-9ec0-290bf0ef9ad0 · outbound

This paper cites Learning interactive real-world simulators.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learning interactive real-world simulators

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.512792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.881947Z digest=sha256:60525d9fd90cf1ba44cc3a6f1f17b5c1ca005fb038b107f6e3b36810429932c6

Observation e98182d0-7005-4f6c-990b-52c5682e3e98 · outbound

This paper cites General Flow as Foundation Affordance for Scalable Robot Learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control General Flow as Foundation Affordance for Scalable Robot Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.886094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.886094Z digest=sha256:b2e606ca0853dc90e99d010eb286947cf603014b461f1f7e2346d5919b413843

Observation ebfb568e-d33e-4cc0-9c8c-e98c7a8b588d · outbound

This paper cites Physdiff: Physics-guided human motion diffusion model.

Learning from Massive Human Videos for Universal Humanoid Pose Control Physdiff: Physics-guided human motion diffusion model

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.496467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.891668Z digest=sha256:5d7d7b7e304b7d080e418d8ea653735fd7e8924a7568e15dcf7efa0bc42aaa50

Observation 1330e2fd-e66e-4280-b902-9ead981309e6 · outbound

This paper cites Generalizable Humanoid Manipulation with 3D Diffusion Policies.

Learning from Massive Human Videos for Universal Humanoid Pose Control Generalizable Humanoid Manipulation with 3D Diffusion Policies

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.896458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.896458Z digest=sha256:70543acbf947008f9a9476439ee88b6a9d9b83b76c4459eb5eb432a11150a55e

Observation 21434a24-904f-4380-a395-71cde8870739 · outbound

This paper cites Generating human motion from textual descriptions with discrete representations.

Learning from Massive Human Videos for Universal Humanoid Pose Control Generating human motion from textual descriptions with discrete representations

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.480757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.902019Z digest=sha256:6a5f49c316565a5dd0f3a4c3d715fd95d86fd85101fb22c573af07f4f6e83cf3

Observation 36b9c2f7-452b-4378-a5a0-5bd034a9005c · outbound

This paper cites Generating human motion from textual descriptions with discrete representations.

Learning from Massive Human Videos for Universal Humanoid Pose Control Generating human motion from textual descriptions with discrete representations

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.465678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.906707Z digest=sha256:3bb39ab49937e9e78ccc9a3b7e50f8850099ab73244d0745d96da3fca6899245

Observation dcfd5f5c-7898-49f9-824d-a563558e11a4 · outbound

This paper cites Motiondif- fuse: Text-driven human motion generation with diffusion model.

Learning from Massive Human Videos for Universal Humanoid Pose Control Motiondif- fuse: Text-driven human motion generation with diffusion model

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.448262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.911733Z digest=sha256:2c3df952ea884c02656bc637e9ea750bc4bc9dc9922346fc69641f9e369df914

Observation e0919a9c-1a30-4b8f-ac8e-5e1529fec4c2 · outbound

This paper cites single person.

Learning from Massive Human Videos for Universal Humanoid Pose Control single person

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.431332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.916944Z digest=sha256:8afc04109992217b94e97ae13fd7e665253b9305b62303355f31da5b693a0e42

Observation 4dae4c7f-b7ae-4019-abd5-e711bcd73e31 · outbound

This paper cites an unresolved cited work.

Learning from Massive Human Videos for Universal Humanoid Pose Control Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:29:47.415315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.923121Z digest=sha256:434aaf4350081c9ce219d3801811ff60e1d7aa57a488d59812ed3181a619e1f1

Observation 54ebceff-50f1-4195-8523-0f0630a45f92 · outbound

This paper cites a man/woman doing something [adverb].

Learning from Massive Human Videos for Universal Humanoid Pose Control a man/woman doing something [adverb]

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.399475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.928184Z digest=sha256:2daa3f4ce89aab5eef0ea86d7934d44f888121a9b110f0008cfaa0cd1f4ac193

Observation 7ecde2e4-bc2f-4941-9a02-b1071c7f43a9 · outbound

This paper cites an unresolved cited work.

Learning from Massive Human Videos for Universal Humanoid Pose Control Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:29:47.382743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.933936Z digest=sha256:2b5084a06cc8dc3ab75c1e8800ad09f6272e8b2a2496172217e7a8f00c0993c6

Observation 70b81db5-3d5d-4a96-8263-430869f7885f · outbound

This paper cites in the video.

Learning from Massive Human Videos for Universal Humanoid Pose Control in the video

Reference 79

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T12:29:47.366443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:29:46.938716Z digest=sha256:934bc22f73e820f2dfd7ce2fdb475de4953b75d01c564fdbcb6b66d8948aafce

Pith citing papers

Observation e95e81a9-ca41-49c5-869e-436bdc0b47a1 · inbound

RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning cites this paper.

RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:12.977121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:11:12.977121Z digest=sha256:27017e2d7d43d362fc5d12c53ec7a3fe248d8de5e28ea560dbc69878a8d1d05e

Observation 76932038-a8d1-4943-ae15-49c66e385006 · inbound

LangWBC: Language-directed Humanoid Whole-Body Control via End-to-end Learning cites this paper.

LangWBC: Language-directed Humanoid Whole-Body Control via End-to-end Learning Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:28.337040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:28.337040Z digest=sha256:75423ae127aac854a041b77cc900246c4da19fd34e589fb9c8811e3ce222f882

Observation 84d8370a-0097-4440-8df1-dbcd5f39b247 · inbound

KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills cites this paper.

KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:00.880531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:00.880531Z digest=sha256:9d23d8929c47ab3715eeb0a42dd63b2222e561950c66a30352e83cdfdcea43c0

Observation df6e441d-3b27-495c-ad3b-49a70f3d5b3d · inbound

GMT: General Motion Tracking for Humanoid Whole-Body Control cites this paper.

GMT: General Motion Tracking for Humanoid Whole-Body Control Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:37.670272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:37.670272Z digest=sha256:a35ab1ab86986f3090965dce1a9fcc919c7d903feba3a2c035d0809e7ada6fcf

Observation b180192d-dceb-43b6-9bd1-e13689839634 · inbound

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation cites this paper.

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:30.082678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:30.082678Z digest=sha256:154f214e5a9b6ee638039cf323fe2c85fdf0e4d5c57e6cd4e7c57d7612bef2d1

Observation 1152f504-9176-47a9-8d57-31cca7cfe022 · inbound

OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation cites this paper.

OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T08:37:20.378298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:37:20.378298Z digest=sha256:28646bce4ca0eb1d0818d410c1dbe542d712c6ddc22e949a39b405e284e45829

Observation 85e8596c-7c13-4273-9e09-2ee668e70951 · inbound

Re$^2$MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement cites this paper.

Re$^2$MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:11.606859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T05:42:58.700609Z digest=sha256:ce5440f3bc2296c8206647e237a77f298b5ccffc653e6c576b7e1a711cd65fd7

Observation a58d7ac7-c965-4497-be50-23df6d32eb8f · inbound

An LLM-Driven Closed-Loop Autonomous Learning Framework for Robots Facing Uncovered Tasks in Open Environments cites this paper.

An LLM-Driven Closed-Loop Autonomous Learning Framework for Robots Facing Uncovered Tasks in Open Environments Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:14.659231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T11:27:34.992950Z digest=sha256:bb48ee8cd1806d5cf0275e56cef12be34b2956e1e6f180017cb643e8ecadb685

Observation 7f5625e8-70d6-4949-9b5e-9488ba635e50 · inbound

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking cites this paper.

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.700049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T10:01:00.409280Z digest=sha256:5d68c7be67f3f9cb6fb169046d4ad13ac76e9380259c8eb0483d3df59ea60114

Observation 67260d33-1e88-4235-aee7-b73dc2155708 · inbound

LIMMT: Less is More for Motion Tracking cites this paper.

LIMMT: Less is More for Motion Tracking Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.550462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T22:13:43.198541Z digest=sha256:0cc12bf36668bff5d59bcda99a1c996abf57cc9e5722d7ace53ca6f3fa9ab4af

Observation 05cf4f94-8c67-41ae-9ab7-c4f45457429a · inbound

ThorArena: Benchmarking Humanoid Physical Interaction with Human Motion-Force Demonstrations cites this paper.

ThorArena: Benchmarking Humanoid Physical Interaction with Human Motion-Force Demonstrations Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T18:15:21.362695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T18:10:36.258268Z digest=sha256:582f4d0ffe4f80479288b7b4398bbe7135bb3605f0b5a6aa6c8a8611d7343895

Observation bbb99c54-894a-46b3-aeec-a6c55834a011 · inbound

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval cites this paper.

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:25.294813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:39:25.294813Z digest=sha256:73ee414c657771135b49681a64ab5a669ce85fafd2b621b8c39feac68686807b