Pith. sign in

Paper Citation Record · LEDGER

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

As of 7 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 70 inbound Pith citation observations for arXiv:2409.16283.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.16283 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T12:17:01.294466Z

measured 131 of 131 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 70 of 70 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:32.082618Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T23:07:47.692433Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact29
  • verified fuzzy28
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4f97b871-d623-43ec-9b5a-a0bc11ef5c90 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:17:01.334605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:9c421fcdccb68ada29f2440ba2b7aa7ab5e7beb5110d7e82b65d7c5f13b560ca

Observation 3b93eb2f-c989-4801-af2e-e190efae2707 · outbound

This paper cites Roboagent: Generalization and efficiency in robot manipulation via semantic augmen- tations and action chunking.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Roboagent: Generalization and efficiency in robot manipulation via semantic augmen- tations and action chunking

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.481825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:87979b224b417d04e43bbf9c75c2ed670223a76eca83fcaba89208203c7c579b

Observation ba186ead-e92f-4874-8457-a98a5ac5c30b · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:17:01.462887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:d8de48f4155c8176aaf919c09998456d057b53fc7de5019d0a937a7d4fafcab1

Observation 1bad00ee-ff0d-48fd-94b7-128f8f959599 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation R3M: A Universal Visual Representation for Robot Manipulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:26:54.148999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:b05464a4d8b9198748ec62bb9f53cc97e4770dbe2232ebfb36687001fd74978c

Observation 915ccc8e-caa9-48fe-99e2-fae756108ea1 · outbound

This paper cites Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:17:01.472248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:e86b986827052e81c4e315acade59765d7deb5faaf623ca6ce6650fae45cff3c

Observation 89622009-47b5-4a5c-9250-4378c5632986 · outbound

This paper cites Masked Visual Pre-training for Motor Control.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Masked Visual Pre-training for Motor Control

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.477837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:f3fb94883fc976e4b25e5f19c3b878404cad5559e6c20aaf84f984daaf72376d

Observation 30448461-fdfb-45f4-9876-5fdca78141ef · outbound

This paper cites Language-Driven Representation Learning for Robotics.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Language-Driven Representation Learning for Robotics

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:17:01.343239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:5ed60b21773ebbabc265f29e51982c83635cde45512a293c0fbe0ddbac2f986b

Observation 1d128787-7f65-4e04-8018-f127d2e73f90 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.505902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:bdbacf1871195018057a9c09ca39685b35817a58a2f25e044986e4cc75d8c4ba

Observation 6e107635-096c-4b27-8752-4df6377de08a · outbound

This paper cites Dinobot: Robot manipula- tion via retrieval and alignment with vision foundation models.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Dinobot: Robot manipula- tion via retrieval and alignment with vision foundation models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.509480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:702450004aa561a3be9260765eb60c49ad6d50054d5dfd98b5bc35329cf19dde

Observation d0a3f88f-dc44-4ec2-908e-15fb29196fa6 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:17:01.351949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:e1ab9f127b238621f24eebf2c34d5a58be922d645de4de58ecffbb4377e24915

Observation fed1f86a-8ef4-4aa4-9221-611e5589cb68 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:54:59.300665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:e87c925a5f86524dd124854a7625b2ffed2f4b65ef84250aff51e38b361f0d70

Observation 7bf6bd01-1c16-4381-a345-bc3d2381430f · outbound

This paper cites Visual affordance prediction for guiding robot exploration.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Visual affordance prediction for guiding robot exploration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.521635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:2833ba1c79e027233fed6ce3e992b07c493627daa4f5417153efb70a5c490b3e

Observation 696d368b-c640-40d0-9210-356759a017ca · outbound

This paper cites Dall-e-bot: Introducing web-scale diffusion models to robotics.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Dall-e-bot: Introducing web-scale diffusion models to robotics

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.525259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:af7f0ee481788dd1ae472b773840c17e8d662d154fb0ce86f9afe673c253555a

Observation d24fed34-43c7-4fbe-8da1-6949036ba947 · outbound

This paper cites Towards generalizable zero-shot manipulation via translating human interaction plans.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Towards generalizable zero-shot manipulation via translating human interaction plans

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.529477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:09f8a37e469056e7fddf27d68bc5b2d41d2c8d6a09be56e15a5c6e8cc2776b6c

Observation a301df75-c8b9-42a5-bdcd-b266e91c134f · outbound

This paper cites Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.361481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:42a2fb5edc04f5da517af2017787048550d8388595778ffceec586d92a709c29

Observation 12b802e2-f91b-4393-b4ce-c66e0fd20f86 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.365911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:74efa06cf1538eeee3b6094e6198aa5ce01afe6a7fe109c97bfa2ef2f30dd2a7

Observation 5b89bc29-c0a9-4561-8604-3da481035a2e · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:17:01.369872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:012dce9ca29f868cb228e10d3c2f6fc8f5fc058a721b3bce548b5e97cf883c6f

Observation 4c9b4171-7598-425f-a432-73f44bacb461 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:1f7a8ae9fe7db0ffb7ad485ff3e34d300d91f9a05addec423c12b310be75bc66

Observation c3296676-0b29-49fd-851d-9668b305ccaf · outbound

This paper cites BootsTAP: Bootstrapped Training for Tracking-Any-Point.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation BootsTAP: Bootstrapped Training for Tracking-Any-Point

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.380972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:9fa1139562f34041031bba4dc205b62c7401e70ceb1b69762600938d01f0b676

Observation 4da3e78a-1ed9-453c-8c9e-e36291f47e62 · outbound

This paper cites One-shot visual imitation learning via meta-learning.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation One-shot visual imitation learning via meta-learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.555998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:c3ac36e01fc3e2d78f0d4de3648f479c76c4f213673f40d33dabe3fc2ec181e4

Observation bedfeb06-d08f-4294-a5d9-e4cbe913fd77 · outbound

This paper cites Visual imitation made easy.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Visual imitation made easy

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.560637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:563797d7a376ed3b852b574f278058e585430673649cd4d20e07861a20b04cc1

Observation 962859c3-b178-4b08-85be-18cc97dda022 · outbound

This paper cites Learning monocular reactive uav control in cluttered natural environments.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Learning monocular reactive uav control in cluttered natural environments

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.565517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:47da756b3dd437ac971bebb4472a70afa258f59b84b7d5bf6c3eb1b736f2b0f3

Observation fbbcf40a-aff2-43b2-9203-c776689e9c3d · outbound

This paper cites End-to-end learning for lane keeping of self-driving cars.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation End-to-end learning for lane keeping of self-driving cars

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.569872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:db688dddc21fd932eb52edc6727cb0fd74f47d5a8c4fe7ee0c29008fb2d96837

Observation 702daf53-1701-42cc-9727-dfda05432924 · outbound

This paper cites Roboturk: A crowdsourcing platform for robotic skill learning through imitation.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Roboturk: A crowdsourcing platform for robotic skill learning through imitation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.574405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:5670c6255a22a9bd2ce5f4c201b3a87a2f5c8a1a315c94a10a0a5439eeeec3d0

Observation 0e53352d-3a35-4546-8242-8512a0c39ee3 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.578336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:9c86504c05f5c1b2f1324b1ac2b06e0b244bb55f6bbe7fff463dacb83fed37a1

Observation 451b5ca9-b40e-4b82-a48b-1e5e81e6a3bd · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Bridgedata v2: A dataset for robot learning at scale

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.582255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:791a4fecdbca5fe08bb33798bf2856c98ed4fcd5a451884a475cbe1804d4bd77

Observation 6a224b93-d8df-4195-9e8f-7ce3ebabfcc0 · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Scalable deep reinforcement learning for vision-based robotic manipulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.586306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:7a4e35eb5924713220e33c1132b96068a9e9ac2a95056bace332c7d56d4755a8

Observation 88f130d4-c547-4af2-884c-8766413511e5 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Ego4d: Around the world in 3,000 hours of egocentric video

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.591657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:367119230bc326833d8955d8717631d39b777f280b1ce29a524662a049ab1483

Observation 1ab5a11d-ce63-4eaf-ad70-27358bb1da6c · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Imagenet: A large-scale hierarchical image database

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.595909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:93acb4aaa6054d66c3f20c0c9114aa0ac8f8367a166f4854aa40172cd8017cc1

Observation c791f4c9-132e-4936-a2f3-ec9ee5ae04c0 · outbound

This paper cites VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:17:01.386002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:00519297c1f04dabc16f382bfa831dd4bab3a05f50a4c79181fff5957e0e847b

Observation 9b659480-c75f-44e1-92b2-175ace08e773 · outbound

This paper cites The Unsurprising Effectiveness of Pre-Trained Vision Models for Control.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation The Unsurprising Effectiveness of Pre-Trained Vision Models for Control

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.390293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:7ca668edee6ee75fbee9a742e9d04cb208ba8c55d39cd4c5e15b981717757b83

Observation 94abe090-7b5a-48b4-9184-c406241bf3a2 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:17:01.393862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:ae2fefa05da061389e43f3c887191161f050f91927ec3a210a2feccfc6a7cf17

Observation cb68f435-9ef9-496f-9a99-9d4e3f324fae · outbound

This paper cites Video as the New Language for Real-World Decision Making.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Video as the New Language for Real-World Decision Making

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.397994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:f1c98798e41e40858e13638e78d845bf144e913c88ae2b20bbe792fc64fe539a

Observation 6e114bad-eacc-4fe1-806d-077627ac3e91 · outbound

This paper cites Lossless Adaptation of Pretrained Vision Models For Robotic Manipulation.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Lossless Adaptation of Pretrained Vision Models For Robotic Manipulation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.401990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:54aae3e9dfa38a7a85d0c6365d29f1852e6f24d860ead4b17ff79c94cde33dad

Observation 9e91d310-6ef7-4c52-8155-d01b87350c2c · outbound

This paper cites On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch Baseline.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch Baseline

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.405426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:26a226523265b7d6ad00f110f77fcddcaef11532b02efe28b6a15cdfb584f902

Observation d4b4610d-f874-4e4d-8bde-0ad1a9df96eb · outbound

This paper cites CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:17:01.409154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:3f415c9d9e42c478d0c80010801a6be7bbc3e5611ad1546457d79d7ac1cfbe25

Observation 2eface9f-0b69-46e6-aab5-ff9407666e0b · outbound

This paper cites GenAug: Retargeting behaviors to unseen situations via Generative Augmentation.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation GenAug: Retargeting behaviors to unseen situations via Generative Augmentation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.412942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:4559bddcb6cb754230d6624e4efff237951ddc23e7192fe519559606ad318a88

Observation 88853033-5351-4e2f-93a6-1eeb3d9c76af · outbound

This paper cites Scaling Robot Learning with Semantically Imagined Experience.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Scaling Robot Learning with Semantically Imagined Experience

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:59:10.716066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:026e7dcee2de3c074325b9352efefface06f4e1b42ab7024cf43061495dc6958

Observation d2c9663e-4c73-4ba7-b74b-28ccb40ff8d0 · outbound

This paper cites Semantically Controllable Augmentations for Generalizable Robot Learning.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Semantically Controllable Augmentations for Generalizable Robot Learning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.419280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:42cb961ba2d4978076c4c13908904f5009d5c816b35bc1a568a2b9ef6753876d

Observation 70bbd1db-041f-4bd2-b4ac-14488ad25a6f · outbound

This paper cites MimicPlay: Long-Horizon Imitation Learning by Watching Human Play.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation MimicPlay: Long-Horizon Imitation Learning by Watching Human Play

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.422183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:a75ef393d5c5266712be0895671b0c6f8a7af605fe7cff2b584e01fa2a1e3c99

Observation 3cf8d654-abd3-400e-9331-bac419315b2b · outbound

This paper cites Avid: Learning multi-stage tasks via pixel- level translation of human videos.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Avid: Learning multi-stage tasks via pixel- level translation of human videos

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.551434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:d5f4767eecc7ae5026cea2e9524b6c37329e04cdd0a2445b2e6a2f3c859be2c0

Observation 3b50430a-ba3b-4f55-b2f4-99b805b23cec · outbound

This paper cites Learning by watching: Physical imitation of manipulation skills from human videos.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Learning by watching: Physical imitation of manipulation skills from human videos

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.485775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:0f6c362a4dca72eb6daa009cb88989fcdfa41d621bbfe4ed17d77b133e1b6b84

Observation 1f54d54d-777c-4a64-807e-a2aad2cf201b · outbound

This paper cites Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.425532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:d056037e134fa4d2b74bc994657e480ce06f7bd9506e63f5a4083cf37cc9fd61

Observation 9d923a03-5578-4a3a-b4c0-ed998836fb2c · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Any-point Trajectory Modeling for Policy Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:32:51.209901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:8676f32fcc70a7c01125beeab7f5b220f17c79224582dcec7464488857ed9c97

Observation c96e3811-4559-41b1-b808-f6cb0002dd61 · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.433049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:d14b644bf381e01a33715cccd9dd4145b8cc30895124200e371c880032e220da

Observation 64fcfebe-f006-4170-85eb-a28f636074d3 · outbound

This paper cites DexMV: Imitation Learning for Dexterous Manipulation from Human Videos.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation DexMV: Imitation Learning for Dexterous Manipulation from Human Videos

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.437100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:1dc7864464e59059b2a70c2a764e915b417aaa92da39548e8f338ba4dcfd435c

Observation 4245b2aa-722e-4f7a-a76b-1ca923fe1e8b · outbound

This paper cites Videodex: Learning dexterity from internet videos.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Videodex: Learning dexterity from internet videos

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.513063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:9972dc1f44d9fdc42153e1e92720db5d16f0ff2ad0558f2fdcc4c2884e684052

Observation 4e1848e7-fb48-4a1f-a269-e490b8e25ee9 · outbound

This paper cites Where2act: From pixels to actions for articulated 3d objects.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Where2act: From pixels to actions for articulated 3d objects

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.516846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:db4f14bdbdec8d75d04dbf7f33fe28ead9e42902a0d6d0d722f7805b321ca80f

Observation 1d3fa5f2-3831-49e3-98b4-7cb33c6eba2e · outbound

This paper cites Human hands as probes for interactive object understanding.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Human hands as probes for interactive object understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.533627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:1d5f924dbfa6f12d171846f7d3d91f0c29cce8a47df4c5b7cb9070a8a5d4b6c4

Observation 91261a19-8a5c-4e5b-b19a-1bf4baf0218c · outbound

This paper cites Affordances from human videos as a versatile repre- sentation for robotics.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Affordances from human videos as a versatile repre- sentation for robotics

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.537571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:22bf6ff06d91eb964ed1e2c5e6854ed26ab39bb36721b92d0a55431787c8efcd

Observation a5966486-76eb-450b-b99d-79fb33993a1a · outbound

This paper cites Joint hand motion and interaction hotspots prediction from egocentric videos.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Joint hand motion and interaction hotspots prediction from egocentric videos

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.542481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:cfcbbf76a88c34e0220f1f284941d28331b7cc771bfc9acca3c273218d024529

Observation 0e72634c-d6a0-4311-8398-3e1dea8e9aec · outbound

This paper cites General Flow as Foundation Affordance for Scalable Robot Learning.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation General Flow as Foundation Affordance for Scalable Robot Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.441350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:6e291956c750df1229f241a9e565222f3b2a1c6bf37ab536ecc2218b8662d81b

Observation becc446e-6b9d-4e36-91f0-d29dac873279 · outbound

This paper cites Human-to-robot imitation in the wild.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Human-to-robot imitation in the wild

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.489583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:9dc9cc420dfc927d0280820032988ee5e2c6672f4ce4ebad5ed590aa5dc76b32

Observation 2b69c09b-b1ef-4666-b09b-fd3893a9cedb · outbound

This paper cites Learning uni- versal policies via text-guided video generation.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Learning uni- versal policies via text-guided video generation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.493353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:cab21321909eb6176700bffc050bba3676126c0e9613db559c304367d61162e7

Observation 7bf51a50-cdb3-450e-b276-1f093362f7a4 · outbound

This paper cites Dreamitate: Real-World Visuomotor Policy Learning via Video Generation.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.445669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:1c287061d2af7e491acd405de690b9a9677283685499291ba608f89a5c0982ce

Observation 900bbbef-4de8-4dec-ab25-df54be801184 · outbound

This paper cites R+X: Retrieval and Execution from Everyday Human Videos.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation R+X: Retrieval and Execution from Everyday Human Videos

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:17:01.449995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:25a09a0dec6b669d881f92e4e5900549aca09cc3ab221f18e6ebfbde81f3b7fc

Observation a06c4e4c-ec49-4a0d-9d53-a46f67e4747c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Flamingo: a visual language model for few-shot learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.546960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:e2f9904bb3490f7b31cbbdae35a6dbfe1c8b459de0718791cef2ef63c10ec4e1

Observation 6d351fc8-c2ce-42c3-b269-b3351e08da37 · outbound

This paper cites Tap-vid: A benchmark for tracking any point in a video.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Tap-vid: A benchmark for tracking any point in a video

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.497655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:8d50156473c96fcf855e4b65e73797a104fe58135823c3e544ba3c939d6480de

Observation 7bb7d616-0b54-4b64-8b57-ec4ccfe51c8a · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Gemini: A Family of Highly Capable Multimodal Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:17:01.453650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:33f9d7b3f171179d2850601052e8da995aed3655766f8f974b17039d2307ec77

Observation 456b8d1e-b1a5-4096-9054-ee170a717df0 · outbound

This paper cites CoTracker: It is Better to Track Together.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation CoTracker: It is Better to Track Together

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.458795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:924f64260c4509b93b420143e9bcee5de5ba53bea89114f941ed683f455756ab

Observation 0876b1eb-637b-4417-b8e9-5e2d9a722a20 · outbound

This paper cites Dense optical tracking: connecting the dots.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Dense optical tracking: connecting the dots

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:17:01.501918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:3ec0e397e90cb9e124179fa6f18c538000cc63963025664d8134dbfe949a8425

Pith citing papers

Observation 7ca60efb-2d6a-426f-a05c-562b94407f37 · inbound

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation cites this paper.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:1310eb7f95e4054f9928ffffc09bb4e92f20b406ff1735e33f26c061113a0401

Observation 08fcb0ee-3113-4928-b7a7-419907738d21 · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:23cbeb3b6b937d8a53b0bb9e63b3dc4cfdedcbf283a29615ac9be74b57301e45

Observation c55d477a-ce5a-4eb9-b509-48ea5ed0074c · inbound

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots cites this paper.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:7cbad9088cd8cfd605704ac8697e2d0ef55c436c96b823b983c1c1c65eb2e061

Observation de8b0878-736a-499b-8dbc-6a840b28d918 · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:21:45.233665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:9974562e61e7189534b15c6bfd5f1f29635bb7f630ee18492e96a63a32c9a9b6

Observation 2f036991-61f9-4386-8a23-b7724c83f23e · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.347709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:64e740fbb8580519eda1bf1d9a6003b4d23082f3ab59eabc146b82bcaef3c7f5

Observation e564a0c1-f300-4ade-b9a8-5fca9b62ce10 · inbound

DreamGen: Unlocking Generalization in Robot Learning through Video World Models cites this paper.

DreamGen: Unlocking Generalization in Robot Learning through Video World Models Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:50:45.515634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T23:50:45.332466Z digest=sha256:4852b3d4864584f0f692283393c56322985fb25eaa60ce414bd9055d5ff5a075

Observation 1d7c1656-b81d-498e-b6b0-42b23703f075 · inbound

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning cites this paper.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.082618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.082618Z digest=sha256:2291a5a204fcb462163c24c312eed3c1ab0671e2ab199f12a06d9ed2d24945dd

Observation e47415ba-09ac-47d7-9104-881ef902dfbf · inbound

DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation cites this paper.

DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:39:28.179162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:39:28.179162Z digest=sha256:02af316a8c69844c13b290f1e8324eb94cc2b91d15aefeadb9428e07a4b3d4d6

Observation 8e2cb8d1-ed23-4764-aa5b-4309cf49ab13 · inbound

WorldEval: World Model as Real-World Robot Policies Evaluator cites this paper.

WorldEval: World Model as Real-World Robot Policies Evaluator Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.746221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.746221Z digest=sha256:7490177cee1827bdae6c97f489b36daac8fc8e59bfad725199e09c5f9bd465b1

Observation 0152c7f7-8fb9-4bd7-8526-17463871aff0 · inbound

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt cites this paper.

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:26.963771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:26.963771Z digest=sha256:fea77a3dc10d508a553b5492830a7897200c218b6cf184c06b8e63f187644066

Observation 4acfe0af-982f-4a83-b4c1-d88260a07428 · inbound

3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model cites this paper.

3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:29.685470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:03:29.685470Z digest=sha256:fa9c799888edf481cd0ec9651b4bea099c52803f1c7006f7d92aa2750691163e

Observation a8a80014-0c0e-4677-8b9d-e979157df557 · inbound

RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation cites this paper.

RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:14.317347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:14.317347Z digest=sha256:0a725f2ab3d4d30a95b7b140f278dffcd535651a5333c0a77c4ae2211847809a

Observation a6364e67-ae71-4d78-943d-ce0914886f8d · inbound

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations cites this paper.

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:37:07.494124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T06:36:13.144868Z digest=sha256:c89cfe4aa9b4e47574dcf79ca4a4d3a4b55f6f07eee9486e66e3e04642d4469d

Observation 68102722-7369-4727-9a9b-da6f7232ef04 · inbound

Geometry-aware 4D Video Generation for Robot Manipulation cites this paper.

Geometry-aware 4D Video Generation for Robot Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.918399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:6d5a5adac8ef58b487a4960b470ada27f205df833557d4a1a915c2261e1bef52

Observation 053b7edf-4cb6-4d7e-b763-9f36948291d2 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:42:41.440781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:750400e5ccb7f15a419b040c53357d913ed06cdec4b587c3fe36b2496b3af62e

Observation 1c0632e1-bbf7-457b-8dad-1a2980ac07c5 · inbound

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation cites this paper.

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:52:57.533297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T03:52:18.984005Z digest=sha256:74ce987e3dfe3b62f7599276f382b6587153e97453d596cf7159548144127d6f

Observation fdbee5e9-1ad2-4191-af9e-e864ff7d47d5 · inbound

Vidar: Embodied Video Diffusion Model for Generalist Manipulation cites this paper.

Vidar: Embodied Video Diffusion Model for Generalist Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:54:28.352175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:54:28.271928Z digest=sha256:3334212e9da50d899fb16bfb2f8bf6de3e7f6d1ad7b404dab867b2504bdc8f68

Observation 78af8e13-775a-43df-af28-0c2d8a3144e8 · inbound

Precise Action-to-Video Generation Through Visual Action Prompts cites this paper.

Precise Action-to-Video Generation Through Visual Action Prompts Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.357011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.357011Z digest=sha256:fd85a5c109bb5b16bda1c1d3a23d2fa6d946190c33eb275581386d8af74388c2

Observation f9efeae9-6188-4694-99fb-6121a36c3eba · inbound

Deep Sensorimotor Control by Imitating Predictive Models of Human Motion cites this paper.

Deep Sensorimotor Control by Imitating Predictive Models of Human Motion Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:54.474371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:22:54.474371Z digest=sha256:fd33efe0f75592f8e62181c27b76b0cead52ab56b647450ca88264b4d306104d

Observation 3444e9b6-0d1d-4710-a6ac-97c920e5c615 · inbound

Ctrl-World: A Controllable Generative World Model for Robot Manipulation cites this paper.

Ctrl-World: A Controllable Generative World Model for Robot Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T01:14:10.261976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T01:14:10.174044Z digest=sha256:f8e56268f0148240e6feefd3a5b782936e7e93e961a1e6d74d2f5bb9e38bc064

Observation bdbbbb16-f417-4aee-87a0-c18cc8b12af6 · inbound

Motus: A Unified Latent Action World Model cites this paper.

Motus: A Unified Latent Action World Model Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:44:36.636455Z digest=sha256:9dd4f4ecf4c702d278ea4253869f9ff39b579d0985bba9d7a496f088b7e67ee0

Observation f14a9dda-1b5a-413e-a890-69e7d6bbbf4e · inbound

Large Video Planner Enables Generalizable Robot Control cites this paper.

Large Video Planner Enables Generalizable Robot Control Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:28:34.008786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T21:26:32.048309Z digest=sha256:6e69c20c73b474dd41ec62a0f79a81c6ad27504c823045e217ff33c1a03f8f06

Observation 9b77c962-01f2-4379-83e6-cbd9c73f76aa · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:08:01.753463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:aa41f2dda998487d157c44c0722a9cc5766f37b47339b333483310c8a611adc0

Observation 28d3a22f-0137-4217-af38-8017209e66c7 · inbound

MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation cites this paper.

MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T02:45:06.654960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:45:06.654960Z digest=sha256:8c782efd03072e1a1bd686f6911adf58d057ee7ad127386af2f62da5b47d4606

Observation 0c6102a6-9dfc-4bfa-8d26-9b49bd797568 · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:c75374c4c102b8274c4ea30886cb77168051367d2e7d103508ac760a3175bdc9

Observation 891af52c-2463-4119-b862-c15f84c4a912 · inbound

FrameVGGT: Coherence-Preserving Memory for Bounded Streaming Geometry cites this paper.

FrameVGGT: Coherence-Preserving Memory for Bounded Streaming Geometry Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-15T13:07:10.012493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:07:10.012493Z digest=sha256:eaaad00586a6c78dae620fa1e6a5f208415ca1477ed5707e97ec99f1de347457

Observation f2d46835-fbf0-4424-a8c4-fcb4f6c7f964 · inbound

Fast-WAM: Do World Action Models Need Test-time Future Imagination? cites this paper.

Fast-WAM: Do World Action Models Need Test-time Future Imagination? Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T01:57:29.753735Z digest=sha256:0f4ea4c3f57a46effe6ec8a78580a21184ea0213c16570cf4aca53fa83437b58

Observation 73d18080-2949-4d96-a4a2-3d0771e32286 · inbound

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model cites this paper.

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T18:54:07.081457Z digest=sha256:dcb6cedb3a40a32b3a980914743c9f4bf12c2b34fe2ad965ba5ee9d6706b923f

Observation 64eeb5d9-16fd-4bc6-80fe-b7227beb37bb · inbound

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data cites this paper.

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T17:02:18.358675Z digest=sha256:d1a9ce8290190b579020eec575721614623de107a1fc8ca4718b19036a8c89bc

Observation d40b12db-4432-4f63-a553-0046a7d12d87 · inbound

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning cites this paper.

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:12:08.970164Z digest=sha256:f7275708441a8b435d0fe64a5e8ba5f30e88083ad19e918966693e2f50e77b98

Observation ac914cc8-773e-430a-ad58-95cb105f6045 · inbound

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence cites this paper.

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T00:03:53.609175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:03:53.609175Z digest=sha256:ae61b32a19de916612bc899a553761010a44620667d3e51763a9ed565ea7b6db

Observation 697cf87c-13a5-452f-bf8d-47e914d4b7c2 · inbound

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations cites this paper.

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:14:24.932972Z digest=sha256:ad5c6465864b7e5805574898c95a96fd336a369ae7d57c81073d985e478954b5

Observation 9ed6ea99-5520-4157-b64b-6cbb56d431dd · inbound

StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement cites this paper.

StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:41:45.903419Z digest=sha256:01cd86ff5824244393517877f5c68983465fab892f237a2cbc957bbf29c8c6dd

Observation b9a8750a-1e4d-453c-9e44-81ca4daf7274 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:02:26.084185Z digest=sha256:571b323a1183152aecce4c437f06c5326e0be5cde71729c8813a0893891ae763

Observation f70af589-066d-4655-9fca-f3b34c5c1e76 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T06:15:32.881140Z digest=sha256:f786775183455a3b3e083f5e2dbe1b77035cdd6f242d71c519d5a95ff3fbc9af

Observation bb78bb12-eb5f-4aa0-a2c1-0b2778fba617 · inbound

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training cites this paper.

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T21:26:26.540403Z digest=sha256:62dd6e723d5423349c864fab097a1e88a85c46734d5d9e12bd2244586deadbdc

Observation 31aa16e6-0fe5-4ffc-8895-a21d9c8e80c9 · inbound

GazeVLA: Learning Human Intention for Robotic Manipulation cites this paper.

GazeVLA: Learning Human Intention for Robotic Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:37:30.784513Z digest=sha256:e289bf7a228f32dd83e874b98a3d7320ac8be2314354e5a95be6a40eb35dcc43

Observation 141045d7-d20e-4e99-be37-c166644b2efb · inbound

Being-H0.7: A Latent World-Action Model from Egocentric Videos cites this paper.

Being-H0.7: A Latent World-Action Model from Egocentric Videos Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T20:48:01.461993Z digest=sha256:4ec9a9057cd643a4655facbfaaf42b821caa10f680abce96a7220837dd1cc715

Observation 4c9a9d0f-25a6-4b62-8685-10ac1c1ab5f6 · inbound

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation cites this paper.

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:28:30.471533Z digest=sha256:eeb2d19ea63d22c1362318a4b4885b67d0682717d1fdbdff27970e898993e174

Observation b1e1f89d-d7d3-4f70-a098-bf0644911523 · inbound

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models cites this paper.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:fdb3f04b249a722d9c5f731457a713ea2a802f8f4aec31d2492bb2e30d718886

Observation f1b5ec9d-cefb-4681-ae82-0a3516090585 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:733efee1877e9a2c429fe806dbf849f6a0a87b375e51657398a92531b5997ba5

Observation 01d0c9a8-0e14-478a-ac8f-656b4e255c40 · inbound

CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL cites this paper.

CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:42:08.454469Z digest=sha256:8a0aec64b87588b29b0cef4df51a0b82641ae99219549362a93818d0a2bfb818

Observation 23be7d0e-1de1-4756-b3e9-aba1039dceae · inbound

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation cites this paper.

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:36:39.753370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:34:03.684055Z digest=sha256:290956ebee6dbec175ba26c6da09a833089442d4dbfe526252c1003bef2dfb8a

Observation 77e6370d-6b84-4122-bc2f-f5b22980e142 · inbound

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation cites this paper.

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T16:54:59.421179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T16:45:15.955954Z digest=sha256:2ec23f0a3891947687aefceeda39413c7c93f7f1400632cb062f6231ab1645c1

Observation 9b91dbfb-c731-47f9-b735-1f89e1ea9c8d · inbound

Point Tracking Improves World Action Models cites this paper.

Point Tracking Improves World Action Models Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-25T03:56:37.114930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T03:51:10.921839Z digest=sha256:804d709edad91cb0351bb549c67f0462dbccd3e2b01f95aa55cd318f9428b53d

Observation 2dcceec9-528f-48cb-95d5-9a3bf11fc57e · inbound

WALL-WM: Carving World Action Modeling at the Event Joints cites this paper.

WALL-WM: Carving World Action Modeling at the Event Joints Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:36:22.454504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:15:14.454649Z digest=sha256:75dc9a3d2118adebd3ec9861d62156fe363487cc0df927296cb043472adfa022

Observation 4a9f0217-97f0-4ce5-87f6-940280aed52f · inbound

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation cites this paper.

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.342412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:28:22.430294Z digest=sha256:d57d3912746943900944a278ee57b2c7dee841cb28f780dcd1e0d3fb695f8b3b

Observation 1a7302a2-5a9c-4a76-a284-65f387db816a · inbound

Dreaming when Necessary: Advancing World Action Models with Adaptive Multi-Modal Reasoning cites this paper.

Dreaming when Necessary: Advancing World Action Models with Adaptive Multi-Modal Reasoning Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.372503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:59:35.783793Z digest=sha256:da76ab15e4e0f5c089a0410fef69a13e283195f960c437b5258a0b63763b46d0

Observation 8a97aa7a-dd40-4df9-b000-dc7efa83f257 · inbound

Light-WAM: Efficient World Action Models with State-Fusion Action Decoding cites this paper.

Light-WAM: Efficient World Action Models with State-Fusion Action Decoding Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:07:23.975665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T19:55:31.100331Z digest=sha256:7065e2f8bb86ea5cc991b2ea633d78f1b0cf1997288ab8a7bdd2cda7f7453d7b

Observation c2f94f4e-2f61-418e-9306-efffbdd95c89 · inbound

$\omega$-EVA: Envision, Verify, and Act with Latent Interactive World Models cites this paper.

$\omega$-EVA: Envision, Verify, and Act with Latent Interactive World Models Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T02:07:33.688051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:10:02.176202Z digest=sha256:a944f0430e6e74b4a59677441efaef90188d99fae354b9d7c49c0f98c5881852

Observation 478088c1-682d-40e2-8e20-817b3da14f7c · inbound

Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations cites this paper.

Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.569488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:31:04.351106Z digest=sha256:c73eb17a5e854a61bbb91ce48e7054d734fb186310451ac074f74a4349fac0cb

Observation 0d3362e4-a8f0-4768-90a5-e6f5bf5c60e6 · inbound

Next Forcing: Causal World Modeling with Multi-Chunk Prediction cites this paper.

Next Forcing: Causal World Modeling with Multi-Chunk Prediction Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.518497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:31:53.905704Z digest=sha256:6099a9cbd62f2dc50e08fd66a3fc698499ccd9de3bfd60d589ee1f70f8d06fbf

Observation 62632859-50c4-46c8-a9a4-a012ea57daeb · inbound

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models cites this paper.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.776438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:5af026423ad814a739ec9ddf6d006796eac5ef79112d4a06fa69acac6131dedf

Observation e2adea06-fa43-4595-8e1d-0cb9c0028732 · inbound

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction cites this paper.

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:09:14.763921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T21:27:47.702578Z digest=sha256:ac64142385f30c4ff174398670cb2bf7272c06663130780e8b7f55d17c55ddb0

Observation 33f6e8f4-b452-4434-bb41-8ebf0bfa61c3 · inbound

MemoryWAM: Efficient World Action Modeling with Persistent Memory cites this paper.

MemoryWAM: Efficient World Action Modeling with Persistent Memory Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-04T04:29:34.769747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:00:43.138441Z digest=sha256:7cdbd0ba4df51543d23e36f9e6e58eb293f8ce1e8ce817e9fb1330c02eeda64a

Observation d5b7daab-af32-4d71-bf43-08a8d4e9ca76 · inbound

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data cites this paper.

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:19:44.730950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T11:47:54.951618Z digest=sha256:ea1486f1d8b8d814085bb99f8dbf83b2e7b4acc4885a97f47f33923ca3e914d4

Observation e95fa53e-ecb2-4148-bcdf-5d04c3342b9e · inbound

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory cites this paper.

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:03:51.959890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T04:58:47.490871Z digest=sha256:cbb6fe50afe20e539e38a1f3cb588354e27194b7b9395cbbb02c8e38b8686a6a

Observation ec6cb3d7-6589-49be-9841-9935670ae0cb · inbound

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory cites this paper.

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T17:14:54.438474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:14:54.438474Z digest=sha256:bf7770b467dcb76390c60bddd5e191d725ec38b6143cfa60748ba574f39f1c19

Observation e63e8435-b33e-451a-8d72-8e755bde7c74 · inbound

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation cites this paper.

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.930724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T04:34:38.286863Z digest=sha256:e4fbed02e8c54dbee1253140ff81a57e913bec8bcede67c397952c6d35e72f82

Observation a898c7f5-e592-43eb-abdd-999f4f0caf33 · inbound

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures cites this paper.

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T19:25:32.471377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-08T19:16:06.847356Z digest=sha256:4e8d386a7f79de084a43a75b59aaadbf9f78250070284b49f7e1f0c216a20e17

Observation 6583d759-8629-4502-8b43-2642b36b575a · inbound

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation cites this paper.

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.178857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-08T02:00:30.551638Z digest=sha256:bfeb0a2d6b252e96221c1a477e466fd43746364a23445a26575c97460b2365a9

Observation c55d4e66-9263-4590-a915-deeff2bf9bda · inbound

Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review cites this paper.

Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 159

Resolution
verified exact
local_arxiv, observed 2026-07-10T23:07:47.707668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T23:01:01.563768Z digest=sha256:807fc19abc1b36532693692b53c90cb302fe341ac5859624f12dde7a22f607db

Observation 605362f1-a845-4062-885d-b6bbb6315e23 · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:16:36.361126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T22:16:31.529359Z digest=sha256:0af00448093db8afc5df3e742fdedc94871b9547554d4a5db2f526a00c469e5d

Observation 3a6679f0-a287-4fa8-a412-e02198d37e52 · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T06:48:14.554799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:48:14.554799Z digest=sha256:2e56a1c8e2010dd6ee72c5cefcab5b09b7d76a073b23c427fdfaa0b8f76f2082

Observation cbaae70f-e6bf-4b66-8b03-9fb849b55596 · inbound

PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration cites this paper.

PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T20:30:40.895214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:30:40.895214Z digest=sha256:8fbfe0cae85e0e4cc527353182835b2b9931da1ce1005e0862148ce54196cdbc

Observation edce3f62-877b-42f1-9843-747fa3f95dd1 · inbound

Robots Acquire Manipulation Skills in Seconds from a Single Human Video cites this paper.

Robots Acquire Manipulation Skills in Seconds from a Single Human Video Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T11:06:14.571661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:06:14.571661Z digest=sha256:97a97aba0b1a0847566055e49a5487ae4a9d895e19866d12019f33f14cb4cfd6

Observation b9323eea-7f0d-48c2-a5d7-918efb23c7a3 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 252

Resolution
unresolved
no resolver link, observed 2026-07-31T08:51:31.805127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:51:31.805127Z digest=sha256:925c26da542d463ee62e3a8f6902bf9716dc207c2434ed23b0c6a3267195c882

Observation 070101f9-48a4-4fbb-83f0-58a652793d82 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 233

Resolution
unresolved
no resolver link, observed 2026-08-04T01:23:12.142924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:23:12.142924Z digest=sha256:f42b1e24625c46f5e28402fae779f843d8660600acb3d5ebe721aecf3303b2e2

Observation 8a9b7339-7976-49c8-9033-5cd791e246db · inbound

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space cites this paper.

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:12.003037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:12.003037Z digest=sha256:1602eee7b7ff3552c8e07c0dd94c8909b09ea0770b60756b0619e0374e24664f

Observation 6a543c22-db01-4fcd-a213-356075d5e2bf · inbound

RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation cites this paper.

RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-05T19:58:40.731403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:58:40.731403Z digest=sha256:60e4e387233f386c95e2d180bab8472bc02bed9c2fe5ec9c6ff7f0314f90bc45