Pith. sign in

Paper Citation Record · LEDGER

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

As of 5 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 100 inbound Pith citation observations for arXiv:2503.14734.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.14734 v2

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:09:10.112304Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 602 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:16:50.626399Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact31
  • verified fuzzy65
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

4
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f4bafd84-306a-4ebd-80cf-94f7275ee836 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:38:47.256647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:5109a72f85bb45b7927b83e47fa6578bf48b23fc9179c0c6a997f7e664c89548

Observation 753fb52d-1e0e-4701-badb-90ce67dd70f3 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:09:24.919963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:a964d1a36d627be7c8d97aa9ab6f28873948dfb7205400f6b54c69f44f009496

Observation 6be0e574-7e2c-423d-88eb-cf57187723b0 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.432128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:d4fecaaf584252ac165ae1459713d6e97cfb305acaf922be8d003ffcad9d1df3

Observation 69618917-18cd-48d1-8f25-4c87ce4c0625 · outbound

This paper cites ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.167687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:fec7d919f4c28fc284617bbe30e8ef5a0ee65df41ab3e7b5b52e1e241cddf517

Observation c092ebf1-0229-444b-819e-eb6ede5e147d · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:30:03.293701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:1bc8f62376ad5dfb5b161944a98ee2a744bcaea87016e52cb275616e07cbc378

Observation d31b20f8-c52c-46cb-af9c-6eba878bddae · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Video pretraining (vpt): Learning to act by watching unlabeled online videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.384415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:28583d379ad05706eaae28d1152da65b6581d318693b7588de5d3d14ba7b6709

Observation c55d477a-ce5a-4eb9-b509-48ea5ed0074c · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:e7d1413c23342955ef49b0da7ca0c7863bc9de8fbcf1559609dd888cebf84f09

Observation 826fad00-fed1-4227-ad27-fdc7db1de9a7 · outbound

This paper cites Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manipulation.arXiv e-prints, pages arXiv–2405.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manipulation.arXiv e-prints, pages arXiv–2405

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.388657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:43196131e48a4d1398fd4c40cf25e574edb8281b38869277f3f9ee038dfc7cf7

Observation 47ff94f7-bf8e-4637-a898-61d70e8e99f9 · outbound

This paper cites Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.390869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:5293f1620bfba4bf36a89b48953b4c5d5720ac3f4d0128f3a36579da7bd57c34

Observation b8f7d2b5-4894-4ee9-b64b-3e88dbed82b0 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-10T19:09:10.181938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:3bb7bfd927d33bc9a30c833a4ee3f860344c69b2bdd975780e25f67dd8416832

Observation 70ac69e9-1a57-4b74-970f-429afd180a11 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots On the Opportunities and Risks of Foundation Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-10T19:09:10.185031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:1e514ae218b8b22ba8011a6bc9a3ff1cee0b180558b160917f1ca7e79eee65ad

Observation e93f4a0f-58bb-45ae-88b0-defb21081e71 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots RT-1: Robotics Transformer for Real-World Control at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:41:14.065009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:883d69197bd1adfe02994f4c687f9cd3adb702d3dfb5edc68896bc18c5f498d4

Observation 75830fd6-8b1a-463a-8de5-bd989145fc16 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:36:04.729348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:c747fbe54842cb1dede4c1e84f9a774ef423a71fedc359c80e990d21acdd1250

Observation e3d01bd8-958a-4bdc-9741-cb334571e0ff · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Do as i can, not as i say: Grounding language in robotic affordances

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.401906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:3f01b8690287b68eb954a149e853b4170ff006c9dbd5c5ddd8224f7c97598e16

Observation a99ec1b2-a8c6-47f3-86d1-8c390ce99149 · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Do as i can, not as i say: Grounding language in robotic affordances

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.404035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:1b8901089b9e30e18475172efa7e526131b3ab73f9f1c0e20fa47a002602dd77

Observation 97e2f392-3c46-44ae-8dc8-54669a1b4ca7 · outbound

This paper cites Video generation models as world simulators.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Video generation models as world simulators

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.406647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:04c3c8af33edb778430205d5385787e2dd1d5a1049e0da8842fb5e64d40d7e58

Observation ea3fe2b8-72a2-47e2-a486-1809c3e68070 · outbound

This paper cites Lerobot: State-of-the-art machine learning for real-world robotics in pytorch.https://github.com/ huggingface/lerobot.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Lerobot: State-of-the-art machine learning for real-world robotics in pytorch.https://github.com/ huggingface/lerobot

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.408808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:05edf82ce1dc5e9222b52006241ed875caa42e73bd3561faeb5e588d9bd48008

Observation c2325beb-aeb7-41bb-ad5d-4ed4d96d3e51 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:3797727d8aa5c08ecbd89fb5407910225c03023711ba3330b0ad396b2a6fe9e5

Observation 6f6ea8e4-0bf0-4d77-894f-9c439d7750d1 · outbound

This paper cites GenAug: Retargeting behaviors to unseen situations via Generative Augmentation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots GenAug: Retargeting behaviors to unseen situations via Generative Augmentation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.198675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:9cbc4bcb5cb8c2380dab2dea1d56797a7961e9097453ec4a52c6de4194dd7a43

Observation 15414586-ee27-417d-8a04-fb78034108eb · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.417053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:0e20e8dbd56e0d5e4e46310a22c8ad0fa3378aaffaee799a77c59e7e8deaa809

Observation 3d7b4b58-e5a2-4cb4-982d-b6244bdb76bf · outbound

This paper cites Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.419194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:63c27673fd1c67d2752cd77b18f0e96487b0fbd9360db1529db81e21958e4871

Observation b2e92f0d-8a95-404f-a617-e173559a7227 · outbound

This paper cites Imitating task and motion planning with visuomotor transformers.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Imitating task and motion planning with visuomotor transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.421130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:d4bddbaaca0b410e251334c49fe6a3eb8f187a45bdf24ff9fb38d6a93c98d123

Observation 2683a10c-23a3-440f-9ce0-b12d8203cb72 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Scaling egocentric vision: The epic-kitchens dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.423406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:a72917d91b39a4840f2093225911f4a53702462e081c1b90bc19032240d0f647

Observation 0c8ce1ce-1d60-4a77-91b5-15e3e0746d06 · outbound

This paper cites Telemoma: A modular and versatile teleoperation system for mobile manipulation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Telemoma: A modular and versatile teleoperation system for mobile manipulation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.425676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:a3f99eef12374dc2104e9cce7693030dc909f5e082d84ea233c9a7463f54a95c

Observation a785531e-2b06-42e6-a472-203fb1be99aa · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots PaLM-E: An Embodied Multimodal Language Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:29:30.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:2a2efab8ca6fa29980f5a62a2f7bb1b7ec5908d4a79c13cbaa8d79a994bb8bea

Observation dbc931e7-f4be-4b23-8d0b-ce8de5b7e79f · outbound

This paper cites Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.429833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:548238f008d1f4ad0a06afa22b261ef7c7cbdb8488e5617673f40fcf9ba6eb13

Observation c8f66632-9cab-44ad-97f5-233bef89f92b · outbound

This paper cites Bridge Data: Boosting Generalization of Robotic Skills with Cross- Domain Datasets.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Bridge Data: Boosting Generalization of Robotic Skills with Cross- Domain Datasets

Reference 28

Resolution
verified exact
doi, observed 2026-05-10T19:09:10.158907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:d11a399789667e183a79c29d7c8fc7cc42fd3760bb042e1279f8b20ceeb5dac0

Observation 61b24645-f8ee-42a3-9c32-413a468d0c85 · outbound

This paper cites Rh20t: A robotic dataset for learning diverse skills in one-shot.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Rh20t: A robotic dataset for learning diverse skills in one-shot

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.434290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:6b9b4400d9852317892802dcccb1fd0802f468fe53c6ff5a7a3118b134256f09

Observation 184fdf03-68d8-4ad1-baae-c67fde80010c · outbound

This paper cites Airexo: Low-cost exoskeletons for learning whole-arm manipulation in the wild.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Airexo: Low-cost exoskeletons for learning whole-arm manipulation in the wild

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.436362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:dee11913137b1b23744f5f61e4eecfbb72cf297a619f2d2844de2c3600144114

Observation 46703490-f5bb-4cf6-9be0-4c961be5584d · outbound

This paper cites Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.438473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:f5c28c503d55a72497b3e5d42a9a88a995f870b7bac42ae668beb3898c24b994

Observation e2be4215-5b84-4633-8a80-963228391349 · outbound

This paper cites SkillMimicGen: Automated Demonstration Generation for Efficient Skill Learning and Deployment.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots SkillMimicGen: Automated Demonstration Generation for Efficient Skill Learning and Deployment

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.208818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:e10db1c296a6c4e2504a688055156e8630e521d20af494ffec63199b4d67ed47

Observation 33703922-be44-408c-8add-2f79577902cd · outbound

This paper cites something something.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots something something

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.442832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:26d91efcccccb19d5c470538d06902bcd34a07849ae3e869506c99cd1027f20f

Observation 8fd911d2-8797-4604-81e8-4beec1aeaa98 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Ego4d: Around the world in 3,000 hours of egocentric video

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.445145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:816ebf322403189c7991afaeffd968e52a786022558441540f29d764a9efe569

Observation ed9309a8-95cc-49eb-8efc-3fe02f04872f · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.447772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:122368faee7c44b1716263848ba4260cba281ddccedba67b8605aa355ae9699c

Observation 4b873187-0de3-4e49-b185-20c2d2734ece · outbound

This paper cites Maniskill2: A unified benchmark for generalizable manipulation skills.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Maniskill2: A unified benchmark for generalizable manipulation skills

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.283913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:6d2f82cab5f3fa38a5c9abedb360e6d6da8d53940ed27daf729f51d8cab4ebeb

Observation 9b1c7c57-c754-4f3c-9e0c-ecf4a1493543 · outbound

This paper cites Scaling up and distilling down: Language-guided robot skill acquisition.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Scaling up and distilling down: Language-guided robot skill acquisition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.287089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:a23d397eceb81480897a5f8e17367d8bb98863a8f4d9d85182195a0fd5556719

Observation 2f3cc9e6-d834-47f1-891d-59a53c4d6759 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots LoRA: Low-rank adaptation of large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.290271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:71c2fa89b460fe7644482616bdeb4d3d8e586771cf202f5feeda9a2c4ef6c4e6

Observation 33a43fc4-2b25-40b0-9899-baf6061cba24 · outbound

This paper cites AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.211966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:464b7a2314081fe2e4e5d1e30b0d92ccd5a468860e9fdf3dfb440a049d1c11d5

Observation 6a01dc03-8611-420d-98bb-99fd93399603 · outbound

This paper cites An embodied generalist agent in 3d world.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots An embodied generalist agent in 3d world

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.296190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:81d11b72dee0c6e07d9312360969a6a9c611e97ec0802e637e7b3741601cb55f

Observation a05ce8ba-9f4c-41c2-a12d-afd3039b3752 · outbound

This paper cites Inner monologue: Embodied reasoning through planning with language models.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Inner monologue: Embodied reasoning through planning with language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.299247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:d24d4600d215f7ee0200dd6b22dd8f5e6ee5018ebd2d84f37173ff2520dcb941

Observation 13c3894b-982a-43e7-baef-6f3bf1fbbb27 · outbound

This paper cites Grounded decoding: Guiding text generation with grounded models for embodied agents.Advances in Neural Information Processing Systems, 36:59636–59661.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Grounded decoding: Guiding text generation with grounded models for embodied agents.Advances in Neural Information Processing Systems, 36:59636–59661

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.302499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:962925edc25ac0597e78bad44f9b02cc0b9b1e7e4371fd086043db3ff504fd01

Observation c60c043b-290b-45de-b783-bc90ea1a6286 · outbound

This paper cites Open teach: A versatile teleoperation system for robotic manipulation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Open teach: A versatile teleoperation system for robotic manipulation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.305646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:1914aff299b115cf4c8be26aad1de40f808d9871620054ca1c7bab21518eb957

Observation 97d94b99-c9ff-493f-bf46-7f6882f401b0 · outbound

This paper cites Rlbench: The robot learning benchmark & learning environment.IEEE Robotics and Automation Letters, 5(2):3019–3026.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Rlbench: The robot learning benchmark & learning environment.IEEE Robotics and Automation Letters, 5(2):3019–3026

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.308953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:ecb429241b2975cd8f377d96644fd87d2118f1d7a4c7eea2f5cf0cbbe8c8a06a

Observation 17174dc9-a584-4840-8d92-67e087ded15c · outbound

This paper cites Vima: robot manipulation with multimodal prompts.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Vima: robot manipulation with multimodal prompts

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.312103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:014c8f67effc0d3580f042fdc673d03c1e0abd5818fdaaa52ef19a8310baf1a0

Observation 7684f82b-a134-457b-b93a-9709f9c5ec44 · outbound

This paper cites Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.315372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:0213d24b525b3835e4992e1f1afd1667a355c04dea9f9e6e15979f760072c636

Observation bda5e247-72a2-4bd4-8d1b-8ddc45133a8f · outbound

This paper cites an unresolved cited work.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-10T19:09:10.318406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:16f6ff7451f6f8a217c10a56a384dd18710532ba9e00e810a6823e5c2a4422c3

Observation 2c708f94-fa56-4930-9978-22f366a2e57c · outbound

This paper cites Language-Driven Representation Learning for Robotics.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Language-Driven Representation Learning for Robotics

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T19:09:10.215775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:1b23deaa3676e959c5bf8ec3e89dc6ff2cb79dfc0dd9f599590da97ddcc1be46

Observation bc327b05-2d15-4abf-873b-d205985f871c · outbound

This paper cites EgoMimic: Scaling imitation learning via egocentric video.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots EgoMimic: Scaling imitation learning via egocentric video

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.323900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:f000269956cab3bb26cf42b26193472118f39cf57584a8885637789c5a7f31a6

Observation 3f89b214-a997-4d25-b69c-14df9599f5f3 · outbound

This paper cites an unresolved cited work.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-10T19:09:10.326477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:386181f6bf3d777c931b6446345dafca7da4b2157ba64b761a6baf822d588196

Observation 5e67bf7b-9dd9-4ddc-bb22-a27dd3124858 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots OpenVLA: An Open-Source Vision-Language-Action Model

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-10T19:09:10.219102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:beb42c5eac166a26e96a2b7691fabe1c6c0094c8a6dc6c223783979acd66759f

Observation fd9bba19-1692-4008-bfa8-67a6719c60d6 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Vision-Language Foundation Models as Effective Robot Imitators

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:919ff2c79a9dfbbd0dc0478b6acb61dccf56f04138b1e0e47f97cabf9b3d7ef5

Observation 643b9fc3-cbf4-43c7-81a6-f09073fdd162 · outbound

This paper cites Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.226821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:b301b676e2afaebf00ef6f362ecbe1d2a96697f2be9f05f912fb23118e9e17aa

Observation 76cbe6dd-82cb-4ea3-9052-a1c21a1c7847 · outbound

This paper cites Code as policies: Language model programs for embodied control.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Code as policies: Language model programs for embodied control

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.337843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:253ea6428454ef927a46aaf3e0825956dd3b9a220a2e080475a48f74288ca051

Observation 5afacf84-7090-4e6c-a2bc-49a3f6576817 · outbound

This paper cites Text2motion: from natural language instructions to feasible plans.Autonomous Robots.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Text2motion: from natural language instructions to feasible plans.Autonomous Robots

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.340508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:c7248f1a925c2213ffc1d025dbc267f7237321c05d2f77d065718e552c5d266c

Observation 6dbbdc15-7333-4684-9d37-4cfc0bd03c4d · outbound

This paper cites Stiv: Scalable text and image conditioned video generation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Stiv: Scalable text and image conditioned video generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.230078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:4bff0aa8d5777544b45d532341812c1a85bed96e7f8e2db1be11fbab7f35b5bd

Observation 37fbcf3f-c60f-4f0f-93d3-2f0638532d9f · outbound

This paper cites Flow matching for generative modeling.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Flow matching for generative modeling

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.344867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:1b1fb669e24093d94d77180fba72d8f695ea7e6ba73dee1294f5c36ad05b621f

Observation 2fc299b1-a8e1-4fcb-9491-3cd5692d57b0 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-10T19:09:10.233471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:fb7e0c0c6136445d6ad752349b0bed122b881e9960f94ae299214ee74be7aa0c

Observation 17d507d2-82fd-4c45-a6b7-f9aaad78f661 · outbound

This paper cites Hoi4d: A 4d egocentric dataset for category-level human-object interaction.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Hoi4d: A 4d egocentric dataset for category-level human-object interaction

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.349048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:39503d118ea8c11da5f53b40feb0c14088c54a5606398333da776575babf3c6e

Observation 694b1496-4f20-4649-b755-8923d5ec6c37 · outbound

This paper cites Interactive language: Talking to robots in real time.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Interactive language: Talking to robots in real time

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.350980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:d64753a5970a5296df6ea0e93893054c3ddb1e652b022b6546161aae46dbcbcb

Observation 555be4ce-0a97-42da-9773-cb49f228bafe · outbound

This paper cites Interactive language: Talking to robots in real time.IEEE Robotics and Automation Letters.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Interactive language: Talking to robots in real time.IEEE Robotics and Automation Letters

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.352854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:24f55aebf2caca85d57a06630b0b98b449b6f47307493fff1f33b77b25ace390

Observation 94c6acde-dc36-4f39-a18a-8cee3b614b53 · outbound

This paper cites CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T19:09:10.236670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:62367552276d3abeb9623bd9a905e50efaeb704a033ca4505a8b8c9759bad15e

Observation 8a9bafc1-4911-41d1-b89e-700922e9b7dd · outbound

This paper cites RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.356627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:24df0782f6ac33e7f7ef88f3e76761a3883d4e76cfa8c29587ad1dc55fa4c61c

Observation 3f9fea75-7a72-4ad8-a0be-ab2b0d410baa · outbound

This paper cites Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulationdatasetthroughhuman reasoningand dexterity.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulationdatasetthroughhuman reasoningand dexterity

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.358661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:2d21b0710c9cc0829a5d29321807fc0b0ce147eb521e595a34524a135715d438

Observation 91d0e019-10df-495b-9205-99520de19b30 · outbound

This paper cites Human-in-the-Loop Imitation Learning using Remote Teleoperation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Human-in-the-Loop Imitation Learning using Remote Teleoperation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.239776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:420930573090612be783fc6768067397ff971e5b2115f2a8e4c91cd10286a2cc

Observation f41cf05f-b940-4deb-b2c0-2017c46fe3b4 · outbound

This paper cites What matters in learning from offline human demonstrations for robot manipulation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots What matters in learning from offline human demonstrations for robot manipulation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.363508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:5d358ccd73ddd637f1cc59601e778be9aefdf3731fe14771c8eaea170ec7b445

Observation ef5061da-b039-4767-89a5-7cba35442710 · outbound

This paper cites Mimicgen: A data generation system for scalable robot learning using human demonstrations.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Mimicgen: A data generation system for scalable robot learning using human demonstrations

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.365771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:0a4260d3d0d89fd5b82c736ecf07e5011f01f281c9c6b6c13bdb5a156f69f026

Observation 0cadad90-e077-4607-8dc9-8b82f9b94acc · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.367921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:8c639897a8072586d3dc81430833ecad1efbc23b65e6d316f25872a5288935bb

Observation 9d746545-6e97-490a-983a-9e6da47e21e9 · outbound

This paper cites Gritsenko, and Neil Houlsby.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Gritsenko, and Neil Houlsby

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.369944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:83633b7d2aaa0bad48fc1453220f63be8ea96174ee21b284b9a26bcd2a2fe4a0

Observation 8d34030f-4ca8-46fc-b4c8-78b05cca223c · outbound

This paper cites Ray: A distributed framework for emerging {AI} applications.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Ray: A distributed framework for emerging {AI} applications

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.372201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:26933aa4a1f8b55219ddeff74b444fdacd21fda0b1375fe97f74c4bfcca08931

Observation 2354638b-c012-43b8-a22f-d844ec2a8702 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots R3M: A Universal Visual Representation for Robot Manipulation

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:26:54.148999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:59b8f4821c796532703bcc33cf72175ad0e6027322eabf59080c7173a367c586

Observation bd47f31b-cb10-4905-aaa0-faf11b81b1d5 · outbound

This paper cites Robocasa: Large-scale simulation of everyday tasks for generalist robots.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Robocasa: Large-scale simulation of everyday tasks for generalist robots

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.376841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:01e598013b423b6949dad925b2e9ab9b4fec4749c7d87ecf96dbb2e59a553399

Observation f563bd26-74be-4526-8bb3-29f275ef2a9a · outbound

This paper cites Osmo platform.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Osmo platform

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.378618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:87741580f8bf4ea9391ae88ff6ab81e95f6abe81873b17d8415acb3290f3ef82

Observation 93ecc2fb-c478-48d9-bc7d-ed652980e185 · outbound

This paper cites Octo: An open-source generalist robot policy.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Octo: An open-source generalist robot policy

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.380569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:c828e72a10d38d2e4c51ef7bfe41cee6541a710dbc36ec4115710b3ab23831a8

Observation 4cd68254-099a-4187-bfc2-8d0db53ae06b · outbound

This paper cites Open X-Embodiment: Robotic learning datasets and RT-X models.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Open X-Embodiment: Robotic learning datasets and RT-X models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.382533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:192b042fd5115561757a380601080c45f9933411f3d75254220bb9b8edf7e3cb

Observation aec10e60-acc8-4f89-bf56-0829233567c4 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.386673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:0a4674a3674146a336f9d5574733419f7f47a7dbdf618be1e9017c1b0dbb2d0d

Observation ed8cacab-8bc1-4501-8b88-30750355c63a · outbound

This paper cites Scalable diffusion models with transformers.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Scalable diffusion models with transformers

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.393070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:476c67b19ed76665faa48e4abb6b86a52f48376825c820aec4e8903f2f9512de

Observation 9b1d2462-baf3-4654-9491-ca86fbaf04df · outbound

This paper cites Motion tracks: A unified representation for human-robot transfer in few-shot imita- tion learning.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Motion tracks: A unified representation for human-robot transfer in few-shot imita- tion learning

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.246567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:cba49ca34aa2e3f26fb28c90c2578bf9cb307dcb37e31e1f2c80508a1d71c8da

Observation e0e03382-e30f-4da4-93b7-ef5ff2975789 · outbound

This paper cites VideoWorld: Exploring Knowledge Learning from Unlabeled Videos.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.250194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:69f85eadcee79341820652784991b5e95606cf245d07df8278cd4c865a4ce500

Observation a503e3e8-4e03-4c24-99af-e8d0e1df7d87 · outbound

This paper cites Sener, D.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Sener, D

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.399642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:2ac72d20641025856085d589f5b51e3f21eff5ba23804d92ea17364a8a53ed52

Observation ffb89af1-7a2c-4ec9-a877-211df9e84202 · outbound

This paper cites Andy Park, Shenli Yuan, Yuke Zhu, , and Luis Sentis.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Andy Park, Shenli Yuan, Yuke Zhu, , and Luis Sentis

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.410794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:e56dc4a8a84ea893f46182322aa027e5da891f9e6dd15e2d76ec5a65b755d8e9

Observation bf713ffa-e035-40e6-a418-593fe9d37080 · outbound

This paper cites Mutex: Learning unified policies from multimodal task specifications.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Mutex: Learning unified policies from multimodal task specifications

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.412857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:4b4b61201067d64a6c3e449ee5e998ef705533de9ff8d2ce24639dab004e18ba

Observation 10d889b6-6c32-405d-b4f6-132c52d8a3ca · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub- pixel convolutional neural network.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Real-time single image and video super-resolution using an efficient sub- pixel convolutional neural network

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.427669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:1cece59ceb9898ee7b0ce47425b29ac1c688c6affa6b421d8cd104865492a8b7

Observation de96d92f-d8b4-4adb-9e5f-4eca4266517d · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Progprompt: Generating situated robot task plans using large language models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.440664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:96c573cdaeee8f658ea49d6454010c4e24aca2431caee2a3ac779900f6654cce

Observation 71c56738-611c-4eca-b206-257df7ffdba9 · outbound

This paper cites Plex: Making the most of the available data for robotic manipulation pretraining.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Plex: Making the most of the available data for robotic manipulation pretraining

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.293135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:a8cf02cc6a4f9f87173b16801e2bd3fc17391b5ba4a43205d3306b0aa8982475

Observation c0ab2512-0553-419c-a184-583898d5ea61 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-10T19:09:10.254462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:ed6142a797348e54eecbf16d02b54c49f4faa3faebabd1713a73acb52b453475

Observation 2fab1635-e7db-4500-acc1-bcab60d51a17 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Bridgedata v2: A dataset for robot learning at scale

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.329282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:a5cabe6cc16ffaecc0df5395d3da07f84c08604ec363ab7e32c158603fb4790d

Observation 5f4e9f3b-6c50-4cc4-8d1c-4afd04e4eae1 · outbound

This paper cites Wan: Open and advanced large-scale video generative models.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Wan: Open and advanced large-scale video generative models

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.332110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:b16a957957c652a3c30a27e13a26e130f77e50235b799548ec71034df6e019a9

Observation 6653a262-1e8d-4e19-a55b-2eb18c5d7267 · outbound

This paper cites Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.335059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:6fda51befa3cb294fd35a3b47b22d1136bdf65a5aa184c130f0a973c7115d912

Observation 30fbfc63-dea0-45b7-a5cb-d549fa5955f0 · outbound

This paper cites Robogen: Towards unleashing infinite data for automated robot learning via generative simulation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Robogen: Towards unleashing infinite data for automated robot learning via generative simulation

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.342811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:c2e24c91a399439468073ec45e7e1cf5270959a29286ca4226ecf6ff654b717d

Observation 16b89cb5-e215-4135-98f7-266ff4dbe388 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:12:26.208161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:3f65892e6d5e1c44e3e546318be770b2d10a8a6eba042e192132101bf219a8f8

Observation e8b47910-cd87-4472-a120-f7da23e6fedb · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:32:05.974462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:f02ad624ccac3956888d43bb6325397a20027d0ee812bcede2478677f99e233f

Observation 50b3a717-2899-4849-b3df-c6afb519210e · outbound

This paper cites Gello: A general, low-cost, and intuitive teleoperation framework for robot manipulators.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Gello: A general, low-cost, and intuitive teleoperation framework for robot manipulators

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.361370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:356add8841adc456d26fe92f95cc45c797a296fe9a985005a2a3cf4086814dd3

Observation c194906c-96b7-4713-9ed1-6e47904fefc2 · outbound

This paper cites Pandora: Towards General World Model with Natural Language Actions and Video States.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Pandora: Towards General World Model with Natural Language Actions and Video States

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.265640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:6554eb6c5c1482a0c339207cd3170f1e2fb15910705a074dc66b4a2c97190ae2

Observation 45457e14-be4a-45f6-b758-70f77ae3bf01 · outbound

This paper cites Magma: A foundation model for multimodal AI agents.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Magma: A foundation model for multimodal AI agents

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.395618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:784f290b83b7420498c9898650c08cba7915e1b35bea6116692246c3f80318cd

Observation 35bc781a-8be6-4fe4-9fd6-c8a32d43416d · outbound

This paper cites Physics-driven data generation for contact-rich manipulation via trajectory optimization.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Physics-driven data generation for contact-rich manipulation via trajectory optimization

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.269393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:2a0f3ce813862f357067771c0beaca30aa518f85982d4b4af303e02d82158825

Observation 9d48cd45-785b-4096-a592-65be6e645af7 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-10T19:09:10.273608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:3bdd9de898254cf7ac82f3996488f51d27bf383ccf36ee603cc4282937e6afcf

Observation 4c4eebcc-c23f-4046-a73d-56d176047856 · outbound

This paper cites Latent action pretraining from videos.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Latent action pretraining from videos

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.346920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:33d68bc30b00898235ae53b1d2415aeb7bcc95bb1838c14d0c7035f7ab28d009

Observation f2486ddc-5fff-414a-abef-aeee946f3192 · outbound

This paper cites Scaling Robot Learning with Semantically Imagined Experience.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Scaling Robot Learning with Semantically Imagined Experience

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:59:10.716066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:fa9d158a01e47331e873c55aa6e337306784a37e9dea1a9984c8ff20d133d803

Observation 94d209e0-ad71-4108-8020-8f5b4f8c890f · outbound

This paper cites Mink: Python inverse kinematics based on MuJoCo, July 2024.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Mink: Python inverse kinematics based on MuJoCo, July 2024

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.374817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:f7ce3723dc037a72d43aad8891049dd6a24671642b0992131e8cdb909077e53b

Observation 7e2c69de-be1f-4825-bee8-61ac8e95bb28 · outbound

This paper cites Deep imitation learning for complex manipulation tasks from virtual reality teleoperation.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Deep imitation learning for complex manipulation tasks from virtual reality teleoperation

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T19:09:10.397764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:5a2f644fb29af0ea935fa8b9a2eb470ac3f63e2fbd2f5753ccd052a546ec0d7a

Pith citing papers

Observation cbb7bf6c-4f98-4772-8676-52a6ae8e77d0 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:25:54.614879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:1b038c899b032c62f940f43455789eabbd3bfe984fcb0b52f21fb04d24a01265

Observation 2ccff38b-29d6-4749-a512-9a3066d82ef5 · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:48:48.984038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:c016e99407951d08b55510ebfe62eee9d78379f089cd9be059ffc1a1e92a1e84

Observation 94e2fc05-08c2-4af4-a242-a63ffe5e5c5b · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:00:48.755917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:be2b20ddfca82241e069dd4ef09cb5313e1d38fee996012286c0e746d812f893

Observation 10b90701-b7af-4b2b-9ebe-eac35d252f1e · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.980064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:251068cef085731a7cd59f89697f473818dda503e59fc798038a7700ccab8690

Observation 78648df8-21a9-4c25-9d8f-4ab662d33c71 · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.175634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:9a9f8fca682d93b702387c23d354676c45517fe32352a575c9ab1dc62c0a5d3e

Observation 07e6296c-bb27-4a18-a1cf-e052f1102a7b · inbound

DreamGen: Unlocking Generalization in Robot Learning through Video World Models cites this paper.

DreamGen: Unlocking Generalization in Robot Learning through Video World Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:50:45.414646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T23:50:45.332466Z digest=sha256:4d7b96d3bd9c39b046fa9386ac8a80bd1394d0fed2ed68212e2a294756c9c269

Observation fb91c658-d002-40e7-818d-ca700cb77e92 · inbound

FLARE: Robot Learning with Implicit World Modeling cites this paper.

FLARE: Robot Learning with Implicit World Modeling GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:59:08.970516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T15:59:08.846629Z digest=sha256:af4e1ce3cb17b1b9569219e7840303343ec3503d25527fa03db941ab19843dfd

Observation 6deb7d8f-5e71-4f60-bb8f-c65f03703998 · inbound

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion cites this paper.

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:52:17.894929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:50:17.902979Z digest=sha256:e73cd696cbc0d9b2242d338cf1f44c3e3715fe7f7680a23c7fbe9daec23add4c

Observation 4d5f39a0-628f-4fc3-982e-27864040e013 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:22:37.140891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:e8f5695960f60d8fde7f623f0eea2854293356457c5fd79bfe99baeae582af1c

Observation 8ad49f91-7081-4c74-a4ad-e2e80af706a9 · inbound

Real-Time Execution of Action Chunking Flow Policies cites this paper.

Real-Time Execution of Action Chunking Flow Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:18:51.758341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T14:18:51.613045Z digest=sha256:19bea2c83b7d2d63b5af098fe5eff8a494e97d2614d04d3f9fa62da0ffcee606

Observation 099f1184-8621-42f2-ac9c-cb920e457fe3 · inbound

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving cites this paper.

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:36:24.363891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T07:36:24.319361Z digest=sha256:c7734485e8444bbda752e26b01144e4355f59907795a805f4eae1c46d5ee773a

Observation 1cac71f1-58df-4132-bd8c-e6bfe8e9cf69 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.658349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:3b55615d18977244f0f65c2229f049de190c708bfd547127bd3a09309c6ec63f

Observation 28f83705-32db-4c33-94e9-6bfcbc3c2b9b · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:55:46.234803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:ae53113b85870bc784ce196c91c84c96db492dc1b3a594e3736e77f421389dd2

Observation 6b143498-a764-47fd-bd3c-0f112b4b53a5 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 265

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:08:35.314257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:caf314a7a98a6d6598218a3ad69e17ae27b08fbed04d8adf2f44f5cbc8adf419

Observation 91ca9e25-72c9-4de8-9590-a2738a024dd9 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:42:41.482810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:8d2c2bd4f451a7734b1b3c866bb3422f2241a321e4e0c9a3703f748205a3ca69

Observation b39b723c-f716-476c-83a4-688fb4f737de · inbound

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation cites this paper.

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:32:56.714378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:32:56.397350Z digest=sha256:e4b0f75d34ce0bcf5ec681d9604c68d41f7ec1438cbee5ce26d8932703141e1e

Observation 9313bc16-81e3-4023-8d7c-1136315d0c5c · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.576715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:c3efc632d96c35fa4e6fa65719427fb0ab9db0dbbe1b0294499bd70d7cbdc149

Observation 3c868a0d-dfe2-4189-96b2-9a056f0fe3f1 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:00.925689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:2dd0784b3830af1e9b28e77bbb375edc39f9b995be21cfeaa197f28398be6bf4

Observation 9a69e291-6c7c-457b-a6a2-dffdc978805d · inbound

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models cites this paper.

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:52:03.130861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T21:52:02.893886Z digest=sha256:eca60b118545d3cf863af29c083ec4b3f83302f78b5cf4c639a3aa748ae7beb2

Observation 8540d009-a800-41fc-9d1e-e06563cf9bbb · inbound

Video Generators are Robot Policies cites this paper.

Video Generators are Robot Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:43:37.242436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T21:43:37.162870Z digest=sha256:06f31f6fbd10f98b8c097895df6086d2024e931b121ea96cad95b296ad2b9cbc

Observation c711b36b-0b69-4d35-9004-88ee17a33681 · inbound

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation cites this paper.

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:28:41.958965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T21:28:41.904725Z digest=sha256:edc3eeb563e399f6d450650e8918690ba446075be4705e44d54c33d06ca91c8e

Observation 48171647-2b2b-4126-b3c1-14667dfd248c · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.626399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.626399Z digest=sha256:75fa7a9fb6a762371a6d81cb4fd79aded8a5822a7cffed316c10c4d23461937b

Observation 00f41550-003c-4909-a09b-59a38ad02de5 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:28:16.195470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:6596845657742df6d918fbbe01d3968e47aa5ed07c01ead0e9c4433c72b57b46

Observation 6c843a89-c11c-4761-a0a2-7657f81e66d1 · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.100523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.100523Z digest=sha256:578f255de09d0f97ec2e8fc2bba1fe0b95c9dbeba65e6ddae937451dbc309bfb

Observation c642b905-9070-4eb9-a0ae-edb82b3a893b · inbound

HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation cites this paper.

HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T15:20:20.854333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:20:20.854333Z digest=sha256:7746959f49ff56d13779695c506a62bf4e1d6df56b57ce00ec5b7d18160cd675

Observation 112021b9-e02d-4ae8-a52c-2eca41505d86 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.249247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.249247Z digest=sha256:4f7edfe073e2f0c06171972354552e515b1a9a3568895ce21a31cb22da98ef2b

Observation 535bbbee-a2e6-40ae-8a36-2b20e04b40b5 · inbound

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies cites this paper.

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T05:48:47.230589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:48:47.230589Z digest=sha256:7a4ca83fea49407e36dd90d9d3edbe1adbd6b05c4b0b1feb745ccad20f115eff

Observation 83c17878-ca85-4ce6-bc81-96fc5e1e803e · inbound

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation cites this paper.

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T05:26:37.119377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:26:37.119377Z digest=sha256:82d2fc8016aa2b0078767b91838beab17cc02ca7f2bf1505c485c557116ac587

Observation 05dc57cf-dd86-42d6-9117-816ec0187372 · inbound

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation cites this paper.

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T20:10:13.585605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:10:13.585605Z digest=sha256:a4da20285fae1c2461b22e0d44a58712a7d67b0625b13e124731419dc5477c7d

Observation 623bb6f6-37c1-47b5-8afd-102391bce1c6 · inbound

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations cites this paper.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.452731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.452731Z digest=sha256:68c5f05991a4d2913c2ea0ca4dc2353386eb36888063ce491308d29cd8ae6490

Observation fd980bf8-ad9f-4c26-b2c2-a672564b4056 · inbound

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models cites this paper.

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:49.331033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:54:49.331033Z digest=sha256:3bfa578b83f861b68653b4db630c819823d051752ca816d0908bc41cd215a47f

Observation b7286cbf-518e-4ace-96dd-24a59d2fd82b · inbound

Contrastive Representation Regularization for Vision-Language-Action Models cites this paper.

Contrastive Representation Regularization for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.363486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.363486Z digest=sha256:b42dc22277911599764fda23c389040e4aabc414a81d3587b38cdcdfab797c74

Observation 0f4400bf-11f6-448c-80b1-8118337a3ae2 · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:11.242201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:11.242201Z digest=sha256:f4e1952fb9d1773ad77f71eb00c3cf8b01fbb4c90632de824f6808fc0519cefa

Observation ef5e1a58-7967-4f27-8231-dc7b6437dfcb · inbound

Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI cites this paper.

Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 230

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:01:14.221694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T09:56:36.716680Z digest=sha256:5a6acf51eb297572055f55c36e06fcbdc36918b9f4bbdede37f5815c43afa787

Observation 21f644ce-f1a9-4671-b9d4-db12a218ecca · inbound

Verifier-free Test-Time Sampling for Vision-Language-Action Models cites this paper.

Verifier-free Test-Time Sampling for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T11:20:11.670683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:20:11.670683Z digest=sha256:d53a7e4c48689ce11cd9a865386c73e89892e9ca8a739b2e396939de8be6b288

Observation 72eb214f-c039-4cb3-abe4-97c14d8543df · inbound

USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots cites this paper.

USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:20:31.773923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T08:20:24.424886Z digest=sha256:38af36905c97a436ee1f0ae454f38910c2ba24120b97b8fbf168a316b7c48ca1

Observation 94f0a72f-e0da-43f1-b8fb-d74e77a670ee · inbound

Reflection-Based Task Adaptation for Self-Improving VLA cites this paper.

Reflection-Based Task Adaptation for Self-Improving VLA GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T07:31:02.992211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T07:28:11.187479Z digest=sha256:a0125d30f2433512fcc3f96ec10322b82d0b1f71a4608fefbe810f03783b3c0f

Observation bfe7f689-70c9-4c08-a5fa-5a44323ac4ee · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:09:39.717164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:63222e705a2e76767ef9bac3b185819b4581e089e5643f3eea424c88edace620

Observation d0a9373e-2ebb-4cd2-b2c6-9c0cf9176b6f · inbound

Co-Evolving Latent Action World Models cites this paper.

Co-Evolving Latent Action World Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T02:52:21.753455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T02:51:21.569303Z digest=sha256:124530c2b2ffc7a26791db36ce813b944e89e2ff7c5f0b88508119625fba2a8f

Observation fd1067cb-fee6-469e-a104-84c13e6da7a7 · inbound

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model cites this paper.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.695032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.695032Z digest=sha256:4851723cb2e6d3f66f81a0e6cbb1d412510422c3fd7e0f94356fa0de2a5de236

Observation 8b76e304-375e-4841-a6bc-bf9387a2be39 · inbound

World Simulation with Video Foundation Models for Physical AI cites this paper.

World Simulation with Video Foundation Models for Physical AI GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.924336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:d75aa93ad26a4b34ef490cdd2cb739a202a68ccd5cf59bd556d7992dfc1df88f

Observation 556ab213-c9ee-4a81-b82d-35b115f4d06b · inbound

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations cites this paper.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:05:33.728637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T01:05:01.553811Z digest=sha256:2299fdf159ab105b0273efd1f0e95f662f777c96d8ee2651a86cf056fda3cef8

Observation 8106047e-be79-48de-a858-39188c005d9e · inbound

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations cites this paper.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:01.100730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:01.100730Z digest=sha256:9f17e57593cfa267710e1621883f417d0f115bf78473caefc1c9d0e49f1ebbd3

Observation 47b7a6f1-0348-440f-91e4-ee42ef558fc6 · inbound

Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning cites this paper.

Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:22:52.813374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T10:22:52.472354Z digest=sha256:4c3854e433b0ef2cef16b74d7dfb1fe5153034341ee811d3816ce5876e80bea9

Observation 39f98808-3696-4ad3-aedb-d756d42438cc · inbound

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control cites this paper.

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:31:31.703503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T12:27:16.448513Z digest=sha256:d399fa84dc0ee2f9b1fbc194f0b8697428f677b5d9c5e7f54375ae1616b740e8

Observation 71b6de69-dd64-47b9-97e4-35478fde678a · inbound

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models cites this paper.

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:10:48.904586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T03:09:09.713822Z digest=sha256:7a08adf0c63ae35ba8ce5fd76c7c0f733128c1f38cad1d1216b155d60b23e81a

Observation 68e5a78f-aba0-4b5e-9cc2-e4fde87ae9fb · inbound

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models cites this paper.

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:09:09.479952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:07:42.311608Z digest=sha256:f8a94972859a7f51574a7ad4e6d2604952ddac2d4e23b943e19172f6831e3b9d

Observation fc84a973-201a-4b46-aa6a-cc2db57606ca · inbound

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention cites this paper.

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:29:09.994156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:28:22.652509Z digest=sha256:24784f88f3c543fdaf42b096e6f2f4ce72b464a0d2312f64713f2786340c32cc

Observation f4c64c22-dc5f-4690-bd3a-de8607a848af · inbound

Mixture of Horizons in Action Chunking cites this paper.

Mixture of Horizons in Action Chunking GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:32:25.624340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:32:25.624340Z digest=sha256:0097c4d5bfee28c0f3b664b408104fc1c34d7e81d12d4c0eb24fce5d8c571da0

Observation 542ecb6d-567a-4b5d-a575-5cf96c6fe10a · inbound

DASIP: Dynamic Test-Time Compute Scaling for Robot Control with Stochastic Interpolant Policies cites this paper.

DASIP: Dynamic Test-Time Compute Scaling for Robot Control with Stochastic Interpolant Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:13:24.640412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:13:24.640412Z digest=sha256:fbfedfcac998c4e8c7f9c68c197195e868a84786f4dfb47e6354b90ac4eb87bf

Observation b5f7ad55-855a-4332-9104-444817a56120 · inbound

LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning cites this paper.

LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T19:50:36.292530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:50:36.292530Z digest=sha256:9555fa491c9e724799ca4280e69db90030822ceef4f1d60b33dec0cc5d2efbc9

Observation 6fefb71f-7ff9-49c4-b05b-b20606d8413c · inbound

Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary cites this paper.

Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:14:04.609119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T05:12:04.200536Z digest=sha256:c28f86a6dd35d553d7b2177535e1eba5f0bd77d53a424b5196cc5c57a6f6401b

Observation f389878e-8078-4b00-95ff-e9d7b1623598 · inbound

SimScale: Learning to Drive via Real-World Simulation at Scale cites this paper.

SimScale: Learning to Drive via Real-World Simulation at Scale GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:34:01.855773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:33:03.629533Z digest=sha256:5323b9829c664b5db440f8a98a94a7558537d416fc6f5ac7bc6dd868a68b18bf

Observation 5069ca53-9669-4972-aaf3-f9d100f0c711 · inbound

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models cites this paper.

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:42.865399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:42.865399Z digest=sha256:a24a8ef5ae6154115291a2f65aedba24ae382e66521f05253210b2a85c69f3a2

Observation 3df1b40c-0d55-494a-9a46-69cc93586f59 · inbound

IGen: Scalable Data Generation for Robot Learning from Open-World Images cites this paper.

IGen: Scalable Data Generation for Robot Learning from Open-World Images GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:58:54.964034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T02:58:36.214948Z digest=sha256:603c36bba83ce5c5316bbef887ff907a71f6e319f4ff0a96e3210f4d988bfe9f

Observation 97814cc7-6b4c-420c-afdf-d27a9efa5888 · inbound

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models cites this paper.

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:01:20.402327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:01:13.910539Z digest=sha256:1c696e0ed7d5b0ad6c5b3d508fe229d6d699609947bded27c821cd96a5f8dfac

Observation 0dba494f-8b02-4b86-ad7f-4e1cda30b2e4 · inbound

AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis cites this paper.

AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T16:48:59.889800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:48:59.889800Z digest=sha256:fe22583d1264c3a3d7447edb8d52ee8ac599c22ae2dbb02180dcd07003051f03

Observation 503b94c5-1ba1-4cca-822d-aad636a2e830 · inbound

Motus: A Unified Latent Action World Model cites this paper.

Motus: A Unified Latent Action World Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:44:36.707880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T18:44:36.636455Z digest=sha256:223864c3cbed29455f180399c2bbe0775b2d795a1e5d50d4d28187ffc669f1d2

Observation 7011809c-aefa-4c7a-8623-91fefa294740 · inbound

EVE: A Generator-Verifier System for Generative Policies cites this paper.

EVE: A Generator-Verifier System for Generative Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T14:10:34.029107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:10:34.029107Z digest=sha256:30e7d82fd162cf7aedafe268b668450be5e0e035e4406e76446b4438efb0a04a

Observation 945a4346-7193-4425-b6f5-36103c137368 · inbound

AstraNav-World: World Model for Foresight Control and Consistency cites this paper.

AstraNav-World: World Model for Foresight Control and Consistency GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:28:20.902027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:23:58.769472Z digest=sha256:f9d098eecb5cfad129c8f7410fdf72b2f9cbb48c44c8a070d001a9e33b8ea3e5

Observation f730debf-8187-427d-967b-d0a5525fe6cd · inbound

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision cites this paper.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.750637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.750637Z digest=sha256:90b2ad6f4e0856a4ea5030f6275290f0e1fa48ff3f1fbc77e66c8d71d1602988

Observation 4b4bea4b-4855-44e5-804e-005e691daeae · inbound

Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding cites this paper.

Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:31:13.359172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:28:35.576661Z digest=sha256:697911cee16a9c17b18b2fe96a83c8d4263c23cf33c2cf91fc2ea57c0726386d

Observation f5a33f8b-4185-4306-ac8d-020c37d3374a · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:41:10.952039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T18:39:59.449746Z digest=sha256:6b7914e5be67d2bd2a697ad0dc393e94f5685441a6884a402567b3b0370db936

Observation fbd85fc0-9aac-487e-8c29-2a9eb13454df · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T13:35:54.033728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:35:54.033728Z digest=sha256:8adf63139b88daf01d9efa631dd8cc2081b958e104600569762fa592aed0bc79

Observation e96ffd6e-49aa-4860-ac73-783a5992f4ba · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:52.654680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:52.654680Z digest=sha256:5321bbd60a0f5b024c4cd4237bb94bf6ddc9adb3fe21f5c1ee31f2ab0e43e294

Observation 815aa5bd-c3f8-494e-a2fc-25a6ab02e8c5 · inbound

CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding cites this paper.

CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T12:37:24.371572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:37:24.371572Z digest=sha256:2fcbf36216f7d997c0785becd5b544cc7fbac42e6a6594514277e741942a57fe

Observation 5c5bb674-010f-47b7-a1f0-d7b01b006d87 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:08:01.692873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:4ed4bfd5e5fb49dfd9393b078b50728e7efbde44790a4160a6ed32ea60997a9b

Observation 27da30cf-1518-4cf2-bf92-3b1271dd183b · inbound

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control cites this paper.

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T09:07:38.745721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:07:38.745721Z digest=sha256:b63cd5fc86dd3235dc107e4d7786a108c4c53b23cff9c4a1a3fd2a87b3ca53fb

Observation 7866a67c-0cba-4284-b84a-50e5afb5c701 · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T14:50:12.846172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:fb32356581b77fb5a615bcb45b7ef6714a57f34eb6801c01178200a09f7d8778

Observation 43f5ec15-ebe2-4b11-a49a-4ca020af8d5f · inbound

A Pragmatic VLA Foundation Model cites this paper.

A Pragmatic VLA Foundation Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:16:59.740280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T21:16:59.712007Z digest=sha256:585e421a9cfabb742827aca8362f7376bf4fb595edd7acc9a47f7b2d2c119409

Observation 5984a3b9-33c5-43c1-b935-1064c51a82d9 · inbound

Causal World Modeling for Robot Control cites this paper.

Causal World Modeling for Robot Control GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T13:53:52.436349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T13:53:52.188890Z digest=sha256:ce6f19acd6bf2338a340b3b76e73bb674391c11d74f44abbe2b6d0ad39c56639

Observation 68c7a14d-b60a-468b-8be3-91b8be394917 · inbound

CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining cites this paper.

CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:27:36.855310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T08:24:44.943709Z digest=sha256:5ab559e2441a4a3e5947d1789018c2075767f666373ff18fd8abda4c52078478

Observation 84793eb5-e971-4812-979b-b483d1d99893 · inbound

Vision-aligned Latent Reasoning for Multi-modal Large Language Model cites this paper.

Vision-aligned Latent Reasoning for Multi-modal Large Language Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:50:44.287592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:48:23.272002Z digest=sha256:048997d1f1da3e019f407b2716362ae8d9a37fdcdc090ff4a265175e52816e0c

Observation 9b1c3508-e64d-47ac-9aaa-c218c5d13ced · inbound

MobileManiBench: Simplifying Model Verification for Mobile Manipulation cites this paper.

MobileManiBench: Simplifying Model Verification for Mobile Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:23:40.786265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:23:40.786265Z digest=sha256:08a05a8d22081f158597f655d163ecb4a37b86ede3be8bd9b3dacb54530cce83

Observation 2f6c2287-5591-46a7-a5bc-adc118c913d2 · inbound

RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training cites this paper.

RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:28.953380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:01:54.696390Z digest=sha256:39ccd4db27949a859e24e354a304353d3a0cf679bb93f014e6929b50c4310a71

Observation ca558d04-ed66-414f-9d4e-36c09a6f5fe3 · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:03.534148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:03.534148Z digest=sha256:d5a50e4b3f95d218c741e5432d1444f5e1ae88ac9da92fc19672b312d15f60ed

Observation f240d912-d0b1-453f-a761-b03b25b523ca · inbound

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos cites this paper.

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T17:02:34.215091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T17:02:33.997887Z digest=sha256:b641877d88a9d4c04f4c351df8ab384d07159daaf2121bbe860a9614e879f3f9

Observation 8415ac7f-1898-46cd-ac80-266340a87624 · inbound

Action-to-Action Flow Matching cites this paper.

Action-to-Action Flow Matching GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:57:29.205410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:53:35.153155Z digest=sha256:fbcf15ae961e7dec9f675afb472a540e8ef43596315391fe2409b9b5f841068c

Observation 000540c8-f2ec-404a-8275-6ab7009aaa50 · inbound

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning cites this paper.

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:04:11.877362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:03:48.795572Z digest=sha256:79d597feab53395e01fd5ae3944d0ce338a064ea3d9ea94090a0e2077dde66fc

Observation 82b770ab-5822-40b1-bffa-6e6249a7d98b · inbound

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation cites this paper.

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:14:10.848559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T13:13:53.818915Z digest=sha256:a482fc7c3e2f03dab0784b3cea6cb083afc9e150c166be5e29bd7b66e9396452

Observation bcaf9a38-3ef0-437d-a261-c38e318b7ec9 · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:37:13.906617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T03:36:09.272019Z digest=sha256:1de1f18e63b72a3009cbc3d75f38387d1544e7afa27abefc6c794863d37d7249

Observation 134c3080-53b2-4137-a4ed-2487c59a5ee3 · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T02:53:14.198220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:53:14.198220Z digest=sha256:4f2e1fbe1fb6357f91235e3bfaf2c2e289c83722040f41a5fb7581a75c9b9ac3

Observation 35710cff-1314-4cd8-87b4-f924276d7f57 · inbound

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning cites this paper.

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.239715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:07:10.387869Z digest=sha256:df9f76339e3bd843c92dd1a4f5fead33f7ee11faaad8cdb2d127391eacc24130

Observation d59c15e5-73f6-4b0b-8d99-3a5a47ed5a22 · inbound

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models cites this paper.

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:02:24.944885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:01:30.803128Z digest=sha256:556cbddb6217a104c7c8e713169658bca96380c8e60e82ab5bb7f6b51f61115c

Observation f5c34b01-c70e-42f4-be36-ca20db7262b7 · inbound

RISE: Self-Improving Robot Policy with Compositional World Model cites this paper.

RISE: Self-Improving Robot Policy with Compositional World Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:30:32.164798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:28:37.997148Z digest=sha256:ac82c081eff62f28aa80ed012c1d7e052bbe608addc95c87fad94a4b46416f8b

Observation 9493940e-b2f7-44a7-82e3-045e4b15b4a0 · inbound

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning cites this paper.

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:12:11.731271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T03:11:52.645633Z digest=sha256:ac0fe45d0ad501ffdd2851ed895602422e74df8b07acbf809992f86398f549a5

Observation 425c4198-a162-47f1-9d01-a60a944e4b62 · inbound

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation cites this paper.

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:11.975943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:11.975943Z digest=sha256:8721bd8b71fe9609b31f7c530e5bc168fff63e2a379494c74d97e862ea0111e3

Observation 2a0f81bc-9fbc-4c27-990e-02e94bea70eb · inbound

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models cites this paper.

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T23:48:55.703806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:48:55.703806Z digest=sha256:9ef3d7f52aa9d1912d553909f15b627eb9e063e6a0bcda1ab3ace1638f67d2c1

Observation 7d441156-200a-4638-8b95-2ec9e131e50f · inbound

Learning Native Continuation for Action Chunking Flow Policies cites this paper.

Learning Native Continuation for Action Chunking Flow Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:40:08.755661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T12:38:26.522838Z digest=sha256:b556fe03d234ab6af92ba18081248702b146310b5d531a623f6bb4e6b9663a2a

Observation 95ae09e7-68eb-4862-b5d5-0417cc39087c · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:18:15.765647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:05d2b08efd2fcbcdec487c16f8abb39c847fc9cf3ac08cf7732850dbf5f1e71d

Observation 3899ce56-48b7-49bf-85cb-4afe9dbc0702 · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:25.839232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:25.839232Z digest=sha256:0ae0e9c6fcbf21bec62bbfdac4cd19063ad44a56f451277fc384ad816977aa18

Observation 91a89579-96fe-4a66-bd99-578a2e942d9b · inbound

VLANeXt: Recipes for Building Strong VLA Models cites this paper.

VLANeXt: Recipes for Building Strong VLA Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.829471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:2b1a583ff1fae588f904ace3816403d85410b314932bd1b1b36827ef07efba16

Observation 7e85009e-bda6-4ad1-b174-313733b79512 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:05:10.082463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T13:04:30.544504Z digest=sha256:b35c4a5da91f3e8575df82e4514bf5561a705418c8108f6880013ab5765063b8

Observation e4356ceb-fefb-4b66-918b-6a8d9d2f8b4d · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:09.848953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:09.848953Z digest=sha256:4271c213f01eca39b1ff4f6945cc9a26d56bb340a68ca5dba3a29ad281a2021b

Observation 7f39e6b8-3e08-463e-bf9a-ab3fcf56ae44 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:20:17.675042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:5f69c13eddfa6444308ac0e60d6234e71410ef151d7e6d714b25d5c4e1a39e8a

Observation 1e9898f8-f9c3-4532-b1d2-2ba1ec23117a · inbound

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models cites this paper.

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:20:17.404389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:20:10.435886Z digest=sha256:a4aec339975f4a4b6b86f6239cc8a092d4eb9ca479cdde8ebbb71a2230012abe

Observation 0438a898-3beb-42a3-bdb3-0b86bdec0bb2 · inbound

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning cites this paper.

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:06:33.824926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:05:30.324309Z digest=sha256:6deeda3e838b686c96cdfd8c9d91c8494d099afb5e890e0d6210ea373edf15aa

Observation acdd75fb-43a0-4431-8566-c45707e268de · inbound

Force Policy: Learning Hybrid Force-Position Control Policy under Interaction Frame for Contact-Rich Manipulation cites this paper.

Force Policy: Learning Hybrid Force-Position Control Policy under Interaction Frame for Contact-Rich Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:36:32.707311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:35:21.153175Z digest=sha256:dbb3356e4a224a6c67ac764a5d81fe5789aa5766c43ee8e0e341189206054455

Observation 2fe17fa6-c417-40b0-9466-0a0b7e554c3e · inbound

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation cites this paper.

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:30:20.930062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T21:22:41.935691Z digest=sha256:a1ccb4dd239a07fd7a17d24739ca89dcf97217613990db5568a80ef52c80537d

Observation 03942d3e-cb62-458f-8f0e-b58f668f99dc · inbound

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models cites this paper.

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T19:35:37.761816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:35:37.761816Z digest=sha256:95545226d9a0e38178f67dc611109dbee081c7d501423138dfc8a07b7d155beb