Pith. sign in

Paper Citation Record · LEDGER

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

As of 11 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 100 inbound Pith citation observations for arXiv:2410.06158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.06158 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:09:33.761708Z

measured 171 of 171 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 169 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:45:08.187972Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact27
  • verified fuzzy39
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

7
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1681862c-8797-498e-ab03-0e708c0f6962 · outbound

This paper cites GPT-4 Technical Report.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:34.019586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:95be07a03dcb551feff3f3e5372478fe625c057a1165b5e9ab271830be885478

Observation 16ce52d4-f074-4d00-a845-28a55ebe854a · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation SAM 2: Segment Anything in Images and Videos

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.820856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:386d909556cc2ce01a92ec0d954681b52a407171b4c59434cd975ae4c72a8a27

Observation 2aa8cec0-f7f5-43ac-8f02-7564bad2625c · outbound

This paper cites Video generation models as world simulators.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Video generation models as world simulators

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.081547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:6c06aa1a8afd7859a3417f12aebac23f5ea722fc1399df0acf68fd5e298c8a4c

Observation 3cd5f690-4395-4629-83e9-e783e58ebc25 · outbound

This paper cites Language Models are Few-Shot Learners.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Language Models are Few-Shot Learners

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T01:09:33.849058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:9f516c85cc503dbf132d4bc8893121bd8a32406113a067fa0cbdff31302051e6

Observation 6a447728-a6e8-4d22-a898-b8a063f217b3 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:32:05.974462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:1d8ac0f9120e7f9ff14383befd05e17ef8c1128524201d94d6a4a74dac503451

Observation 0e737ffb-8f5f-4b6f-b9fe-5826499686e2 · outbound

This paper cites Learning transferable visual models from natural language supervision.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning transferable visual models from natural language supervision

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.085913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:2878b17c49785d05d3de25f6f2278f94fb32a97428c8cbc1b0a7e12a9282388d

Observation 441ec216-e095-4645-8a32-b697ef14b739 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Taming transformers for high-resolution image synthesis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.092286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:8b83339bd7b55b07b9657885446f78bc2a301b5f0e1aa3265388d0c570079fc1

Observation 89f8363d-7a4f-4465-b62a-7aa8f8e927b2 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.098267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:6258bd2fb52e9bb7f806ee49ba9b58adf4ebd51986358544a3f1fb1c841ed47c

Observation 3a38fe7a-fa78-497a-a924-d4d98fc45e07 · outbound

This paper cites Ego4D: Around the world in 3,000 hours of egocentric video.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Ego4D: Around the world in 3,000 hours of egocentric video

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.103362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:04459b9f393442dd860c26e915afa67d7a56b2e7f5eed4bd2d69bb444eb42864

Observation 49508fdd-42b0-4dbd-a2c8-22c579393ce0 · outbound

This paper cites something something.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation something something

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.110280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:f9f00bcded092461c2a3a7c14c601fbf459c7eb00542f6af2e3737f7a76ea9ab

Observation 62463d2e-1f17-4f8e-83a2-9a55965958c6 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Scaling egocentric vision: The epic-kitchens dataset

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.115219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:39619152fbe23573345e7432ff083f2efaa3a8c493e55b19e9a2966dd0c92105

Observation 5ad23283-0ada-4c2a-8427-860ad863e79b · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation A Short Note on the Kinetics-700 Human Action Dataset

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:33.954658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:c40146225d4c8c535cc9ed873153a2fe15593f02927ff8231813e4914b7e7717

Observation 58c428cc-2796-450b-87ae-949709600d7d · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation MediaPipe: A Framework for Building Perception Pipelines

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:14:23.315480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:bd95e36ea25b6feb7a5974ee30462ff1dff208abdf79164aa716a1deaa720407

Observation 69dab474-df5c-42f4-b93d-986272b955b6 · outbound

This paper cites Open-Sora: Democratizing efficient video production for all, March 2024.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Open-Sora: Democratizing efficient video production for all, March 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.119558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:9efa13e132688099c62be2592d21063a76c1e16c1ddcc191954895a75a5b10b5

Observation 7ae26f0a-6d32-4375-a17a-a4d574c2675a · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.831021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:40aa6ed81ddc0f221bf392bd3fe1b078ac4d2751b5ecdaf07ce0759b9a637a04

Observation 1dd81452-8926-44e4-bfac-7b938c8fb77e · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Bridgedata v2: A dataset for robot learning at scale

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.125337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:0a1d8ac580b7301268682aeee6a5433651afc9ff11722820e9ad6fc843dee37b

Observation 8e5054e3-9deb-4a16-b147-42a3735798a1 · outbound

This paper cites Learning structured output representation using deep conditional generative models.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning structured output representation using deep conditional generative models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.132704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:96e1797bba8324a25d869fcedf641baf2f2e97b798499d88bc6c7871d6a615a6

Observation 8af91c3b-b87b-4a5f-ad58-80efcb4a903b · outbound

This paper cites Auto-Encoding Variational Bayes.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Auto-Encoding Variational Bayes

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.940957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:1ff6222e5a8c585f91a96f54c6a59a9e23a8d9d0abb807c97d4a2ea136438f5c

Observation c69ba143-c4b4-43bb-a989-26311cd4081d · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.945626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:297d190c2b3b139fa04d4bcd5bfb46b0ce9ae75eda68828837d6a6310da3a755

Observation 0736680e-b7f6-4c5e-ab8c-d16a24b5d3b5 · outbound

This paper cites MOMA-Force: Visual-force imitation for real-world mobile manipulation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation MOMA-Force: Visual-force imitation for real-world mobile manipulation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.136176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:a99661b5b2ab792a5ecd8fad2aa57595df469e56705acb4574030e2e8bde0698

Observation 30f4e127-9ee0-4025-b01f-6237c9b772de · outbound

This paper cites CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.139669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:30eeb51cbd4a03aef7643ea91b402f4936c8e0372976956ff3ca0cf68893b9da

Observation a84e9232-eb80-4d27-8f98-29d337ffe866 · outbound

This paper cites Denoising diffusion probabilistic models.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Denoising diffusion probabilistic models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.143091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:6ad4ea0cbf087023e28e85114f0f0374cb5c6c4c4cac47584eb3af72bc24e8ae

Observation 7bd6f53c-6754-4429-8698-0fc0cd3a90f2 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.147499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:d9706a4aef45adf00605cad65ca94d08d86432d4e4159994e70dbec8bfdea010

Observation ceb5b34c-6da8-4628-be1b-a4342c81dea2 · outbound

This paper cites Segment anything.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Segment anything

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.152913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:6b6580133e225edefe883332e1d69a47d56e7ef033a07aa8632cfff9fa95cf78

Observation a85f75e1-5a6a-4787-b844-1042c2ff57bd · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Latte: Latent Diffusion Transformer for Video Generation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:45:35.835056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:dc4809b9aceb0108b79c482fd5a2641009435b184d471fc5a7941a4caacf5317

Observation 2b89f9b8-447e-4a74-9298-61be67248f89 · outbound

This paper cites RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:09:33.866487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:fdb420ecfe11d575c6d7e35f47e44fdd91df8d51b637bcfd301af886ee5d18be

Observation 68e95ac0-7a3f-4dc2-8a90-d819b0095180 · outbound

This paper cites What matters in language conditioned robotic imitation learning over unstructured data.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation What matters in language conditioned robotic imitation learning over unstructured data

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.160186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:a61c916a08ba74697e67411fd44fe967d15ebbfdc4b99bbf6b3b910bc7ea7d81

Observation 826d4c58-8d89-4f86-a768-72cedeec22af · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:2b4e94ca4afad118bc2c2bd86a9dcbb50b4904bc86276c7a6789543674a7134c

Observation cb94fd1d-303c-4ee1-8144-d70b041977cc · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.896519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:86afebb1b005758ae51fc58f68f66c60ea48c1d7e755c3d258319fd7e575a357

Observation 7510ea83-c361-43ef-86c2-79f0af46a2b1 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation VIMA: General Robot Manipulation with Multimodal Prompts

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:09:33.914168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:031910e7cdf1777642528e3d3cebefd70c329f0129af7e5a7d9228db820f5080

Observation c8908412-3de5-4aca-b777-a135c72f5127 · outbound

This paper cites Language Conditioned Imitation Learning over Unstructured Data.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Language Conditioned Imitation Learning over Unstructured Data

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:33.936383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:ed8dabdd9315936ce7795a7bbd040da47152c1faf5e95a65e664b87ba3a966eb

Observation c4364eba-7c13-4166-a29e-4b2e0a3cd779 · outbound

This paper cites BC-Z: Zero-shot task generalization with robotic imitation learning.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation BC-Z: Zero-shot task generalization with robotic imitation learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.164295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:96df5cc396bfc3e22a199e8c556bdce709c7ef5a5fa66688899ed64159d1ba31

Observation 2a989c10-d26a-4ea3-9150-3cca8f4962e5 · outbound

This paper cites Multimodal diffusion transformer: Learning versatile behavior from multimodal goals.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Multimodal diffusion transformer: Learning versatile behavior from multimodal goals

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.167537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:c8ef43cb170d7e7eb7e54f80874b1ff41d918a12f451a4d11caa5e4bd388265b

Observation d10340df-1f50-49c3-94d8-33acd6fe79f8 · outbound

This paper cites Scaling up and distilling down: Language-guided robot skill acquisition.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Scaling up and distilling down: Language-guided robot skill acquisition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.174033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:129d5ea15f58d17c59f195fc69e44253b1d4cd59944afcd0cfca47347f94b147

Observation 2cb7219b-d27a-4c9f-a5b5-672bc7919744 · outbound

This paper cites CLIPort: What and where pathways for robotic manipulation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation CLIPort: What and where pathways for robotic manipulation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.177773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:100cdbf3c9b93d476788242088093961483c14c7f632e7ae7f54d5b9833583de

Observation cd9a79bb-4583-43f2-836b-b93cbef79ce4 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Octo: An Open-Source Generalist Robot Policy

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.977111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:00c8a038c74a2bfef4d341b4b5846521a99bd057b7bec00843514055a48bfc0a

Observation e1489ac1-b3cd-4d82-9466-25100e2a6139 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:54:59.300665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:4de5350872d72cc4072c962b865110859234dda9a1aa67e60bcd90527f5fbe18

Observation 48865a7a-ac69-4281-9795-6bd3731ce0a6 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:34.000529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:411d5eaa9053556e90458e19524e58a538c083f78053d14f6559c99e1fa8c5b0

Observation c3716f0a-25b9-4000-a3aa-11cb4b2ea967 · outbound

This paper cites A Generalist Agent.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation A Generalist Agent

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:24:50.082909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:285d5fa82ab81b8c698e2e15c11cd5eee5af44476b240a44ae973a502d5afa0a

Observation bc03956f-270b-43af-ad25-83488161bcde · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.815304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:59f84da7f24660d5512e41fbadf72d9e584dbfb07d6118f2206e76373b93f238

Observation a3139217-4873-46d5-a46d-7d0d627c9184 · outbound

This paper cites Perceiver-actor: A multi-task transformer for robotic manipulation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Perceiver-actor: A multi-task transformer for robotic manipulation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.180949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:e1d48e48a9fb2fd57d064af640f327f67435c20bdde8c45718763973281ce9f4

Observation ba602364-84f6-4aae-b4a4-dd2fe0a72f88 · outbound

This paper cites Chained- diffuser: Unifying trajectory diffusion and keypose prediction for robotic manipulation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Chained- diffuser: Unifying trajectory diffusion and keypose prediction for robotic manipulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.184083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:58442cad981b1833ad1fb555a29a4ee86e9a017fdbc681a75ae4b6c5541e97b1

Observation 5515dcbf-7e13-45eb-b594-7229737c80da · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:00:05.246520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:8a2288d09ec544b738b29a06586be16acb49e03ebf09e887caf1337e33418852

Observation 2d56cbc6-9735-45cf-b5fc-ff4a9e50c181 · outbound

This paper cites Act3D: 3d feature field transformers for multi-task robotic manipulation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Act3D: 3d feature field transformers for multi-task robotic manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.186857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:aa3a088d4b39f82351ae2d52cb16d1f5da83bcd0f7030c465b76c8aa2240a7e9

Observation 3ebd26d9-a903-4019-9991-e3adf16ae4ab · outbound

This paper cites RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:33.855141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:6fc52e297accc64090b5fbc508e0937e299281bbbdfb755087e5abbe83ebca26

Observation 26de1763-2ed0-4fb5-b271-f593569bf819 · outbound

This paper cites Transporters with visual foresight for solving unseen rearrangement tasks.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Transporters with visual foresight for solving unseen rearrangement tasks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.190351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:0a49f3c29b056e70362ea30cd18127a90d74e1acb81640efd638429f9b73296f

Observation fbdcaae4-5cf7-4f00-8db8-9ed792447446 · outbound

This paper cites Learning to rearrange deformable cables, fabrics, and bags with goal-conditioned transporter networks.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning to rearrange deformable cables, fabrics, and bags with goal-conditioned transporter networks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.193652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:11817761a995e3cd42d3d805febc2ba18eb1ffab919f7741fe00a13adaeeeaa9

Observation 8a22ef28-6291-4268-bfac-2d71b9463da4 · outbound

This paper cites Goal-conditioned end-to-end visuomotor control for versatile skill primitives.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Goal-conditioned end-to-end visuomotor control for versatile skill primitives

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.197390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:c93362a7a3682295e75a66551691b545e5e649dc65ff582773e19e564564dca2

Observation f082fad4-1c65-4480-92dc-e973cfee7a05 · outbound

This paper cites Wish you were here: Hindsight Goal Selection for long-horizon dexterous manipulation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Wish you were here: Hindsight Goal Selection for long-horizon dexterous manipulation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:33.888003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:75c4df17dcae4948b8b865fecef93d2996ae25ebe3bcf13639e39808dd8d91c2

Observation 74d8fc61-c806-4f7b-b7ed-78afa32c4953 · outbound

This paper cites Masked autoen- coders are scalable vision learners.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Masked autoen- coders are scalable vision learners

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.200372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:644e551df00f43ace0ddf3ed2786337e2ea4fea7cb093ba7005c03da1c9e68ea

Observation 6f82857e-191b-41d7-aa8e-df746ffaef48 · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.203484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:b1343424d679594659fc865b44b36da4581f71d7005e41859f88a9c66796a1c1

Observation 3401a4f1-c1f1-4fbe-8cb2-3e8106ed70f5 · outbound

This paper cites Masked Visual Pre-training for Motor Control.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Masked Visual Pre-training for Motor Control

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:33.901022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:cacb931f3626e2e9e201ed2a48080fc37b80c904a2f09b33ccbf5f0bb7134e89

Observation 98f0a839-81df-4218-8816-f37ed5680be1 · outbound

This paper cites Language-Driven Representation Learning for Robotics.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Language-Driven Representation Learning for Robotics

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:09:33.906792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:4188082ffc60c403885f53c65b55c35ca05f1405ea993edab804892af384f925

Observation 268d886d-ad23-4abc-8dc6-ac4a19ae3005 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation R3M: A Universal Visual Representation for Robot Manipulation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:26:54.148999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:0c8fa038ae0c82f8962bbc7dceb9ff018355a71a458f5e68c07042360bd6578e

Observation db68736c-cf24-4b92-af13-937ee2d967f4 · outbound

This paper cites Robot learning with sensorimotor pre-training.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Robot learning with sensorimotor pre-training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.206954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:59efbe36f8c6fc35b126c3c90a0c45825aa6d0eec2d1648149e11c63cfb8e315

Observation c2841de8-36dc-4eb0-b673-db594337255f · outbound

This paper cites Mastering Diverse Domains through World Models.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Mastering Diverse Domains through World Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.918442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:6e0690adf6e6d181c51350ebbc51b7fa9b58e2a32c040d97cf5134563cf4e5ed

Observation a5371c3c-0c55-465e-93e9-6f5bcb0331eb · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Any-point Trajectory Modeling for Policy Learning

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:32:51.209901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:70110bad8b57b4f8b3a4d721570d20b7f86526a1896b8b005d863714aa6473c3

Observation a9f476a2-de5b-4155-a0e9-025b556cc244 · outbound

This paper cites Learning Interactive Real-World Simulators.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning Interactive Real-World Simulators

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:15:18.951679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:f6ba1d2d8be250800f96604921c8b1ce47a678c7e0eee8f47552f19bfbb2fdbf

Observation ce25e3f6-fcc5-4004-9897-dfd748bb997c · outbound

This paper cites Masked world models for visual control.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Masked world models for visual control

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.210241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:50cf8c99df7a47b6af7627c9d3ab86a85a690c3ffdf997f5f975803aed8e8deb

Observation 34b07f7d-af36-4ccb-a4ee-5b8e2d31a340 · outbound

This paper cites Real- world robot learning with masked visual pre-training.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Real- world robot learning with masked visual pre-training

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.213940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:6d5855416d3f694f26de8d9d0a4a7c7a170b094b218f7191c62fb532e738a5ff

Observation 9847743c-6f68-4f7f-bec6-8fc2151c589f · outbound

This paper cites Exploring vi- sual pre-training for robot manipulation: Datasets, models and methods.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Exploring vi- sual pre-training for robot manipulation: Datasets, models and methods

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.028783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:dcade542c74b3f1bc8065849326d6d623e93d77eb7388945653d757d7fcfd2a1

Observation aa706a92-cc41-4c30-aafa-6693eed27395 · outbound

This paper cites Curl: Contrastive unsupervised representations for reinforcement learning.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Curl: Contrastive unsupervised representations for reinforcement learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.038585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:276f34423de58415ab1fc1e7ed8369416e27002247058079579d5fed35282e25

Observation e4ba7868-7030-4311-b156-befb91c8dc81 · outbound

This paper cites Time-contrastive networks: Self-supervised learning from video.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Time-contrastive networks: Self-supervised learning from video

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.045549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:9fcc11944ea5de0f7ac3f9979f6122330e2accf715db98fa5aa378249bd50ca0

Observation 575ad86d-0aed-44c6-8388-5c8a5117ae7d · outbound

This paper cites World Models.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation World Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.964061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:344aac23c51b6fa2c842489c5d332c3dcb2b8ec38c149b763f9681e77409f70e

Observation c5c33d96-5d94-4782-952a-f1ce9fb841e4 · outbound

This paper cites Video prediction models as rewards for reinforcement learning.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Video prediction models as rewards for reinforcement learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.052219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:14570e0eb6a02c462414455eba8777edf08b96466d6de006c513b28b34bba270

Observation c46571f0-e670-4d3e-82dd-85e9c7d8cb3c · outbound

This paper cites Learning universal policies via text-guided video generation.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning universal policies via text-guided video generation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.059156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:353ea1b7041d7ce1d8c43cda6cd47f6acfab94b64a37cac1120fe27c2ebf6abc

Observation 8b5e8a73-9bd0-4b75-8a42-e95b0725678f · outbound

This paper cites Video Language Planning.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Video Language Planning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:33.995353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:440c5292d4a2636dd3ec38289ed1fd7d01f13d4ba8a778c1d996d8c506f43e9e

Observation 17551b90-c5e5-4645-bbab-33b1f217ad9d · outbound

This paper cites Deep visual foresight for planning robot motion.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Deep visual foresight for planning robot motion

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.063399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:57ba054dbc7cc4655448281715182353eefe66d8ec3c102624f77f135aafcca2

Observation d20b0b42-f77d-4a59-96a1-a9d30b90dd33 · outbound

This paper cites MaskViT: Masked Visual Pre-Training for Video Prediction.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:09:34.008304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:53b366c045bbdab5ef32baac205ab541288aa27a49db65a6da7ca18ba6c327c4

Observation 90527fe9-f2ea-4b38-ac1c-f7fe7b9e5a1b · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Video pretraining (vpt): Learning to act by watching unlabeled online videos

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.071744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:e82015505c96312e89be4365db59d8d3759d5da43a7aabf7d6468e31c325ac47

Observation 0ddd4cde-02bf-444a-94e3-b5646e55156b · outbound

This paper cites SpawnNet: Learning generalizable visuomotor skills from pre-trained network.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation SpawnNet: Learning generalizable visuomotor skills from pre-trained network

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T01:09:34.076330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:ee7b814e33097b661da9cff2cdba159675fbdafbd432e5d059b4897f6ee9d4d8

Pith citing papers

Observation 939f2fcf-40a1-4df2-8026-e1512e961037 · inbound

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation cites this paper.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.455226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:49f4795fb1e3176c4364b3ce53f12df66ae0e53d593e2f44ff76ff4129a24b87

Observation 71da2607-7bd9-4001-9cb3-c3a03890cc6a · inbound

RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning cites this paper.

RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:08.187972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:08.187972Z digest=sha256:f1b7dbd23734b4bd63652631ef347e8f1e8d51a2f1587a5e57f07659079fb531

Observation fa15a85a-a35c-42c0-844f-70c1165c7a25 · inbound

RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation cites this paper.

RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:14:17.902796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T22:14:17.798964Z digest=sha256:94df9a4680e9352c59af7bf899db57fdd20a97e5c82998c0632a158a75cb7a90

Observation 9418149c-13e5-4fdc-84d4-6477d619c478 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.712601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:f48b1018947deded85b2c77baf5561b4218ba850a3ef54372b34c6d601866adf

Observation 0d1d4298-c3d0-4f38-b305-c9920570d619 · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:251918ae9d6351ada04568381089e3f4a3c86b525bb5c2d6433e1503e115f444

Observation fc482c4c-3273-44e8-85fd-cce4d26bac5b · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:3f1b7e3b9cfc14cda4021ca402442e1cf32a0d1dda403e00f3f53800a4aaefee

Observation 604bd5dc-71fd-48c2-8d4b-e9ac1705d730 · inbound

Universal Actions for Enhanced Embodied Foundation Models cites this paper.

Universal Actions for Enhanced Embodied Foundation Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:29:39.712076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:29:39.712076Z digest=sha256:067597cf97b0afe93fdc4260f6ec1db9fad9eae259f758b4977f9c349e3cadbd

Observation bbe0baee-a21d-480a-ba10-a40440c467dc · inbound

Generative Physical AI in Vision: A Survey cites this paper.

Generative Physical AI in Vision: A Survey GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 253

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:00.991440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:53:00.991440Z digest=sha256:ff9d0c6c8e8dc4d2403a99ab78d0127d6df905764283d111535ce8b8219ec425

Observation a654762b-eda9-4afb-b248-92a2a0f8a95a · inbound

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model cites this paper.

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:12:19.883534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T06:12:19.643111Z digest=sha256:19d225049f4724faf1774acb84f1ccc0fd0a341ce68fe49a7777fabc33c55669

Observation a6ddb56b-d85b-4a29-a34b-bd2e55e2a666 · inbound

Trajectory World Models for Heterogeneous Environments cites this paper.

Trajectory World Models for Heterogeneous Environments GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-09T15:36:48.634944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:36:48.634944Z digest=sha256:7557ae8233b437239df82f81927ab9a3704d7bcc76123fe8f2bff201bdee78a8

Observation eb92019f-2fa3-4549-9c74-070073dbd8d9 · inbound

Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation cites this paper.

Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T10:56:57.571437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:56:57.571437Z digest=sha256:9f954187a6360fe5d8f8c3b045b16aceff286f9d525a04f78865b990a4de5897

Observation d8444895-aa39-440a-81df-09e0df18a715 · inbound

RoboBERT: An End-to-end Multimodal Robotic Manipulation Model cites this paper.

RoboBERT: An End-to-end Multimodal Robotic Manipulation Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:17.892459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:35:17.892459Z digest=sha256:3de78805655f2f6edaa5057f430ebe50551add374508dccead046f39be3b8b96

Observation 02dd43f7-c83b-4d07-8567-a89caba4d63b · inbound

Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation cites this paper.

Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T00:05:05.608402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:05:05.608402Z digest=sha256:07e36740f6f9828b9a0214a3dcc163f236b408c58e26a30a4ef72ded7531d12f

Observation c2325beb-aeb7-41bb-ad5d-4ed4d96d3e51 · inbound

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots cites this paper.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:e10ab8450f2ac677200d839cdeae6d05bee909c702bbb3dbeff8be683cfda81f

Observation 49d0c939-2b4a-4ed6-a0f2-31cb3d6ca508 · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:21:45.217988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:6f0097da4112f7a9db9b16bb293cd826f63b333d2552369b4dff1a173ce98e52

Observation c6e23427-a530-41bb-980a-4fb7e5c805d6 · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.341346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:fd7d7a5e86ff5398964b181da37e5370350e4449520d6ab97011ad2e4aee44e7

Observation c98000b6-9851-4c6b-b8eb-892195b29495 · inbound

DreamGen: Unlocking Generalization in Robot Learning through Video World Models cites this paper.

DreamGen: Unlocking Generalization in Robot Learning through Video World Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:50:45.527387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T23:50:45.332466Z digest=sha256:cc52f219a71c12244e4db9303ff487801b36df2ebee6a5655b06150dcddef193

Observation 371321d0-6505-4559-9a2d-8c866389aa0a · inbound

Toward Embodied AGI: A Review of Embodied AI and the Road Ahead cites this paper.

Toward Embodied AGI: A Review of Embodied AI and the Road Ahead GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:07.249927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:07.249927Z digest=sha256:7f7f0f75e4d7268fa0a0781a6d01de830b8200e2df1fcbabdb4896944848aabe

Observation dd1f52a4-92e0-4d3f-9706-7bd8a79e4e69 · inbound

FLARE: Robot Learning with Implicit World Modeling cites this paper.

FLARE: Robot Learning with Implicit World Modeling GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:59:09.002921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T15:59:08.846629Z digest=sha256:2144c5b042de89991d9a5eba0cbd7b1d523fc2f0d669b4c60546803e50e361e7

Observation 24bbbd95-4594-40e9-8c1b-be99106b09b0 · inbound

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning cites this paper.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:32.880004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:32.880004Z digest=sha256:ec0b7e1347ee992cd4ce8195d33a21203a580b831caa2c6eb7b556d2c3802186

Observation b38532cb-ba0d-403a-a6c3-4d68fef0db36 · inbound

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models cites this paper.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:59.173764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:59.173764Z digest=sha256:ba2aadac867627a2c8830166aac34f6266d5804bd86d6a1d19120046dd087a9f

Observation 95e8f96c-6bd4-4f52-ad50-8481b3b626d0 · inbound

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks cites this paper.

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:50.945418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:50.945418Z digest=sha256:ec58aa723b9522cec3c07dfce92a9c42ad1df72d35139bc1e9662cf0014603d4

Observation 46dac9a9-e6ef-4c9b-9cad-84e4d1079bce · inbound

SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models cites this paper.

SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:37.624131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:37.624131Z digest=sha256:84959d09048728fccc8240273751f753aa3ba8c3fe897fabb285ec69094db80d

Observation 01f3c103-744e-4601-8a31-d7da21894728 · inbound

Real-Time Execution of Action Chunking Flow Policies cites this paper.

Real-Time Execution of Action Chunking Flow Policies GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:18:51.785756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T14:18:51.613045Z digest=sha256:45cbb0cd25c2ef1f4ad6f1c71cea5c40cc00f8e6bb43335aa101de0b2ea4002e

Observation 6be3951c-9154-4f5b-a981-bf788e1bd5ae · inbound

Scaling Laws of Motion Forecasting and Planning -- Technical Report cites this paper.

Scaling Laws of Motion Forecasting and Planning -- Technical Report GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:23:30.399877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:23:30.399877Z digest=sha256:1701f6ef1f894476afe43a7b978158cbc3db1bbf751853b38cc8e8752fd18657

Observation 9d7ab69f-7459-4c38-8c22-54fbacc7952c · inbound

Analytic Task Scheduler: Recursive Least Squares Based Method for Continual Learning in Embodied Foundation Models cites this paper.

Analytic Task Scheduler: Recursive Least Squares Based Method for Continual Learning in Embodied Foundation Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:52.057995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:50:52.057995Z digest=sha256:a1b920cd058da03698092612dc3e2ed31ebe2547494b7936ec6e9303fe6d1e82

Observation bfe90d72-9fec-4ea2-9b31-73ae96cec333 · inbound

ReSim: Reliable World Simulation for Autonomous Driving cites this paper.

ReSim: Reliable World Simulation for Autonomous Driving GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:32:16.666959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T09:28:00.597160Z digest=sha256:b15e71b8658d05621fc28a0b61a2ad65776b0671c4fde824ac8c87dee77c72d4

Observation 7ec13d52-da31-4ccb-9917-c8efb31e6db8 · inbound

AnchorDP3: 3D Affordance Guided Sparse Diffusion Policy for Robotic Manipulation cites this paper.

AnchorDP3: 3D Affordance Guided Sparse Diffusion Policy for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:17.292943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:17.292943Z digest=sha256:a135d09b044ec1ff5e0695ee0a6137e1f2678b181f0bfe0ace94161f64dfaf35

Observation 74e98e85-7c9f-4150-862e-1c6478143aa8 · inbound

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models cites this paper.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.111915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.111915Z digest=sha256:a69f3852dfbdd6aba4a0b7b081a6962963d39266a2cba442eed5799034176302

Observation 326aedce-c969-49eb-a43f-14627f5503f2 · inbound

WorldVLA: Towards Autoregressive Action World Model cites this paper.

WorldVLA: Towards Autoregressive Action World Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T22:57:07.883617Z digest=sha256:3923c9615e2615f08f4fcfc000dc42a0e204d8c90aa271ecd4763c3d2f5317a8

Observation c5e4d021-8b86-4a0f-91c0-7a05e508d76b · inbound

A Survey: Learning Embodied Intelligence from Physical Simulators and World Models cites this paper.

A Survey: Learning Embodied Intelligence from Physical Simulators and World Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-06T21:09:05.729636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:09:05.729636Z digest=sha256:69ddb989168f82ac18796a84e0db7da0f93a29bf17b54beaaaf0a77ee24eb590

Observation d43c688d-cb8a-4c24-815a-9141f9a3bea3 · inbound

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers cites this paper.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.569402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.569402Z digest=sha256:6a0d7cfcdaec29541fab2ed268c5c2da416d139ebdc78414def9fdc93ff83d80

Observation 88eee43d-281d-4d61-b201-fb9be1880757 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 257

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:08:35.273510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:4cf16f0c49cb94b7bf5a44822b91aef022f42d300267a8c2d41fcfd8dfb887ca

Observation 8eeac56f-b0ab-46a0-ac26-360dd6c6c23f · inbound

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation cites this paper.

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:30.281065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:30.281065Z digest=sha256:64aa337a527ed2d958a87d307a726c74b05b3b5279a1798e6886014bc7537b4e

Observation 53ce3f98-e0ed-4b1c-8840-5b64511bde7f · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:42:41.436739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:e67c4945a51c0e955347df11aea53e3a8ecfa8546a0d81a393ef32bcba42c6fb

Observation 0f6cb723-6265-49bd-96ab-e3c0fb3182a9 · inbound

Is Diversity All You Need for Scalable Robotic Manipulation? cites this paper.

Is Diversity All You Need for Scalable Robotic Manipulation? GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:10.940685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:10.940685Z digest=sha256:4a8662d2d84a86c26aff9d9fa4b6472e5a82aa34f06c12b13d9ba3c234c9d6b5

Observation 127bd5ac-8728-4d62-8040-0de9ecd378db · inbound

EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow cites this paper.

EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:30.472359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:30.472359Z digest=sha256:28872351724dc70312c7cf09adec69f5b5511edbf684d608c15f6b8aa3e810f5

Observation 410b8cb8-e57d-4839-9281-21db122a0a85 · inbound

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation cites this paper.

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:52:57.544944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T03:52:18.984005Z digest=sha256:90a53dd69384b680c7e901e4cacb07f8bf0d3e10d37d354ff410b62e15ec77aa

Observation 3fd71bcd-4126-4639-a489-df9b1984b346 · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.687590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:490383c2189b09e85b08e66bbcb15d859296b45f08930e05bba8ff42cb53513f

Observation 368ff7d1-664b-4688-9812-1998d46451f2 · inbound

VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback cites this paper.

VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:34.079434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:34.079434Z digest=sha256:2f3a3be8a982d6eff25fa327cb0df96e7a079fb5c2fcd90ac1579a3dab1c97dc

Observation 586916b1-9c11-4d27-9b30-2cb73d4aff15 · inbound

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems cites this paper.

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:39:37.917854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:39:37.917854Z digest=sha256:801d83f0af5d26b415739c3e9a4f7646ff80db14b4273d1ea8e8487be5aafa4c

Observation a754cb24-dfa2-4b2d-a1d8-0a9e118bdbbc · inbound

Video Generators are Robot Policies cites this paper.

Video Generators are Robot Policies GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:43:37.301332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T21:43:37.162870Z digest=sha256:9b88563e4a2c39ecb6f51be3800df13590afd51d23e02694f3efdfb6b6ddcd0d

Observation c99be82a-8dec-4b14-9ead-6d31a1e9b659 · inbound

GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions cites this paper.

GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T21:59:01.047936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:59:01.047936Z digest=sha256:44ba83a12784b4e86776f96558e1b737526ee8fe81a22f5f5fd944bb2015c4ca

Observation b7bcd2b4-7d39-453c-8bd1-836226b29e87 · inbound

Time delay as the origin of oscillations in anodic Si electrodissolution cites this paper.

Time delay as the origin of oscillations in anodic Si electrodissolution GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T20:50:28.898609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:50:28.898609Z digest=sha256:99dcc3a177839cce882dcca07e7d1a9c6fc18416f861ed93cfb1b3d4bac168c6

Observation 1f28b881-4803-4a69-a153-860e9e1cc6ee · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.575554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.575554Z digest=sha256:7b0919604efaebf7e199c22c0798f4245e8a93345d3cdb66c7e16e4dcf36671c

Observation b69a492f-c3ad-4efd-8d90-333674e1472f · inbound

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation cites this paper.

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:42.481513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:42.481513Z digest=sha256:a92419ca457651c9a48a33d5b06153eb0f78128f45939ab2feca234bc424eef9

Observation 80d21dc2-8b78-46d1-acae-9c4e192cad86 · inbound

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds cites this paper.

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T11:41:25.735930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:41:25.735930Z digest=sha256:9d9e734b916dcca8cd77dbafc5caff239a6052036a69a4674ca7320b8e964013

Observation e07360b3-d48d-48bb-ada4-0936878f2ed4 · inbound

ANNIE: Be Careful of Your Robots cites this paper.

ANNIE: Be Careful of Your Robots GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:00:38.085472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:00:38.085472Z digest=sha256:26f59e84a3b8682332ec8d2766c1f9465aaa5c28ecc86873136a997473b27129

Observation 05fd20a6-dd03-4bbc-9684-6b345a15f7fa · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:11.692188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:11.692188Z digest=sha256:29ff78ff70118e1453ce38e628c614bd10470f21e0d4c9003d3599e1c5b9c30d

Observation 9a6b3d53-f35f-4db1-94f0-717918725f95 · inbound

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation cites this paper.

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:51:09.094307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T08:46:51.278024Z digest=sha256:3f8a6ae7fe19d83bd6646d1ffad82635308cc5a620b780343cf77cce6fb7f85e

Observation c303c05a-5946-46f5-9ef9-266b8c1e57c2 · inbound

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models cites this paper.

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:41.373666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:34:41.373666Z digest=sha256:0645faa9b4c8d97231e27e4bad94c149b5dc979f7cb3221fb96189f7e82e51e7

Observation b5eee9cb-bc7b-4867-aa39-3256c67d8e44 · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:08.084135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:08.084135Z digest=sha256:119703a253593412c55fa085005a6784d5c1056efebcaa8633eabd9ba2956c95

Observation 6df8e2d7-5b34-4efc-bbd4-3b7d601b151c · inbound

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model cites this paper.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:45.901236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:45.901236Z digest=sha256:35a214d4694cd373e23820b9ad3e63ee0aa869138a75382c51515c1ae001e909

Observation 29b1e0fd-6614-488a-87d7-1198111f46ab · inbound

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations cites this paper.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:05:33.801038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T01:05:01.553811Z digest=sha256:fcfbd0a52cea456284e5936deee376c879d2a752b1e9f7c16ba372e00a4c39ec

Observation a6b38cc4-616c-4007-9599-7213a9cbc1d7 · inbound

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations cites this paper.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:02.038513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:02.038513Z digest=sha256:dbcff6d49b1452e95a87981f6828fc57668e46df349067c4b59fdb39f19ac0a9

Observation 9cc3191a-57ef-440e-928c-cd04e9b565ed · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:50.487718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:50.487718Z digest=sha256:e0cea239bc5852cc87859da75a837eefb445a3333d352052cb86f838fa1067ee

Observation 9a048972-d77f-4733-99e5-186f1e0bdedc · inbound

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models cites this paper.

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:42.853298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:42.853298Z digest=sha256:fc6e350cb031828b7008d9b7875e3d82c3f7491e0028033944a091339c443cab

Observation 9f5a139d-437e-4269-8eae-14788252d343 · inbound

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models cites this paper.

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:01:20.391335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T23:01:13.910539Z digest=sha256:a082b7cea9965abf38724d99942f89d9cad18651982fb37e4f29b99f8b60eb37

Observation d4dd7a13-477e-4674-b918-b269d633eca4 · inbound

AstraNav-World: World Model for Foresight Control and Consistency cites this paper.

AstraNav-World: World Model for Foresight Control and Consistency GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:28:20.760419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T19:23:58.769472Z digest=sha256:30e4d79436d24297f872eb9c2acf41e45ba30f0c16a71d860600261d2ec6ea11

Observation 485c0a13-a128-414f-a4cd-726f972065b4 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:08:01.714240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:19ebf6bf441429817c61b0d73243885f4cf99c240ec057487d9b74f2951a3eb0

Observation e838f07f-1b58-47ff-8154-a40dfbe6f4d1 · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:9a4ef388a4b7d5bde3763ba79546b86d8b4a8f77afe7f7e531211f4fafa0d39a

Observation 1473e9a7-a850-4d5e-9eb7-2eabd6112025 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:20:17.653988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:412e13830ec2c840005de152ee63c24b69e65c339ebe73186264494b09da3f5b

Observation f064819c-76c3-44ea-82c3-1336c4c75be5 · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.300469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.300469Z digest=sha256:9f9ebe2dbecc92ef008cfb0a14cda5bb780dc9566333d41f0ac2c2d438304c9c

Observation ea1abddc-40df-49e0-a190-6f5c6410a8b7 · inbound

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation cites this paper.

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:30:20.887158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T21:22:41.935691Z digest=sha256:0349bc28180969a49e09fac6cc14c655c32f6b47be32c2e698b52d8b84c27360

Observation 525b26c0-04ae-4b24-aa52-120a0c2d6713 · inbound

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation cites this paper.

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:49:54.536283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T09:49:03.333757Z digest=sha256:694d1ce3e9df07b789ec8f086a079c3338ca226a04610e95dfcacc9f0007571d

Observation 224bb51b-4ca3-4e17-b486-7da2341838a3 · inbound

Fast-WAM: Do World Action Models Need Test-time Future Imagination? cites this paper.

Fast-WAM: Do World Action Models Need Test-time Future Imagination? GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:57:30.135077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T01:57:29.753735Z digest=sha256:9f1839f7c6cc52ac08122c80505994c58b4318d70c85b01517fbc6088e9ee88b

Observation 31c6fc7b-9d62-4de1-92cf-f095e6b3c5cc · inbound

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding cites this paper.

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T17:54:04.264928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:54:04.264928Z digest=sha256:f52b2ee2daabe338fd52b0f2767b5292b44bc07c5f520992a73601216df8ee99

Observation 2c92e281-012e-4baf-8358-cc1ca07863d0 · inbound

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry cites this paper.

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T14:05:26.303000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:05:26.303000Z digest=sha256:b6cae128b6f5f2674792978d42d6e8cfe53e37516ca55302e3b6fc6939377249

Observation 64aca153-c709-4421-8759-fe5fe7bc754a · inbound

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model cites this paper.

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:58:08.868507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T18:54:07.081457Z digest=sha256:2fc2adb0865f1e7828b24bf6ebf1d1ac088c0ca6f0c9a994ea0eeafd40d4e95f

Observation 8a7c097c-feda-44a4-949d-af339ece0be9 · inbound

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data cites this paper.

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T17:03:01.191166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T17:02:18.358675Z digest=sha256:4c51fb091602bb2e57c54c680ae685c85ec60944b1136c9434bed866999143ec

Observation 0b05b596-6a27-44ba-9aa9-ca3441554943 · inbound

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning cites this paper.

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:12:08.970164Z digest=sha256:db52d8a0b52dcd26fc6473c98e3821cd6cebed111e6e730eb06d42b26c424f0d

Observation 1b020702-8eba-4d81-97f0-0122d24b5898 · inbound

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence cites this paper.

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T00:03:53.609175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:03:53.609175Z digest=sha256:ef6c810762d583e890448df2fa20f5c09b6470d0f5bd8c84aa1eb4b20b0a8f63

Observation 3af0af38-378a-4cf0-9628-184a4d5cc607 · inbound

SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds cites this paper.

SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T17:13:04.878923Z digest=sha256:94b6075df306aa109012a6054b9255f1fd0d53b50ad0d5d8a7727433631c1f88

Observation 9da4db36-6c76-4929-b26a-199d50fa8edf · inbound

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis cites this paper.

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:16:31.378588Z digest=sha256:a2557465e63f56b9046616cc50680a5602236ed4b53a77e1abfb5607ec384a49

Observation 467ae8fd-660d-4e7d-921a-f3becf43a469 · inbound

Device-Conditioned Neural Architecture Search for Efficient Robotic Manipulation cites this paper.

Device-Conditioned Neural Architecture Search for Efficient Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:55:17.595703Z digest=sha256:ea01998c5954aeb0539962beb898530a876a7cddd1db1b839b98bf0ee2ba6562

Observation 7745e842-42d2-4364-8def-827e50012805 · inbound

ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation cites this paper.

ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:11:29.558490Z digest=sha256:02a7bbfe5e559774cf64e82b6fd0abad0b999a3a72f79994218b48b873916fbf

Observation bd0a1c43-44db-47e6-97de-f46d4502d24d · inbound

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap cites this paper.

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:48:08.135538Z digest=sha256:d59734d23ad9e4c2c26870e79c692bb05bd5e12e3b27060bed3288356fc17f37

Observation c96b2652-88f5-40c4-ba0c-abfc23f58041 · inbound

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities cites this paper.

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T11:42:34.409651Z digest=sha256:cfafc3c1c17e369edfc4ea205a6e0492bda2f0ccdafbbfdb806b5dbb994a84c8

Observation a0b35e68-5302-41bf-9c25-4a6ca57c5e99 · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:7613f6aff76666ab20f4270f5e2ba18e27bfb64943c70f6c877b25dcaf7e1080

Observation 5e9a00a5-00b6-481e-b19c-fb0b53e84d95 · inbound

M100: An Orchestrated Dataflow Architecture Powering General AI Computing cites this paper.

M100: An Orchestrated Dataflow Architecture Powering General AI Computing GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T05:50:06.986471Z digest=sha256:7d3f05b5aded1f32734b5ab5b8397c67f60c96ffbb231c81471310e8d7a8a353

Observation 0c50f298-35b7-4fca-bdd2-a363c430bc6a · inbound

StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement cites this paper.

StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T04:41:45.903419Z digest=sha256:7beec7ed4965636bd76b8b5478a409d23e81ec4f0180a8697b4e468f03b8be7c

Observation 5f4596ca-8680-4a9a-9359-7e116c10c0c4 · inbound

Cortex 2.0: Grounding World Models in Real-World Industrial Deployment cites this paper.

Cortex 2.0: Grounding World Models in Real-World Industrial Deployment GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T00:52:05.348704Z digest=sha256:a66e6362865011f1539699ec5950fcc9f76428f61b67b78e3a7081e25359178f

Observation dca35df8-43a5-4d03-a98c-fdff7984f6d8 · inbound

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors cites this paper.

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T22:07:24.208555Z digest=sha256:8d7f3260155c11ca778a5a877dde0fa3c8abe679800b2720febc203cee2a1ed7

Observation 7a9cd68f-5c53-44dd-97c5-3e1e97644997 · inbound

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors cites this paper.

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T18:39:49.173828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:39:49.173828Z digest=sha256:484871b78c6e7c2545f08fbee2abe0589f201e30bf02743c7af790774e4b168a

Observation d41e0369-b592-467f-9508-0e51d3ae119a · inbound

GazeVLA: Learning Human Intention for Robotic Manipulation cites this paper.

GazeVLA: Learning Human Intention for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T11:37:30.784513Z digest=sha256:005245f91cb2a1ba28c24c683b8435e23bdd314032d3343c2ea0a531ae3710ee

Observation 4d3ac599-5a6f-4ecd-ade8-ec4dfe477813 · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T02:51:27.662262Z digest=sha256:1eae7f637c4e00e1d4103a8959a4caf0ee5d4fb73312143a606e1b500325516c

Observation 33dcb4f0-4aaf-445b-bc5c-634516bcc4b5 · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:21:29.167222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T11:16:58.104663Z digest=sha256:79be0d67e3769da67fe78dabbef76214cbef914c156ab8a5df8c4714b2028371

Observation 125432a8-ce2d-4373-83d4-7786ef498517 · inbound

Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models cites this paper.

Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T15:42:10.954090Z digest=sha256:46d43546768369233ebc274d149ae9ba4f68251cb11db470c459980148e41401

Observation af2fc661-27fc-46ad-96db-31a4c56b5248 · inbound

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation cites this paper.

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:46:26.015629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T13:50:23.835653Z digest=sha256:069f9d10ee2b1eb7a9eb0c8f5e46d17bf21bf5a44aaa56e0ede1c9a15edd0d53

Observation d58a3187-12f5-46a6-a5a5-36cb28827e9f · inbound

Being-H0.7: A Latent World-Action Model from Egocentric Videos cites this paper.

Being-H0.7: A Latent World-Action Model from Egocentric Videos GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T20:48:01.461993Z digest=sha256:07a1202442388726ac98c5351768f0e175148646ad24ec895230f5bc21ceb462

Observation 511d6429-44c9-434d-8bfc-5dc5c18018df · inbound

From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs cites this paper.

From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T17:01:38.481471Z digest=sha256:e6332bf21f3baf0e89b051a1c895b027de9471194918f36c941bb7f191b713fe

Observation 677de21c-38e8-4801-8e43-6609a3c233cc · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:51:30.705838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T16:06:10.595164Z digest=sha256:b0b646a116b50e2ba23015c739791d43345517bf76f4ed1121782d4a1c75d91c

Observation 5722ee95-773c-44ba-ac19-8edaaad46dfc · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T19:02:19.756092Z digest=sha256:bf912ed9dea4257281ad7d582f961ffac60f7f0c8589cf101e5890fa9ac172c8

Observation e192821e-1e3a-4637-bea8-c83f788243d8 · inbound

NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models cites this paper.

NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:57:03.927659Z digest=sha256:efedcd82819b22477075bf7dc232db454053b5315422690fb93d69b4b9ec4d83

Observation 12c7b0d7-79ca-4ce5-8bc5-8cab853b66d4 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T03:39:41.090350Z digest=sha256:215d14a921d224d6aa44dfe6cd2c5c233accf3c44ff731d33baad779db8ab670

Observation 257ef8f8-8173-4a0a-a913-448483894112 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:26:29.149977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T02:53:54.608425Z digest=sha256:811b93b3675a839d0cc118a88de7ae51a557cf56d00c14e41621879aba286c80

Observation a244e96d-25a2-4338-b7c4-95d901b66493 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:19:49.716782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:16:48.180290Z digest=sha256:cb45fae1a6aa8efdef867e97b2cf10df02021632029a65ccae48417ed7b18116

Observation 704d9921-473c-491f-9ffc-c523aa957017 · inbound

LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models cites this paper.

LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:36:27.577339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:07:17.590690Z digest=sha256:6cfe3e39742a92639b17d2fd7aa4d0a2d568167ee0e92d72f371069a70b16e46

Observation 02e5f680-6634-470e-b1a3-2165d8d2bdfb · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:26:27.035564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:14:54.885244Z digest=sha256:e4faad1fdc4b02bae00101cc841ed625bcaa0b310dce1a535df5603e0865790d

Observation 081ecbda-fb28-4cec-a851-fbb5a04d2201 · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:17:59.644023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T21:14:56.501485Z digest=sha256:98e96653383c4892eba54ef3a81fe1dbac467d4f7b8d869a42ed5d8bcfcce396