Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T01:09:33.761708Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 100 inbound Pith citation observations for arXiv:2410.06158.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T01:09:33.761708Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:45:08.187972Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
71 of 71 outbound references displayed
External citation measurements
7
pith, observed 2026-08-05T02:28:24.338817Z
Observation 1681862c-8797-498e-ab03-0e708c0f6962 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 16ce52d4-f074-4d00-a845-28a55ebe854a · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation SAM 2: Segment Anything in Images and Videos
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2aa8cec0-f7f5-43ac-8f02-7564bad2625c · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Video generation models as world simulators
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3cd5f690-4395-4629-83e9-e783e58ebc25 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Language Models are Few-Shot Learners
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a447728-a6e8-4d22-a898-b8a063f217b3 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e737ffb-8f5f-4b6f-b9fe-5826499686e2 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning transferable visual models from natural language supervision
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 441ec216-e095-4645-8a32-b697ef14b739 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Taming transformers for high-resolution image synthesis
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 89f8363d-7a4f-4465-b62a-7aa8f8e927b2 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3a38fe7a-fa78-497a-a924-d4d98fc45e07 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Ego4D: Around the world in 3,000 hours of egocentric video
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49508fdd-42b0-4dbd-a2c8-22c579393ce0 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation something something
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 62463d2e-1f17-4f8e-83a2-9a55965958c6 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Scaling egocentric vision: The epic-kitchens dataset
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5ad23283-0ada-4c2a-8427-860ad863e79b · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation A Short Note on the Kinetics-700 Human Action Dataset
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 58c428cc-2796-450b-87ae-949709600d7d · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation MediaPipe: A Framework for Building Perception Pipelines
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 69dab474-df5c-42f4-b93d-986272b955b6 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Open-Sora: Democratizing efficient video production for all, March 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7ae26f0a-6d32-4375-a17a-a4d574c2675a · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RT-1: Robotics Transformer for Real-World Control at Scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1dd81452-8926-44e4-bfac-7b938c8fb77e · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Bridgedata v2: A dataset for robot learning at scale
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8e5054e3-9deb-4a16-b147-42a3735798a1 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning structured output representation using deep conditional generative models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8af91c3b-b87b-4a5f-ad58-80efcb4a903b · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Auto-Encoding Variational Bayes
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c69ba143-c4b4-43bb-a989-26311cd4081d · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0736680e-b7f6-4c5e-ab8c-d16a24b5d3b5 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation MOMA-Force: Visual-force imitation for real-world mobile manipulation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 30f4e127-9ee0-4025-b01f-6237c9b772de · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a84e9232-eb80-4d27-8f98-29d337ffe866 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Denoising diffusion probabilistic models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7bd6f53c-6754-4429-8698-0fc0cd3a90f2 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ceb5b34c-6da8-4628-be1b-a4342c81dea2 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Segment anything
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a85f75e1-5a6a-4787-b844-1042c2ff57bd · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Latte: Latent Diffusion Transformer for Video Generation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2b89f9b8-447e-4a74-9298-61be67248f89 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 68e95ac0-7a3f-4dc2-8a90-d819b0095180 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation What matters in language conditioned robotic imitation learning over unstructured data
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 826d4c58-8d89-4f86-a768-72cedeec22af · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Vision-Language Foundation Models as Effective Robot Imitators
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb94fd1d-303c-4ee1-8144-d70b041977cc · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7510ea83-c361-43ef-86c2-79f0af46a2b1 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation VIMA: General Robot Manipulation with Multimodal Prompts
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8908412-3de5-4aca-b777-a135c72f5127 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Language Conditioned Imitation Learning over Unstructured Data
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c4364eba-7c13-4166-a29e-4b2e0a3cd779 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation BC-Z: Zero-shot task generalization with robotic imitation learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2a989c10-d26a-4ea3-9150-3cca8f4962e5 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d10340df-1f50-49c3-94d8-33acd6fe79f8 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Scaling up and distilling down: Language-guided robot skill acquisition
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2cb7219b-d27a-4c9f-a5b5-672bc7919744 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation CLIPort: What and where pathways for robotic manipulation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cd9a79bb-4583-43f2-836b-b93cbef79ce4 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Octo: An Open-Source Generalist Robot Policy
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e1489ac1-b3cd-4d82-9466-25100e2a6139 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 48865a7a-ac69-4281-9795-6bd3731ce0a6 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c3716f0a-25b9-4000-a3aa-11cb4b2ea967 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation A Generalist Agent
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bc03956f-270b-43af-ad25-83488161bcde · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation OpenVLA: An Open-Source Vision-Language-Action Model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a3139217-4873-46d5-a46d-7d0d627c9184 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Perceiver-actor: A multi-task transformer for robotic manipulation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba602364-84f6-4aae-b4a4-dd2fe0a72f88 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Chained- diffuser: Unifying trajectory diffusion and keypose prediction for robotic manipulation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5515dcbf-7e13-45eb-b594-7229737c80da · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2d56cbc6-9735-45cf-b5fc-ff4a9e50c181 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Act3D: 3d feature field transformers for multi-task robotic manipulation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3ebd26d9-a903-4019-9991-e3adf16ae4ab · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 26de1763-2ed0-4fb5-b271-f593569bf819 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Transporters with visual foresight for solving unseen rearrangement tasks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fbdcaae4-5cf7-4f00-8db8-9ed792447446 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning to rearrange deformable cables, fabrics, and bags with goal-conditioned transporter networks
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8a22ef28-6291-4268-bfac-2d71b9463da4 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Goal-conditioned end-to-end visuomotor control for versatile skill primitives
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f082fad4-1c65-4480-92dc-e973cfee7a05 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Wish you were here: Hindsight Goal Selection for long-horizon dexterous manipulation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74d8fc61-c806-4f7b-b7ed-78afa32c4953 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Masked autoen- coders are scalable vision learners
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f82857e-191b-41d7-aa8e-df746ffaef48 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3401a4f1-c1f1-4fbe-8cb2-3e8106ed70f5 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Masked Visual Pre-training for Motor Control
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 98f0a839-81df-4218-8816-f37ed5680be1 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Language-Driven Representation Learning for Robotics
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 268d886d-ad23-4abc-8dc6-ac4a19ae3005 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation R3M: A Universal Visual Representation for Robot Manipulation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation db68736c-cf24-4b92-af13-937ee2d967f4 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Robot learning with sensorimotor pre-training
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c2841de8-36dc-4eb0-b673-db594337255f · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Mastering Diverse Domains through World Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a5371c3c-0c55-465e-93e9-6f5bcb0331eb · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Any-point Trajectory Modeling for Policy Learning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a9f476a2-de5b-4155-a0e9-025b556cc244 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning Interactive Real-World Simulators
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ce25e3f6-fcc5-4004-9897-dfd748bb997c · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Masked world models for visual control
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 34b07f7d-af36-4ccb-a4ee-5b8e2d31a340 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Real- world robot learning with masked visual pre-training
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9847743c-6f68-4f7f-bec6-8fc2151c589f · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Exploring vi- sual pre-training for robot manipulation: Datasets, models and methods
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aa706a92-cc41-4c30-aafa-6693eed27395 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Curl: Contrastive unsupervised representations for reinforcement learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e4ba7868-7030-4311-b156-befb91c8dc81 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Time-contrastive networks: Self-supervised learning from video
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 575ad86d-0aed-44c6-8388-5c8a5117ae7d · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation World Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5c33d96-5d94-4782-952a-f1ce9fb841e4 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Video prediction models as rewards for reinforcement learning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c46571f0-e670-4d3e-82dd-85e9c7d8cb3c · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Learning universal policies via text-guided video generation
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8b5e8a73-9bd0-4b75-8a42-e95b0725678f · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Video Language Planning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 17551b90-c5e5-4645-bbab-33b1f217ad9d · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Deep visual foresight for planning robot motion
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d20b0b42-f77d-4a59-96a1-a9d30b90dd33 · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation MaskViT: Masked Visual Pre-Training for Video Prediction
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 90527fe9-f2ea-4b38-ac1c-f7fe7b9e5a1b · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Video pretraining (vpt): Learning to act by watching unlabeled online videos
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0ddd4cde-02bf-444a-94e3-b5646e55156b · outbound
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation SpawnNet: Learning generalizable visuomotor skills from pre-trained network
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 939f2fcf-40a1-4df2-8026-e1512e961037 · inbound
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 71da2607-7bd9-4001-9cb3-c3a03890cc6a · inbound
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa15a85a-a35c-42c0-844f-70c1165c7a25 · inbound
RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9418149c-13e5-4fdc-84d4-6477d619c478 · inbound
What Matters in Building Vision-Language-Action Models for Generalist Robots GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0d1d4298-c3d0-4f38-b305-c9920570d619 · inbound
Cosmos World Foundation Model Platform for Physical AI GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fc482c4c-3273-44e8-85fd-cce4d26bac5b · inbound
FAST: Efficient Action Tokenization for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 604bd5dc-71fd-48c2-8d4b-e9ac1705d730 · inbound
Universal Actions for Enhanced Embodied Foundation Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe0baee-a21d-480a-ba10-a40440c467dc · inbound
Generative Physical AI in Vision: A Survey GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 253
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a654762b-eda9-4afb-b248-92a2a0f8a95a · inbound
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a6ddb56b-d85b-4a29-a34b-bd2e55e2a666 · inbound
Trajectory World Models for Heterogeneous Environments GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb92019f-2fa3-4549-9c74-070073dbd8d9 · inbound
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8444895-aa39-440a-81df-09e0df18a715 · inbound
RoboBERT: An End-to-end Multimodal Robotic Manipulation Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02dd43f7-c83b-4d07-8567-a89caba4d63b · inbound
Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2325beb-aeb7-41bb-ad5d-4ed4d96d3e51 · inbound
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49d0c939-2b4a-4ed6-a0f2-31cb3d6ca508 · inbound
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c6e23427-a530-41bb-980a-4fb7e5c805d6 · inbound
GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c98000b6-9851-4c6b-b8eb-892195b29495 · inbound
DreamGen: Unlocking Generalization in Robot Learning through Video World Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 371321d0-6505-4559-9a2d-8c866389aa0a · inbound
Toward Embodied AGI: A Review of Embodied AI and the Road Ahead GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd1f52a4-92e0-4d3f-9706-7bd8a79e4e69 · inbound
FLARE: Robot Learning with Implicit World Modeling GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 24bbbd95-4594-40e9-8c1b-be99106b09b0 · inbound
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b38532cb-ba0d-403a-a6c3-4d68fef0db36 · inbound
Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e8f96c-6bd4-4f52-ad50-8481b3b626d0 · inbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46dac9a9-e6ef-4c9b-9cad-84e4d1079bce · inbound
SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f3c103-744e-4601-8a31-d7da21894728 · inbound
Real-Time Execution of Action Chunking Flow Policies GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6be3951c-9154-4f5b-a981-bf788e1bd5ae · inbound
Scaling Laws of Motion Forecasting and Planning -- Technical Report GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d7ab69f-7459-4c38-8c22-54fbacc7952c · inbound
Analytic Task Scheduler: Recursive Least Squares Based Method for Continual Learning in Embodied Foundation Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfe90d72-9fec-4ea2-9b31-73ae96cec333 · inbound
ReSim: Reliable World Simulation for Autonomous Driving GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7ec13d52-da31-4ccb-9917-c8efb31e6db8 · inbound
AnchorDP3: 3D Affordance Guided Sparse Diffusion Policy for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74e98e85-7c9f-4150-862e-1c6478143aa8 · inbound
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 326aedce-c969-49eb-a43f-14627f5503f2 · inbound
WorldVLA: Towards Autoregressive Action World Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5e4d021-8b86-4a0f-91c0-7a05e508d76b · inbound
A Survey: Learning Embodied Intelligence from Physical Simulators and World Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 178
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d43c688d-cb8a-4c24-815a-9141f9a3bea3 · inbound
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88eee43d-281d-4d61-b201-fb9be1880757 · inbound
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 257
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8eeac56f-b0ab-46a0-ac26-360dd6c6c23f · inbound
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53ce3f98-e0ed-4b1c-8840-5b64511bde7f · inbound
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f6cb723-6265-49bd-96ab-e3c0fb3182a9 · inbound
Is Diversity All You Need for Scalable Robotic Manipulation? GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 127bd5ac-8728-4d62-8040-0de9ecd378db · inbound
EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 410b8cb8-e57d-4839-9281-21db122a0a85 · inbound
AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3fd71bcd-4126-4639-a489-df9b1984b346 · inbound
GR-3 Technical Report GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 368ff7d1-664b-4688-9812-1998d46451f2 · inbound
VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 586916b1-9c11-4d27-9b30-2cb73d4aff15 · inbound
Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a754cb24-dfa2-4b2d-a1d8-0a9e118bdbbc · inbound
Video Generators are Robot Policies GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c99be82a-8dec-4b14-9ead-6d31a1e9b659 · inbound
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7bcd2b4-7d39-453c-8bd1-836226b29e87 · inbound
Time delay as the origin of oscillations in anodic Si electrodissolution GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f28b881-4803-4a69-a153-860e9e1cc6ee · inbound
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69a492f-c3ad-4efd-8d90-333674e1472f · inbound
Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80d21dc2-8b78-46d1-acae-9c4e192cad86 · inbound
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e07360b3-d48d-48bb-ada4-0936878f2ed4 · inbound
ANNIE: Be Careful of Your Robots GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05fd20a6-dd03-4bbc-9684-6b345a15f7fa · inbound
GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a6b3d53-f35f-4db1-94f0-717918725f95 · inbound
R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c303c05a-5946-46f5-9ef9-266b8c1e57c2 · inbound
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5eee9cb-bc7b-4867-aa39-3256c67d8e44 · inbound
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6df8e2d7-5b34-4efc-bbd4-3b7d601b151c · inbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b1e0fd-6614-488a-87d7-1198111f46ab · inbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a6b38cc4-616c-4007-9599-7213a9cbc1d7 · inbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cc3191a-57ef-440e-928c-cd04e9b565ed · inbound
RynnVLA-002: A Unified Vision-Language-Action and World Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a048972-d77f-4733-99e5-186f1e0bdedc · inbound
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f5a139d-437e-4269-8eae-14788252d343 · inbound
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d4dd7a13-477e-4674-b918-b269d633eca4 · inbound
AstraNav-World: World Model for Foresight Control and Consistency GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 485c0a13-a128-414f-a4cd-726f972065b4 · inbound
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e838f07f-1b58-47ff-8154-a40dfbe6f4d1 · inbound
World Action Models are Zero-shot Policies GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1473e9a7-a850-4d5e-9eb7-2eabd6112025 · inbound
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f064819c-76c3-44ea-82c3-1336c4c75be5 · inbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea1abddc-40df-49e0-a190-6f5c6410a8b7 · inbound
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 525b26c0-04ae-4b24-aa52-120a0c2d6713 · inbound
Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 224bb51b-4ca3-4e17-b486-7da2341838a3 · inbound
Fast-WAM: Do World Action Models Need Test-time Future Imagination? GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 31c6fc7b-9d62-4de1-92cf-f095e6b3c5cc · inbound
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c92e281-012e-4baf-8358-cc1ca07863d0 · inbound
World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64aca153-c709-4421-8759-fe5fe7bc754a · inbound
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8a7c097c-feda-44a4-949d-af339ece0be9 · inbound
From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0b05b596-6a27-44ba-9aa9-ca3441554943 · inbound
ViVa: A Video-Generative Value Model for Robot Reinforcement Learning GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b020702-8eba-4d81-97f0-0122d24b5898 · inbound
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3af0af38-378a-4cf0-9628-184a4d5cc607 · inbound
SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9da4db36-6c76-4929-b26a-199d50fa8edf · inbound
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 467ae8fd-660d-4e7d-921a-f3becf43a469 · inbound
Device-Conditioned Neural Architecture Search for Efficient Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7745e842-42d2-4364-8def-827e50012805 · inbound
ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd0a1c43-44db-47e6-97de-f46d4502d24d · inbound
Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c96b2652-88f5-40c4-ba0c-abfc23f58041 · inbound
${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a0b35e68-5302-41bf-9c25-4a6ca57c5e99 · inbound
Human Cognition in Machines: A Unified Perspective of World Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e9a00a5-00b6-481e-b19c-fb0b53e84d95 · inbound
M100: An Orchestrated Dataflow Architecture Powering General AI Computing GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0c50f298-35b7-4fca-bdd2-a363c430bc6a · inbound
StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f4596ca-8680-4a9a-9359-7e116c10c0c4 · inbound
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dca35df8-43a5-4d03-a98c-fdff7984f6d8 · inbound
CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7a9cd68f-5c53-44dd-97c5-3e1e97644997 · inbound
CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d41e0369-b592-467f-9508-0e51d3ae119a · inbound
GazeVLA: Learning Human Intention for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4d3ac599-5a6f-4ecd-ade8-ec4dfe477813 · inbound
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 33dcb4f0-4aaf-445b-bc5c-634516bcc4b5 · inbound
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 125432a8-ce2d-4373-83d4-7786ef498517 · inbound
Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af2fc661-27fc-46ad-96db-31a4c56b5248 · inbound
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d58a3187-12f5-46a6-a5a5-36cb28827e9f · inbound
Being-H0.7: A Latent World-Action Model from Egocentric Videos GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 511d6429-44c9-434d-8bfc-5dc5c18018df · inbound
From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 677de21c-38e8-4801-8e43-6609a3c233cc · inbound
RLDX-1 Technical Report GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5722ee95-773c-44ba-ac19-8edaaad46dfc · inbound
RLDX-1 Technical Report GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e192821e-1e3a-4637-bea8-c83f788243d8 · inbound
NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 12c7b0d7-79ca-4ce5-8bc5-8cab853b66d4 · inbound
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 257ef8f8-8173-4a0a-a913-448483894112 · inbound
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a244e96d-25a2-4338-b7c4-95d901b66493 · inbound
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 704d9921-473c-491f-9ffc-c523aa957017 · inbound
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 02e5f680-6634-470e-b1a3-2165d8d2bdfb · inbound
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 081ecbda-fb28-4cec-a851-fbb5a04d2201 · inbound
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.