Pith. sign in

Paper Citation Record · LEDGER

GEM: Generative Supervision Helps Embodied Intelligence

As of 6 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 3 inbound Pith citation observations for arXiv:2605.28548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28548 v1

Coverage vector

measured 100 of 102 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T13:38:27.263726Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:39:40.704498Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T07:14:45.365277Z

Reference resolution

100 of 102 outbound references displayed

  • verified exact69
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d149a499-859e-4355-865e-1ae78eac85b5 · outbound

This paper cites write newline.

GEM: Generative Supervision Helps Embodied Intelligence write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:10f99f2e56e931b044cb6d1021c4d45ce431429ad13ff3be74212313649f95b4

Observation cce4394d-aa34-44a0-9a72-ee1efba4bc1b · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

GEM: Generative Supervision Helps Embodied Intelligence Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.951334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:7b2977e89c1dd7f1041d0ef6504ee3b83a51bcc3fdd52c54be5562c00766ca0d

Observation d2c07c8c-e579-45ae-a33b-f17058acbf10 · outbound

This paper cites Qwen3-VL Technical Report.

GEM: Generative Supervision Helps Embodied Intelligence Qwen3-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.946776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:99824c3ea909f5170e27e68e42982059724a5b3bf0784c5b4c3fb5a731536f4a

Observation 51404052-19ab-47d6-8d16-ab7e07d4cbd9 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

GEM: Generative Supervision Helps Embodied Intelligence ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.804096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:1b4f7247174b72e63f524668c520749df7e5665e9cc986af6037c405f219c627

Observation f3327a7f-ee00-4e74-aa6b-7bc6423df2ea · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

GEM: Generative Supervision Helps Embodied Intelligence PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.793790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:bb488e26a0653e3a3c76eaeace50a3583075c866f8fedb31ccdca4c852f9fb1c

Observation c46c33a2-bf7f-4032-9c7d-164324bdf8bf · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

GEM: Generative Supervision Helps Embodied Intelligence RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.921818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:f9e0753228400b3755d416e3854f3eeca79959c72e2cd52303d863ac67cace2f

Observation b18ed60c-d513-41b6-a475-f2592cdaa16f · outbound

This paper cites Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.

GEM: Generative Supervision Helps Embodied Intelligence Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:134de99cab09f3f8e506edca974f3132f6e6b0eaa25710cf221ad4df78ba04bb

Observation dd2c1c0c-3251-4f5f-ac73-e9c70abd3172 · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.

GEM: Generative Supervision Helps Embodied Intelligence Scaling spatial intelligence with multimodal foundation models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.949546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:4c81f43d40388cbf7d72804aead6b7f711a319dcbdae6a7f2b566000a8ac7d65

Observation 8d3e1211-b1f8-489c-8ced-33474cd2eabf · outbound

This paper cites SAM 3: Segment Anything with Concepts.

GEM: Generative Supervision Helps Embodied Intelligence SAM 3: Segment Anything with Concepts

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.876255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:f8517af6ed07cfb83d115109d4538c329904e23e75d65234fd7d34718aea5c2d

Observation 4f799f29-9b9b-4aa7-be31-067e8ecb854f · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

GEM: Generative Supervision Helps Embodied Intelligence WorldVLA: Towards Autoregressive Action World Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.924772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:3f796a36b092e00168e99731a4d928931372630691e033e170369f084b02194e

Observation aab9c73a-15fd-4a9c-8148-179cb121f737 · outbound

This paper cites GR-3 Technical Report.

GEM: Generative Supervision Helps Embodied Intelligence GR-3 Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.764033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:2c932c49a422129ed8c897f2a885e9becc4097b4df1c3669f4e1c22d53ae6a0c

Observation c92fe02a-1a3d-4205-8329-bffe61f6edda · outbound

This paper cites Blip3o-next: Next frontier of native image generation.

GEM: Generative Supervision Helps Embodied Intelligence Blip3o-next: Next frontier of native image generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.792569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:5c7cfee69aeb696982d3bdebc1f9c4a039972e2e9d6dc883db453e9a96215d69

Observation 78ecdf25-e94e-4225-b0d2-a12583a15c8c · outbound

This paper cites Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets.

GEM: Generative Supervision Helps Embodied Intelligence Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:43:28.939856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:db113060c232b671755b4cff8d01656bd6bd92140db4626b0e315d1b9a874fc5

Observation 3580ea63-40f7-4af6-bebd-f53a390412f7 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

GEM: Generative Supervision Helps Embodied Intelligence Diffusion policy: Visuomotor policy learning via action diffusion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:b3e9ea06467d5498795e2b91d35e1442bd1227ec1a3762aee91aca8446d12669

Observation a8b5c485-6f35-45d2-a448-feb30664d723 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

GEM: Generative Supervision Helps Embodied Intelligence Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.927253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:1dd0cbf63d0f7e92c2923fd1a129082058cd8103cf7933dcfb3002858fd0a6de

Observation aaaee761-e61f-4c03-b94c-22b714a04762 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

GEM: Generative Supervision Helps Embodied Intelligence Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:8404c422460183c85f5dc0a42688cd537bba95d74fb12be9b5fe457cff00835a

Observation ad7f14f4-521a-46a0-89d9-5dc61ec3a577 · outbound

This paper cites Rynnbrain: Open embodied foundation models.

GEM: Generative Supervision Helps Embodied Intelligence Rynnbrain: Open embodied foundation models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.777758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:790c9fba79d9448cad7c9add2f510b3d22956158824c96a5450c38e1078e34a2

Observation 9ba5e188-954c-4619-a305-f0fa0f162471 · outbound

This paper cites Embspatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models.

GEM: Generative Supervision Helps Embodied Intelligence Embspatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:0340025686595667ddf2ea4a01cfa9205b5355fc96c3b86bfd586965901b69ba

Observation 5ad13a23-9cba-4100-a90f-dd043cf2f2db · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

GEM: Generative Supervision Helps Embodied Intelligence VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.785956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:77e2d4ebe1222522931eed3b10f03c98d613d6fe1421a523e0b31e90a58ed5d5

Observation c74c67d6-8a92-4a42-9dd2-cd7b5946ed9c · outbound

This paper cites Visuospatial Cognitive Assistant.

GEM: Generative Supervision Helps Embodied Intelligence Visuospatial Cognitive Assistant

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.944298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:cf6fe55aff582700bc02d0d21b64a32e1234e4a1091a1999b233892a454f1dab

Observation 45701357-72f3-4fe9-8998-35f72aa9c7b5 · outbound

This paper cites Roboafford++: A generative ai-enhanced dataset for multimodal affordance learning in robotic manipulation and navigation.

GEM: Generative Supervision Helps Embodied Intelligence Roboafford++: A generative ai-enhanced dataset for multimodal affordance learning in robotic manipulation and navigation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.954054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:f6dc7b89c79d84e51d0f442b52032c58eebca546b6a9232b9927a21296b505d3

Observation 5b6cf150-4ea8-42a7-b59e-46072d6ddae4 · outbound

This paper cites MiMo-Embodied: X-Embodied Foundation Model Technical Report.

GEM: Generative Supervision Helps Embodied Intelligence MiMo-Embodied: X-Embodied Foundation Model Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.934583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:1c10aaf23db6a3b83b64b6d66b1a846740c7bce2b56c1a988f176c888cfcb36d

Observation 0dfcbdc7-a68a-4b8e-a4bf-32a0476fa631 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

GEM: Generative Supervision Helps Embodied Intelligence Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.944572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:e5a17de0210da1726789d9c00762c1823c2d7005cfd023b0d6aaeda891edeab8

Observation 1896c951-cbf5-49ad-868c-2121fd568b56 · outbound

This paper cites ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning.

GEM: Generative Supervision Helps Embodied Intelligence ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.917162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:16caed6d0231366c962d1ded26e200ea96c4b932eb042b24db2802acee96b6b5

Observation 89ad9c6d-6cf2-4437-a452-ea8ef2027d7a · outbound

This paper cites GPT-4o System Card.

GEM: Generative Supervision Helps Embodied Intelligence GPT-4o System Card

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.914189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:d41b501a28ff1cbff1bc447f117a443e003b5b1631c07b7a86962618fe13c3b1

Observation 8ec73933-0dda-4e21-b120-cb70c1d9a08a · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

GEM: Generative Supervision Helps Embodied Intelligence $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.908931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:465af509f6b564726ed4027d0ce377f14c8bd453498b9b8a2378b5940a23354a

Observation 9fe3db1d-6a66-423e-87de-c84b1baa9d07 · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

GEM: Generative Supervision Helps Embodied Intelligence RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.926574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:93f812d7ae767cf485e9732428170f09f9f540742285621c6d79d148caad600a

Observation ed181f75-1356-4880-aaa9-4b5863f53358 · outbound

This paper cites RynnVLA-001: Using human demonstrations to improve robot manipulation.

GEM: Generative Supervision Helps Embodied Intelligence RynnVLA-001: Using human demonstrations to improve robot manipulation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.929111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:e933b415221471099002fc1a94b498e7f32e04f1b46bc0479d588873c99da091

Observation f0a74d72-77bf-4c37-8041-1b363cbf9d68 · outbound

This paper cites Cotracker3: Simpler and better point tracking by pseudo-labelling real videos.

GEM: Generative Supervision Helps Embodied Intelligence Cotracker3: Simpler and better point tracking by pseudo-labelling real videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:356a254fa9d71d433544939fbb91e1284fd17459f0d47fa959826921cccc45f3

Observation 22f5facc-aa40-4113-923e-9cfb0b8ea02c · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

GEM: Generative Supervision Helps Embodied Intelligence DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.951915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:f9c98f792ccc53b790c4a6a4940488cc8184c46bbe894bf2cb9548ea3bbfe84d

Observation d3ef739d-bc24-4664-9722-4212ff047edc · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

GEM: Generative Supervision Helps Embodied Intelligence OpenVLA: An Open-Source Vision-Language-Action Model

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.924107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:1872797458579ecc270250dde750df37e0ffbe5e6a18ccf51aed6ae1abb318e7

Observation cae21675-b328-41cd-ae0b-ac355fa4d93f · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

GEM: Generative Supervision Helps Embodied Intelligence MolmoAct: Action Reasoning Models that can Reason in Space

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.895201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:aff147c25ca99cfa6f5f274dd32683e0a34fa58e4fcc47aeec3d10d42f5d73f7

Observation 24a59ded-67bd-4f9f-b8ef-40577202b133 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GEM: Generative Supervision Helps Embodied Intelligence LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.954255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:a777228ad63875f9f05c8d658520885e0f1acb53ac3efff6778710b13ef5cf5f

Observation aea6106e-fbe8-4c8f-ac8b-7c57f1422c16 · outbound

This paper cites Pointvla: Injecting the 3d world into vision-language-action models.

GEM: Generative Supervision Helps Embodied Intelligence Pointvla: Injecting the 3d world into vision-language-action models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:9c4e2563e3db02cf7e067e0815eb8f36c32b0f25d62b908768a3e88152c2f703

Observation f29ec315-13fb-4534-8dc7-ca5e719354a8 · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision- language-action model.

GEM: Generative Supervision Helps Embodied Intelligence Spatial forcing: Implicit spatial representation alignment for vision- language-action model

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.936462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:abe1d7f54bfc49b979ef54462d6f4222762e12ea1328874e06004e9400ae0c08

Observation 7638faef-1280-4847-bdcb-39420c9c4093 · outbound

This paper cites 3ds-vla: A 3d spatial-aware vision language action model for robust multi-task manipulation.

GEM: Generative Supervision Helps Embodied Intelligence 3ds-vla: A 3d spatial-aware vision language action model for robust multi-task manipulation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:02988521a02da82c2676860e92bdd62325077a0b48d6de42fe411bc4db72d987

Observation 32c9a440-5ccd-4b66-8b71-731a5e88279d · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

GEM: Generative Supervision Helps Embodied Intelligence Vision-Language Foundation Models as Effective Robot Imitators

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.884284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:c5e701ab3c7651750b413ff9ce9f169ba88e793e0085e970fe5675c31217aa74

Observation 420d55aa-e54b-4445-a22b-2486fb8596f8 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

GEM: Generative Supervision Helps Embodied Intelligence Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.901126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:f764c7ec9374a16f7d2c4cec860d087c56956d774cc54cfdc54c5c2af269d640

Observation 63fa33e6-d0df-481d-aac7-8cbcefdcc2cf · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

GEM: Generative Supervision Helps Embodied Intelligence Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.933921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:97e9a3da67415ad5a41acc441edaf6e06598f9b64016067d586d1ce41cfc3ab6

Observation 8316b8f0-9ee0-41c0-8158-a3d2f1a04e95 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

GEM: Generative Supervision Helps Embodied Intelligence Depth Anything 3: Recovering the Visual Space from Any Views

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.849859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:06a84bbfe45d681ce47264bf18367b76c02934df5522130470b46086f5bd154d

Observation e06f79e8-75bb-4d92-a64f-e90d4b851b1b · outbound

This paper cites Flow Matching for Generative Modeling.

GEM: Generative Supervision Helps Embodied Intelligence Flow Matching for Generative Modeling

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.941625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:71fa9d3ceb393f21dc8403349f5e4b9ca671a99a684f3dea3620058267e84a02

Observation ed31c78d-0064-421b-a293-ca77b70ef5a4 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

GEM: Generative Supervision Helps Embodied Intelligence Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:1b0246ebb456cecb75880f1458d98f709b74471b1b2a5c2ddf336ada15088139

Observation 9bf277dd-3037-4213-8f4a-cc6cbbf0e714 · outbound

This paper cites Visual instruction tuning.

GEM: Generative Supervision Helps Embodied Intelligence Visual instruction tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:0633c0717ca28af4c725b0dd9d117a598b8de9252ccd6b2e40292a3c6e274863

Observation f4134f73-13cc-4725-a6e5-1fcc1a83f232 · outbound

This paper cites Towards generalist robot policies: What matters in building vision-language-action models.

GEM: Generative Supervision Helps Embodied Intelligence Towards generalist robot policies: What matters in building vision-language-action models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:5c96cd6b98ae205c84762e26866e03f0cb61e2a24c531f706a146af1447e0b45

Observation bda20188-d9fd-4f43-a762-7b0433f1aecf · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

GEM: Generative Supervision Helps Embodied Intelligence HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.798947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:ffe66f1eee9e69acebfe7714a88ccc14608bfb73c63b2d53b21ca1e922726868

Observation 011333c3-bc43-4a5e-92f2-0153eb018680 · outbound

This paper cites arXiv preprint arXiv:2602.03310 (2026).

GEM: Generative Supervision Helps Embodied Intelligence arXiv preprint arXiv:2602.03310 (2026)

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:43:28.795239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:45d2ae18542f5ef914c6a0081b1b50ac42e2e8eee1f0e36c74b5fe16ad58ef8a

Observation 0396a53c-9495-4734-841d-a9f0cd4f4001 · outbound

This paper cites Vl-grasp: a 6-dof interactive grasp policy for language-oriented objects in cluttered indoor scenes.

GEM: Generative Supervision Helps Embodied Intelligence Vl-grasp: a 6-dof interactive grasp policy for language-oriented objects in cluttered indoor scenes

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:9445fcf009ac1f03b56a5840b36d0779b1aa666719331d982b9501ee2c57858c

Observation b4b13735-5d8f-4332-935e-4cce4db729cf · outbound

This paper cites Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces.

GEM: Generative Supervision Helps Embodied Intelligence Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.937015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:83598813efd6ff4e89eaeeeaa6d859ef2d700faddfbabb41c3a8d0a3d099ae1a

Observation a8ae2393-4e79-48f5-9916-c6b8ee7bc8ee · outbound

This paper cites F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions.

GEM: Generative Supervision Helps Embodied Intelligence F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.946603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:a2a0b6540078013b08d97242b76fb82961b8096bc3447b362acd3668aad2f114

Observation 6b80c074-06b8-4843-ba0a-74075e53344d · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

GEM: Generative Supervision Helps Embodied Intelligence SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.881710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:96627e63f9de419fffb535459fb35597338c479373700e0ce92d6c85844d1680

Observation 7443e8b9-874b-4ec0-beb1-b359a0d40b0a · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

GEM: Generative Supervision Helps Embodied Intelligence Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:887de6e25ab357a55eb9d94830365c32210108694f7fb37401e37cf7b3f0e762

Observation 757facc2-130e-43c5-9987-6323a51dced0 · outbound

This paper cites Scalable diffusion models with transformers.

GEM: Generative Supervision Helps Embodied Intelligence Scalable diffusion models with transformers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:cd5580ec86628a684102268eb4e8e3d8107cda4a5d3012edade4c500ced251d3

Observation 23c0f6ae-9ec2-43f2-a155-ee34fd3bed08 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

GEM: Generative Supervision Helps Embodied Intelligence FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.893229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:e416c0a542848665f4b784dd47a1eab4c69c53609a38ca218f9f1618d3f34341

Observation 0a39b032-3823-4d47-af39-e073a2dd5721 · outbound

This paper cites Eo-1: Interleaved vision- text-action pretraining for general robot control.

GEM: Generative Supervision Helps Embodied Intelligence Eo-1: Interleaved vision- text-action pretraining for general robot control

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.922358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:c5967138e34c107c2190b2f74efee92f4d9e5030ceb8e029cf8dbe1c0b7ed00d

Observation 8e88c1db-c1f0-4e51-91bc-3de251b11927 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

GEM: Generative Supervision Helps Embodied Intelligence SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.867061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:7b57dcffa4a3123bf93a971e69790d209ba48d2e36b7d75d33cc0fbe9b6e0258

Observation 0ec6e0d1-103e-464c-bff1-b5be53ddfba4 · outbound

This paper cites Paco: Parts and attributes of common objects.

GEM: Generative Supervision Helps Embodied Intelligence Paco: Parts and attributes of common objects

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:7d0bbd5237da078c11bf01a29b234020b87891d06591d269ba9f2ed12c21d962

Observation 2c21c9b8-9b0e-4001-ab59-24bdc7c3523a · outbound

This paper cites an unresolved cited work.

GEM: Generative Supervision Helps Embodied Intelligence Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:e32d2aad963c9571b0a6c4e0b37828fe3d3185f707c8b7736835b91f58c576e5

Observation 3676bcb4-76ed-49c6-b56f-5cf19048cee2 · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

GEM: Generative Supervision Helps Embodied Intelligence Robovqa: Multimodal long-horizon reasoning for robotics

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:e45f78ad66b5c8bb89253c2391cfca561eae1d3b30e0b9bf36e7eccef232731e

Observation e491a828-d36f-47ab-96ef-c40910d87bdf · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

GEM: Generative Supervision Helps Embodied Intelligence DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.861506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:4462e83c7b84c03978ddbce3f5527e6de433d4120e5e34d9e9d9da6977aaa557

Observation bf9b9fd9-d1c2-42f6-b1a6-247efb96cba1 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.

GEM: Generative Supervision Helps Embodied Intelligence Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:0feb6caebcebfb97dd282f598493db2be1f73214c57b74515180431610b123cf

Observation 9fb6d44f-2a71-4c89-bed3-bc563272dbff · outbound

This paper cites ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.

GEM: Generative Supervision Helps Embodied Intelligence ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.942277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:8cf93c827227f79dec8ebbfc2d8206a54ae1e7c7058c52b7a0e1cf1dd0e81e7f

Observation 59befd15-bfbc-4da6-859e-0d76b8179f58 · outbound

This paper cites Starvla: A lego-like codebase for vision-language-action model developing.

GEM: Generative Supervision Helps Embodied Intelligence Starvla: A lego-like codebase for vision-language-action model developing

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:bd3eb5ff73e1fafa3ee30ce20f92ad3e1731a5fdd4632f88d6fc1a468fa24aae

Observation 477cffff-638c-47d8-b17e-458e07374898 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

GEM: Generative Supervision Helps Embodied Intelligence Gemini Robotics: Bringing AI into the Physical World

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.869499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:dda67c62cd009b1d0ec629d7cc6e61523eda3bc1b115dfb6cd139d01cbff61da

Observation 00e6adc2-62a9-43ec-9b09-071d2d2a4756 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

GEM: Generative Supervision Helps Embodied Intelligence Octo: An Open-Source Generalist Robot Policy

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.874260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:0c20742c4922aa8b81784c0c9c94ea2a48698052bfab25f151ac4865288de35e

Observation cf16ddec-9cb7-43db-97d2-1bb8a77c7e27 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

GEM: Generative Supervision Helps Embodied Intelligence Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:b83074448d237c84dc62893b05198561fc343ce0b346eddc9e0292319411a5c7

Observation 77592b1e-6166-47b0-805a-2b002ec7e1d4 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

GEM: Generative Supervision Helps Embodied Intelligence InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.914029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:27201e03a0f501cadb2555c6dafde45f4241978277491e6dabe8f7a7644ecb07

Observation c087c14f-7309-451f-ab3e-31e832e2427d · outbound

This paper cites Unified Vision-Language-Action Model.

GEM: Generative Supervision Helps Embodied Intelligence Unified Vision-Language-Action Model

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.871097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:cc62bdfd77d770f0ece714327a909754cc4e68917a9f5f136015113fdb31df04

Observation cf21edbb-3892-4a0a-94ef-bd6f289db748 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

GEM: Generative Supervision Helps Embodied Intelligence Chain-of-thought prompting elicits reasoning in large language models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:186897a2aa13dfa7f6f4cf7e5c704c72ed5558eb7b57e0320a3cbde07b59dc55

Observation 7749cdb2-e03e-48d6-b7d8-fec1d4aade5c · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.

GEM: Generative Supervision Helps Embodied Intelligence Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:d6b9035dcd8d06e4779aede263a23fbf2237bf674b21f190908a080bebbc12b2

Observation 882414c5-d819-46f9-98bd-130c6ef31df2 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

GEM: Generative Supervision Helps Embodied Intelligence OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.850081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:08845255b633b4d60854c7a33c36602ade90e2c5dc4180308ad2ae5c19b272cf

Observation c9c7379b-66ff-4fa3-b08d-7033d905fc6e · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

GEM: Generative Supervision Helps Embodied Intelligence RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.858519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:050f623f2e2a03c55b62cd1a164d1435673409cc8cdcad96f57751afb203d019

Observation 0f6dca17-7b46-466b-b44b-f633467122c8 · outbound

This paper cites RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation.

GEM: Generative Supervision Helps Embodied Intelligence RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.949108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:c3a02bf08f3f2f5452253b3b2380ec79db9f74cc2b2208b61bfffe46fe8daf03

Observation 4481ce55-7ea7-476b-b11e-eadcdd9f5432 · outbound

This paper cites A Pragmatic VLA Foundation Model.

GEM: Generative Supervision Helps Embodied Intelligence A Pragmatic VLA Foundation Model

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.845036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:5b38c99a7f240b30f32bb25a1ffbfd965d66db9cf66cc5881ed9b2e2d4e1f6a0

Observation 690dcfed-bd83-4f91-9d08-52d4b0bdc4a9 · outbound

This paper cites HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents.

GEM: Generative Supervision Helps Embodied Intelligence HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.847594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:43f9981e9b48270d869d827a12058b0b6e9fdd7877e01075ca79fc86c899ad0e

Observation fef89dcd-5141-4252-aef4-4732d826f87f · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

GEM: Generative Supervision Helps Embodied Intelligence SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.852429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:2f42f3974db3734329eea23020ba1d35469d70f962cb0e9530a39e2f0418b396

Observation 1ebaed09-99d6-469e-9d12-4cdbcbef8e66 · outbound

This paper cites Qwen3 Technical Report.

GEM: Generative Supervision Helps Embodied Intelligence Qwen3 Technical Report

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.855673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:9657d567c3f9daf50867bce39cc5629834c9c8a7c093b3e1019e54772256101e

Observation 2841ee5a-1afe-4dee-b816-c0eceeb941d2 · outbound

This paper cites Vlaser: Vision-language-action model with synergistic embodied reasoning.arXiv preprint arXiv:2510.11027, 2025b.

GEM: Generative Supervision Helps Embodied Intelligence Vlaser: Vision-language-action model with synergistic embodied reasoning.arXiv preprint arXiv:2510.11027, 2025b

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.842607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:091eae132a291ac7e90c5ac47a4e7053f0628be68c254f38cab0655dc6dd8b2c

Observation 2a47a3e9-d31b-4b46-bae8-a1979ec1f024 · outbound

This paper cites Magma: A foundation model for multimodal ai agents.

GEM: Generative Supervision Helps Embodied Intelligence Magma: A foundation model for multimodal ai agents

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:daa47eafa8ab7dcc1ffb47b2ca6c56ac9ebd84e32962fc63bb547a23fc12c110

Observation 901bbc8e-98cc-4570-a9e3-b2c672d4005b · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

GEM: Generative Supervision Helps Embodied Intelligence Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.855009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:c53ace512064a47b86ffa0d99a27f3d3ea394f2836ee8cb4e71acdea92d500c6

Observation 2a18153b-99ac-4dbf-9a69-f2b2503746d5 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

GEM: Generative Supervision Helps Embodied Intelligence Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:43:28.873798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:ca03fc723ea8a771976531ca24bcf90089369310109cc262405310099cc4b4c6

Observation 9e1f3388-9e16-4361-bcd6-c0826a546caf · outbound

This paper cites Cambrian-S: Towards Spatial Supersensing in Video.

GEM: Generative Supervision Helps Embodied Intelligence Cambrian-S: Towards Spatial Supersensing in Video

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.833403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:9d636eb3154ef115e9467559d7a16393bf8309a1fd16f9e5af10f707c9df43f9

Observation cc10bdb9-2a05-4871-bab4-2bd054ce30db · outbound

This paper cites MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence.

GEM: Generative Supervision Helps Embodied Intelligence MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.844628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:03aa5cc6866e8170dbf9a087361c404f2235c5abc752330efca2bb59e996a999

Observation 34de4d10-d593-4454-b3e4-34752fbae90a · outbound

This paper cites ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding.

GEM: Generative Supervision Helps Embodied Intelligence ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.911677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:90561f54447c82223cb64c647bf9138315d602c2e1503a783022134e8019d597

Observation 18754250-5f34-488e-bf35-4046fb1c6787 · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d indoor scenes.

GEM: Generative Supervision Helps Embodied Intelligence Scannet++: A high-fidelity dataset of 3d indoor scenes

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:0be9ba99b5287d7112f42b767a22763bec92c0c6a363e59b215e3e4815cf97e9

Observation bd951bdd-e0a6-4387-9914-8c31179266e7 · outbound

This paper cites Spatial mental modeling from limited views.

GEM: Generative Supervision Helps Embodied Intelligence Spatial mental modeling from limited views

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:267589a2daf7794e2d8071d30bad4794ca2676deb866aa5eadc9de16bf20922f

Observation 4fe380c9-921a-40f9-a01a-1e5fbc9e33eb · outbound

This paper cites Depthvla: Enhancing vision-language-action models with depth-aware spatial reasoning.

GEM: Generative Supervision Helps Embodied Intelligence Depthvla: Enhancing vision-language-action models with depth-aware spatial reasoning

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.837226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:a7fca416890fd507ae85bd80ce72bd1a062fc607b3a297a1dbf8a2dbab512d10

Observation 16fc1a9d-0a96-405d-833a-d9b5e6981623 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

GEM: Generative Supervision Helps Embodied Intelligence RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.820593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:3977d653f20d7ca8d9dea9f9138f4bfe14cf58b5a171e76d5962cb54fedecd0f

Observation ac05f819-a2e2-4e5f-a106-eb4f53d22e91 · outbound

This paper cites Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation.

GEM: Generative Supervision Helps Embodied Intelligence Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.817755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:6d4735c220835382f65078bdf82ed13632ba1199e11335b2191cb1e49285ad53

Observation d81748c7-be38-4348-8b05-e7dd10731347 · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

GEM: Generative Supervision Helps Embodied Intelligence 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.815179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:fd2e7199cc5184708c2999dc71b0afbbc959bccf6a23069d55e725cce6d51435

Observation f39df5ec-2370-4c23-ad8f-fc59df839686 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.

GEM: Generative Supervision Helps Embodied Intelligence From flatland to space: Teaching vision-language models to perceive and reason in 3d

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.791564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:580518aa1fde9ad270eb3f43f02423ed387c4184c924fbbf6e84969c54ecb664

Observation 0e610237-103c-4e29-85ab-bfb040883cc3 · outbound

This paper cites UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent.

GEM: Generative Supervision Helps Embodied Intelligence UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.812492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:143a2b0f2a2a33c4e2487e1b911b503047a7a742348f18d7f67261b50b2d4f75

Observation a51b2dcf-2af1-4390-b5cb-3398cac4ac65 · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

GEM: Generative Supervision Helps Embodied Intelligence DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.138463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:544eb26eff7c7bf2800628c5ba9ec01664c84b87d18ab0b5794dd02cd48cddff

Observation 0df22255-3ef6-4892-a4ba-11ee48feff11 · outbound

This paper cites Zhang, C.

GEM: Generative Supervision Helps Embodied Intelligence Zhang, C

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:43:28.820997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:09b5afbe0e7b6f35191ef26f6467883dd4f3f68a9c7c22dfea2ece6e5c58fd55

Observation fe4372dc-fa7b-47b5-8963-6475ddf37df7 · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.

GEM: Generative Supervision Helps Embodied Intelligence Cot-vla: Visual chain-of-thought reasoning for vision-language-action models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:2c1b6b603f484cf6b643c3893087878a64d8ee4bf41dac0c8386701795616c92

Observation fe1324de-d757-4bcd-8f4d-16d37117c13d · outbound

This paper cites FlexiDreamer: Single Image-to-3D Generation with FlexiCubes.

GEM: Generative Supervision Helps Embodied Intelligence FlexiDreamer: Single Image-to-3D Generation with FlexiCubes

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.809289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:808df21c6643d5dbf060b8356e61a1242f8c414b181431dcd9e17c0b9e3d3be5

Observation 87a8e034-e975-4b8f-a6ad-596d92d895fe · outbound

This paper cites Deepmesh: Auto-regressive artist-mesh creation with reinforcement learning.

GEM: Generative Supervision Helps Embodied Intelligence Deepmesh: Auto-regressive artist-mesh creation with reinforcement learning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-06-29T13:38:27.263726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:3ef45db1bb1b25d9a8aa0ad07bf86993f6dc3d24546b4d967a510d3750f88f2d

Observation cedfbd86-02c0-40c5-8fa6-fba9b19e4bdc · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

GEM: Generative Supervision Helps Embodied Intelligence 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.812321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:0101392e5a0718903fafc9a3df8c14093dd6c690d411c0d6115853787d75a340

Observation 50c0f55d-bc2f-4a19-807f-bc9bcee0951a · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

GEM: Generative Supervision Helps Embodied Intelligence TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.803323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:d194cfaa036a5652870092d032ddffe4d4fb694e1bbae52714e9919e28094e9f

Observation 5438bedc-2204-4b28-8b6d-4a5f7b2dcc1d · outbound

This paper cites RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics.

GEM: Generative Supervision Helps Embodied Intelligence RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.823705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:567fdb3687cd35f1a4e96908677153d0645ae7d9d4a457c50b11e47446391f5d

Observation 60ac761d-65ae-454b-a709-412f26e1fbbe · outbound

This paper cites Open3D: A Modern Library for 3D Data Processing.

GEM: Generative Supervision Helps Embodied Intelligence Open3D: A Modern Library for 3D Data Processing

Reference 103

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.832119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:fcb2d3e17d19bf18240252fa4e3e832579b9dbc475417e4b30d7229484a01458

Pith citing papers

Observation 994de202-8e44-4e4a-8417-b3dd6f3266d0 · inbound

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models cites this paper.

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models GEM: Generative Supervision Helps Embodied Intelligence

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T19:33:54.716433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T04:41:28.564104Z digest=sha256:e3cdec73ebefa88300987c7c334a6ae01374a918dce730271cf200eeeb5bf237

Observation 5f4a89b6-e6f7-4190-8252-542136a69e98 · inbound

From Foundation to Application: Improving VLA Models in Practice cites this paper.

From Foundation to Application: Improving VLA Models in Practice GEM: Generative Supervision Helps Embodied Intelligence

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.366955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:116c9f10f68d0fdc828d86ee21ddd0a54c505408eb3f9fd0a6152f6bc324b521

Observation 9a094702-8da5-4b8d-b8e7-8868f3660ee5 · inbound

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text cites this paper.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GEM: Generative Supervision Helps Embodied Intelligence

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.704498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.704498Z digest=sha256:ebd2a55a4de58b9b5f1f4de16b1f7f3a3dcc27e98d6e74a7759d744a36f9c7ab