Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T05:09:21.028373Z
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2605.10485.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T05:09:21.028373Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T01:34:31.166730Z
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 288c0cb6-07f0-4c68-9fa7-dce3b86da573 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bf4bf8f1-22f4-4bcd-b2cd-4fd870ece669 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 86a7a6e7-2c61-4028-be60-8a39c949787a · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models π0.5: A vision- language-action model with open-world generalization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f48eba0b-4134-4598-8b20-2b2756793505 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models π0: A vision-language-action flow model for general robot control
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9fa21a8b-2f14-49c7-be3f-87dac521f47d · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 77d23117-d408-43d4-9d1a-9757cd1f4a01 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d3f023ae-bd48-4772-a659-21a03d8fb49e · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Knowledge distillation with the reused teacher classifier
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a70c30b5-eb5f-4749-a389-89f8d8610c2f · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 52b97257-e3af-4ce2-b985-68a1af823548 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 6bc36c34-3ce5-459c-8525-608760fe5bb3 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 454005b6-74b7-4281-a547-e1eeef9b7a20 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Objaverse: A universe of annotated 3d objects
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0c08331c-e246-4fa6-9091-a023d0744d34 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Rvt: Robotic view transformer for 3d object manipulation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d886c025-cf7e-4a93-8481-fb7bd91d3c31 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models arXiv preprint arXiv:2512.09619 (2025)
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 715b4cce-4370-4acc-9fc9-1b6d8a1b2c7a · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Lora: Low-rank adaptation of large language models.Iclr, 1(2):3
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 940efacc-92b9-400c-bc45-90cea4912d07 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models An Embodied Generalist Agent in 3D World
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ba798ace-4dc1-4890-a2af-bf2dcc5632aa · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Mllms need 3d-aware representation supervision for scene understanding.arXiv e-prints, pages arXiv–2506
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation eab724c8-8dae-4cc4-9b6d-888b28e2c9f6 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models What’s “up” with vision-language models? investigating their struggle with spatial reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7ab7566c-37b2-44ca-a372-912b59a37155 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b45b0076-b0c6-4a4e-a783-2dddc83475aa · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models 3d gaussian splatting for real-time radiance field rendering.ACM Trans
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation df3b2a92-fde4-4ee7-8998-fa1d832ede1a · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3bea8396-6ecf-4c27-922f-6191b3b69979 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2506cb1c-e9dd-46af-bc32-65ff85fbb04c · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models A review of robot learning for manip- ulation: Challenges, representations, and algorithms.Journal of machine learning research, 22(30):1–82
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8175d114-beaa-4b83-bc1a-5511238af15d · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models A review of spatial reasoning and interaction for real-world robotics.Advanced Robotics, 31(5):222–242
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e5835fe9-69b1-4615-a6eb-99a27f7d540b · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Pointvla: Injecting the 3d world into vision-language-action models.IEEE Robotics and Automation Letters, 11(3):2506–2513
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2a8b4649-1302-4a36-8aca-c3dedbe5e7d3 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Spatial forcing: Implicit spatial representation alignment for vision- language-action model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1f7d639f-3222-424c-9e96-3549df67ffae · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation aeb4f825-3c86-4c40-a6c6-ced04b25825f · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation dbe476cc-2772-48f6-b10c-9808912ccfa6 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5f0b34c4-c348-44be-8e12-221c47f87d3d · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models DINOv2: Learning Robust Visual Features without Supervision
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 734766e6-bd05-4d34-bfbe-5cc6ed4bc686 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b1798aea-963e-4d10-9c5e-9539590e036e · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Film: Visual reasoning with a general conditioning layer
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d5e96b2a-3d85-4d7c-a0e7-fc27ade84f9d · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7304605e-475b-4f2f-ae88-4d10a8a2145c · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Perceiver-actor: A multi-task transformer for robotic manipulation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation df22501b-42a1-44d2-b133-cf43852717d7 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f5979ade-eb3a-4b41-93fe-c6b66ec3e213 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Rocket: Residual-oriented multi-layer alignment for spatially- aware vision-language-action models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 27f97e89-9f07-455e-b991-4f3b9dbc9460 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 553f3816-b02f-45ec-b513-2f8cf2bab091 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 57f4e4bb-af80-4afe-b382-23f7831ed237 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Bridgedata v2: A dataset for robot learning at scale
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 689ab1b1-4cb2-410b-b924-eac27ae39780 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Vggt: Visual geometry grounded transformer
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 482c5687-db62-43d7-8405-4cb4c7ae365d · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Depth anything: Unleashing the power of large-scale unlabeled data
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 903892b5-6838-4477-a974-7a5381729265 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Depth anything v2.Advances in Neural Information Processing Systems, 37:21875– 21911
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bbe0a506-6e8f-4564-9f9f-dce213bb1c31 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Scannet++: A high- fidelity dataset of 3d indoor scenes
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 97c48c6d-4df8-4c1a-86fd-39b946f8bd03 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 18578f66-026b-4a86-9dd7-05bb52f171a7 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Improving 2d feature representations by 3d-aware fine-tuning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9b5d40de-7725-4c42-a780-9c37402bdccf · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Sigmoid loss for language image pre-training
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d462be3f-9b37-4448-b91f-275895efa924 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a8584f01-c829-465d-920b-0c1b98325f7b · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Rt-2: Vision-language-action models transfer web knowledge to robotic control
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ae111cc0-b8df-4cbf-b2d2-735e49b36e8f · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models This task requires precise spatial perception to locate the screen and hinge, as well as smooth and controlled motion to avoid damaging the articulated structure during contact
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f9c5205d-ad76-4eef-adde-3a561af8b304 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models This task requires accurate object localization and a smooth transfer trajectory to ensure stable grasping and precise placement without dropping the object
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c17223a3-08d7-41d7-842e-a7adb38e5216 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 556c2877-d691-4ec2-a8da-7b4e56801933 · outbound
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9637b991-567d-4564-8175-23d732540a7d · inbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.