Pith. sign in

Paper Citation Record · LEDGER

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

As of 3 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2605.10485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10485 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T05:09:21.028373Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:34:31.166730Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact19
  • verified fuzzy30
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 288c0cb6-07f0-4c68-9fa7-dce3b86da573 · outbound

This paper cites 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.923919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:8db39a5fe092ea813572618bf91943b81f19fb2d4eeec972d190b167ca635064

Observation bf4bf8f1-22f4-4bcd-b2cd-4fd870ece669 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:36:24.912557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:fda2f2ddb47915decf5bd99490bed9f2f4deff5ef1fd5ae5f47a4dc568f0a816

Observation 86a7a6e7-2c61-4028-be60-8a39c949787a · outbound

This paper cites π0.5: A vision- language-action model with open-world generalization.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models π0.5: A vision- language-action model with open-world generalization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.065834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:3df381d0b924ecc29f3da023551e695ed75c97cbbac60181e804cb0a25bdffc2

Observation f48eba0b-4134-4598-8b20-2b2756793505 · outbound

This paper cites π0: A vision-language-action flow model for general robot control.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models π0: A vision-language-action flow model for general robot control

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.069312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:f9ea6d1fa36b6e74fda864db1858ec4d8ea079502b4c68386d4669b62359a578

Observation 9fa21a8b-2f14-49c7-be3f-87dac521f47d · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:36:24.888491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:6cbb9de9f542c124360a99239ae7b39a6d8b5fe4471da8009deaeab68f37bd7b

Observation 77d23117-d408-43d4-9d1a-9757cd1f4a01 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.061909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:4c5f80b03fa40a23b3161415fe27f5e6fe263d9c3aed8d7e5f9024338932759c

Observation d3f023ae-bd48-4772-a659-21a03d8fb49e · outbound

This paper cites Knowledge distillation with the reused teacher classifier.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Knowledge distillation with the reused teacher classifier

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.071961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:8574dab82b59ce16f2b9834138e6a441737e466609e7247d032b95f186f98ea0

Observation a70c30b5-eb5f-4749-a389-89f8d8610c2f · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:36:24.899722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:5875df0e5fb22c75332e867358ae93e82b925259f224e4773639d30f8f9ed140

Observation 52b97257-e3af-4ce2-b985-68a1af823548 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.871267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:9ad2ead6721dd11640f81b0eaf206f9166649ad42a37e2f5a958bb44681e808e

Observation 6bc36c34-3ce5-459c-8525-608760fe5bb3 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.074517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:3cd0164fd2df5b20a577d81a68517dca2228fa8f5736e486419e50115bee17af

Observation 454005b6-74b7-4281-a547-e1eeef9b7a20 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Objaverse: A universe of annotated 3d objects

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.148408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:2ef3d6839c5e592dd29c27c4cccdb36312cae61d54f722147a96ef8182963428

Observation 0c08331c-e246-4fa6-9091-a023d0744d34 · outbound

This paper cites Rvt: Robotic view transformer for 3d object manipulation.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Rvt: Robotic view transformer for 3d object manipulation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.164911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:7054e0ade0900533697b88a1ded85cb278505cdb3be6a02c7d0cce3f7c2a547b

Observation d886c025-cf7e-4a93-8481-fb7bd91d3c31 · outbound

This paper cites arXiv preprint arXiv:2512.09619 (2025).

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models arXiv preprint arXiv:2512.09619 (2025)

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.193107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:8a26cfa99e23fdf6b5e56311953402eeb9a07f13a1977c63a487cb95ff8ea115

Observation 715b4cce-4370-4acc-9fc9-1b6d8a1b2c7a · outbound

This paper cites Lora: Low-rank adaptation of large language models.Iclr, 1(2):3.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Lora: Low-rank adaptation of large language models.Iclr, 1(2):3

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.153632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:a087b45ef03f460579081739fd92c0ee143a737c70ae0c63a4b2df86446a0073

Observation 940efacc-92b9-400c-bc45-90cea4912d07 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models An Embodied Generalist Agent in 3D World

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:22:18.773673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:4389cd224050c200d9f6933173d4db7cc0edb363917048b88ccae6e8ebcc4169

Observation ba798ace-4dc1-4890-a2af-bf2dcc5632aa · outbound

This paper cites Mllms need 3d-aware representation supervision for scene understanding.arXiv e-prints, pages arXiv–2506.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Mllms need 3d-aware representation supervision for scene understanding.arXiv e-prints, pages arXiv–2506

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.137843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:cb47563bb5c7a884caff2c7f0775d9eec9eb71f3abbc46d2eb489c18f17f491e

Observation eab724c8-8dae-4cc4-9b6d-888b28e2c9f6 · outbound

This paper cites What’s “up” with vision-language models? investigating their struggle with spatial reasoning.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models What’s “up” with vision-language models? investigating their struggle with spatial reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.101141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:5b6fe08f3751239eae265a3912b9dedeb25790578e182eafd4098e01b3189581

Observation 7ab7566c-37b2-44ca-a372-912b59a37155 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.104776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:6a193659dbc60fc8d466d148ae1e5af759ea4f57ca8790f4c06c96aecd6c4627

Observation b45b0076-b0c6-4a4e-a783-2dddc83475aa · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.ACM Trans.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models 3d gaussian splatting for real-time radiance field rendering.ACM Trans

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.108503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ff9f4a6490c406681460e5efff2548cf152327846127efb9238661546a002627

Observation df3b2a92-fde4-4ee7-8998-fa1d832ede1a · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:31:26.205311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ce9f6193fd23a0bda795b52f89afa779dfd95a3275a5d005aa2c786fa907d54f

Observation 3bea8396-6ecf-4c27-922f-6191b3b69979 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:31:26.139493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ea7cc5bfb312d68cc62a4922f8fea5cc37f811347966291fb681902c014fc672

Observation 2506cb1c-e9dd-46af-bc32-65ff85fbb04c · outbound

This paper cites A review of robot learning for manip- ulation: Challenges, representations, and algorithms.Journal of machine learning research, 22(30):1–82.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models A review of robot learning for manip- ulation: Challenges, representations, and algorithms.Journal of machine learning research, 22(30):1–82

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.119351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:f051ea47fe75f927473f0f5eea97f17629adc42f4aa4754a3d949362b6bddce7

Observation 8175d114-beaa-4b83-bc1a-5511238af15d · outbound

This paper cites A review of spatial reasoning and interaction for real-world robotics.Advanced Robotics, 31(5):222–242.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models A review of spatial reasoning and interaction for real-world robotics.Advanced Robotics, 31(5):222–242

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.123395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:2209b18427d82b1e0c077e70b2e34b37de67ee7f8ec34d6e7bb7f8c122d697c9

Observation e5835fe9-69b1-4615-a6eb-99a27f7d540b · outbound

This paper cites Pointvla: Injecting the 3d world into vision-language-action models.IEEE Robotics and Automation Letters, 11(3):2506–2513.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Pointvla: Injecting the 3d world into vision-language-action models.IEEE Robotics and Automation Letters, 11(3):2506–2513

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.141966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:495ce3ee7735aec947cba970097c0297a2ca88855f7f3bab0ec3962f705266d4

Observation 2a8b4649-1302-4a36-8aca-c3dedbe5e7d3 · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision- language-action model.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Spatial forcing: Implicit spatial representation alignment for vision- language-action model

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.188928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:a680a28fa4518e453d3f4779b6208fd5c6c8ed6528ab238d01f9872302bc7221

Observation 1f7d639f-3222-424c-9e96-3549df67ffae · outbound

This paper cites Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.174886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ed0e93ec48e7475dd36241b6835d87d4476353f0434408c4c65eebf4d38546d7

Observation aeb4f825-3c86-4c40-a6c6-ced04b25825f · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.093663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:bb75d6019758e730dc247b09283e1b2193e63aca19977860f398a4d30ce78e3a

Observation dbe476cc-2772-48f6-b10c-9808912ccfa6 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:31:26.219409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:28116831b080b3005d8704c6f3418284eae75fbf35c1bba3a16f18bb62ff1a12

Observation 5f0b34c4-c348-44be-8e12-221c47f87d3d · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models DINOv2: Learning Robust Visual Features without Supervision

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:36:24.928992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:086e4e29217dc2df732e229a8c317d34e42cea8c1e77d3a4a5c21e2571f3541f

Observation 734766e6-bd05-4d34-bfbe-5cc6ed4bc686 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.145253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ee45aee874fd3d18d3e7e601ec19bd25ec5f2b65c5d1938a0bb92cdffc9ce651

Observation b1798aea-963e-4d10-9c5e-9539590e036e · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Film: Visual reasoning with a general conditioning layer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.157223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:7f9162ae38081bafac79e2e17f6414fef8d1058127fa9bda642009a33eb3e937

Observation d5e96b2a-3d85-4d7c-a0e7-fc27ade84f9d · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:12:22.898827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ecfdf6166d698cba78adcb8eea403666d5ffa41fce26cc8190cdc76db8e2f555

Observation 7304605e-475b-4f2f-ae88-4d10a8a2145c · outbound

This paper cites Perceiver-actor: A multi-task transformer for robotic manipulation.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Perceiver-actor: A multi-task transformer for robotic manipulation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.126906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:b6904087381348c24ba1c781e114b99325c5cd81510e5c0b06b77f555ee57242

Observation df22501b-42a1-44d2-b133-cf43852717d7 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.150931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:d36235f161dea4e0b6f80d1c6f9467009524bc229609dbe02df729bdf6ab73ea

Observation f5979ade-eb3a-4b41-93fe-c6b66ec3e213 · outbound

This paper cites Rocket: Residual-oriented multi-layer alignment for spatially- aware vision-language-action models.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Rocket: Residual-oriented multi-layer alignment for spatially- aware vision-language-action models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.168690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:2e4453b1a249c5a3bedcf6264f52907606e012bc8f9ab5c29b8143c1a5c0821e

Observation 27f97e89-9f07-455e-b991-4f3b9dbc9460 · outbound

This paper cites GeoVLA: Empowering 3D Representations in Vision-Language-Action Models.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.179579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:81197ad0bba02de533b3f6a860ef466c3083b20baaa5ce75cc08a5c87bd0f33c

Observation 553f3816-b02f-45ec-b513-2f8cf2bab091 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:31:26.154335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:c87ad8c8d1b6fe2c3e93d91d6b8d938e09f30bc6071848dbaa20799c5cfc3be1

Observation 57f4e4bb-af80-4afe-b382-23f7831ed237 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Bridgedata v2: A dataset for robot learning at scale

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.130223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:5525e7fb849dd4a2cb4ab409586f5b16f3b4c6798b97b0d41a13601b856dc6ed

Observation 689ab1b1-4cb2-410b-b924-eac27ae39780 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Vggt: Visual geometry grounded transformer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.090295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:3f936dfa93fff19ff7de99c913cae97122e0802435d6dfb72c59dcb8e50e9550

Observation 482c5687-db62-43d7-8405-4cb4c7ae365d · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Depth anything: Unleashing the power of large-scale unlabeled data

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.096839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:be47ef4aed87c896acf7e7ceda89955044012cd826dab3d2de5cc86f33587216

Observation 903892b5-6838-4477-a974-7a5381729265 · outbound

This paper cites Depth anything v2.Advances in Neural Information Processing Systems, 37:21875– 21911.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Depth anything v2.Advances in Neural Information Processing Systems, 37:21875– 21911

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.112063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:051a5ff6bc1a754d82eaf8b35edb869d1a38043cacbbab4bc8faf0458383f751

Observation bbe0a506-6e8f-4564-9f9f-dce213bb1c31 · outbound

This paper cites Scannet++: A high- fidelity dataset of 3d indoor scenes.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Scannet++: A high- fidelity dataset of 3d indoor scenes

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.133373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:21b3fafa1a729547efc7c3bf2f5daed11f3e797659fea5ff1c71622ce2aa4b9d

Observation 97c48c6d-4df8-4c1a-86fd-39b946f8bd03 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.143646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:5dab4113543c5dbb111f37083007aaf1fba50a246710d07bc6ff97cecf0d83fa

Observation 18578f66-026b-4a86-9dd7-05bb52f171a7 · outbound

This paper cites Improving 2d feature representations by 3d-aware fine-tuning.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Improving 2d feature representations by 3d-aware fine-tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.115308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:edc2d4bb07f9dbdaadbc5632c54a764c7f2a454812f3db875eff8dc103085da5

Observation 9b5d40de-7725-4c42-a780-9c37402bdccf · outbound

This paper cites Sigmoid loss for language image pre-training.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Sigmoid loss for language image pre-training

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.159779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:8e2237cf21134a4b0d846b89ffdd15b1b5e08d4c095b7a72c85d9d07380ec4fa

Observation d462be3f-9b37-4448-b91f-275895efa924 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:5a010f9d20a3ef6402acf0fa57c9232dbe54d51d43db5774f5e149409dab8f6f

Observation a8584f01-c829-465d-920b-0c1b98325f7b · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.162351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:f14ef6297f1f6a72d14e597b13e4aa7ef16ce56d1686ab74616ecb069e155e0c

Observation ae111cc0-b8df-4cbf-b2d2-735e49b36e8f · outbound

This paper cites This task requires precise spatial perception to locate the screen and hinge, as well as smooth and controlled motion to avoid damaging the articulated structure during contact.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models This task requires precise spatial perception to locate the screen and hinge, as well as smooth and controlled motion to avoid damaging the articulated structure during contact

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.083821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:3aa2714b4da072e914cc541f08ce34eee38e3740eeb252fbdaaba2c0964fe9b9

Observation f9c5205d-ad76-4eef-adde-3a561af8b304 · outbound

This paper cites This task requires accurate object localization and a smooth transfer trajectory to ensure stable grasping and precise placement without dropping the object.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models This task requires accurate object localization and a smooth transfer trajectory to ensure stable grasping and precise placement without dropping the object

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.087208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:1409ad9e54a88d4011cfe3a0d0146178381add05cd0cc49c68cfdd441271b911

Observation c17223a3-08d7-41d7-842e-a7adb38e5216 · outbound

This paper cites an unresolved cited work.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-12T12:11:33.080104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:9852d719ac836a40003087838325ff9098ce3c0d9b0de3f3e6e6a160f5896ea6

Observation 556c2877-d691-4ec2-a8da-7b4e56801933 · outbound

This paper cites an unresolved cited work.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-12T12:11:33.077238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:e4994fb0cdbffddac4e6ddfeae3052403c66bdfb1de8596a7e3dae2932cf520a

Pith citing papers

Observation 9637b991-567d-4564-8175-23d732540a7d · inbound

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models cites this paper.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.166730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.166730Z digest=sha256:9821df7b518beabf596431e3e8769192a29ffcf44f393b4c8bfc7cd6fa5e6272