Pith. sign in

Paper Citation Record · LEDGER

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

As of 9 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 36 inbound Pith citation observations for arXiv:2508.09071.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09071 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:16:50.651526Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:19:46.803171Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:38:43.940925Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 044b3559-17b0-4d81-81da-b5245e869a54 · outbound

This paper cites Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.631795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.631795Z digest=sha256:f0a6168c03c6f188aca073d576d53f0128d7920d36d18cea0dae0ca825ce8dbb

Observation 3f75680b-eb5e-4213-93c0-4d7397703433 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.637420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.637420Z digest=sha256:7202c66d5d9fa2aef64b3a750f68e7e4bec65ff6eadf42e5d274008f175d5835

Observation 19dc6192-36bc-4fbf-995e-a2ea50decd90 · outbound

This paper cites Flow Matching for Generative Modeling.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Flow Matching for Generative Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.640176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.640176Z digest=sha256:906432caf7ca1ad383668f87d1b1b1d50d68a2c78699c47579ce0e0ebbd92c96

Observation 0953712d-9267-4135-882f-c3bfcbf75c5f · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.642960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.642960Z digest=sha256:fb29ed9eff91eab8a1392dd3329976843e82c17ecb297f426d5d948e4a0a5e1d

Observation 4a3bf52f-d29f-42f7-a82d-f45e792260a3 · outbound

This paper cites an unresolved cited work.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T21:16:50.739395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:16:50.651526Z digest=sha256:768bbea021bccf20ce96bfb6c35b55fafd029f341435aa8d79be17f34ddcb39e

Observation cc1f00d7-649d-4957-920c-df764718e823 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.647280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.647280Z digest=sha256:fb0184b641cf0b53a6d5a3f00f7b2acc59e95d0a0a84457d0c31f60a17980b4b

Observation 8fb6f571-f63e-4466-985d-9a52c58cb6d2 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.649376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.649376Z digest=sha256:ca589589bf2adc933c948b97e6604eeab404ad949b61738d020336bb1b3a47e9

Observation 8779decc-64d5-427b-a144-f8ce6795cb08 · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.629382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.629382Z digest=sha256:5538d6bb614ebc6e29ce3afe187ffd641a1c9c631674e8e4880b2d81db1d16a7

Observation 845a691d-df31-4f89-9d08-7449379b87df · outbound

This paper cites Qwen Technical Report.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Qwen Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.623102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.623102Z digest=sha256:f64d436eb7a49224b6678cd6ee2c806a6292332849c2461baa221c3ecfaef3af

Observation 75752dd1-39ad-493c-888b-7bad0658224e · outbound

This paper cites DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.645267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.645267Z digest=sha256:79fd188b9ba3842a4433e57f8b578dbbbd872634aab3d29a1149163df6e7de2c

Observation f9c6fa5d-81b5-4426-873a-e0eae9c2790e · outbound

This paper cites PointVLA: Injecting the 3D World into Vision-Language-Action Models.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.635037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.635037Z digest=sha256:b0c42a444c5ea68e873d8ea5c07aa454551646f26842fbb5e30bd2dde9ab34e1

Observation 48171647-2b2b-4126-b3c1-14667dfd248c · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.626399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.626399Z digest=sha256:f872cb0588b73db52588f9952869337b66ad96a7f8e66c18dac70864c177f45e

Pith citing papers

Observation 62c69a0d-4cd7-467b-b9d5-3b8d2080edb3 · inbound

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:43:24.471673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:43:24.417901Z digest=sha256:86ae3848c07d9ff7d76d92eae7f6922a9e37c845ca068106dd1e0b5a8ce1a847

Observation a014a169-5699-4e65-b9f4-7d9f9f2c16aa · inbound

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization cites this paper.

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:20:02.012837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T06:20:01.885711Z digest=sha256:f63c541b4d0ac051ffae140c98fd80f14e3b14a4a6dfac25f5d3ecd833afae9d

Observation 1e1e29c0-0145-4860-8940-ad6e07cf9bed · inbound

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization cites this paper.

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:55.445345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:38:55.445345Z digest=sha256:7cc20eae10c9c6deb742dbfb98e591c05fe44b10cc8389a1e2a3da0535f674e9

Observation 8d85dd0d-8aee-4bdf-875b-303f0b8d2231 · inbound

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models cites this paper.

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:44.272979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:34:44.272979Z digest=sha256:157ca6cb0865bc33272a5d49636206180919ab7ed3f013b3ece6e6b9f2f0f9b2

Observation 71d97c6e-4e7d-4b10-b9ff-f83eedfed65a · inbound

A Pragmatic VLA Foundation Model cites this paper.

A Pragmatic VLA Foundation Model GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:16:59.751177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T21:16:59.712007Z digest=sha256:f5e01f9fac0d2abf94feb0dde528a9465fa05239d275d6db3f820ee4360784e0

Observation 83ff0ae9-ccf5-42c4-9dee-05e79c17ddda · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:00:48.382348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:21:45.365156Z digest=sha256:eff453f07b9751337507adcfd76bae3f197503c4d03656968aeb65555f940c68

Observation 7d9bab16-e82e-41f8-93dc-9dcfc1d2a85d · inbound

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment cites this paper.

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:55:52.326722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:27:12.286456Z digest=sha256:20e93ddfe2889f28ec7e7c122e7361aae1933f378a8f7f4b3458873d12edbefe

Observation 48bd2dfc-665b-498f-8670-62cde3646a6f · inbound

Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models cites this paper.

Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:01.810013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:20:47.418874Z digest=sha256:cf0afe0e78d282f519dd2169dc5df0b1e2d62234ad3dc2eb0b97f7f50500ec24

Observation 3c4e72ad-2694-47f3-8c08-22e5abc0a37c · inbound

R3D: Revisiting 3D Policy Learning cites this paper.

R3D: Revisiting 3D Policy Learning GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:00:04.200200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T10:57:54.273960Z digest=sha256:214ae62e542a83d91a16ac06443132b24e68492703da7728909c33f295cb7da6

Observation e8d9e6ef-2437-443e-abcf-a7a7f98caced · inbound

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance cites this paper.

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:29:47.564969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:04:16.073230Z digest=sha256:e18b613769fa03e9931e2b600140c436e540bc38b54dec2725c1ed8b3827972b

Observation 19e63f52-6734-438b-a66a-4cbef4048519 · inbound

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising cites this paper.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.281607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T10:40:04.767657Z digest=sha256:37c954b381daad860a259c6d29debfd44d14ae98efdd8a698925ab52dd1976c4

Observation d57c6f33-91f7-46fa-97f4-0e8527a07f2e · inbound

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising cites this paper.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:12.310771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:20:04.435904Z digest=sha256:980d398dfe371aec612b3fe452111fc90ea612dd777d9cc450fcae5422109fb9

Observation fa371804-0d82-4488-a3c2-42452d784ade · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.398109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:54912f97e257b52114b06743e0826833887acb639319fa763ec6b0e4a53a1756

Observation 27f97e89-9f07-455e-b991-4f3b9dbc9460 · inbound

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models cites this paper.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.179579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ce38a12f21376ead954efd4b58bd73aa1d9def4cc80f9495511a8342f736eaec

Observation 4fcfcc17-bca0-4983-90e1-8e69f46bb3db · inbound

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation cites this paper.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.495377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:b05412647cc519a7a433f1cfdbd3cdb8c7a336eabc84e3a987d91f9c99c4a546

Observation 230cea23-6c8d-4fe1-8421-db1277ed9a20 · inbound

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction cites this paper.

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:52:17.039700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T04:50:07.792757Z digest=sha256:e1014ee7583afb6f12f2a84d22cfd76602a17a584d3476d88c156e605537d14b

Observation af347835-7fd8-4fa8-b65a-10715cf232bb · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:57:12.743980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T03:57:10.338617Z digest=sha256:1ef01b0146c875f96391e0eb7ff47e93bba3496e2655f5a8479a069db0a3575f

Observation 02fd77d4-1ee8-499e-9500-af46074635e6 · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.741680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:0d113bb2e57ee7f86a4d2b6245dc9ccea700317398a7d1d0a3b1f57fd0198069

Observation ff736a69-2844-4053-9402-17fddd5d9379 · inbound

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction cites this paper.

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-21T03:39:29.563473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T03:36:23.117616Z digest=sha256:111527a5fd0381f676a97774ca6404107597b244eade8119744282c356bd8b89

Observation cdff4942-d152-4ee1-bdc1-672ff08d89e0 · inbound

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation cites this paper.

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:13:15.294796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T06:57:41.245418Z digest=sha256:d4eaafefa60ef591329ae170cbd0f853a59a21306eaafc8ea565a8ae2671511c

Observation 774ff56d-cb86-4531-ac0a-e37f7093b0b8 · inbound

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning cites this paper.

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:20.364540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T14:41:50.084254Z digest=sha256:68930da329c1cef635ef33761fd557bdec043cac6053e5dcfe4889ed38b40bf7

Observation f163582c-b0d7-43c9-905a-83044c62c510 · inbound

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models cites this paper.

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.193553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T10:04:10.627420Z digest=sha256:a4ef935943ec80a5b30f94e0e4246754f3a2a34fb43c61a54dffa7c89b6d7849

Observation 3fcff667-0757-4581-994a-7d30ee88679a · inbound

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training cites this paper.

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:36:44.079773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:20:40.998337Z digest=sha256:1a542257002e19bf379a87b0416063068c0b1a63e73c03d02e849bf38b54c81a

Observation 3921b61e-cb8c-4e63-85b0-ca7ec9126dd7 · inbound

Robots Need More than VLA and World Models cites this paper.

Robots Need More than VLA and World Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:46:59.303478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T01:01:33.530167Z digest=sha256:9448a5848446cbaa05d3b0821e59adbdca8dcc124e404b5139a9a0d145f4a377

Observation 8e4205c0-1ddb-4ee1-8255-a97cdf6b1ef5 · inbound

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model cites this paper.

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.707491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:24:59.678266Z digest=sha256:ea1b3eaddc4625fe27a5422388f2ae81308f61dcee2bd02e573a3750f004f659

Observation 9143f34f-a725-4658-85f3-38421e58cb5e · inbound

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models cites this paper.

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.294172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:13:19.238535Z digest=sha256:2a5933c0e5e6b489da36b27430f4b894b6c9c7a3f2d6d9622d91dfbc8f071099

Observation deb79e72-96fd-4fb6-a5db-a938de7c19f2 · inbound

VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models cites this paper.

VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:17:41.839593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T12:47:35.314486Z digest=sha256:f7ee518ba275cda1bafd15ea03dbd5d7820e960c8deed51768686a85b1b842a5

Observation a0f12d5f-9ade-463e-ac63-4ee602e7259a · inbound

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation cites this paper.

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:33.009592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:37:21.516021Z digest=sha256:f31c33195ba480e03d3d8839d1e58e1075e68e840d9f56c6bfeaa6a4d36ff626

Observation 51bb3c66-f216-4640-a45b-a17b4c375d58 · inbound

Geometric Action Model for Robot Policy Learning cites this paper.

Geometric Action Model for Robot Policy Learning GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:43.942419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:59:17.983035Z digest=sha256:ae4dc2048434e4dde307d2fedcb6be00ff2c285b7a907cc71389e574f2e624a5

Observation d36d26c3-0f0d-41e6-8210-ae9cac506e7e · inbound

Learning 4D Geometric Priors for Inference-Efficient World Action Models cites this paper.

Learning 4D Geometric Priors for Inference-Efficient World Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-11T14:23:57.266710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T14:23:57.266710Z digest=sha256:29320b9f8402fc6d53c461b5ce43150b4479d3df85bd8f565a2fc01334292b58

Observation f7dcc86c-a1b0-4b2b-a6fb-29884bb9807e · inbound

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models cites this paper.

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T05:11:57.092685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:11:57.092685Z digest=sha256:1cb3a2fb6ef4018ec42d73a7290d114563009837b891933f400599b2c9bcaa3c

Observation 856e1487-93bc-4916-9ee1-2bf9a0884165 · inbound

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation cites this paper.

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:26.629789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:26.629789Z digest=sha256:3b6e8b3c57eab7cf145af0acbc96bfb30ec696515c893fe389d5da42b2126752

Observation 049d2ea7-780b-4e41-95b4-8a945fbce450 · inbound

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models cites this paper.

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T01:10:40.118599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:10:40.118599Z digest=sha256:7be12f106aecbde949e98a5a2a00222027d0cf45ee3ea4bebdb803a1c472ce5a

Observation 7dd3cd96-9da4-4c72-b559-74a2723c3c96 · inbound

VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification cites this paper.

VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:35.555813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:21:35.555813Z digest=sha256:d04c1d93250fea8d6bbf0a5f98e9f54389b924d1ddb4fe6d4ddc5445f4d59728

Observation db459a88-3610-46fd-9aaa-6187dbaf2234 · inbound

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models cites this paper.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.803171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.803171Z digest=sha256:16a56c8a423d4dcab8cb7340dd98b8336401d5ede39d2709623f70b445abefab

Observation b7df52d0-2bea-422c-a38b-49f06c97be77 · inbound

DreamWAM: Beyond RGB Future Prediction for World Action Models cites this paper.

DreamWAM: Beyond RGB Future Prediction for World Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:27.996941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:59:27.996941Z digest=sha256:bce44dfa5d62f414dc87f3583b9e14824ce4b6c0da30e835789c24eefb6a2511