Pith. sign in

Paper Citation Record · LEDGER

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2608.04633.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04633 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:19:50.356389Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26f9b6b8-c9c3-4892-ab74-7105435e5b9f · outbound

This paper cites Zitkovich, T.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Zitkovich, T

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.396151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.396151Z digest=sha256:971b6b5847ec746a3ec6e3f32db10f6ef84f70c51e1389623839c34a2595618c

Observation 3f472e05-b48c-4833-8843-c2850544695d · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.446123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.446123Z digest=sha256:5f10c1233c07397b07a7d7af6f462be0e7bcd1ab79fc41fefbeba6d6ee3b6aa2

Observation b922b83f-7d15-4061-8911-9e7cd6d05b4e · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.520992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.520992Z digest=sha256:e4c74173d52dd8d7d0e601e9c8c36471be6d39ae636a55989e31b428fbd4c4cb

Observation e68b6461-d7a1-4051-8dcd-6f2b3aea963b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.577353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.577353Z digest=sha256:d19658da2abfcb26eb74c419ee98dea1723b344dbbd9a00d89e14c2c0f9ff1ac

Observation 5194fd60-de51-48a8-aec8-e97196c28576 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.863591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:46.671171Z digest=sha256:5014f44899fcff9b7ef695fc852441c2f8a1a2cd96f710f3ad4f390ccf14fe98

Observation db459a88-3610-46fd-9aaa-6187dbaf2234 · outbound

This paper cites GeoVLA: Empowering 3D Representations in Vision-Language-Action Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.803171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.803171Z digest=sha256:bd25bcae88be309445bdc5d8bcfd8dc8adcdd5b28db942140766c847e6cfda8b

Observation 9cd5a458-dec3-4ee2-b750-754efc5b8c04 · outbound

This paper cites Singh, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Singh, A

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.923578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.923578Z digest=sha256:22283284a34b9346c6a669760e0d8a3c334de6e98cbfa06132075d4b39d53d1f

Observation 92770e72-495c-422f-923d-8c5a48bad184 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.003040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.003040Z digest=sha256:f4874d7b10da204847b9b7db7cd7c2c8ea85f126a15f64941b010eb3bcd92914

Observation 13e9efc7-a1bd-46c6-8d49-60b75febcc95 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.078147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.078147Z digest=sha256:c0428be1495710ad8a859938f7013f615309967aef4fd2b5c91b3ecb23b81926

Observation ab78dc4e-908e-443f-8b81-016d8a0891fd · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.636357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:47.155194Z digest=sha256:fe686d22edbe5852b57aebcaa0afdfb0047dcfefba9e50233a7e85250ea70f4a

Observation 1e5c079e-cc74-487e-944a-c312cd13ca68 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.208817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.208817Z digest=sha256:3d2d4718ea7c2be4d070964d8406d736bcb1f2493785b9b2268e2b26ae8b5722

Observation 1f34d831-2310-4095-a423-78be06bfecb4 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.275991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.275991Z digest=sha256:0a923cae7161a78def55d2f382fe03bd17073ed3691fd21424b35026521d949b

Observation 677ce944-fbae-4bbe-babe-38852f0f9f5e · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.408288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:47.359540Z digest=sha256:be53324796917266b61ea26699550bf523d319bbb3a1be3cb77d6517e95f9a6a

Observation 6238bce0-fabf-42d2-a31d-6f4a6730eef1 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.403646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.403646Z digest=sha256:21be5c333f1ac1e75ed6fe3de94162f19403286200c0e9e0b5e8f81bf5c46b6d

Observation def2df4a-59f4-4fa3-9f90-ec4a4a19695f · outbound

This paper cites O’Neill, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models O’Neill, A

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:19:52.215031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:47.510694Z digest=sha256:ceb25bf15ffe51fb5967214d2685646b70832bba2eba536649406cd628d2d15e

Observation 567b9100-3b8f-489b-9b99-5978f5f252dd · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.572949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.572949Z digest=sha256:30331ded21ee635c519bf4dfd9aff88c4e7702cd7dfe8cb745dd185e33c4a6c0

Observation a909a4b9-f254-43ae-83a7-e22beb324557 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.653884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.653884Z digest=sha256:c231b130563951ad11acb74d107293d289486a85a6d7ea299f80c758bc6f9b12

Observation 29970e42-475e-44e3-9b81-74697dd96599 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.739039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.739039Z digest=sha256:c2aa93841e1b3e8db9e5aa0f5d3c44a8697d0704906e1bf3b7dd31b407faf577

Observation 22e3fb0a-6958-4a39-a54f-fa6f9591b62a · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.832939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.832939Z digest=sha256:4671c8326a937fdbde0b405a7f145bb97872368cef4bdaa11396e35ac1867d84

Observation d55763d2-e79f-4c5a-aa60-919c89b1877a · outbound

This paper cites Shridhar, L.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Shridhar, L

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.896250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.896250Z digest=sha256:cfb7654cf427ff9ee2b5e2d2d34374c5aa6c4f8db594986ab226dcd188ca5e7a

Observation 5947d0b1-08dc-443b-a0ca-54017e52ee19 · outbound

This paper cites Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.935248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.935248Z digest=sha256:41168e080a9d55dca957964c4bbf2c5552a35b20e2ad52c9ef892595f0b36b34

Observation 43fa7b85-4d46-44ac-9aa4-378678dda7dc · outbound

This paper cites Goyal, J.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Goyal, J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.011334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.011334Z digest=sha256:ea59a4ffd6168acbe4a14e47c15f7de6fd05e4587a5a632ac8558f6ca82bce86

Observation 3da0a491-2e47-4f1f-95f8-48f090164ce9 · outbound

This paper cites RVT-2: Learning Precise Manipulation from Few Demonstrations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models RVT-2: Learning Precise Manipulation from Few Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.095191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.095191Z digest=sha256:6831cb5e828a45bf54d633d6e41ce8b951d101c77e3578f3a0ce6b0704d71715

Observation e1e07334-0fca-48ef-b91e-2331ec311f7d · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.187917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.187917Z digest=sha256:3efa7885d95e35cfbddc4da165b74563e56babdc6572f24bcb6566a01def4ee6

Observation b1fd0a23-d6f6-41c7-a9e3-d7f7e487a763 · outbound

This paper cites Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.268018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.268018Z digest=sha256:51530ff10cff784683084f945740ac44c4cdfc8480cb01e9e5ff1c2a2b4f4afb

Observation 752060ec-438d-4d3c-ad9e-06c242fde276 · outbound

This paper cites StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.331513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.331513Z digest=sha256:1b8cfe448520883dd47f34602c886ac7eae836921d37e3bbd2ccf42c7a234df6

Observation c046748c-8091-4aee-bbdd-93269d49b7a9 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.983253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:48.393864Z digest=sha256:4ef74ba0ba0a7a15f7c38ceea0e43d13a383e070adff7ecd8cc3da6632489259

Observation cc182fc0-0a71-483b-a2dc-10046dc23220 · outbound

This paper cites QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.478639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.478639Z digest=sha256:50f23ffac04645966c62e7f9e240d10cc9f14dbfaf961b0dc0db10bc968b6666

Observation 25f249e9-f40e-49e6-b0ef-f99092c85327 · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.534851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.534851Z digest=sha256:241253c276329aa82276fa2bfce9ffa702dfe9e8e956c84cfa92d8638448c88b

Observation 767aec74-c80d-4bca-8b52-692215445dc4 · outbound

This paper cites PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.590609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.590609Z digest=sha256:ab70a7f0f105646b750e9b8cf0026ebfa7e134c060feb554383eeb8ee099f188

Observation a8669714-812a-40b7-8883-a0c45c288110 · outbound

This paper cites Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.667454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.667454Z digest=sha256:bb382585de077a32b832ec647b056b98541ec9503efff8e5ec099bf457fa12ba

Observation cd9d14f8-9612-487f-8ec8-2cde133f0f58 · outbound

This paper cites Peebles and S.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Peebles and S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.767277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.767277Z digest=sha256:cd21f4a89bf3e29002568a197ca590eace8386154ec996ea4d7497a5fd6875c6

Observation 06563ac3-47a8-4429-baa6-b30e995206c1 · outbound

This paper cites Radford, J.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Radford, J

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.843593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.843593Z digest=sha256:b7845ea23de6e105ce79da564f4e56766829afac89c8f78171eaa06e95530feb

Observation d672b3e2-c805-4b58-81a8-5925afb5e535 · outbound

This paper cites Denoising Diffusion Implicit Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Denoising Diffusion Implicit Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.975765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.975765Z digest=sha256:0b2ca36d462c7e77dc4d1a2a889822f56b09338f656fe72eddbe9b6921a231bc

Observation cc9f1646-a4c2-418a-b34c-61d6444f36c1 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.097273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.097273Z digest=sha256:b5eaf6bbd257c53ba9c28c1855732f31626626e9809740e9732dd8441d9906e0

Observation 23102af3-4a0e-486f-b583-ba540c05c04d · outbound

This paper cites Rombach, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Rombach, A

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:19:51.738404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:49.191361Z digest=sha256:4fbae9cdeef333f6b4c3b188571db6fd847ea12f5bbc16e426efd0e53cd0da96

Observation 1c149296-e834-440a-89e7-1cfb36079b83 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.316462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.316462Z digest=sha256:52db9e1ad1e82d1cbdde3e71c5c4f8cdfacdd70610da2c499b8c4ec48ffade71

Observation 1a4a7068-291a-4154-a9ff-248e47bae85a · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.424569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.424569Z digest=sha256:8faf4404e83c31cfc2dc3163954ad4725c1e43db3639413b2300021a415ca1d5

Observation 37bcff39-8ec9-46f8-838f-803038cf6f9b · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.501211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:49.557016Z digest=sha256:197abce3365c8081f4356e2de07b3a182123bb95366398d8af21b21a849198a5

Observation d46dfc98-2a71-4997-be37-2ba63028c1af · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.656649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.656649Z digest=sha256:fdb7089044da0f5d7a17c6e3a8637e3824498853b3a3c27961e8ac9f4f6687cf

Observation ef3bd510-365b-4b64-a3ab-bce873a810f5 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.758226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.758226Z digest=sha256:8c50d686bf265c5aaafcc78b6274668ac3220fbcbe138b231da50d11a8dfb67e

Observation c004a486-b24f-41c0-8c5f-aa26d8d09415 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.893880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.893880Z digest=sha256:06ff38621bc634da66fa49fe650adad144463d34102ad9ccd906dd7dca25d70f

Observation 58340e5a-1563-48b3-8529-a965af4e892c · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.283184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:49.988923Z digest=sha256:3a72473ad88de06a1f5e94478074b1255eae0cf8abfd93bbba26a53cdd7a90d1

Observation 479ab36c-5a63-4a51-90f5-cdbe56196a57 · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.092976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.092976Z digest=sha256:64e22f5e5f85ef14783feee2beb537ca2ed16bc3bff43a9b8bb4d57efaea1a77

Observation 6e930a4c-d7e0-4273-b5c3-18d9d8fce342 · outbound

This paper cites Kirillov, E.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Kirillov, E

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.211537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.211537Z digest=sha256:22578ff66be33e53abe68f472d8dcc46fc8f501f8a741a36aacfe95d17e50028

Observation c57b9e74-d5e7-4767-8b1f-30d0aa36d0b2 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.356389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.356389Z digest=sha256:b97f386dd13e310dc21e8423913d0bd34e28a4efc9fe1ea7aeb1a5f6a4fd6b93

Pith citing papers

No inbound Pith citation observations are available.