Pith. sign in

Paper Citation Record · LEDGER

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2608.04633.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04633 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:19:50.356389Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26f9b6b8-c9c3-4892-ab74-7105435e5b9f · outbound

This paper cites Zitkovich, T.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Zitkovich, T

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.396151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.396151Z digest=sha256:745c6c4eb0081e7926061e64f52ef673347c53d24b6976fdd930227dea278be3

Observation 3f472e05-b48c-4833-8843-c2850544695d · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.446123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.446123Z digest=sha256:23c17781fdc4e3843438c74cef3bc3c0e9ae678d76c3ce61a0815635385be4c0

Observation b922b83f-7d15-4061-8911-9e7cd6d05b4e · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.520992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.520992Z digest=sha256:74f9086e97ca3af492b4cf32dec366cd4afb0b7a3220fa49ae8e7b3526b76d97

Observation e68b6461-d7a1-4051-8dcd-6f2b3aea963b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.577353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.577353Z digest=sha256:5f2a62fee81203ed3d92e23b4899ec71e55538a06f50b2dcd24c04a2edb06869

Observation 5194fd60-de51-48a8-aec8-e97196c28576 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.863591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:46.671171Z digest=sha256:9bdc044925af5be3ec7df648b97cedb3e4f7697efeaa4d7fec35945301b56f41

Observation db459a88-3610-46fd-9aaa-6187dbaf2234 · outbound

This paper cites GeoVLA: Empowering 3D Representations in Vision-Language-Action Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.803171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.803171Z digest=sha256:c9a75c346843208fac003d153e10442578ff5cc17dd6da6af1754cb71671c4ed

Observation 9cd5a458-dec3-4ee2-b750-754efc5b8c04 · outbound

This paper cites Singh, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Singh, A

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.923578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.923578Z digest=sha256:ed046317403bf89d1eb851bdeeaf541fdfafcd898504f69b84960fd56acf443b

Observation 92770e72-495c-422f-923d-8c5a48bad184 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.003040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.003040Z digest=sha256:eb65a6bb6278f780b94cc6ba5d6aabbcf1f903b1657853835741782d9f1f274e

Observation 13e9efc7-a1bd-46c6-8d49-60b75febcc95 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.078147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.078147Z digest=sha256:7ec81c3894bf35cbf119e5ec150324bd53f0af1dd3bc9bf845db1fe7782c712f

Observation ab78dc4e-908e-443f-8b81-016d8a0891fd · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.636357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:47.155194Z digest=sha256:53cfbd5aae80b9e7587b653581a6d502dda99e678db84b522b3638612c24992c

Observation 1e5c079e-cc74-487e-944a-c312cd13ca68 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.208817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.208817Z digest=sha256:76cff7caa8bc68a2dfcd2be8c3c85c751edf75bc21059103e8f98bfb2c601fe9

Observation 1f34d831-2310-4095-a423-78be06bfecb4 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.275991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.275991Z digest=sha256:0ddcb0334b9a199fed8b8811eb03c9460fea30953487ce561b5928577f965e89

Observation 677ce944-fbae-4bbe-babe-38852f0f9f5e · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.408288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:47.359540Z digest=sha256:c24e5830ee38628e371a4d465dead965358ea55fd1263141a81bbe0f8911d5ea

Observation 6238bce0-fabf-42d2-a31d-6f4a6730eef1 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.403646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.403646Z digest=sha256:250f1e57e55c8337860adc35ac49d49981a2398baf9e5acbfdb0c662c34d6d7a

Observation def2df4a-59f4-4fa3-9f90-ec4a4a19695f · outbound

This paper cites O’Neill, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models O’Neill, A

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:19:52.215031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:47.510694Z digest=sha256:1d222f23d6aaf77931252555dcf44a67fbb3d2e4d72bcba9813ef122a37ed3ae

Observation 567b9100-3b8f-489b-9b99-5978f5f252dd · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.572949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.572949Z digest=sha256:ad49365b05f6b5a20c41eb73bb0256df1b310567635f866c26cdcb8889888fef

Observation a909a4b9-f254-43ae-83a7-e22beb324557 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.653884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.653884Z digest=sha256:d8cffedc2859fb98af030dc97945353e62160d6d4f184581141cc9aae661021e

Observation 29970e42-475e-44e3-9b81-74697dd96599 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.739039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.739039Z digest=sha256:e5b0f8469db3d83a6e170bb2e5874270056b690ea665263f6ba89b5258448be2

Observation 22e3fb0a-6958-4a39-a54f-fa6f9591b62a · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.832939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.832939Z digest=sha256:b0cc1ae716973222524e5742b67574e849fd672d33bab4b478282aa2a1d5ca33

Observation d55763d2-e79f-4c5a-aa60-919c89b1877a · outbound

This paper cites Shridhar, L.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Shridhar, L

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.896250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.896250Z digest=sha256:7c28c5f20b06938beccfb61387f1a24ed2f553ae45c9c54fbbe3cd16c289fcb7

Observation 5947d0b1-08dc-443b-a0ca-54017e52ee19 · outbound

This paper cites Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.935248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.935248Z digest=sha256:293c6252e57697fd0f10ccc5b547ded33f66d8dd135b164b183ef8dfd090d388

Observation 43fa7b85-4d46-44ac-9aa4-378678dda7dc · outbound

This paper cites Goyal, J.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Goyal, J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.011334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.011334Z digest=sha256:abd2631b6af069b48ef1c5b0085a39d970c808eedaacd1904fddace2d9ebe28f

Observation 3da0a491-2e47-4f1f-95f8-48f090164ce9 · outbound

This paper cites RVT-2: Learning Precise Manipulation from Few Demonstrations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models RVT-2: Learning Precise Manipulation from Few Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.095191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.095191Z digest=sha256:89d5d84edf23486ac8fbadc0531fda40b33e869abcff628993693e76fd12b9b5

Observation e1e07334-0fca-48ef-b91e-2331ec311f7d · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.187917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.187917Z digest=sha256:040aec35587a1c9ab115a794fbadcebacbcd47ddb98c843bba7c60887ba074bf

Observation b1fd0a23-d6f6-41c7-a9e3-d7f7e487a763 · outbound

This paper cites Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.268018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.268018Z digest=sha256:5411412aa2c1f423aed0a6f5bd10363fdaced4f4c870ec6936234fd700597bce

Observation 752060ec-438d-4d3c-ad9e-06c242fde276 · outbound

This paper cites StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.331513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.331513Z digest=sha256:3aac75971ef18ab09248053b328fef43ceb7f6f5082ff48b7e72bd893322f3a5

Observation c046748c-8091-4aee-bbdd-93269d49b7a9 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.983253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:48.393864Z digest=sha256:e6be2063055a78bf3ab98e781d8ae76b6ce29ce0a80f468abb550ce3c7587d9d

Observation cc182fc0-0a71-483b-a2dc-10046dc23220 · outbound

This paper cites QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.478639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.478639Z digest=sha256:92c59c27f393265c9c427e8c0ebab1b43b9e3fbb17828517c01774061a2804dc

Observation 25f249e9-f40e-49e6-b0ef-f99092c85327 · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.534851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.534851Z digest=sha256:03d39af2b54e65cd17da4c49a7495fcb8887615b9eec63be490ad4b5519efdec

Observation 767aec74-c80d-4bca-8b52-692215445dc4 · outbound

This paper cites PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.590609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.590609Z digest=sha256:b9dc017894d1b4d91786187e8335eba0241191f87e3b2ec433b2cc31a0fb0c6a

Observation a8669714-812a-40b7-8883-a0c45c288110 · outbound

This paper cites Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.667454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.667454Z digest=sha256:f4a65880cf1373cf2d08be4596b6e408a46ffc1f82b23f1b7d5b3d769c608d1d

Observation cd9d14f8-9612-487f-8ec8-2cde133f0f58 · outbound

This paper cites Peebles and S.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Peebles and S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.767277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.767277Z digest=sha256:169997375957f2b3fef8d3921ba11d5c9287e7849d9e80ed6717d6d2bb0ae04b

Observation 06563ac3-47a8-4429-baa6-b30e995206c1 · outbound

This paper cites Radford, J.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Radford, J

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.843593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.843593Z digest=sha256:d1ef54866ca0add8c67ec5676264258891a21d4513f08e352a9e906dc02a373b

Observation d672b3e2-c805-4b58-81a8-5925afb5e535 · outbound

This paper cites Denoising Diffusion Implicit Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Denoising Diffusion Implicit Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.975765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.975765Z digest=sha256:0a823c17c8a938f7c5c865124a7900e8add25ec270532a1a63fd1c2a7347d6d5

Observation cc9f1646-a4c2-418a-b34c-61d6444f36c1 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.097273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.097273Z digest=sha256:baf528ab5891367404743a1938f27d1d59b4bfc4836c4c4512f96f0f5d68b599

Observation 23102af3-4a0e-486f-b583-ba540c05c04d · outbound

This paper cites Rombach, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Rombach, A

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:19:51.738404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:49.191361Z digest=sha256:bd4363255264abac9a48f435ddb0df80d4cba6bc78a2f42e16601045140fa417

Observation 1c149296-e834-440a-89e7-1cfb36079b83 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.316462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.316462Z digest=sha256:c1e368d404061ce61767e94939fe7f8b7b169074089bcc30c83bbdbf086437be

Observation 1a4a7068-291a-4154-a9ff-248e47bae85a · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.424569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.424569Z digest=sha256:a415126989a3a41598adac95dfd2c0f02232560b0051009467397252cbf0f7e0

Observation 37bcff39-8ec9-46f8-838f-803038cf6f9b · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.501211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:49.557016Z digest=sha256:5a770cec9ee244bd8ac0e968c79ff2cbb061b0dc056f3e18b8fd3a512324d531

Observation d46dfc98-2a71-4997-be37-2ba63028c1af · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.656649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.656649Z digest=sha256:ecba48f6a5103e1fae2b6ab828543248d2e55015c2a5ed1a60c781d8a7cf7241

Observation ef3bd510-365b-4b64-a3ab-bce873a810f5 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.758226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.758226Z digest=sha256:0a8d4efeb3fa749d1086853d121ca4424e7a0c77dc94bb584384ae66c2ae51b4

Observation c004a486-b24f-41c0-8c5f-aa26d8d09415 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.893880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.893880Z digest=sha256:76ecbc72fdc35fa3db95c7c6edbb6c161cd02f45140742f24db8d8790841edc9

Observation 58340e5a-1563-48b3-8529-a965af4e892c · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.283184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:19:49.988923Z digest=sha256:eea313cc2b16c1fdff71eacdbb90bfa4f102a784bb1bd66cacd71e4f241aacc7

Observation 479ab36c-5a63-4a51-90f5-cdbe56196a57 · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.092976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.092976Z digest=sha256:da65886012f611067a9779416294df0358f118515ba0c4fef8a738185a17fa8d

Observation 6e930a4c-d7e0-4273-b5c3-18d9d8fce342 · outbound

This paper cites Kirillov, E.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Kirillov, E

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.211537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.211537Z digest=sha256:195b8b00b9d4f7f2f3e7bb11fa2b0b60bc9b9c73617a97cd17471bba056620a9

Observation c57b9e74-d5e7-4767-8b1f-30d0aa36d0b2 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.356389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.356389Z digest=sha256:a9fff7303f6f0d7153dfe958d02edcff5be562e10ad8a4616bdc20402cbb7486

Pith citing papers

No inbound Pith citation observations are available.