Pith. sign in

Paper Citation Record · LEDGER

How Should Vision-Language-Action Models Use Proprioceptive State?

As of 9 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2608.03052.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03052 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T01:01:44.151987Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afc253b5-1f6b-43c1-bb74-6511dc44c14d · outbound

This paper cites HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation.

How Should Vision-Language-Action Models Use Proprioceptive State? HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.071866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.071866Z digest=sha256:a5cbe13d320aff1cfeb107fba318951721f4020b572ff3c33d0461bcc0b7b982

Observation 8a3b5760-2bc7-46e0-a334-855dcc767c0e · outbound

This paper cites RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies.

How Should Vision-Language-Action Models Use Proprioceptive State? RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.084075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.084075Z digest=sha256:2515a20fd4a365647e04145f7695fe9926078264a3a3a571a4f959a865f925f6

Observation cc10f7a7-1860-46b7-8577-790329b85c56 · outbound

This paper cites Davies, Y.

How Should Vision-Language-Action Models Use Proprioceptive State? Davies, Y

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-08T01:01:44.989531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T01:01:44.087910Z digest=sha256:7b34a2ccbc637aafafa0bd51f75b4216984cba6f1431f486a26dfc8172c71ad8

Observation af6f9f4e-56ad-43f8-817a-cb06c02ca0cd · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

How Should Vision-Language-Action Models Use Proprioceptive State? $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.094721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.094721Z digest=sha256:c38b5e9bf14a50711d1fde10414300ff91f19829879c9e6e992f678d5fe5249b

Observation 203c76f1-4b45-4a9f-ac77-aeb8ac7954d1 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

How Should Vision-Language-Action Models Use Proprioceptive State? OpenVLA: An Open-Source Vision-Language-Action Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.098166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.098166Z digest=sha256:7736fbef100c9bce81ddd8509e9a8f198730a6e602f95fc7799700ddae1d0801

Observation c83c44b3-fcb4-4aba-8afe-dc0c84019d49 · outbound

This paper cites 2602.12032.

How Should Vision-Language-Action Models Use Proprioceptive State? 2602.12032

Reference 11

Resolution
verified exact
doi, observed 2026-08-08T01:01:44.528761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T01:01:44.107978Z digest=sha256:99046990bac87cdecf10efb031766e0cce638bf37410f79b90366fd9b95c5c4a

Observation 0d5b8cb8-3eeb-498b-9b0a-d553c6637214 · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:01:45.035967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T01:01:44.111421Z digest=sha256:f8676d8a6e8caeebfb85fc6955876f08685c660cd643f6bbcce194a6b156efa4

Observation d9ce61a8-bbd2-4c14-8313-63116db4ecd5 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

How Should Vision-Language-Action Models Use Proprioceptive State? MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.115129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.115129Z digest=sha256:0a2b210bf184410d0eeab63365e5abeabcb07774dbfceb087034020a4aa7a910

Observation e54940ea-f858-4887-9295-6bd8d2cf92cc · outbound

This paper cites Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies.

How Should Vision-Language-Action Models Use Proprioceptive State? Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.118690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.118690Z digest=sha256:538696daadd3bb716cfe5fd5ca71d43ab791bc1b3fcf7cdc2274c31f5956ca0c

Observation 8c5ad290-f9e0-4386-bc10-c1d1fc7b6ee2 · outbound

This paper cites Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies.

How Should Vision-Language-Action Models Use Proprioceptive State? Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.122151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.122151Z digest=sha256:095f216acf54d3f14160cc4f77533b8ce4565aaa898f3e3b66760fc28c0b877d

Observation c3e3f441-ea11-4b9d-870b-847031eba6c7 · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 16

Resolution
verified exact
doi, observed 2026-08-08T01:01:44.433068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T01:01:44.125411Z digest=sha256:35b8998add573a9fae3818b40801b10a40a8725b061bf30f1f582ae245ebce41

Observation 35f2b077-8ae6-4598-8b4e-67b33e60a9da · outbound

This paper cites Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models.

How Should Vision-Language-Action Models Use Proprioceptive State? Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-08T01:01:44.721229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T01:01:44.132304Z digest=sha256:27ee984801c3ae8043c6c25e43f00321cbe3e90ad855cbcc926887253ba92a7a

Observation 30f870f2-dca5-419c-9ab7-5c38ad0279b5 · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.135594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.135594Z digest=sha256:0d3bb04c0d6288602eeba9b65b3e56e4dd745b5cb9f62cda21b0263599b460be

Observation 463d4e7e-96de-4c03-b6ea-a74878bc71d5 · outbound

This paper cites 2601.14133.

How Should Vision-Language-Action Models Use Proprioceptive State? 2601.14133

Reference 20

Resolution
verified exact
doi, observed 2026-08-08T01:01:44.358840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T01:01:44.138844Z digest=sha256:722ef9d9494d2e2e61613dcc90a680c5d0d72e1c3563d530fea19d446b05223f

Observation 66820ca3-7f3c-4dce-9585-fba23ab4730d · outbound

This paper cites Zhang, J.

How Should Vision-Language-Action Models Use Proprioceptive State? Zhang, J

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.142251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.142251Z digest=sha256:edfeb277120e0683818b88495e9e3a1605d331e4da5f36ca00820791f167a333

Observation 6d897209-cf3b-4e60-a0cb-24de3d6d59ec · outbound

This paper cites URL https://doi.org/10.

How Should Vision-Language-Action Models Use Proprioceptive State? URL https://doi.org/10

Reference 22

Resolution
verified exact
doi, observed 2026-08-08T01:01:44.284230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T01:01:44.145383Z digest=sha256:ab86f2c2085af5eeea4ac3c22fa9f5ea9d6e12fc13fa30d75ebc9a08f21ee9a2

Observation 5d2a6ede-04cb-4506-b734-b0bf1f0fd6a4 · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.148650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.148650Z digest=sha256:6cef6943fe5e5ef020b44f9805f774ef8b93288d3372e83fd0789db9c0b94abe

Observation 85208fef-0207-490f-8633-3851c3c81abe · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.091348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.091348Z digest=sha256:45a874af8842f7564f585436f6d8f83c33479d52d931541fb513d49000fb2f02

Observation 87fdd124-616f-4096-9bdc-3ba1d7801856 · outbound

This paper cites ROSA: Harnessing Robot States for Vision-Language and Action Alignment.

How Should Vision-Language-Action Models Use Proprioceptive State? ROSA: Harnessing Robot States for Vision-Language and Action Alignment

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-08T01:01:44.735748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T01:01:44.128884Z digest=sha256:a86ab7c7ab438b5aad8d3f17e917259caaa1178260545484e3463550305844ca

Observation b93173d7-4de2-45eb-9873-d8c5239c2712 · outbound

This paper cites Whenwouldvision- proprioception policies fail in robotic manipulation? CoRR, abs/2602.12032,.

How Should Vision-Language-Action Models Use Proprioceptive State? Whenwouldvision- proprioception policies fail in robotic manipulation? CoRR, abs/2602.12032,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.104786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.104786Z digest=sha256:8215762147fc7918c85927c455c4bdd413d16b43be91280137849786e0ced64c

Observation c42ebbaa-e9d2-4ae7-919b-8806c2da7a10 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

How Should Vision-Language-Action Models Use Proprioceptive State? Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.101602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.101602Z digest=sha256:ee5511f4e5fe53b62be2094f6fe6785dc33b6f777c1120f3dd1e753ba4edc698

Observation 854015ee-14eb-4e0b-8f34-60d21ed41dfd · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

How Should Vision-Language-Action Models Use Proprioceptive State? $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.080292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.080292Z digest=sha256:7dc38fe6d8e3a054a365741770f028d58952e64c91a79dae422220255a6baab8

Observation e72c9459-5697-4a3c-a778-220795f9843f · outbound

This paper cites HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation.

How Should Vision-Language-Action Models Use Proprioceptive State? HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.076699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.076699Z digest=sha256:7d74bd1519440402a8f0fa035c6071622d7fb6674328b3c617e08bbc651008b7

Observation 808705d3-d049-4661-9671-42e1387f8e10 · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 4096

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:01:45.026043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T01:01:44.151987Z digest=sha256:d99b810601063c246f17607293d6c36c4f4361909211b78bc0263ca30cd38a1c

Pith citing papers

No inbound Pith citation observations are available.