Pith. sign in

Paper Citation Record · LEDGER

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

As of 10 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2602.06575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.06575 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:56:05.890768Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T01:01:44.122151Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:07:27.134549Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e79c0e1f-5ab5-4967-bf62-3c2d1efe2343 · outbound

This paper cites write newline.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:03.390527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:03.390527Z digest=sha256:8744f49ed9d82412864877521c11ffafa6b891420c4b501be0b07b9e7ea1c525

Observation ca558d04-ed66-414f-9d4e-36c09a6f5fe3 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:03.534148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:03.534148Z digest=sha256:d5a50e4b3f95d218c741e5432d1444f5e1ae88ac9da92fc19672b312d15f60ed

Observation 866479ce-3577-489f-b0ec-00297cb223a3 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:03.641804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:03.641804Z digest=sha256:56b32071a69130650e6ac2e827597875446f53fdaf641bb8e23bdb48d5d73f0b

Observation ce5024b7-fe2a-47fe-90ea-ebddaf07717a · outbound

This paper cites R., Finn, C., Kumar, A., and Levine, S.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies R., Finn, C., Kumar, A., and Levine, S

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:03.742124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:03.742124Z digest=sha256:c7fdd772f040152523d2cdaeccce19c20c9931ddd2327e3a9734f377588adeb2

Observation 5a02d83f-7c01-4b53-8c32-d1abd795b7eb · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Diffusion policy: Visuomotor policy learning via action diffusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:03.850627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:03.850627Z digest=sha256:eb7818f1f27cb65ab8266a46375d8b95284f913e6fe84329871e0c718a4b8f44

Observation df2cc7c7-9f52-487f-90cd-13cfab484e06 · outbound

This paper cites The Ingredients for Robotic Diffusion Transformers.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies The Ingredients for Robotic Diffusion Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:03.958826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:03.958826Z digest=sha256:8e59ea207f0096eae3d3ea1616476caba8f24f7a1efca6550de3e61eb8a4775f

Observation 2eacf72c-5def-46fe-b13f-20757e186e7d · outbound

This paper cites U., Akram, W., Saoud, L.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies U., Akram, W., Saoud, L

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.064798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.064798Z digest=sha256:1ca62bc93ec8fa827f13cd70f1a8ebdfb78ea7256fb47e2b58bd54c05117da69

Observation fea78ca6-3221-45b2-8b6f-f2d5b3c5c48a · outbound

This paper cites Latent Theory of Mind: A Decentralized Diffusion Architecture for Cooperative Manipulation.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Latent Theory of Mind: A Decentralized Diffusion Architecture for Cooperative Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.120185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.120185Z digest=sha256:80e444892d908f50e200f2716d87d5f024ed947667f24b55967be43e2cefb2b9

Observation 85b205a8-ea66-405b-94bf-339714a66c7e · outbound

This paper cites Demystifying Diffusion Policies: Action Memorization and Simple Lookup Table Alternatives.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Demystifying Diffusion Policies: Action Memorization and Simple Lookup Table Alternatives

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.228948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.228948Z digest=sha256:96be4e9f2776fc19831ae58b5f23583b6e18513edd16835f32ff99c9186262d3

Observation c4026ade-74dc-4276-8eaf-1bf3f6496f0c · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.338230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.338230Z digest=sha256:4ddd4f6e58370ad8bcb4b5aec4c7559e5a0e0a9aad04886f00e04444ef3d8bb5

Observation cfd835a7-5f16-4087-a785-e8ed73593688 · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Video prediction policy: A generalist robot policy with predictive visual representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.397799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.397799Z digest=sha256:dbcc5a648a0fb3b3aba2259114c1253fa60bdd3f3611d71e8374dbc73192d145

Observation 90fb4eaf-421f-4151-95a3-dd0dd3999e1c · outbound

This paper cites Otter: A vision-language-action model with text-aware visual feature extraction.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Otter: A vision-language-action model with text-aware visual feature extraction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.501586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.501586Z digest=sha256:dbf3b26a2bec060015a57dde613e749f1653370425469efa37715848baa66536

Observation 822fa9cc-27c0-42aa-804c-b12f3cc2c5b0 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.604747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.604747Z digest=sha256:57e0e1985366181c85e7904d603cc987250e5a3dda5632f8cf964b4be379ab5b

Observation 68cc964f-af0e-415a-acf5-9449d98324b9 · outbound

This paper cites The better you learn, the smarter you prune: Towards efficient vision-language-action models via differentiable token pruning.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies The better you learn, the smarter you prune: Towards efficient vision-language-action models via differentiable token pruning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.715989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.715989Z digest=sha256:2dc1e0d04605195e4a802c7cb4ddf90b8ae34a101bb4f7c00ddb41cd11f461ac

Observation 0b503360-8bb5-4e77-bb98-2bb75e829740 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.772125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.772125Z digest=sha256:47415a8acaa7c7746c65f39ff830fc573e6b590ca78ee67fbfca9ba94a016835

Observation ad7a0ac2-e012-4c8d-a8dd-4998de472333 · outbound

This paper cites J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.879074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.879074Z digest=sha256:23bc0d50d97426e1616a861b7dfa65080264ffd5585f989851e9ddbcced4a22f

Observation 2cb37b32-e59c-4e28-904d-aa6761045a37 · outbound

This paper cites CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.939069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.939069Z digest=sha256:cc2ed101a95aa80f270cfcb91ef52780d118cea484f2ed0ae3c4bee763c5e41d

Observation 6214642d-950d-46ac-a6d1-7c53307f2485 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.993667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.993667Z digest=sha256:d7fc38f888bf8e46d1f73509b10e97bddbd154bcf807059ae6ed48eba2bed663

Observation 2fca8ba7-75e2-492c-90b4-926308054eb2 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Vision-Language Foundation Models as Effective Robot Imitators

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.024414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.024414Z digest=sha256:86d13f80f0e250b920d5498d4cf9614cdae4da360c4ab8566ba8363943e708f7

Observation 5c5d9faa-2b20-47fd-883d-0ef0312d7808 · outbound

This paper cites Flow Matching for Generative Modeling.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Flow Matching for Generative Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.083076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.083076Z digest=sha256:efaf6dd87dd74f2434de897601232cc81f43502b2ea8e28be5a6e0cbe484c22e

Observation 6547da2f-54b9-4ce2-a21a-70f92dbd2a79 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.142388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.142388Z digest=sha256:1c10df5c79b4615fe1c38f6b805c564ade0f13f5d2066c321676dbdecb59eb55

Observation cc3bae6d-825b-4b3f-a816-02c0e955849d · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.199824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.199824Z digest=sha256:324e7aa441e4701b29459710320b8bcdd92e7aff6a08ab45b5a7caeb133f9e68

Observation 6a4daf34-bf94-4cfc-9186-dee12cc36bda · outbound

This paper cites and Xie, S.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies and Xie, S

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.315533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.315533Z digest=sha256:810600d7c946a36e6e7607d77bf03c9b4f69cccc5c1b3105a8c1428ccb3aa469

Observation 82f5e85d-e02a-40d5-a581-3a169114d44e · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.379652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.379652Z digest=sha256:efe689000baf69d08848891741a343b3b35d796c4a2260e357e43aadfeecedce

Observation 15fbe715-5e6c-4768-9716-468af86a2bc5 · outbound

This paper cites Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.396182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.396182Z digest=sha256:f6baf61ba8717b7a917779c99911a38f30689073148a6bb6ffea30a289df4be1

Observation e99d5310-da24-401b-bfe8-fdbc90974028 · outbound

This paper cites E., Otto, F., and Lioutikov, R.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies E., Otto, F., and Lioutikov, R

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.459011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.459011Z digest=sha256:51f21eff4a233f0274af755db0bdd598d8a8d0cdc2345c496fc962984977210e

Observation a4b86630-081a-4e87-b794-675d41d62a52 · outbound

This paper cites TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.496346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.496346Z digest=sha256:55c73f7f9a84bf0eb6a354e29d0c4e8932075184ab07a1a7ae6748843fb822f6

Observation 54748125-b756-47b8-9bf5-fe9c211be117 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.558783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.558783Z digest=sha256:8af99e56c7d3e2685d5451002323355fe56dc3b9b58c8ac0b43539fad2a15f67

Observation 87a6f7a6-44de-44f9-8bce-769e41b39213 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.618437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.618437Z digest=sha256:6d1b4185a8fdc777c53269c2e14752380f676d59418c5ffc3241163912f1a996

Observation 537af961-80dc-4af7-90f5-d7c32efc8a9b · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipulation.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Unleashing large-scale video generative pre-training for visual robot manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.689402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.689402Z digest=sha256:ec220eaf41af1923b8165037ce143b0d00e316eb89e4e6cf1b04bfe059e95d9a

Observation ddf23d8c-a1e1-482a-8f98-b654a88a821d · outbound

This paper cites Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.729625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.729625Z digest=sha256:7b6ff27e5938696f898bde3fabfa330af7b06384ebef2b7dcc7770f226d2621f

Observation 31281768-7a8c-4a6f-8d55-0dd8cdc96c43 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.890768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.890768Z digest=sha256:b616c75b87a786771170108f38c019a736fdea8975340c2b36e97d60baee9ab2

Pith citing papers

Observation bc31dfb5-636f-475a-9eb3-7fda664f660e · inbound

When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA cites this paper.

When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:53.217702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:26:47.393632Z digest=sha256:35718c546efea40f678ddf4878d1bbda799bc3c614620466925bae2f438ae9be

Observation e54940ea-f858-4887-9295-6bd8d2cf92cc · inbound

How Should Vision-Language-Action Models Use Proprioceptive State? cites this paper.

How Should Vision-Language-Action Models Use Proprioceptive State? Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.118690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.118690Z digest=sha256:538696daadd3bb716cfe5fd5ca71d43ab791bc1b3fcf7cdc2274c31f5956ca0c

Observation 8c5ad290-f9e0-4386-bc10-c1d1fc7b6ee2 · inbound

How Should Vision-Language-Action Models Use Proprioceptive State? cites this paper.

How Should Vision-Language-Action Models Use Proprioceptive State? Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.122151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.122151Z digest=sha256:095f216acf54d3f14160cc4f77533b8ce4565aaa898f3e3b66760fc28c0b877d