Pith. sign in

Paper Citation Record · LEDGER

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 18 inbound Pith citation observations for arXiv:2505.21200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21200 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:43:04.537493Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:36:11.929173Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 42200279-e32f-4538-8d46-af1268e082e5 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:58.620579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:58.620579Z digest=sha256:abc6aa1cf84d832cabbe6b0e380ccb6fa433c1d631f80cbdc4cdb5b1ba89f032

Observation 220d2467-d9a8-4f82-af99-bfdd8cdf092a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:58.759097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:58.759097Z digest=sha256:64f8171d3c8b1ee72eae8dcc684dc5a50fbc5d6f458bb849f09f7ee49bf54e17

Observation f040c164-af62-4b84-a709-3bca321d088b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:58.929949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:58.929949Z digest=sha256:02313dafdaa1931972d977d53a59c1e1669bec334fe09d41c66acc0221710459

Observation b33ab496-7c76-4034-bbbe-d09dfbaccd17 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:59.076821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:59.076821Z digest=sha256:85991b859d6972d699075636c5a64b1eecf1515b204b6ac74edc0955d228aab1

Observation b38532cb-ba0d-403a-a6c3-4d68fef0db36 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:59.173764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:59.173764Z digest=sha256:ba2aadac867627a2c8830166aac34f6266d5804bd86d6a1d19120046dd087a9f

Observation 262519e2-caa4-4cbf-b62e-00dbd37ca820 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.768722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:42:59.237915Z digest=sha256:010f6fbff9acaed9af939d3e2ac02fdf73b8eaad0e0df2cbc1a643e733f75246

Observation 5552c08d-1b32-43ac-aee0-2431f622ec04 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.651756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:42:59.492720Z digest=sha256:36037a01d901992d255c150414e0ffd6a8d4feed24a0fffdd223e58ca0ae6512

Observation ef69b40b-65a0-4648-8698-b9ac0637225c · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.466170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:42:59.647436Z digest=sha256:c5fc5e0034cf62f9404227496e1dcd210dfe3f79d5c762ef9ae5819e45c8919e

Observation 2e7eda50-524d-4190-a527-f414ba586b84 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.302972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:42:59.712646Z digest=sha256:b48c3bae21bc26aa41aadbb63f60b19f09867ba7d3778253f894874d8f8e216a

Observation 14bfbcab-b152-44de-b9fe-5e4af94d01e5 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:59.879271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:59.879271Z digest=sha256:e76d78b922f750d8392961bb4205611cb4a71828d00c9d23d41d75313dbf44e1

Observation af051aea-0e19-4b7f-96ae-f55f54d1c16c · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.008350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.008350Z digest=sha256:6a1bdcd6aef60f22f8b26650fb46d17ca65fc91c6d150566af357afb24b9f33c

Observation 8df6e184-76dd-4822-a79d-9717c2273671 · outbound

This paper cites Manipulate-Anything: Automating Real-World Robots using Vision-Language Models.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.230985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.230985Z digest=sha256:c8680ce14179756a31137b9aa5fe61176961e4525c4ad93c0b184e092283b334

Observation e6e1485c-b068-44a3-b2bd-e5fe14d36143 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.431971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.431971Z digest=sha256:2be3fa72126c83baad5938d73826d72690a21283ad36cd895e02255ff6463d72

Observation 55fea985-5ab4-45d7-ae53-ef1212805730 · outbound

This paper cites Diffusion Transformer Policy.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Diffusion Transformer Policy

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.597211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.597211Z digest=sha256:debc9de4be4acaea8e9b1b5898ffb3c3fdfb116e8900a82956164bb18955e1a8

Observation 9e645849-c1fb-48f3-96b6-e46d5bcf4fe1 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.735532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.735532Z digest=sha256:815faa55bb1fcb6c300505f61a570c3eb0ad1dded402ec6e781bcbf10cfb90eb

Observation 5f2091d7-02ca-4580-a627-db9cc2403b6e · outbound

This paper cites Jiang, X.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Jiang, X

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:06.120893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:00.879965Z digest=sha256:3bd91294fe4c9caa75f22e5c6973ca488d7cf9f109bf57dbb078817351b0efb4

Observation 9d404b8d-14f8-44d9-86af-fd0f71ed1923 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.083162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.083162Z digest=sha256:cd793aad6f9a21060a7296e09c83c7470abacdc42e066aa023fdf486a3045fd6

Observation 4050d5a9-723a-4ca2-aeaf-cabc24b5d3b5 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.279474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.279474Z digest=sha256:482187dee1f7714a7fcdcabe07533375795e0c24f0ff310b68ffa1caf54b892f

Observation 9b4c248c-1307-4904-93bb-016fe5ee8e86 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.351486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.351486Z digest=sha256:6660588cc90d493c5f390b8db140a5582d98fb4c21485c30e658e46362d3e106

Observation 664d0ca2-60ca-40e3-9239-9fe0ecb9f715 · outbound

This paper cites Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:43:05.160857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:01.359705Z digest=sha256:6839533fe478901886a6bf295d3944d1421c8e00bd4f01772747060708a7becc

Observation f06cc0f4-5215-4d09-866f-98dd11e095e7 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.480814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.480814Z digest=sha256:d2f842f5c86afce2b2b13e0f1a1f3dab8aa87ca0b82f427ecf282a54588eda73

Observation 34ebeda1-0fb0-4921-8ac1-4c0902d3e029 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.638742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.638742Z digest=sha256:1e26ef33013dfcf1b2839b28df844ca89cf23ea0c84ae16b86d404e85d9e121c

Observation a517b294-5f40-42d8-90fe-a0e1c789012c · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.760587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.760587Z digest=sha256:7ce8fb90d33783680e9c367a0d23985a945103670205efe2ac84abd1351e6180

Observation 83926484-51c0-4457-93df-48ae6ccede47 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.916518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.916518Z digest=sha256:efff609563c87646e4b578bd943c3dfa12607b04dd80aa789e8b65ea1ba397a9

Observation d0bc70a6-f547-44cb-8a61-00668dc6b4f8 · outbound

This paper cites RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.033013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.033013Z digest=sha256:26b7fd46211ffb3346fb6061850bd62d3b44a71ccbd7c47112eb0d5f12fe6204

Observation 8e3fbcf7-60c1-4579-bac9-cf9dbd6b82d4 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.178483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.178483Z digest=sha256:b3c4b5c2aceb5fec6423b6783d8f8e8265672b36bc7160e51739b7827d75ce0c

Observation 6ba2346a-6f98-45c9-a41c-80c1c81e0659 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.340523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.340523Z digest=sha256:c389f425e3544f5db8ebca07839f986033b3d979653c1372ddd6427887d5bcfb

Observation bede8346-9c03-4645-b231-b8d865439650 · outbound

This paper cites Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.506958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.506958Z digest=sha256:a679912e8dad2a6abfdfbc74100ba646c90c60dff881a268b9d941416c8a4576

Observation 7b4e0d74-7f2c-4ab2-8c46-f9f0c3947d2c · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models R3M: A Universal Visual Representation for Robot Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.629364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.629364Z digest=sha256:adabf860a809f075a67d5df8af49b3696fc616c89d1ddde5436df901f4465543

Observation 47a2d80b-83af-4431-8303-138fafda7b9b · outbound

This paper cites Oquab, T.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Oquab, T

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.739636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.739636Z digest=sha256:6d2ba0dc1d76e731e1f9ce6539f4be047998479d4660481a00d437ca14afd77c

Observation be8158a6-1d9b-4d35-af86-66d8b09a0150 · outbound

This paper cites O’Neill, A.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models O’Neill, A

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.744927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.744927Z digest=sha256:bad22672e19859acac9df416a03a8e448999f5710b8844ae811f266638d07e73

Observation d8b20428-4c1d-403e-9387-b86e1b507580 · outbound

This paper cites Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.750457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.750457Z digest=sha256:0ff65a8893494a4822f131ffe0dd6a3649a31557ba64e35773055ca87ca352cd

Observation bae20548-c5b9-4c76-a6e5-58533845c6b4 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.850761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.850761Z digest=sha256:ed84f8a856e2fc826c60f1d546fcb692b1c9aa270ba299456ae9744f3eab2529

Observation 5cc97333-604b-4e29-9a59-121840267127 · outbound

This paper cites Singh, V.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Singh, V

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:05.924852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:03.002933Z digest=sha256:ab10c5ae173db07ed53a951e0fc83879893d916b93fada59d443a7b6e233e82b

Observation bc2b18de-deae-4f08-ad3f-df8ae4ddf3ed · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.201022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.201022Z digest=sha256:bc95d7d380a7a71dee863d514bd7b8c9a830f8535f47edf32bfc0cb6df47c0ca

Observation 849551f8-8e3d-41bb-8fd4-fde8f3c680d5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.323929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.323929Z digest=sha256:a9aaa570e16bb89817cbadc2baac6484989e2e98f5acb5c5cc4615897c877060

Observation d9b01b24-1d17-4365-b681-446034d39b78 · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.455619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.455619Z digest=sha256:22e46ccb96653a698322f7061e7eb869cfa074860dd3fdfca41b00f7d9adb177

Observation 0fc3c282-6fd8-4637-930c-58cea05178ab · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.601909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.601909Z digest=sha256:456d18810e24be6768ea317fae3a4f6a37c1091f18a4b9ecfee46e6b46df6063

Observation d9482747-4466-46a6-ae62-9b5a24e150ee · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.749499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.749499Z digest=sha256:e3d3676e6d987b26c5e81218f8d10b9251e0b54eabb3da2bb54ce5b36ad81761

Observation f8fc4f28-c74a-48c7-881f-8a7e7109521e · outbound

This paper cites DNAct: Diffusion Guided Multi-Task 3D Policy Learning.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models DNAct: Diffusion Guided Multi-Task 3D Policy Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.817109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.817109Z digest=sha256:6e78e6972757204c1ced0e8543cdd3239880a4537587b1c7b1c83e02e27b8bbd

Observation abffa18c-c2c1-47ef-8a27-82f0e2f3eccf · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.910286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.910286Z digest=sha256:92e34abf056cfc05ba2b059f96a5cb3af81c5557b60b281eeafe40a067111f23

Observation 8c50a432-9f19-4125-981f-ee77a805da2e · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:05.796712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:04.018067Z digest=sha256:d2ced3704429a264065aa9094517718c312e57d54ee9b5313ff9a825c3b9762e

Observation c769917f-87c4-458b-bc29-fa86fccecb79 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:04.147405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:04.147405Z digest=sha256:716571f0728a922347c89cc238a5a4e12f55fa9fa8f930c4c7315b4b41450b28

Observation 4e812d10-ea7d-45f4-80f4-3e28783cfb9e · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:04.258730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:04.258730Z digest=sha256:578d269cac71c30d02bc7db7b19a7471a2646d14beb57acf6c06c577ce2e75ea

Observation 87e0c8c6-c14a-4a5f-9e62-32f80e215ed9 · outbound

This paper cites Zhang, Z.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Zhang, Z

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:05.638348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:04.323508Z digest=sha256:d8b38117ac257c4d7cdb38894da3ac2c6192da3da187858e2a8d497d378d6239

Observation 549af06b-7b60-4c27-9aa4-e09d009bb627 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:04.407246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:04.407246Z digest=sha256:e557033b7ccacc30634b70875e5892b3f9311b36a6248c49c0c8785f7f2984de

Observation 5de09fb8-0c6e-4ec0-b728-13a0d27c0ed1 · outbound

This paper cites Zitkovich, T.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Zitkovich, T

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:05.434319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:04.537493Z digest=sha256:8ad6252c79e37e53db92990e37f9784fe8f894b546d16fa960c319e1b042f6aa

Pith citing papers

Observation 27e0a52e-db45-4355-a05b-7d132211f3f4 · inbound

RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models cites this paper.

RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:36:11.929173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:36:11.929173Z digest=sha256:897920cc5e3c21da923f55a452719beee2b6f28f479707b9fbfd77fe11105ec9

Observation e9ad02eb-0de1-4f98-89bb-7bfaae0e40a2 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:28:16.251395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:c5cf17dea66b868c0ad8cfcf09749a6009e919c14e5ec277e5b39dfc91537cee

Observation 5cc907b8-c95a-409f-af39-091b8910feed · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:12.680018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:12.680018Z digest=sha256:199449bfba4056f936ce0c01ca69fbeb143ef86585d3f83820c432491668183e

Observation c7669fec-08b0-4352-868d-3ff350b648b3 · inbound

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models cites this paper.

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:09:09.433198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T06:07:42.311608Z digest=sha256:d43875d322bf40faef9c07f56260cda400edc98a53e75c2b1f93859297f47940

Observation d75ee6f8-1593-40a5-aff7-9f64f760f4ba · inbound

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control cites this paper.

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T09:07:39.309018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:07:39.309018Z digest=sha256:05427ddd88e21bdf25eb1123978fcfe53d177b96f0d21d9663556b877c09da05

Observation cdeed0c5-557e-4128-8df2-49523da42a74 · inbound

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models cites this paper.

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:40:14.434166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:39:58.095112Z digest=sha256:6203729b37ceb8669eb83df16f8694c74d2b13be0283a47663b13706e15c84a8

Observation 566736a1-5b5a-44d2-9456-127d74258f25 · inbound

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism cites this paper.

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:50:03.934133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:46:12.134869Z digest=sha256:fe8702b2fbf16c5d5ebf482a9c0555acc40ee69ec0f8cc1126fc39ffff721ef5

Observation ebda3e5c-1bb5-4920-babb-7c268ddb6849 · inbound

VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success cites this paper.

VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:25:50.142804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:33:26.708966Z digest=sha256:ccc8c620c6f106c95c6e921dce1dee9716719117e1ba5e306b734b5bf3b1e6c7

Observation a65763df-9030-4369-a569-943e7c27f4f5 · inbound

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models cites this paper.

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:22:06.687892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T02:19:39.604140Z digest=sha256:26bc098b422e7247db4c0354dd0ab3d516e258646802ef7089a500dab9bd8b79

Observation 2a3220a9-f94f-44ad-8e8d-4501e7d6aa96 · inbound

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models cites this paper.

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:15:07.104557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T06:13:12.442859Z digest=sha256:e31b060d95308393ec976d8d326b73a74535f2f4ccf57e77a082d5e798f8ac3d

Observation 346d9813-74f7-4774-a93f-75f41c53f5bf · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:47:37.220705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T18:43:08.029165Z digest=sha256:000e4f75831fcedf984e08758d499b2480606bd1ca01877386c27476bdfce404

Observation 012a3ec7-7c05-4075-a2b5-c77670f3a936 · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:45.902573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:42:05.592911Z digest=sha256:24aa0d357c07c678636e0b74dccbc5ca9f284d80909acc53c04e1aeb004055e3

Observation 207fda53-dc32-48a8-a34f-e553e199d3a8 · inbound

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models cites this paper.

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:13.041269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:16:27.229603Z digest=sha256:3ab803549ad214bcfd7960cb4279a9dd7832a8ecedf701b3b778f04490122848

Observation 43146e1b-4c59-4bc5-a8d3-d2e9ed35e07b · inbound

GeoSem-WAM: Geometry- and Semantic-Aware World Action Models cites this paper.

GeoSem-WAM: Geometry- and Semantic-Aware World Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:16:34.619734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:11:10.246045Z digest=sha256:32b5325219f020fd6f4ac167786ec74465c531cfb4081f2283e6e3ace13f815b

Observation 10a0d5dd-2dfe-427d-b614-7934d9e4435a · inbound

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models cites this paper.

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:47.661540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:57:01.861091Z digest=sha256:4f949a5901ef047925247f839626689f64719cde3eb23cb1ab9ebaa7aaac19a3

Observation 8bdac25c-b34c-4d14-943d-f0123da50301 · inbound

Improving Robotic Imitation Learning via Trajectory Standardization cites this paper.

Improving Robotic Imitation Learning via Trajectory Standardization Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:45.564349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:37:26.864436Z digest=sha256:548c534d5d4b54fccd171414e954c42b5906cb409847c8fbaceaae704520676e

Observation c89b7deb-90ab-4cc3-b7c5-789c8b0edc79 · inbound

Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? cites this paper.

Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:52.300386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T04:56:00.919208Z digest=sha256:764c04c8d4d3f031cdc5bb0e59cca33617009cbf83d601f0ecaf68955d129b34

Observation 6ee57c48-dd03-412e-ba93-5ae81aa47fbb · inbound

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation cites this paper.

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.864902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:07:45.792132Z digest=sha256:d8cdd58ffbd870c7a5c6919e5a142f2f772688cdb4b2bd9baeba91be4f588211