Pith. sign in

Paper Citation Record · LEDGER

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

As of 16 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 15 inbound Pith citation observations for arXiv:2507.01016.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01016 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:07:26.130586Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:40:42.849572Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:19:43.485742Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0bb710e7-e1dd-4a4a-9df3-36d5f420dc2e · outbound

This paper cites GPT-4 Technical Report.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.122914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.122914Z digest=sha256:bb0373efc5d46bbc17db53f03ecbdedea423984c48b9b8d1f1109df14a3af833

Observation d2b560d4-2189-4451-aa56-58d7370ba347 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.189706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.189706Z digest=sha256:d0192b0f2f9bd55c2731ffa7a53e3579fe3cc3a6efe377249ebcefd6e412da9d

Observation 8cb96dad-3e34-4e3b-be2a-3df86d9bd178 · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Flamingo: A visual language model for few-shot learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.458950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:22.275615Z digest=sha256:bb81f47be6fbc558e3b793a19e270486b59c5a28cb7c85df404da2dbf44f6074

Observation bab2aed8-d8df-48d5-b6d8-f8f3ca6e8650 · outbound

This paper cites Minivla: A better vla with a smaller footprint.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Minivla: A better vla with a smaller footprint

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.333788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:22.363654Z digest=sha256:875eeb0f583dc6e426df9f45f307d04e1dbcaf8304f5366c61427c8413866149

Observation 9b752740-a23b-42b6-a8c4-1e9ce8e0b41b · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.451956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.451956Z digest=sha256:3529b685ac33191b8af148c585a121fb556621c5d5f024a28dd04b0c7283fbc9

Observation 6df64edd-df2d-43d0-b50b-7a9991989922 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.506991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.506991Z digest=sha256:91b992634323ee048aa3ac6a8719be0095165401fd345fa1b8fb38848ac318a9

Observation d43c688d-cb8a-4c24-815a-9141f9a3bea3 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.569402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.569402Z digest=sha256:859b74542eec6d3e9610e0c9c628ac7135a7454f59f15a726ab87a7c4eb0c726

Observation 3a2658ac-b097-4a57-99f2-2d45dbf88426 · outbound

This paper cites Anyvlm: Unified vision-language model for any robot morphology.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Anyvlm: Unified vision-language model for any robot morphology

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.147454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:22.651487Z digest=sha256:79de22c31b160dbf17dbee8d85909ef3651727f8f771e41d4eeacf9c78eb7024

Observation d4f6ece7-02f1-4998-ac13-45b3652590c6 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.706473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.706473Z digest=sha256:ade7d33711c30d818e04aa056ac65f4a66477eb028642ebf5c8545bf188d2d41

Observation 3827fc8f-d211-4da5-b32d-9e48b3d2b97c · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.759775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.759775Z digest=sha256:2e6bf1073fd85b2259cec756cb0a26b413a7409ae4d8602b8e97a09332b37a31

Observation 7457aed0-8bc8-4c67-8070-4a723c9fbbd1 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffu- sion.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Diffusion policy: Visuomotor policy learning via action diffu- sion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.002060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:22.864369Z digest=sha256:786c6c0aa153919216710345cb151208c332fe9d35045ae6f285b89bc99cc39d

Observation 49ca3dd1-337e-4860-8823-47335f20c2c8 · outbound

This paper cites Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.903469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.903469Z digest=sha256:d07d33969b420a725656c1504eff7a3c563f760cbd0dd300edacdf228701cc06

Observation f6a7de39-98df-4cdc-836c-6243442d69a2 · outbound

This paper cites Palm-e: an embod- ied multimodal language model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Palm-e: an embod- ied multimodal language model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.918616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:22.959736Z digest=sha256:7e212769b305014e30f1bfc91a12d8c7f4104efba4961fe4c1bc4cb2cc2cd074

Observation 97837d45-4a19-4c9f-a7d8-8072d999b00b · outbound

This paper cites Exploiting llm quantization.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Exploiting llm quantization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.833382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:23.014734Z digest=sha256:e51cdcb3a837e2b54e7c0af5b052da9a31be3ddbb918563706c7d14a80371c1c

Observation f2dbaa8c-661a-4e26-ba7c-87affb1f7875 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Taming transformers for high-resolution image synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.731198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:23.128367Z digest=sha256:2a3945fc274b36c33e309899d396d0c231c3639450191d7e128d0484225f54a4

Observation c5c4ba28-880a-40a5-b658-21c9594fb4d5 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.240619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.240619Z digest=sha256:0716293123f0e4f63267e060e75173cf7011f70157bf5f24feb34424af1d309f

Observation 96be1bc3-b35f-431b-b1a5-c6304326c03f · outbound

This paper cites A new algorithm for data compression.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers A new algorithm for data compression

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.620501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:23.348202Z digest=sha256:6690eea846d6aedbcbe88b9f2ed30bbc8c86ed4822f2df903e633943645383b6

Observation 743aa2af-c4ba-4d72-96f2-2b3cd9565335 · outbound

This paper cites Act3d: 3d feature field transformers for multi-task robotic manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Act3d: 3d feature field transformers for multi-task robotic manipulation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.531754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:23.457029Z digest=sha256:ba60df55f27c82e0c422ecd78da0c208326afc0dc84ee780837ee3e3d81a800d

Observation 85db7f8c-2df1-49c7-836d-3b681e7bcffb · outbound

This paper cites Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.545915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.545915Z digest=sha256:289c6a209e765d55bd686ff5424ff08ddb063d5b382d37e8b4c02489a65eff62

Observation b1115345-e57e-4ec5-a187-8b9c1ed1f47e · outbound

This paper cites Deep Reinforcement Learning in Parameterized Action Space.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Deep Reinforcement Learning in Parameterized Action Space

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.686753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.686753Z digest=sha256:3dfee1d44349c4363ec992d868e4ee6b34404271b5e374689e6707f107a5343a

Observation 8e9902a1-93fd-498e-b383-c9e4935b72b7 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.792753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.792753Z digest=sha256:7e0bf70d3df368989543a5f4e9ae2024206abb7e58ec72c2cd8241af5adc0a75

Observation 13dca03f-3454-44ce-8b2e-858468f84538 · outbound

This paper cites Rlbench: The robot learning benchmark & learning environment.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Rlbench: The robot learning benchmark & learning environment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.442533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:23.874297Z digest=sha256:4b4ccbb715839b9686bc8995d533a7d37e621b7c2c83a38cba49fca787012e94

Observation 398ef956-5728-4acf-aeae-20f2aa4fcf33 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Pyramidal flow matching for efficient video generative modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.982056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.982056Z digest=sha256:a5dcd9c82e88bf091deb638190a4e67f511548d87129a4cb646d9132306dee9f

Observation 60582574-bb70-4fc5-abd2-93782b851e18 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers OpenVLA: An Open-Source Vision-Language-Action Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.072732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.072732Z digest=sha256:af148dc42bef215cc2ccb3d133e585a27dd358226c0776a2523f3e573b543428

Observation aa408865-f62b-4e67-8fa0-68500444dcf2 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.169056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.169056Z digest=sha256:7b61e3ef53badcf7cb9d573ef55eb4fc69629994b71ba225c3160370373c0456

Observation ab9bdf39-151d-42fc-93a5-b7188949b3f1 · outbound

This paper cites Behavior Generation with Latent Actions.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Behavior Generation with Latent Actions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.226708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.226708Z digest=sha256:7c9f66723007e759ece556ea75459577cc254ffd1a05c361b59bae2df333b8a9

Observation 12ce03b6-eb68-4fdd-b832-d35b9d18ff2f · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.295662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.295662Z digest=sha256:7875323fbb556cedb11ba9c6aa7d0ef54929ca869d242279ac1d8d5aafa82be7

Observation ff2f85e6-844b-4225-8780-158474d501a5 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.365470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.365470Z digest=sha256:d0a18ccd866e4217a0b4076bf5f76f1e5ac5a46a64fc3f0f4869d4fb0867be38

Observation c61f6e63-fbab-4d01-a09e-f871f4665d8d · outbound

This paper cites Language Models are Few-Shot Learners.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Language Models are Few-Shot Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.445625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.445625Z digest=sha256:71707dc9be471373ce77e16ace6e66848f5c848c3c30ee3089e8b5886862d29b

Observation 2dc1c629-12aa-4ffc-b9c5-24dd9705b76a · outbound

This paper cites Quest: Self-supervised skill abstractions for learning continuous control.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Quest: Self-supervised skill abstractions for learning continuous control

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.322251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:24.514228Z digest=sha256:0f3b8e4bd637140812e73c0aa47b5cb860cc5afe4f9b97219cd0782d1a3ee2c2

Observation 5beecf7a-5133-43f5-a2de-2873d75a10cf · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.596296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.596296Z digest=sha256:357126130d99b8fc8df975b5488e42adaa36fa1b50b61a0510089e09f52bd377

Observation 2b13a5cb-123d-45f7-a071-a5d63ffe1527 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.226869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:24.653961Z digest=sha256:46a9823e520a11b422712bb4b40c00818329a57864b488796ad22c3b69923d68

Observation 0c87a25f-e31b-4a1c-bcc9-b8cbf4155947 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.690665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.690665Z digest=sha256:e36e7c9d088995be6ddc2afac8b633c6d8478b31d6b3a358323e09f4014c4d67

Observation 49f8508b-3bab-41b6-9b32-6d4ba7b90ecb · outbound

This paper cites Language models are unsuper- vised multitask learners.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Language models are unsuper- vised multitask learners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.133600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:24.774942Z digest=sha256:16751a5c91225b367761a5013374c88e60b87f970ff9a2c325d3c3ffd358b15b

Observation 3a536382-8929-4bb4-a7a0-e36e558a26ef · outbound

This paper cites Scalable Image Tokenization with Index Backpropagation Quantization.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Scalable Image Tokenization with Index Backpropagation Quantization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.825964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.825964Z digest=sha256:b160d27877f7a82e6de600b1e9f10e969c763801c3e5d1764d9a6fc03ad004b4

Observation 7366d8dd-f706-4cfe-b1e1-b8382b733708 · outbound

This paper cites LLM Pruning and Distillation in Practice: The Minitron Approach.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers LLM Pruning and Distillation in Practice: The Minitron Approach

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.915904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.915904Z digest=sha256:a1096b9e3bce1fdd2ef678be171cf62fb5af16062cc9b7b1ee898453ca02ca14

Observation 4db8a6af-59f3-4159-a494-9b86e285c1fe · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.973912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.973912Z digest=sha256:6927b27dbd2adcabcfdf6e2c0c9dcb55d95748f32a9a90fde9d10d666df5d975

Observation 8554926e-dcd0-4b2a-93ab-00b3c145a3de · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Octo: An Open-Source Generalist Robot Policy

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.030945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.030945Z digest=sha256:9f693a3dc7329c16ba6c9630432b40a63df0dc7e9b6045afde87d9faad0feded

Observation d9d6a6d4-90fa-462c-9991-fda81792a0fb · outbound

This paper cites Neural discrete representation learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Neural discrete representation learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.105749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.105749Z digest=sha256:a774e8ce9416eb5db7a49680fee738967b2192b8c59ba39e59fa94216e54415d

Observation 9c69885d-71c4-4ad6-b823-a8a55f78fdc6 · outbound

This paper cites Chatgpt for robotics: Design principles and model abilities.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Chatgpt for robotics: Design principles and model abilities

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.044323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:25.151700Z digest=sha256:750505f8a673efc017c17b5ba2f1008b36e462ad649c7b2da2be56405b3e2dfa

Observation bd82c91d-c9d7-45fc-a6a3-c4788a484fdb · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Any-point Trajectory Modeling for Policy Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.244285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.244285Z digest=sha256:b8bb61cfa5bd03cbd951ceb76bb980f41ae1b71346bcfaa32f8165f80cba95af

Observation fd872960-f77d-4bab-9005-d9e24eec506c · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.323052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.323052Z digest=sha256:c28c346af768067dd9e6243fe774b5eeb0319ec3c87fd9408bee765fbcd4db23

Observation 748b3e90-2048-49ad-a580-f499ac889134 · outbound

This paper cites Transferring Foundation Models for Generalizable Robotic Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Transferring Foundation Models for Generalizable Robotic Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.362504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.362504Z digest=sha256:b1c7d52fc542611d69ca19c2e9ff46d31be53d1a4abd3d83f40ea77260feb7d8

Observation da1bc365-02d9-4d08-a4b0-9b76864969d4 · outbound

This paper cites Spatiotemporal Predictive Pre-training for Robotic Motor Control.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Spatiotemporal Predictive Pre-training for Robotic Motor Control

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.446793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.446793Z digest=sha256:8888a54b51aab4ec74d1206bc4d5ffc58d2f36b83ff09dbea3eb6601da307eba

Observation e1d8b0a8-5381-4bc9-bd51-22bf5b3618a4 · outbound

This paper cites Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.504405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.504405Z digest=sha256:06bb759a42d0243ffa365b3140d9ec683a7d7099ac451580c7c7f3ddb757bcc3

Observation 5a69ef97-ae9a-40d9-9334-e9c04c16722c · outbound

This paper cites Latent Action Pretraining from Videos.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Latent Action Pretraining from Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.536516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.536516Z digest=sha256:512bca0bbad44cc95f864f9faded5cecd3c836fa9b72794f9fc6395c6c67ce7c

Observation 1b81bb15-fd8f-409e-9e5f-108a76b3b3ba · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.595946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.595946Z digest=sha256:6e43d11c3a304b465a7a04f10e9e5fe53fd09ea5b8ab83863e5daa29e0550459

Observation a4a5ee28-9536-481e-b8cc-fca2f56b8fbf · outbound

This paper cites Soundstream: An end-to- end neural audio codec.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Soundstream: An end-to- end neural audio codec

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:26.923590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:25.664555Z digest=sha256:6cc28cfc290c34c4d9d69f0343a04ed8213f717d89969fd9b038a773690526c2

Observation 296c81bb-282e-4713-b691-debf1cd58ac1 · outbound

This paper cites Moviedreamer: Hierarchical generation for coherent long visual sequence.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Moviedreamer: Hierarchical generation for coherent long visual sequence

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.702842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.702842Z digest=sha256:acaed6ca8a26e464bda16ff86d152c98fe2ed2b76564fef42dfe9c45f94d6bde

Observation 850984de-93ef-4ff3-ab41-3585a9bf2cea · outbound

This paper cites Diception: A generalist diffusion model for visual perceptual tasks.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Diception: A generalist diffusion model for visual perceptual tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.791210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.791210Z digest=sha256:b970fe1765977452beacf854ed43f58e06e62eecc993896642ddecc29529fdbd

Observation 9b04ccf3-2d15-4542-855f-c4cece5049b7 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.850810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.850810Z digest=sha256:b6a48e4b472a69ecd0ac5673558739a7c283a12841fbbc0fa85032888759e410

Observation ef8e333e-3baf-4787-8619-ad432f15f447 · outbound

This paper cites VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.936596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.936596Z digest=sha256:510522dbfd243b1761748c8852a525bdc899ef9e589e317293a38732b19c420c

Observation 50648549-db24-47ed-9cd1-4360a95c363b · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.998032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.998032Z digest=sha256:9b0d0014daaf143512089739fc139936a41b275c264d24e22319c9cc346e3cf4

Observation cb1aa5e6-43e7-4995-b644-82f7e86dbef5 · outbound

This paper cites Point cloud matters: Rethinking the impact of different observation spaces on robot learn- ing.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Point cloud matters: Rethinking the impact of different observation spaces on robot learn- ing

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:26.828145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:07:26.049381Z digest=sha256:4883fd782d2d2634bed6ced1cffef307d265e83c842a1e02c98fe85622df11a1

Observation d873b553-398d-4f19-8b8c-2d544b0d7e47 · outbound

This paper cites SPA: 3D Spatial-Awareness Enables Effective Embodied Representation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers SPA: 3D Spatial-Awareness Enables Effective Embodied Representation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:26.130586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:26.130586Z digest=sha256:3059548f81bcae3001e06697f93f0a9c38b2e7ac0e9f1097de0b8ddf42bbedd2

Pith citing papers

Observation 95b0d24c-17d7-4bd4-b034-f64e3f093f85 · inbound

ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks cites this paper.

ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T17:40:42.849572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:40:42.849572Z digest=sha256:a151358354133a980cbd41683efa99414797d1f3108f371c27de0a322a1e54a4

Observation 3a81d664-dbb7-49fc-8697-3ce1f15be229 · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:53.042021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:53.042021Z digest=sha256:54a31426c23584448da6439422d4733bc290a6666f8d068a795305916028088c

Observation 4f2fbe16-6b13-4c90-9544-c2ded5458320 · inbound

Continually Evolving Skill Knowledge in Vision Language Action Model cites this paper.

Continually Evolving Skill Knowledge in Vision Language Action Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:04:09.248500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T06:02:31.638120Z digest=sha256:48d23cad7f1b468ab2a7e8b5119439af5c7c7ba6b24cb02eb25bac40d9e2dffe

Observation 5a469e98-446a-4ff9-8737-5725a9922559 · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:48.923259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:48.923259Z digest=sha256:724f9f7acd1a0a2255196545e3d1ab12539f1589d283dd8e1f1a389c1d16ff65

Observation d3391853-c917-4460-8253-b556712ed4f2 · inbound

Learning Native Continuation for Action Chunking Flow Policies cites this paper.

Learning Native Continuation for Action Chunking Flow Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:40:08.704943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T12:38:26.522838Z digest=sha256:86781628358335fa0f170c5beb6706864f13b5526c3354a3277e5b52d47d7e9c

Observation 2b267576-52e1-4a81-b10e-890416b67a99 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:05:10.051106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T13:04:30.544504Z digest=sha256:6be5f9c9d1641c98ca1f69bee823b19274b4342fb40c2c49401565c7042fcf8e

Observation 408a266b-fe64-4d6d-a97c-0ced98d15a1f · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.978308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.978308Z digest=sha256:8cd1f61e77514e154f8d5122a16443ccf2b55dce1bccfb9b5db81762e3ebd987

Observation 97714748-41f3-43a9-b146-f3ad49458b1d · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:57:12.827282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T03:57:10.338617Z digest=sha256:c5bf430524706ae1d0c8f8dde9855744d8cc926ef659a43b8858ab5960be14b5

Observation 6a8e3e34-85de-437e-998b-533c131989ae · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.671785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:5e9544e3bb5dbad4bc3302812993c7e628aa3f001d10a62077a2ba3c63e6d236

Observation eb409383-3e7e-4ba8-9645-4bef12279702 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:06:08.666110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T06:05:52.461348Z digest=sha256:72707292fc4e812e51f4f9a56653ae4f9c235507937e26f6fa5ac41772e7a8ed

Observation d7a5fd1e-bd22-4e94-bdcd-8398898ae269 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.265667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T16:52:05.565650Z digest=sha256:66a81ea30f51d8bec2c8acc9c8f73a2e773b08c937f4b96b669729ade9e49991

Observation 3df18ac4-53bf-401e-94a0-bec8be992c4c · inbound

NAC: Neural Action Codec for Vision-Language-Action Models cites this paper.

NAC: Neural Action Codec for Vision-Language-Action Models VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.911642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T13:59:53.484305Z digest=sha256:aa1445982e093400b7c9b9307f978934364e9bb3d594e8894e25d2a67237ad81

Observation 29b1bc9b-80ac-436d-a26e-b6ec5eceff9d · inbound

ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation cites this paper.

ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:19:43.487289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T10:14:42.111428Z digest=sha256:1e5e93993cb70224e1d5f1a3a01929b1e53fec79f96649e66f8a173c107b7f76

Observation 64cce292-40e0-4212-8839-749231c2e2d0 · inbound

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation cites this paper.

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T05:40:47.306935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:40:47.306935Z digest=sha256:9ae394d05548cd2862ab37079fd34569f43efa277e05eba4c9a597fb69a68823

Observation f734bcda-218e-4a07-8473-f5a4d39067db · inbound

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots cites this paper.

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:01:26.608329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:01:26.608329Z digest=sha256:e22c0b5fe3e52d8d9e2a849d85c533e6cc2affd3c924c54f244d43c07c8bece7