Pith. sign in

Paper Citation Record · LEDGER

FAST: Efficient Action Tokenization for Vision-Language-Action Models

As of 5 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 100 inbound Pith citation observations for arXiv:2501.09747.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09747 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T08:52:31.686474Z

measured 173 of 173 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 312 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:38:42.966861Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact34
  • verified fuzzy33
  • unresolved1
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch3

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation bae3ca87-2444-40dc-b1c3-d16c9b308870 · outbound

This paper cites Dis- crete cosine transform.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Dis- crete cosine transform

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.405513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:e13ee48f07dcabb622362b954f7df29a6dc27e5ed2aa4d9a3a6525ac8bd1ee43

Observation e0d2a305-fe01-4fb1-88d8-8aa637f3478b · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:31.959920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:42cb191c9054fbc55da442500ea2285b71673fa9ceeabaf1fbda4197e24d8dc2

Observation 2c402382-742d-4233-a846-ace70612bfdd · outbound

This paper cites Minivla: A better vla with a smaller footprint.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Minivla: A better vla with a smaller footprint

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.386911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:27763c8180676356042c8a584ee8624d1654ca744e79dd94d473bc0664919664

Observation 7c4b5521-4ff1-435b-bf32-e5ac0ff4fe6a · outbound

This paper cites RT-H: Action Hierarchies Using Language.

FAST: Efficient Action Tokenization for Vision-Language-Action Models RT-H: Action Hierarchies Using Language

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:53:27.953855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:39b608d44b984908e590ebd2789fb3f3697ae4d033299082cbc0ab1cff8586f1

Observation bc76d6e0-6103-431e-84e1-79558827d5e8 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

FAST: Efficient Action Tokenization for Vision-Language-Action Models PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.987351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:cc62235b44d9d3f6aa6ea0813bfd3878322137d8adf1f92da0d1b081ec5695ab

Observation 4a6108e2-a318-4763-9cb0-824cf8f55913 · outbound

This paper cites Roboagent: Generalization and efficiency in robot manip- ulation via semantic augmentations and action chunking.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Roboagent: Generalization and efficiency in robot manip- ulation via semantic augmentations and action chunking

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.367864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:0b7cd81506e0431254bb184c88ae59848caf0cb6938f00588efaf11480ebd84b

Observation 0600c0bc-a69d-4080-8a38-1d103ef5b0e3 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

FAST: Efficient Action Tokenization for Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:31.885570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:296b3ff6eece545905f6f3a8bf9d8a7c2c6ca24ca563de2981017e88be651181

Observation 4ed60879-415d-45e0-b15b-cdbf7d349585 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

FAST: Efficient Action Tokenization for Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:31.895348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:89db29ba99e82651ccb99fe59bde86075cc36f68c6ab1c74c509fbe07538beaa

Observation e0b41df9-e0d6-4477-9494-b451fa7e1ab9 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

FAST: Efficient Action Tokenization for Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:31.904673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:4794a96fad2bff6af62e92322f4cb3ab9e8bb52e4adcc0ed4ea9cffef3d45524

Observation fc482c4c-3273-44e8-85fd-cce4d26bac5b · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

FAST: Efficient Action Tokenization for Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:09:34.215368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:07e0a046a6b83b1cee134311765b538774d2918d1b4e2cf56fc3b279490884db

Observation 0fd8edf6-6bcf-447c-a786-d6886f7d1814 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

FAST: Efficient Action Tokenization for Vision-Language-Action Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:52:31.926052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:6b6750632bb2ce9cccbf195ed66bb073ee664809c5c19a4ae05b7bc33b5bc3b4

Observation 836504b2-6e5d-48a7-9dc3-fa941533a835 · outbound

This paper cites NaVILA: Legged Robot Vision-Language-Action Model for Navigation.

FAST: Efficient Action Tokenization for Vision-Language-Action Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:31.934448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:0c199c505a32886fc7cf673ea3bf5e74a4ea40896e27b9c357d3497f97e84abb

Observation e63e21ac-9087-4699-806b-214448dc7f00 · outbound

This paper cites Open-TeleVision: Teleoperation with Immersive Active Visual Feedback.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Open-TeleVision: Teleoperation with Immersive Active Visual Feedback

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:31.940615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:8feb3a5889989979f630d7cdf2857ffda9557dbf77aca504b3bdba820be5c633

Observation e854caf7-f3ec-4203-b09a-61cc28501a40 · outbound

This paper cites Dif- fusion policy: Visuomotor policy learning via action diffusion.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Dif- fusion policy: Visuomotor policy learning via action diffusion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.276352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:fffe954738d3feadbd7da0c6be1f0e79634c00a7686801b3a1115f2694f66b47

Observation b526abab-ecb1-49be-b41f-bd5255d464d2 · outbound

This paper cites Universal manipulation interface: In- the-wild robot teaching without in-the-wild robots.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Universal manipulation interface: In- the-wild robot teaching without in-the-wild robots

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.294575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:2396981d8eefb19b45447b145997b9d810a6dd5b4b873b6b380890f2981d9aeb

Observation f542c32f-c3fa-4e53-993b-2248a17a4a93 · outbound

This paper cites An algorithm for the machine calculation of complex fourier series.

FAST: Efficient Action Tokenization for Vision-Language-Action Models An algorithm for the machine calculation of complex fourier series

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.311544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:7d28349ffea519dbd3c128056e05239bdfd018c6cdf30165d1b4a3abb570cbac

Observation c62ca3ab-b25a-4109-9967-2a984199dc1f · outbound

This paper cites Keypoint action tokens enable in-context imitation learning in robotics.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Keypoint action tokens enable in-context imitation learning in robotics

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.315767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:6d4c94742136ce3892afd432fbaf30b0e004347e04e18145c7794c9b5daa62e1

Observation 2ce0f7c7-e664-495d-ad2e-176afc93a187 · outbound

This paper cites Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.319470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:8afe00b4999927bfeea43aab3b5bc9da890dfeb0e4820c05d214229d26f86380

Observation a54155e5-b5b8-4145-9dc9-ca23203faeab · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

FAST: Efficient Action Tokenization for Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:31.964936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:e16e616f490bbd26ef6e33452b04def27d3cbf4fdb5b87bde23d515da514ef08

Observation e4f9e417-9a48-4503-bd8c-7f4dbf7b5d84 · outbound

This paper cites Tam- ing transformers for high-resolution image synthesis.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Tam- ing transformers for high-resolution image synthesis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.330179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:cdb491cf951fc0b45f449ddf9109eace149c72cc41d73fd06267ffe7c985cc37

Observation d99d2fd7-e9c8-47e6-b7f8-9b1e24cab9ab · outbound

This paper cites Qi, Yin Zhou, Zoey Yang, Aur’elien Chouard, Pei Sun, Jiquan Ngiam, Vijay Vasudevan, Alexander McCauley, Jonathon Shlens, and Dragomir Anguelov.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Qi, Yin Zhou, Zoey Yang, Aur’elien Chouard, Pei Sun, Jiquan Ngiam, Vijay Vasudevan, Alexander McCauley, Jonathon Shlens, and Dragomir Anguelov

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.333571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:c72873ad46bfd509d4a6e6535aa6b13cde1635ee9921f2ff4a713c480f616591

Observation 0aee9769-5477-46f4-aeba-c69181368476 · outbound

This paper cites Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.337276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:753665d362c8caec4bfbb07718d411f903b1ff8d5b08d7ac0dbe224e1bea6ad3

Observation 74ea73da-19d1-4b11-8546-fccdad4d667c · outbound

This paper cites Moka: Open-world robotic manipulation through mark-based visual prompting.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Moka: Open-world robotic manipulation through mark-based visual prompting

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.346454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:7ba7dcc9917b51a2d098a9833185e14346e023f5ed009c65278166125ff455b8

Observation 1de2e5bb-5534-4c44-add6-9afae8b7ec89 · outbound

This paper cites Humanplus: Humanoid shadowing and imitation from humans.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Humanplus: Humanoid shadowing and imitation from humans

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.350837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:8af9c907926d4faf6e8d65ceb5fcba6cc768d2b0a5f127dad9b8472578438289

Observation 139b7318-649a-4636-997b-a44d25034a95 · outbound

This paper cites A new algorithm for data compression.

FAST: Efficient Action Tokenization for Vision-Language-Action Models A new algorithm for data compression

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.355125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:082b635e07ab76e251a595ee938d9b279a9cc3217f6162161eab5d58bad5982c

Observation 9323ea76-cf37-4d3c-88f2-e2d062098c6d · outbound

This paper cites Multilingual Language Processing From Bytes.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Multilingual Language Processing From Bytes

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:56:20.888260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:97445e2ce6ed5dc6b06e2c41f039acf3a2a41014c325ca8b3ae46a9f7b31cab4

Observation 1887f98b-1110-4843-aaf3-33a3be1c7998 · outbound

This paper cites Singing voice graph modeling for singfake detection.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Singing voice graph modeling for singfake detection

Reference 29

Resolution
malformed identifier
doi_truncated, observed 2026-05-11T08:52:31.815851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:9b0a5ade6a22915854dbcfb1a785ecf6e5e4f884b67449d921144c35e67e8853

Observation d374ff46-e885-4174-8c12-8a6ca347d6e5 · outbound

This paper cites Bridging the Human to Robot Dexterity Gap through Object-Oriented Rewards.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Bridging the Human to Robot Dexterity Gap through Object-Oriented Rewards

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:31.976627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:8c9897600161a92bbd4de1cd96204924b4c9a2af7d639fb5766a5fa662bf0d60

Observation f3ff5b79-7cd0-47cf-b94b-fc97f1c58d63 · outbound

This paper cites UMI on legs: Making manipulation policies mo- bile with manipulation-centric whole-body controllers.

FAST: Efficient Action Tokenization for Vision-Language-Action Models UMI on legs: Making manipulation policies mo- bile with manipulation-centric whole-body controllers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.376947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:ee37e0bcba456632ffc282e7f7290e602156938232d5104cc62f34b7785d021c

Observation 1951f2e3-cbfb-4a6f-bd29-26b41b2ad0f8 · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

FAST: Efficient Action Tokenization for Vision-Language-Action Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:25:18.388921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:b5eda0e22636849d94f3515e91c44e042b22828b2649ce1ac56d0111efa4b990

Observation 930bdf6b-cfc2-4353-b922-9366a8ef917a · outbound

This paper cites an unresolved cited work.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Unresolved cited work

Reference 33

Resolution
verified exact
doi, observed 2026-05-11T08:52:31.831709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:a1ff85ef24e7765201e2ee24e1d6fc7f7b53d3c5a33af8cfdedf680fbdcc22f6

Observation 70bc0a3e-e3ef-4560-86b2-de67a24c6e60 · outbound

This paper cites Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:31.993969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:b9f433c32e87e611c145a97a7b84f9b6a4222480637dad7e03e10727939861c4

Observation ca0b98e2-7f33-4f5e-b31a-8f295255e97a · outbound

This paper cites DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning.

FAST: Efficient Action Tokenization for Vision-Language-Action Models DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:52:32.000886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:cb6fb03dac5e3f52c97240bb4cfac0fa1e222d27b75212a6325a811e3bc28408

Observation 17831756-63ec-4a95-9488-fbbc611cb21f · outbound

This paper cites Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.008245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:49a3bbab976a8bec95b83b4ca2ae095f6d74054e859d7a79b37c21a594071c08

Observation 5d086ce1-2226-47e6-bb46-f50689720485 · outbound

This paper cites Pris- matic vlms: Investigating the design space of visually- conditioned language models.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Pris- matic vlms: Investigating the design space of visually- conditioned language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.412738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:ad38b68d9a852eb62883f8a0ec1f20fb9f3054300110484747089e4c90d17985

Observation d11cc28c-9a22-44b3-b2f5-50a8e4853b89 · outbound

This paper cites an unresolved cited work.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-11T08:52:32.420266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:e912e83e2be2a0430eff23cb0edd75c5864cac85e86238bd29eb61f9c9472921

Observation 88da5c10-2c5c-450b-adec-a4abfa1220ff · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

FAST: Efficient Action Tokenization for Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:32.015560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:dbbe1d9cfe29058d382b5e66a094e5d6691ad1512f3cbb4d03af6db2634eeb42

Observation 01f6c475-9287-4e2a-950d-bbd4ac6a2fbe · outbound

This paper cites Action chunking as conditional policy compression.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Action chunking as conditional policy compression

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.429969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:1801ea425a16592094cd3d0dd7214a370a32de810927bc9d29addc2c15933b7c

Observation 5603573c-d0d0-49bb-aada-44aa7a7651f6 · outbound

This paper cites Behavior Generation with Latent Actions.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Behavior Generation with Latent Actions

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:52:32.024830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:c137e43d28e2545c98a4de8961f297ffdb675c49379bccddc6e01cf91035d375

Observation 773e604a-216c-4d02-8af2-93adc7485486 · outbound

This paper cites Learning Visuotactile Skills with Two Multifingered Hands.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Learning Visuotactile Skills with Two Multifingered Hands

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.029968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:b422967af6f454c48890ccdb0f88a13caf5b9f0ec4a2c7cd9f3cbdb0fe85cf97

Observation 5d35112b-3e22-4330-a9ef-823018900679 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.192350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:b5e059763727d81eb4f268d5d70abc7059ff5d0316496a9940219372b47b199c

Observation f40da83b-c9a4-4e13-b4c3-1df1f3ac1ec0 · outbound

This paper cites Visual instruction tuning.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Visual instruction tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.199076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:02a603656b24454d0bc84c92c39b45e9f8bfd46ea8b5ca01d57b11da3ccf9031

Observation 56bfbb44-3817-4bd2-ac75-64ef67e9f81a · outbound

This paper cites Decoupled Weight Decay Regularization.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Decoupled Weight Decay Regularization

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:32.037835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:a6b850f7a132cb2b81afb5571ea8faab273b1ee3bcb6989cbf1759878e74c968

Observation 98c3cb93-c186-4589-95dd-d5c95f24ccfc · outbound

This paper cites Serl: A software suite for sample-efficient robotic reinforcement learning.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Serl: A software suite for sample-efficient robotic reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.218349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:a0f93ab695888c19c4ec38b24282eeddb20cbf64463e091bbd452b03b7110841

Observation 87371680-dead-4f95-92a2-3784f2bcfecb · outbound

This paper cites Roboturk: A crowdsourcing platform for robotic skill learning through imitation.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Roboturk: A crowdsourcing platform for robotic skill learning through imitation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.237348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:56085907ce0843b21330ee55c95e51a1a14c00efbefe004066fb3765e05da231

Observation 559e37a9-ac78-4c21-a367-b1a0e981aad3 · outbound

This paper cites Finite scalar quantization: Vq- vae made simple.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Finite scalar quantization: Vq- vae made simple

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.246351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:d29c11383545eae3d094567e015439b5284a62a1c4a9120f806d886875038706

Observation ca6e0a4b-3e52-4a6f-aca0-5036622043c5 · outbound

This paper cites QueST: Self-Supervised Skill Abstractions for Learning Continuous Control.

FAST: Efficient Action Tokenization for Vision-Language-Action Models QueST: Self-Supervised Skill Abstractions for Learning Continuous Control

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.044659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:841e5df8226d5fd43dd2c1062fb484680307c7cb6328f61c3fbab4a74eee9096

Observation c9591172-98da-4c6b-981a-ac2101d7d1e0 · outbound

This paper cites Pivot: Iterative visual prompting elicits actionable knowledge for vlms.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Pivot: Iterative visual prompting elicits actionable knowledge for vlms

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.268353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:3e98f7936ef2dfba4d7d7fdb3fd2f44ec52085b4d7d91cbe54b5b32e58278226

Observation a516bdb4-14b9-47b2-86c5-abd31b7ec9f3 · outbound

This paper cites Octo: An open-source generalist robot policy.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Octo: An open-source generalist robot policy

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.326348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:a0e01fb84df8ba217814ec34b7bc256474081b7ba6a22a7d5273bf4cc6f9711e

Observation 36b4bd67-6172-4190-ad77-550a7bb25457 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:23:25.503833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:74adc837e103839997215c3d7a9ae228d52678743f7b6b43084070c36cd8de85

Observation 4a4e7967-2e5d-4712-818d-959e767d0ece · outbound

This paper cites Byte latent transformer: Patches scale better than tokens.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Byte latent transformer: Patches scale better than tokens

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.373299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:6f7c0104fd4f2f3875de76506d5cfa0dd1dbba36ba8a2cfb5798eee47b9b31ca

Observation 788e0cc3-a08a-4a69-b4d0-b08df8f32c1e · outbound

This paper cites In-Hand Object Rotation via Rapid Motor Adaptation.

FAST: Efficient Action Tokenization for Vision-Language-Action Models In-Hand Object Rotation via Rapid Motor Adaptation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.060952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:86d2a78cfa1f7ff93db00602f3b9b0eed24fdfe041cc845de6f280623fb751de

Observation 5ab9d296-bab5-404f-8188-319f34c3b3e7 · outbound

This paper cites Language models are unsupervised multitask learners.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Language models are unsupervised multitask learners

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.391035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:0d4c575abac8e6bc6de82c2655d5cef576b6ac05edd0fa0fc063d48d53b2673c

Observation 1f02c590-d58a-4d6e-84e0-6815ecfb47a5 · outbound

This paper cites A generalist agent.

FAST: Efficient Action Tokenization for Vision-Language-Action Models A generalist agent

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.398230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:3123d80d84bc74cae4e8968f4328a256b0b2a056062c7f61b6c446de3895b818

Observation 30d41e1b-2759-45e0-acd8-0a23b752066e · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Neural Machine Translation of Rare Words with Subword Units

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:15:20.428834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:532fad1323d378ce6277c2966010ab88f9bf7c5ec1e9e40131ccdeb26fddf894

Observation 5ce49807-3e8e-46a9-87c7-5ab79d51082a · outbound

This paper cites Hand-Object Interaction Pretraining from Videos.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Hand-Object Interaction Pretraining from Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.070856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:9471a7c3ef2200f8de808917dc81b1cb2ae25827960d51641fb9fed8b2ae7d5c

Observation 136b9880-42d6-40fa-ac61-9d3b0c7f47a5 · outbound

This paper cites Neural discrete representation learning.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Neural discrete representation learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.180700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:e43dda09249c8ac42a228dc9a9e58676e1d1928e7ebefa987d186c3e42d84281

Observation 7ac2cf8b-b9e1-4fdb-81cd-fbd253706109 · outbound

This paper cites Neural Discrete Representation Learning.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Neural Discrete Representation Learning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.079345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:abcfb8959ff6ed979d8c9c72164939471b4f3630028219a6e1c853dec9f0e2a9

Observation 5d493849-b62b-42ef-9b98-f67b871fe005 · outbound

This paper cites BridgeData v2: A dataset for robot learning at scale.

FAST: Efficient Action Tokenization for Vision-Language-Action Models BridgeData v2: A dataset for robot learning at scale

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.204338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:c20d16ec472f994010b0c4ad38f56e84499e7db2eddaea8905875d296908f4af

Observation 62f4316b-b1af-4955-81aa-85293122d50e · outbound

This paper cites The jpeg still picture compression standard.

FAST: Efficient Action Tokenization for Vision-Language-Action Models The jpeg still picture compression standard

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.252985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:e4a0dc41c0edfc213d711d2439c7d686f51f0de813bbe0398067827ac953ce9e

Observation 0ef39012-e3bb-40f7-88b8-b72e348e123f · outbound

This paper cites Scaling proprioceptive-visual learning with hetero- geneous pre-trained transformers.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Scaling proprioceptive-visual learning with hetero- geneous pre-trained transformers

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.362588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:bd783de2e02cd6146643c891b9877a61702164e7f5d5441f8ad497a2b44f005b

Observation c5216e5f-b39a-4f44-83d4-f866473f92fc · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

FAST: Efficient Action Tokenization for Vision-Language-Action Models TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:12:26.208161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:ac58cff772a0e6b5794014faaccde311706df0dbdda5709265aed0875ec486ca

Observation b1eb95a1-be23-421e-96f6-83b3bedfc13b · outbound

This paper cites ElasticTok: Adaptive Tokenization for Image and Video.

FAST: Efficient Action Tokenization for Vision-Language-Action Models ElasticTok: Adaptive Tokenization for Image and Video

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.096052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:73eb2a6728f5200c1e78970670155b6e7d9070b3c18fa755ede0412961b8a37a

Observation 57cfbfcb-8c20-4ce7-95b9-c5e853cb1f70 · outbound

This paper cites Latent Action Pretraining from Videos.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.103559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:a048582f8e7563cd0d42fc57b2381a95152ebc64518402a94f43cf92ad64b676

Observation b8ac0df3-9642-4073-aa10-fa8153dde81c · outbound

This paper cites MAGVIT: Masked Generative Video Transformer.

FAST: Efficient Action Tokenization for Vision-Language-Action Models MAGVIT: Masked Generative Video Transformer

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.112802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:2a216b3d5c12566f11aacd0178becb1b3cbd681e89c6473c6e16ed64857b2273

Observation 4f83859b-9f94-42aa-8527-aa8ca9156103 · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Robotic control via embodied chain-of-thought reasoning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.424610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:e65e56b1dba2854fbd934848929956b30d1346b140fb4da6155971528e9cf526

Observation 79955a72-0acc-47f7-a02a-8a740e4e8185 · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec.

FAST: Efficient Action Tokenization for Vision-Language-Action Models SoundStream: An End-to-End Neural Audio Codec

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.118854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:03e804a5300da8dc31ac12331b1d07ef47105ad3e194f187e4567a1b6aeefe61

Observation c6a7f4cc-755b-43a4-b328-2f3589bc4b7f · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:32.132745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:511bbc61e2413113242a73761c2a7b394197fbec593ebdc5db007fb9d4c8e5cc

Observation c58671ea-6b25-4182-a6a0-aada05b60156 · outbound

This paper cites ALOHA Unleashed: A Simple Recipe for Robot Dexterity.

FAST: Efficient Action Tokenization for Vision-Language-Action Models ALOHA Unleashed: A Simple Recipe for Robot Dexterity

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.146345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:94306753e4635169d05404a4aa167def7d1717ed654b177d09e72a3df2835baa

Observation 74032957-6ba3-4cf5-b1ba-8fb2cc4d88bd · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

FAST: Efficient Action Tokenization for Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:eeb050ac054b4ecd3b2336bd147d211e4f4d4a4bdc029218c7e2dc347f71ab10

Observation 1b72a23b-1c38-46e7-b57a-0998e8f618de · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

FAST: Efficient Action Tokenization for Vision-Language-Action Models TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:27:23.210785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:212538de0d9e692a326d8ffa07bd82ff6f57a6e1f826ddba6b81d3a3fa3b5bd3

Observation 6e150253-6d71-4ff0-ba0b-2ab9035b8ed9 · outbound

This paper cites Autonomous im- provement of instruction following skills via foundation models.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Autonomous im- provement of instruction following skills via foundation models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T08:52:32.381099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:1d295e682a349871789c455ef1a1f0b02387348a099f2b735009f5b50a01f426

Observation b6d45ae7-00c2-46b6-840f-b972d00b3a48 · outbound

This paper cites Compression of individ- ual sequences via variable-rate coding.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Compression of individ- ual sequences via variable-rate coding

Reference 76

Resolution
malformed identifier
raw_fallback, observed 2026-05-11T08:52:32.185722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:73f11a97840b986498d45d6920e793f8ab521010af6e58fc5a396d2890e97172

Pith citing papers

Observation 3c17340d-0043-4f36-a7e6-03d4a668d1f0 · inbound

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model cites this paper.

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:12:19.840291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T06:12:19.643111Z digest=sha256:c881fa5cfd93aa3e732289f173cd2156cf3a8945e1d64c1b041cde2c293e821a

Observation c177e2c1-4370-407b-86d3-c32632cc9b07 · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:48:49.064333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:61c88091c1498363410425de61afe3268e180dbde7d325dafd48468e5e1e762d

Observation 432d5666-6801-4c84-88bc-4e89fa89b527 · inbound

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models cites this paper.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.233298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:1b15058201f006fc8fbe1519c9a12a8cb1f5417ba3866ada9399c32ed7402dfd

Observation e1190dcc-6d96-4cf3-96ff-2a9be3cf739c · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.433811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:c66d800ea835808fb8d76b9431045e55a72997cc225b7bce7572af253fdfee01

Observation f78da1bd-03f3-4eb1-9027-821ab3f6391b · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:32:22.553145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:b5c3d132e276eaeb44f2aa4dd20fceb9dafce4aefd8264e728e8f38d87ab6901

Observation 0b31a95a-8506-4798-b973-785e19acce9d · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:00:48.842942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:fa2eaf2ff59b17fc22cd7689c86ecf9c7fc1ab729d28680d67332937519a4f36

Observation 20008719-ae60-4e23-baf5-b47dee0db4db · inbound

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks cites this paper.

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T15:53:29.476407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T15:53:29.412890Z digest=sha256:80322894e522190aa09cf2e956613b486bbaeec437d3c37cb97556b41d355965

Observation 74dd22fd-a4a0-41d6-a649-ec3271b9cf9c · inbound

Policy Contrastive Decoding for Robotic Foundation Models cites this paper.

Policy Contrastive Decoding for Robotic Foundation Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:11:38.557820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T14:09:48.762737Z digest=sha256:6c269dbaa1a10889a452d741b0f9d0d6018e5e795f40611babb146fac4c0ea54

Observation a6e56e9e-e96c-40e7-83bf-857197eb6c07 · inbound

FLARE: Robot Learning with Implicit World Modeling cites this paper.

FLARE: Robot Learning with Implicit World Modeling FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:59:08.897738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T15:59:08.846629Z digest=sha256:e882bfb9f42a1f8983236156fc3dc7c8b70215077ab68ab5c576b6be483501e6

Observation a523c01d-a6c5-48eb-a79e-506e999fe7c4 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:25:47.254494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:02481ff07565bd1ecbc4289d06fa0b4d3c80993b02fbb645843c8de36b5f0e0b

Observation 6990683a-0047-4f84-bd40-d34b0f706960 · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:55:40.496669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:4febdae1b295626d42f66a6c56f62ca0ce0daf321d6083b71e87547e8be4b504

Observation 14c9ae98-a376-491f-8152-377b49fbaa14 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:22:37.500199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:f05e8119f74dbb34ffc0b890287ef771ea499fef11321c45fb61f66051dfbb86

Observation e8512285-7216-41cf-9287-43949bdc2555 · inbound

Real-Time Execution of Action Chunking Flow Policies cites this paper.

Real-Time Execution of Action Chunking Flow Policies FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:18:51.648607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T14:18:51.613045Z digest=sha256:34a7316fb4c6abf12607f71c3b6511334a1719237659d2849df46c5d58546e8e

Observation 553badf2-5ebb-4667-9690-cfadccfa0be3 · inbound

AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning cites this paper.

AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:46:44.180673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:46:43.955825Z digest=sha256:103e3772b288363400103872ae77e7e262704a1d03eca588bb14bd1215854658

Observation 4330f603-7bf6-4bb5-93a0-158667a91944 · inbound

WorldVLA: Towards Autoregressive Action World Model cites this paper.

WorldVLA: Towards Autoregressive Action World Model FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T22:57:08.224546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T22:57:07.883617Z digest=sha256:ef65c940c822f291995c5c90f15c65bfcd68a181d57fbbe3ff5ab5d439a99336

Observation a189148d-5bb7-4167-b865-d57482b0d284 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 126

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:08:35.401768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:6ce8130dc326dfff625105f95131d6206a0b1ae5fa9c76f714d0c56fdfaf761f

Observation aa644d30-0dab-443b-9c0f-b9cda949ddd7 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:42:41.493612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:f88b767e213a27afe797ad06d6d0d66311cdbf23e93521a847fac2fa0d39547e

Observation ddd62206-4c0b-46d1-95b6-60a2f4ee14bc · inbound

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation cites this paper.

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:52:57.504899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:52:18.984005Z digest=sha256:e04fb802745bc57499fba0339b0c79570f3d5382c42d3f2af895835a5f01d7f4

Observation 5130156d-0f88-40e8-bed6-72e2c2798d39 · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.517032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:b26d3513d181d2c68276e83e58b1b4d07cd66dbe6758acc16582f066bb76ee34

Observation df4247d9-d154-4a04-b7b0-c36570efb474 · inbound

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models cites this paper.

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:52:03.141334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T21:52:02.893886Z digest=sha256:ea46ed52a718230f78128e784d724269533d453041c159fde72cc464d95f859b

Observation 957ef696-5bca-4bed-9f3c-ac4dd53f3b93 · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:45.915628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:45.915628Z digest=sha256:5335cb4e82df3ad960cca74e7c9c2d226acb43dbe156b95fb89c56d9619a3f97

Observation bd557c63-484d-40f1-856b-a2b5012b244e · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:42.966861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:42.966861Z digest=sha256:6a0995908dc66d494b881905dffff56005fcc8eb9d42f643b4f5eb8eb3bfbc10

Observation 61a8e01b-d9bf-4685-9cb5-d77a6979d8e9 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:48.143380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:48.143380Z digest=sha256:943334e2c46b6037d8b74298d86b425f6b3369c2060acdd28aa374145a573682

Observation 4e33c551-7fab-42c5-bc95-867a4a125250 · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.189100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.189100Z digest=sha256:7821c6a7831d3b3538800a6fbe600a34b1779fc99edf1161d5562f70c2a01a47

Observation 4c7d7538-4911-4813-b2db-7f81eeaad77d · inbound

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:43:24.519745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:43:24.417901Z digest=sha256:6212195bdfb4d0628068a9204cfbe03bdf023eaaca78e0e2021ffcff82336918

Observation ecbac68c-3f86-4700-a0b6-05c1aa1d1bf1 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.725953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.725953Z digest=sha256:2a9b5da00f3a91c33406d5169c79e03f90ab38ca2d054b2d60870490dbd3a4ea

Observation 99d155fb-1462-4ecf-86ab-dfdaf40e4073 · inbound

Mechanistic interpretability for steering vision-language-action models cites this paper.

Mechanistic interpretability for steering vision-language-action models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T13:50:17.405514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:50:17.405514Z digest=sha256:69565360b3f58bfb2bb5662416db38e8dafc9c818a3da7ad9f57fd2aa96378ca

Observation d71fa393-e759-4444-b8d8-e12ec46b78aa · inbound

Galaxea Open-World Dataset and G0 Dual-System VLA Model cites this paper.

Galaxea Open-World Dataset and G0 Dual-System VLA Model FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:31:09.933298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:31:09.933298Z digest=sha256:30a241fb61b4cdb000295addf3da7784031d18c252c721c9e8792245529a7ff2

Observation 87117c1b-5aed-4d9f-a546-1a2317f3df49 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.300757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.300757Z digest=sha256:5fb568699b7055d054119b5204e31e8540abc74025d9bbe2d4d701b1ae9ccbb4

Observation 439805ee-919d-4ebb-ad27-10464cb11206 · inbound

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies cites this paper.

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T05:48:48.107487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:48:48.107487Z digest=sha256:60eda20671879281ea656b293220a5bee4ef737b3fdb597d4502fc084ab9e0e7

Observation 5f748370-c075-4b89-bc33-0ce4e1404976 · inbound

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions cites this paper.

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:42:47.062488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:42:46.978148Z digest=sha256:ac6d7a5980af346bd698f0cd04b009d5708efefe7792d7a4025226bb42c90a73

Observation 46dee853-bb34-49b8-b5b4-1db200593787 · inbound

TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models cites this paper.

TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T21:29:09.427398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:29:09.427398Z digest=sha256:4fdbac76c10a3e62c512c2b5424046121a8c27eb2e1802d0fc85242d3d735542

Observation cff1b3eb-0556-428f-9437-8e359a293b07 · inbound

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation cites this paper.

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:10:13.383967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:10:13.383967Z digest=sha256:a8e9ea476e383f1f7695aaba0fecc46215e03642ece26d7e48dd0a1d67a4f9aa

Observation 90e33f91-4bd9-483b-87a4-aa566715af39 · inbound

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue cites this paper.

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T16:17:25.976904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:17:25.976904Z digest=sha256:092bb793d63e832c33f24cce8d7b37e84a059bea0e8d07bfe3aa811bfdfb2d44

Observation 2e7e6cae-fae1-4465-9325-d976b76eb960 · inbound

VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation cites this paper.

VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:51:25.745416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T13:47:30.943307Z digest=sha256:562a62e09725c95214e85e011b206972ce41fa76a5b29ea3f1122299f9ca8dfe

Observation 193d796d-07f2-4d6b-aee5-002dfeeb6fd2 · inbound

World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training cites this paper.

World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:51:23.523541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T12:48:32.123998Z digest=sha256:0171b2c005e9a466bf0317b2e6ad33dc868699a93d88339692eeda4d1aa172cd

Observation 0de6fb67-aa8b-45f4-8da8-a62bb926f755 · inbound

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations cites this paper.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.522955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.522955Z digest=sha256:78bb89746747bc0935c4eac9409f1df539e4e62cbcd0d2d9b6528a432374fdfe

Observation 287b5301-89a9-43ae-9f5d-424e6a6fde02 · inbound

INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models cites this paper.

INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:59:12.684224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:59:12.684224Z digest=sha256:5cba387298cb4a74d57c12808277ca7af65166112b81fcb10a0d86fe3358a3a9

Observation ebc4c4b4-72dd-4335-bef4-e9d2bf973cff · inbound

Contrastive Representation Regularization for Vision-Language-Action Models cites this paper.

Contrastive Representation Regularization for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.056099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.056099Z digest=sha256:ed7a19c34c8fe3b460cfcc0a72d3d9c9d5614166639566a9eee3e548bc1cc387

Observation 79e7768f-f511-4db8-a540-5ded61693a00 · inbound

Verifier-free Test-Time Sampling for Vision-Language-Action Models cites this paper.

Verifier-free Test-Time Sampling for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:20:12.458075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:20:12.458075Z digest=sha256:8e6cde08081fbbb24eb82e02b409b81ed67be06d50de251fccf945bf5d1a3668

Observation 62bff911-0d1e-4e43-bc97-ef5521b9c79c · inbound

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation cites this paper.

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:36.206030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:36.206030Z digest=sha256:7d681e710cd413b99345a40165d927fd20f45795648dbb681bc02a9d819e7911

Observation edf9e0ab-06cc-4ebb-9615-0da54e91ebae · inbound

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models cites this paper.

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T10:24:57.896357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:24:57.896357Z digest=sha256:16f2ba7263d0026f3e14efcae156d011a74f23ea844897e44bc568429550cf3a

Observation e68e9b60-6f35-4c6c-8d4f-e658c700d96d · inbound

Ctrl-World: A Controllable Generative World Model for Robot Manipulation cites this paper.

Ctrl-World: A Controllable Generative World Model for Robot Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-16T01:14:10.514253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T01:14:10.174044Z digest=sha256:e365b67dcebfe3194fc974c1b233986a7816c7c545a4334d47b0ac28734c7869

Observation f0ab2c21-24d0-405a-a98e-ed8f47ca68e0 · inbound

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving cites this paper.

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:48:01.082599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:48:00.943591Z digest=sha256:1de1345e6075338c19317b201872e98e8a33f4cd6878043a080701c21d41048f

Observation 39461b7e-6354-4ea3-aa17-8eaaa2a13d30 · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:09:39.861416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:f429efbb91c2a0e5e8679c6190d9eff1b0c8b51cd9fc338ee430d94e98092750

Observation a1a160eb-83f9-4afc-b7b9-b8bc5f6f352d · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:14.854816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:14.854816Z digest=sha256:3fa80a732d39c04f576c2fbbedddb5ea508a44191a8feb10c5a651e2596b1544

Observation 72f02667-d3bd-4e7a-9ff9-d0d050a2785e · inbound

Speculative Coupled Decoding for Training-Free Lossless Acceleration of Autoregressive Visual Generation cites this paper.

Speculative Coupled Decoding for Training-Free Lossless Acceleration of Autoregressive Visual Generation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:20:48.684648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:20:35.021120Z digest=sha256:eaf3015e119d3a932fec981296f7068217c291015a58e0ce5ebce2be49c44b06

Observation 90750e1c-e854-4de1-8a6f-05946be50bdd · inbound

LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation cites this paper.

LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:35:27.906366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:34:55.240907Z digest=sha256:9f5edb077a698722571c806ccd260f4a52607ece58a81755d064aac614483233

Observation 0b6dfa7c-d79a-441e-b759-72be7aa7a3ba · inbound

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation cites this paper.

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:25:32.989265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:22:41.611893Z digest=sha256:32ca49335cf96a4a3b51092091e2806b02dede4cec84f814d3d2e7ca32e228e4

Observation 00c6d782-afb3-41f0-96ca-d67350f92265 · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:30:18.279960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:f5cb34185627c2db95c2f5c4c0349e06ee0069f2538426ac7b35cd2de62c4f10

Observation 5e0fcc0a-341c-4338-bf45-cb8dd4e1cd8f · inbound

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding cites this paper.

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:04.607173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:21:12.375936Z digest=sha256:4a8766b6878e05609b3685bfa31984c23f7f1fcf88d843627ce5a348ffce5baa

Observation dc291951-9c29-49b9-b360-0ed466e93021 · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:52.110751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:52.110751Z digest=sha256:500a255be735cd27a686b2d6cb6453b2a150182c58f3001cf72753b468e9f8d2

Observation 2dd1631e-8d20-4d39-a5ff-adaca0501b58 · inbound

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention cites this paper.

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:29:09.932906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:28:22.652509Z digest=sha256:34684ef1b102228c27b57bf7aa9424da5222680ec5a9956b4c7bd5829c5f1711

Observation 28178c42-aeb1-40bc-8033-01a9d1997d73 · inbound

Mixture of Horizons in Action Chunking cites this paper.

Mixture of Horizons in Action Chunking FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:32:28.585704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:32:28.585704Z digest=sha256:6576c2872ca8c7d803c421c701d97d0a0f332c767082d8cdf7bb3311ff25271d

Observation a8884fcb-8c89-4f58-b5b6-ee2dd654999b · inbound

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision cites this paper.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.802389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.802389Z digest=sha256:449b6619a94a660f2c9e3f407426088a1948234acb240c99e278278a251e45dc

Observation 1e649f92-2d80-4e7b-b914-4925b27d864e · inbound

Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding cites this paper.

Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:31:13.375451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:28:35.576661Z digest=sha256:45f7e0170f9c755d350daa0c3c4f0579efed7f71dc2c6babe5d780e96dcecf4b

Observation c65aa1f7-b389-4017-92a7-465d9e25a3f6 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:26.092436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:26.092436Z digest=sha256:27cd7471d2f014ce577bfdd210db9a27f6076fefe3ab512453cc3cfdf5a3d49e

Observation 0065bab7-4218-4845-b536-8c89d63e06e6 · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:47.926004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:47.926004Z digest=sha256:b04202284593304bb66980b2b8bfac41b5f7fb12a1c60d27f0c6618e7ba5384a

Observation 0d790f23-2631-4d05-b551-82bdf4dfa8d7 · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.846355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.846355Z digest=sha256:24122faccf7e4fe3365d948a075eb6f4066031af12411bbce1f6c0171de36acc

Observation 063ca814-7ea0-4958-ac76-c7a1005d8d41 · inbound

Stable Language Guidance for Vision-Language-Action Models cites this paper.

Stable Language Guidance for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:23:05.294401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T16:21:18.449987Z digest=sha256:c754f90c662a153a7900f2b18c3789a27e3fa25616b17b7f337e04cb0a1672c5

Observation 2ac26e8b-edd2-4aed-bd47-aa1b694e3992 · inbound

DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning cites this paper.

DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:07:50.693252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:05:08.804472Z digest=sha256:80a95f3d8bc52c9fba1bd60fe7c3243a42b00dd29a9de893211837fa9c0d2c90

Observation f6e8c49b-1f4f-41e9-9f9b-b548a1219b0d · inbound

Supervised Mixture-of-Experts for Surgical Grasping and Retraction cites this paper.

Supervised Mixture-of-Experts for Surgical Grasping and Retraction FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:30:48.168225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T09:30:17.251223Z digest=sha256:2d7a7dee0834006963fa2cb56e718e75aaff0b2df56cde1cba41df36ca33864b

Observation 9c8d4159-35f2-4fb4-848f-eb7f6e40c2e7 · inbound

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources cites this paper.

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T06:50:22.904022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:50:22.904022Z digest=sha256:a0874a5b708c1977408c70410599600cb3b2913b6f19b1c7ae1457b0880e8838

Observation ca02669e-59e7-4cf2-a841-51dbd3fae137 · inbound

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning cites this paper.

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.391086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:07:10.387869Z digest=sha256:250f13ad279aa4626f15d2b7b2ecf3b7f2c1a1038cf7df9ee555762db41f1543

Observation ae4d161f-6ec2-4318-80e5-328138cf44c1 · inbound

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning cites this paper.

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:12:11.859871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T03:11:52.645633Z digest=sha256:7208632afe08899cb4b2b617a5df808473325edf1d448061a849bd883cca59d2

Observation b3be4d23-8927-472f-9f06-f2782764af10 · inbound

Learning Native Continuation for Action Chunking Flow Policies cites this paper.

Learning Native Continuation for Action Chunking Flow Policies FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:40:08.745158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T12:38:26.522838Z digest=sha256:9eaed4e5632ee19ecc20fbd0039a434e0b8743548fe842014cf1ca291951417c

Observation d38a70cd-7f3f-4f91-ab79-10108d364c10 · inbound

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control cites this paper.

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:06:42.961391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:05:39.797848Z digest=sha256:c9715e22b3c6c2c5ef0d51f5a6f5f038ed9ae26108be1dbe5c2566d3ca856d8c

Observation efdb2e1d-c66e-4073-a413-99ed4ad9b070 · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:28.981293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:28.981293Z digest=sha256:c7fe1524a56c9d4f60e0b113103b739b55f19bc5a4f56a7e96fa1cea30be0fa6

Observation c2198c2b-e888-4d1f-b497-b45c68da367f · inbound

VLANeXt: Recipes for Building Strong VLA Models cites this paper.

VLANeXt: Recipes for Building Strong VLA Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.805192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:a3106ab351b6a933c5f4c891906f5671bb30e83eed59455a93b81d833b425f41

Observation 185ec468-d8ee-497a-bc6c-2be3224e419c · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:05:09.935677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T13:04:30.544504Z digest=sha256:5c10471263731fde8d9d8eca1ec641f1ab4cec0efebed1b4fde57c25c844ad33

Observation 58dc68ba-727b-4dc8-a9e4-3b0ed97a74f7 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:13.747288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:13.747288Z digest=sha256:e5691421506c5a83cd4fb43aabc754f1bb7e7ed1f8e88e32fd4fcfea1db8e15b

Observation 23e97b66-5306-47be-8230-9a484a30460b · inbound

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation cites this paper.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.167016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:a5ed6f49403560dfc4992adf6859e0378bf317f35524c6afaf5921e853ab5dbb

Observation 9a2440b7-8291-444f-a112-09c142415ee9 · inbound

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models cites this paper.

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:20:17.298081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:20:10.435886Z digest=sha256:adce303149d0fd49c8e83aba8ed6260324059efdcd992f629b23d13fc5d87b87

Observation 47cc75fc-f237-437a-b673-4da9cd804bdd · inbound

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation cites this paper.

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:30:20.975440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T21:22:41.935691Z digest=sha256:f186cd39f7681e48d8db90570efd2d3766470cbdbe7989b262725fb6220968ec

Observation ac54e185-26f9-448e-932c-55ea746034b7 · inbound

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models cites this paper.

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:40:14.593189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:39:58.095112Z digest=sha256:c4eff339b5a373e733b1c2d3dda3c493853b1da51d24b905b42d14d85c061337

Observation 9e779ea9-0eea-4086-b075-03cba721c6c7 · inbound

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models cites this paper.

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T19:35:38.385806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:35:38.385806Z digest=sha256:6c51fda1b35165f7251bc3798fec4ce2c737eab3298b130c2292c57964a68dad

Observation ab770e7a-bba9-4b6e-b0b4-0a257bf42ddf · inbound

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons cites this paper.

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:00:12.661658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T17:59:52.365630Z digest=sha256:527ed2ce06363f15b0d86b8d1cbd0430af27956952453f8f101082d813de01d8

Observation 5b0b7430-c71b-4e2b-be27-aa5080bb5a42 · inbound

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models cites this paper.

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:50:37.181369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:50:02.670780Z digest=sha256:e55f9a2f0bafb138af4a9e45e6db4e8c6909880b4ba4d9bb056495b19deaca1f

Observation ff145d29-385c-4ae7-a8bb-486a5b596d7e · inbound

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models cites this paper.

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:25:30.903500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T11:24:45.578472Z digest=sha256:35b16cfde4e2e8f7249f1649a6dbe7fde0b5697d6049297daf269f6a184b108b

Observation 80c3c169-1c8d-4e58-a86c-25698f7a17ba · inbound

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism cites this paper.

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T11:50:03.870101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T11:46:12.134869Z digest=sha256:7fbadf77e03c5241d218fb4c605c166182690a20bd226adc0ea91adbfdb3f0af

Observation dc73b356-41ad-4121-83b3-3aaaf8f56e93 · inbound

Towards Generalizable Robotic Manipulation in Dynamic Environments cites this paper.

Towards Generalizable Robotic Manipulation in Dynamic Environments FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T09:49:54.326780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:49:26.446868Z digest=sha256:ba1eee68495ccdddbc1871ecc156c32f6c1105d7ba1b2e16bf7f2407e233904b

Observation a856bf57-2118-4f9b-98fb-4d89109fc848 · inbound

Towards Generalizable Robotic Manipulation in Dynamic Environments cites this paper.

Towards Generalizable Robotic Manipulation in Dynamic Environments FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T00:13:51.580603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T00:13:51.580603Z digest=sha256:12d959d05276fb7b94ada57d0f5d0f79c13daa31472bd1a8b630cc835e81208d

Observation 2cea80e3-bb25-4814-a574-cc457e8bee22 · inbound

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness cites this paper.

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:19:54.404995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:15:50.963123Z digest=sha256:c9c6c56031199b96f9d398a583a7d02502bdc9c14fdbbf0ad9e0197e44d98c5d

Observation f4d48a6c-510e-4923-9cfa-1be29594313c · inbound

Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic Control cites this paper.

Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic Control FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:39:52.213026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T08:39:03.661087Z digest=sha256:dcf0edf3b9e16d1253ff34ab3401c87430a2d9a872d4ad8793b994bab8d180ec

Observation 4dcfb5fd-2318-4d8f-8512-5ccfaf5b04fb · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T08:05:15.181408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T08:02:13.188363Z digest=sha256:6ed9cef496ca6916daa5fef9591caf2827b47587bfae9c6bd37ce4173eb63c11

Observation 2004ca27-2f34-428f-a72d-21042e7c79e6 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T10:50:00.975469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T10:48:56.280105Z digest=sha256:c4639050700afac8ed4ba221f37953ebde698583e769ed53ddf53c0ea0a17b20

Observation fc8b8e5b-b726-4976-9e72-b1e8451f21fd · inbound

LaMP: Learning Vision-Language-Action Policy with 3D Scene Flow as Latent Motion Prior cites this paper.

LaMP: Learning Vision-Language-Action Policy with 3D Scene Flow as Latent Motion Prior FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T18:17:45.786494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:17:45.786494Z digest=sha256:2ec34036964c9a6d0284f4d22c7cc49ae6188edf9632740821745ef9b03cf415

Observation bccb7a51-0e66-4057-89eb-0d29352d9c81 · inbound

Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance cites this paper.

Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:28:23.102611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:27:35.706107Z digest=sha256:681a3905d9093cf630e30fd64b72213549df88a73bb66505663ff3581accedc0

Observation 51321e68-5091-4560-9106-27ef63b7aff6 · inbound

DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching cites this paper.

DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:59:33.932755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T22:59:21.473359Z digest=sha256:e316924495010ada902d7f0c742be2d7ac41b8f3e2d883d9b8648ee3fd35d344

Observation e024f619-ddaa-41a3-a3c6-f57b528a210e · inbound

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA cites this paper.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.310458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:6f53532339adf07a2ee137fcd6d64ea8134ff8d746fab7eff202fd07a62081d3

Observation bcbff7ca-3b7d-463a-9102-b1afe759f26b · inbound

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model cites this paper.

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:58:08.914108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T18:54:07.081457Z digest=sha256:0fcef989ac33a7340b6e5736c2dd775c4341d7ae81544443b0c8b501437feece

Observation 1226feed-329d-45ac-919c-5dd5640eb28d · inbound

The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling cites this paper.

The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:53:08.378413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T18:49:07.376350Z digest=sha256:cecb15a0cde737cd3cb99dda47f9d16e5ca5b32692a113bf9a05dc6ef566af86

Observation 6167e5e1-afb0-4904-8d8b-fe4b664ff974 · inbound

Hierarchical Planning with Latent World Models cites this paper.

Hierarchical Planning with Latent World Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:18:13.786304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T20:13:58.991298Z digest=sha256:31932d3a818c3a968d559571596e189457ecf4ecf795062cc3ab569a6fc5ffda

Observation 68b4e89c-26f5-4668-8238-97e961dcdcac · inbound

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models cites this paper.

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:08:01.513460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T16:57:52.595165Z digest=sha256:81f7a9f7af23b0b6f2ec5e0e841d09b6853de3e415a222046ae5073db5adaf19

Observation 08970cd1-31fa-481d-b91e-e80f4dcba141 · inbound

Hierarchical SVG Tokenization: Learning Compact Visual Programs for Scalable Vector Graphics Modeling cites this paper.

Hierarchical SVG Tokenization: Learning Compact Visual Programs for Scalable Vector Graphics Modeling FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:52:32.433811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:33.951884Z digest=sha256:46fc8f6fc1b10d030476c5809fa9c293a5ceea0d260cbfa563517cb7e7846c27

Observation 349dd5ee-56fe-419a-b699-aa196cc02fc1 · inbound

VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success cites this paper.

VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.433811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:33:26.708966Z digest=sha256:26ada1251b1051457971ea63c36b6a6e0c906734b945883125214eff99e2494d

Observation 2136fea0-43eb-4595-9015-dac234295ab2 · inbound

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model cites this paper.

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.433811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:31:23.255452Z digest=sha256:9399424b7a551d0c15f0431d5a0b1bb2c6a9561fdf4063fb6a68388a258eee71

Observation 7c5f000f-ae8e-4009-b3e7-b2eafafa7314 · inbound

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly cites this paper.

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.433811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:24:42.800233Z digest=sha256:bee1964b70216018eb38d47a9789489a6afca906e7684526a0fbe05e4f2e5726

Observation bdd9e340-079f-4dbb-9eef-f6cce84d9422 · inbound

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly cites this paper.

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T23:37:01.875768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:37:01.875768Z digest=sha256:bc9da86d70159b253b5c84c7606a7358f3753a1a816cf6b187232c901e8e52cc

Observation 73d523bc-9056-4ba4-af8a-e9afb3b666d4 · inbound

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis cites this paper.

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.433811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:16:31.378588Z digest=sha256:c3d63f249ca3bae121982974d0e622730897e6cb13a796a695d4529becbb213b