Pith. sign in

Paper Citation Record · LEDGER

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 14 inbound Pith citation observations for arXiv:2506.09930.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09930 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:42:06.146713Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:39:40.204306Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c27d9800-1868-4362-86fc-6cdad0e5b7b0 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models PaliGemma: A versatile 3B VLM for transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.448199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.448199Z digest=sha256:5d264f59b0c050509dd21555ac01e773d66023aa71403ac33570414249b1845c

Observation 84b7dd7e-7714-4b6d-9c2a-fd9a277ec88b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.456303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.456303Z digest=sha256:d2ce1201453210df29e78609bb78eb2fd32b2344900970bf12bca16041badd53

Observation 5a4bcf23-5d93-4151-9601-6a5045fa80a3 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.462652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.462652Z digest=sha256:6ff29af1af9cd585fc9e3751fb6728b5ba4b76227eec95432c9246c205c32543

Observation 4b6f62db-9986-4bc5-904f-cb8758868817 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.469820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.469820Z digest=sha256:6d5d5d4036525a1e97f94722155fb8bd19dc3221cf2f3dbdb5273ef0c970b63e

Observation cc1c29cf-2742-466e-9b41-3ef9f52be5be · outbound

This paper cites Language models are few-shot learn- ers.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Language models are few-shot learn- ers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:10.920972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.476254Z digest=sha256:bda5b7549ad42499e24951ffe9a76c2164c47397dd4c49ddd437e447fe37132d

Observation 91e20cde-cd8d-4722-8c14-d6fb38401d96 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.486242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.486242Z digest=sha256:f0f6f80814bb5e8e61f01629fc40c022dad9af02f99c2726f219c0061c8b091b

Observation 341cf704-859f-40e6-a05b-41851bd41b86 · outbound

This paper cites Pali: A jointly-scaled multilingual language-image model.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Pali: A jointly-scaled multilingual language-image model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:10.761426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.490855Z digest=sha256:ccc8e06f5f18e4a9d7755bca313a000cd7705b61578424ea94d9240958b395de

Observation c747118b-51f4-45ef-9502-97fea4230882 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:10.608283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.494634Z digest=sha256:2ffb8f5d2c554abbda31ec0d5ee216018104af48233af61e525906a621f66102

Observation bc69081c-67bc-45d0-8d79-3631fa2d1980 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.500130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.500130Z digest=sha256:9602a26732d864f2d33100b7d090d626aecc1b28cb5d172b1fc2b64682cc60d8

Observation 0d963e5b-7976-44ec-a70f-bf6c8bd322ad · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models BERT: Pre-training of deep bidirectional transformers for language understanding

Reference 10

Resolution
verified exact
doi, observed 2026-08-07T04:42:06.329783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.509005Z digest=sha256:6217bd2dfdedb2348f6a340cd0780ccb7da9d429585e8de5f955972a6009cb25

Observation 7f056c45-23d6-4dce-b2af-3fa2239fa884 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.517492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.517492Z digest=sha256:8ef66ca2d50f9eb29be56a0b93b69684f51d68571da5c7781bcfc57f9fcffe8a

Observation ba37c4f1-fd83-4a19-aea7-bc0a26fc0e0e · outbound

This paper cites Poor performance of openvla on bridge.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Poor performance of openvla on bridge

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:10.418414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.526157Z digest=sha256:7c41b34ffcdbb09361eb644fcf8c9c86f79bee035afbe49d39794632c1717fae

Observation 44f349b3-9c3e-4c41-8ad6-84a9a674b104 · outbound

This paper cites Gemini, 2023.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Gemini, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:10.208868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.534702Z digest=sha256:7a9ee2dbf4d9d114ff98e13fa16b4036e9d70b78384d7288d9318559d706bb65

Observation dd51dde9-1c76-4a58-a908-dce00f64396b · outbound

This paper cites Maniskill2: A unified benchmark for generalizable manipulation skills.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Maniskill2: A unified benchmark for generalizable manipulation skills

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:09.898690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.542407Z digest=sha256:8af123e464f4bb9fe111d5d131b6854596f968eb92fe5e12ba96cdca7534bcfd

Observation 96adb65c-2cc8-4e41-8c87-0cedfbba00c6 · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.549671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.549671Z digest=sha256:3009e5b91214f0ca6cbc4dc9697bdc4eb8dcd7d80f18ac815bf6944eaa4b5d24

Observation 0532ff28-dd99-4a5c-8c30-6d2d8e0bca1c · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.558677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.558677Z digest=sha256:8d93d75407da69ba4764c680e4485ad2eeca2ec652e22b833d9abca52efc54a5

Observation 9cbc4315-164b-4170-b50f-3d0196a42f86 · outbound

This paper cites an unresolved cited work.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:09.627086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.566759Z digest=sha256:50923d06136e4711e2be7ce0316a774abaf0f817d11ae0f1c15a7c0f6ef4d371

Observation 72427009-033b-452a-a48c-2ea4e9ea4ae3 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:09.364393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.575347Z digest=sha256:29ac7d1c7ba939af93170c62fc582239e91b210603e148d66e2ae71daa8b0856

Observation 6077cad2-4627-44b2-84c5-80f9bd9a672f · outbound

This paper cites Open- vla: An open-source vision-language-action model.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Open- vla: An open-source vision-language-action model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:09.173991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.650556Z digest=sha256:a175ec8dde29e8339bcb5a9a246b728eda7824c6125b8e03907a4e815fe85a18

Observation 2c25dea5-686f-4947-be56-be29b27d9155 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.738170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.738170Z digest=sha256:4f7ad1a68f66d5ab9237899a46c078812ef99907f628f549830cc06e6b949eb2

Observation 8a5f7fbd-95d5-4e82-9b39-1ffe23eaf7ce · outbound

This paper cites Flow Matching for Generative Modeling.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Flow Matching for Generative Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.781269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.781269Z digest=sha256:6990bbff38a30f2664ff4eb4a4a1cc53fc9da08c7254241f20335ca6b2812a9a

Observation 5366ed24-bd34-4772-91b0-b58102627c61 · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:03.904872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:03.904872Z digest=sha256:b98b0fe74b1dd720d8780ea29555496b8bc8dbd4d8308a7ebe9c466b7240693d

Observation acd3e924-dcb6-4e59-84dd-32d9efb42b3f · outbound

This paper cites Visual instruction tuning.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:08.961317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:03.997313Z digest=sha256:db78cf681f31f2e2973792a97b3cdcfd28505e64b9c26b0eccf30a855dc76a17

Observation 3a825f77-8c62-4827-a63d-fae94d00ac9d · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:04.036897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:04.036897Z digest=sha256:90461c2d5244b357ccd91c500db713057d68125d4088b0a8e4ed4853386e882c

Observation ea21fac9-c7c1-4fc8-80f2-d937895f41d5 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:04.128560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:04.128560Z digest=sha256:d75de74d801c0680734eae54eb0620baa6b5f1093d9b48455df3d156877a0f49

Observation f7d4a076-07f3-4d94-bb82-cfa8c353b254 · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:04.189941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:04.189941Z digest=sha256:b875d54ca65fe8f2dafc914a1372497d576fe9f64e34d5692e5599a0a220c47a

Observation 156eb4df-4266-4f88-a898-01374b2dda86 · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:08.731297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:04.283011Z digest=sha256:07b3f1a709b0bbbc3cba9cb8e0b4a5e6f13e030401c75b9dc6ba14031ca683bf

Observation 5c66e6c8-1674-4f14-8461-7089d1f0e2c4 · outbound

This paper cites RoboManipBaselines, December.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models RoboManipBaselines, December

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:08.515642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:04.352564Z digest=sha256:50bea183787e7385b5cd89282259e0982c0f551f2eeb0962d4d259a10060fefc

Observation 6bf085e6-01f9-4da8-8d34-e10d1caae89b · outbound

This paper cites Octo: An open-source generalist robot policy.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Octo: An open-source generalist robot policy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:08.177530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:04.509031Z digest=sha256:e3454716357eb6ea7e0029cfed3a6b94debca9931c3007bfe7cd6a66efe972da

Observation 5a7a5c19-55b3-4716-97aa-d25409227988 · outbound

This paper cites gpt4o, 2024.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models gpt4o, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:07.948686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:04.590014Z digest=sha256:fc862b3866f47bb47d2282bd8db16d465a677ba43817ddd3feb17fd85461ad63

Observation 16c5c5a6-c9ea-457d-b02c-a96bcbe279f6 · outbound

This paper cites GPT-4o System Card.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models GPT-4o System Card

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:04.670921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:04.670921Z digest=sha256:e1ad4c2bde08fa968aa6a8b36e39667a1e1ae64dceb46c6b8898eda939cf6f1c

Observation 9bc60aae-10b8-4c33-9b05-dd9d98415e27 · outbound

This paper cites an unresolved cited work.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:07.813039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:04.742455Z digest=sha256:2d3c754e8e9fdc937fb2bbc1d393c171c28265d1eba6082193f7045442823591

Observation 66c08d11-835d-4732-b7d2-f8eac72be357 · outbound

This paper cites Transfer between Modalities with MetaQueries.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Transfer between Modalities with MetaQueries

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:04.839068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:04.839068Z digest=sha256:147ce7e2ff9ae79090bc78f7325157361147787e2c5ea2400b4fac385b88317f

Observation d4f464ad-4aed-4261-9491-5e8dfbcfba1f · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:04.937068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:04.937068Z digest=sha256:6551d4591102bea9cd542d4c7853f4341eee781e1130a83e3f1ef2744a498830

Observation a7599869-dcfc-4e08-b4bc-b0bbdd366888 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:05.044108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:05.044108Z digest=sha256:b8a2a1d3acb6f311fe50638aa48b2f0de9a13c676c4e8e2998269b3df3d36ea9

Observation 63153e01-3e4f-4536-9664-d8ae44c94a4b · outbound

This paper cites Learning transferable visual models from natural language supervision.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Learning transferable visual models from natural language supervision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:07.641362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:05.159411Z digest=sha256:bc6134c7bbc7cf4d0c59456ae8e753f77b4d55a83cedd9ec09612f44888d357d

Observation 928ae829-0606-4ba7-8e14-fcba1b06b291 · outbound

This paper cites an unresolved cited work.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:07.475381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:05.232657Z digest=sha256:3c98f674e9250c1c547b2e4e3af9fdc6f43a060cf50b305e314ddd9740a0ee08

Observation 952a6701-c7ec-479c-9ea3-1305e37ac407 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Gemini Robotics: Bringing AI into the Physical World

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:05.314334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:05.314334Z digest=sha256:e64bcd6d0f4851e8da8b14dc98ff46fa53c0ee70489012c7824a6aa289c11b0c

Observation a804a151-5955-4284-84e7-9cc7f52a29c2 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:07.323345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:05.424420Z digest=sha256:c0c2fadad87f5ea0b2b6da4037f13b6cf3b3a95f0cdaf19dedf7b6181f789929

Observation 1f7feb02-84ed-478e-b985-553503aef950 · outbound

This paper cites Metamorph: Multimodal under- standing and generation via instruction tuning, 2024.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Metamorph: Multimodal under- standing and generation via instruction tuning, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:07.180134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:05.541456Z digest=sha256:ebf6aae678582b83a262f67dba1986f471f6186339fc0d6c572e312870bb70e0

Observation 2b2fa300-de27-470b-9a94-91a669883642 · outbound

This paper cites Vlm see, robot do: Human demo video to robot action plan via vision language model.arXiv preprint arXiv:2410.08792, 2024.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Vlm see, robot do: Human demo video to robot action plan via vision language model.arXiv preprint arXiv:2410.08792, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:05.639347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:05.639347Z digest=sha256:e7abf4f9c7f5e7ca2d4c11568814a1cdaf6d1ac9a9e640f0d6ea62853141f7d7

Observation 581770ba-6bfe-444b-ad65-cdb3dfebdfae · outbound

This paper cites Decomposing the generalization gap in imitation learning for visual robotic manipulation.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Decomposing the generalization gap in imitation learning for visual robotic manipulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:07.012189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:05.716768Z digest=sha256:22149c6fb04d9fda5b6160e9feb7cd30ea4536e823c31f0573b9b4cc7d9a6fb8

Observation 3fc7aefc-af35-4ad5-a4bf-727f9b6e9183 · outbound

This paper cites Magma: A Foundation Model for Multimodal AI Agents.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Magma: A Foundation Model for Multimodal AI Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:05.831412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:05.831412Z digest=sha256:52df248bb65731e09b1ada75a476a0f49afcdac6fd8eb544e3dd58f5dcd7618b

Observation 95b2b9ed-2033-42bb-9e5c-a4c1aca5bdbe · outbound

This paper cites Sigmoid loss for language image pre-training.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Sigmoid loss for language image pre-training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:05.914924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:05.914924Z digest=sha256:43505c3ab748d20b4794ab98616b44e0830b50c03cda1e3fa0a066a9ea09f7a0

Observation 0479ad34-5cb8-45d6-b6a2-9d79d522827f · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:06.021627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:06.021627Z digest=sha256:4220102a5fd74e6508d809f4f45a9eefe95aa95289333f642c608cc4327775a2

Observation b74ebdc7-4d61-4f68-980e-a90cdf7e78fc · outbound

This paper cites carrot on plate.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models carrot on plate

Reference 46

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T04:42:06.813547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:06.146713Z digest=sha256:31f7cd6d0b79c1746fc3f0ff0ddd91f74ae13b2260880332a1d729693f8895f2

Observation d042169b-3098-4228-bdd8-69403383f2f0 · outbound

This paper cites an unresolved cited work.

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:08.365026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:42:04.440709Z digest=sha256:0aa715d106fc2a3e163281535b855178bf5dae13f6e628bae3aec092e1d0ff8e

Pith citing papers

Observation 07cc7d3e-2ee1-483a-9247-4815e3e92ae3 · inbound

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization cites this paper.

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:20:01.937735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T06:20:01.885711Z digest=sha256:694fa7ea86c141d9d34cfbdaba22b2d8ce0014b54c7955190c7bef81e0e07983

Observation bab77f63-a90c-4106-a0b7-b3e759dc711e · inbound

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization cites this paper.

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:53.471843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:38:53.471843Z digest=sha256:ddc6acc6beb1dbec55d50324bd6937fc10932aaba251d203e76697fec89f62eb

Observation 95e87e2d-283c-4f7d-bf50-7687c3e09ee7 · inbound

Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot cites this paper.

Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:03:12.269733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T18:02:11.393346Z digest=sha256:9fe2bf62f4a8da07e5fa5ba875df0c54ef7271fd8e8591953b3a56b3a1bf57d4

Observation 60f0d20b-097b-459f-8ff6-a89a095e4604 · inbound

Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot cites this paper.

Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T12:41:53.551056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:41:53.551056Z digest=sha256:4d249c74c3bd8a68ee73bd16a41eb97f4e3ba5932b28eb4499c922605e1cf839

Observation e144412d-e29f-4851-9400-08dafe29f882 · inbound

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation cites this paper.

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:35:38.600524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T18:21:00.089755Z digest=sha256:307d710b53ffcf8bcba166fd504f81acd7866fbe32f8d16d78cdd55d0934497a

Observation 124dc2f3-cafd-47c4-be9a-8f050d0bbdf5 · inbound

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models cites this paper.

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:20.664056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T14:39:27.174418Z digest=sha256:81c7d8961199ce87d3757a4b0481c76720f2a42a575f0d971f859a772c44769e

Observation 19ff866d-7b68-434a-99ed-c057353da94d · inbound

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation cites this paper.

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:56:59.992400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T00:47:21.667833Z digest=sha256:73662aca51253c2ed2823ff35a4914718998d04b9ffb59f7e5223b6601981d3f

Observation afcccc75-7d19-4c99-b990-e9cd5036660f · inbound

APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies cites this paper.

APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.730672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T09:49:56.894300Z digest=sha256:3d511faa66cec8783ea18f2ad1de4b23499c5101f29b3fa87296c79ea369565a

Observation 42673a39-889b-48e5-933b-039500523b22 · inbound

MANGO: Automated Multi-Agent Test Oracle Generation for Vision-Language-Action Models cites this paper.

MANGO: Automated Multi-Agent Test Oracle Generation for Vision-Language-Action Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:00:01.458115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T23:07:27.176553Z digest=sha256:d5e5869f27c7a4aa0d534564cf7fff6c872894869ee4e5712150ee950d930dc7

Observation 36726a43-7f1c-4f05-83f9-08bbb7e3535c · inbound

EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories cites this paper.

EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:25.123843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-02T22:46:28.245243Z digest=sha256:c6a6e924ae4a5c829c8ff9d79319b3c34449c17d76fc33cf3e186f4652e9a12b

Observation 7274a360-5f3c-4098-acd1-5d1578617171 · inbound

Robots Acquire Manipulation Skills in Seconds from a Single Human Video cites this paper.

Robots Acquire Manipulation Skills in Seconds from a Single Human Video From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:06:14.393672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:06:14.393672Z digest=sha256:e047ea793e1b562ade27bc6486a3fc4d38919581a2518359bb33c8e267a57f8e

Observation f466b2ab-6682-4105-9fb0-844986fa41c4 · inbound

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models cites this paper.

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-30T14:45:01.896517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T14:45:01.896517Z digest=sha256:510fd333f8c13a2451bf37f10b883f1d2e32650d9816889ed0d73cdc6d99649c

Observation 82189999-d2a1-4aa9-8fbd-d1272d3f771b · inbound

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models cites this paper.

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T10:23:13.435458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:23:13.435458Z digest=sha256:e3fb42c5b3a2abcda90ae14427ed2b0acf09776672120dd2ef41d4309d6c31ab

Observation 78dbd86f-6df9-4060-87f0-f9c36e3b9296 · inbound

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models cites this paper.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.204306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.204306Z digest=sha256:3ffb9492eeac366766d166ded654b38db284b7d9c04fbcd73b3875db6ab4fdca