Pith. sign in

Paper Citation Record · LEDGER

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

As of 19 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 14 inbound Pith citation observations for arXiv:2506.17561.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17561 v1

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:10:27.648196Z

measured 114 of 114 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:38:41.712173Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 109 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved99
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9ec5de35-6096-4d2f-8296-10f60ef2fd82 · outbound

This paper cites GPT-4 Technical Report.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.003333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.003333Z digest=sha256:84272060abbd86366c19c79371a541d6cbf0406b7c4aaf665a5e88ed082a42af

Observation 5d926090-8c7b-46aa-a059-5a883583aeea · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.009878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.009878Z digest=sha256:8d38e9f6a21bd3f1a8dd78f55ccb310707e3c738dc8dabafdb2a0168fe1fb995

Observation 9d03ec14-9f3e-48c2-b990-72f4cc939298 · outbound

This paper cites Dexart: Benchmarking generalizable dexterous manipulation with articulated objects.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Dexart: Benchmarking generalizable dexterous manipulation with articulated objects

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.015062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.015062Z digest=sha256:351f2df93746637e9a195acf1ce270a6139f7547e4e2c24e90599ba082bab2dd

Observation 3b7bbc82-118a-485e-9d98-338d4d1fac28 · outbound

This paper cites Minivla: A better vla with a smaller footprint.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Minivla: A better vla with a smaller footprint

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.019970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.019970Z digest=sha256:1b669f4a61a7dd63d2f6373f86edb28866d15f3b957deadd7514eb28d5d69614

Observation 6aa62eb3-b3cc-410f-809d-fd521ce12374 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RT-H: Action Hierarchies Using Language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.025891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.025891Z digest=sha256:4b35da6e4f3cd3668ebbf11e5884a06be489cc74ce1fb5c423b44a1fd71ed3f1

Observation bf2ae482-fef2-4bab-ad87-67c9bd34d633 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models PaliGemma: A versatile 3B VLM for transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.031239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.031239Z digest=sha256:a0230c6360a179ba7332a2afe292952fd5cd74544a8a836c12c02739b5b6b99f

Observation 35b47c6d-5c46-4310-b54f-4e5b814b9559 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.037102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.037102Z digest=sha256:f4bafca569a713081b7e8baa9556791002d21c750ddceba1b2195e69fc025c29

Observation f6ad6bc2-204a-4beb-8e66-71771cb81664 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.042074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.042074Z digest=sha256:be4084f46c4bf41d582e825a11e1284d5213fd73edb9b10802668ccce24235c7

Observation 5d421add-9cf8-4ff3-8c02-89de5e8cd8b5 · outbound

This paper cites MIT press, 1982.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models MIT press, 1982

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.048240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.048240Z digest=sha256:26bba6c6eeb5d33d2220b0abc3e7297afc500c9e218d56750a6e7308acbd0971

Observation e0f215ab-bfcb-460d-8122-9fb53d629c09 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.053827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.053827Z digest=sha256:a10b9b7aec3281bca3b6affc5123bbf4ca57796ac6f2f463c1a4896866dabc52

Observation ae157066-4690-4f23-9fe7-7273fe336801 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.058990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.058990Z digest=sha256:be827abf5b5c7fea0a904c679840b0b8dc65ebc0a60b326173ff49956c79506f

Observation 4c2fa5b2-242f-4fb2-9b04-f0fc4d1e03f9 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.063851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.063851Z digest=sha256:d6fc24bb28d4e30da89b6635a1809ead7ea793b1259d7647ecfbb4056bcfae71

Observation eaf6eb3d-16b3-43c7-a976-d831098c78f1 · outbound

This paper cites Univla: Learning to act anywhere with task-centric latent actions.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Univla: Learning to act anywhere with task-centric latent actions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.068940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.068940Z digest=sha256:61d118b80f229bf435a07c6a08655b3df14dfc4c2ffb4469ba1a728ffa9b4ee9

Observation 8d42b62d-b0fc-4115-a016-0d3b712ff8e8 · outbound

This paper cites MetaFold: Language-Guided Multi-Category Garment Folding Framework via Trajectory Generation and Foundation Model.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models MetaFold: Language-Guided Multi-Category Garment Folding Framework via Trajectory Generation and Foundation Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.073785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.073785Z digest=sha256:f839c678d6f4ab60f7e6cdf2d4a11012d57fcf052bd463a761d65726184e72cf

Observation fcbf8451-f69e-4dd3-85ef-88806e85856a · outbound

This paper cites Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.078702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.078702Z digest=sha256:d1bc412471ff4fe68057e337772840fb1eb3e74298a4687e625b42ee62fa7da2

Observation 4ddca0e3-d259-40f1-a1f5-800e9a5cb0ce · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.083480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.083480Z digest=sha256:aa22387c721e09fafd52be626dbae9071362176430c4e031baa642de3622c705

Observation 720b0bd8-8bf7-44c7-a166-e0a64d9670ce · outbound

This paper cites GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D Manipulation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D Manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.088478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.088478Z digest=sha256:6e380ba3882327eb172bbbba57203fde5e33bcb6a2926a5f4e711942f91f4295

Observation 2c7cd1f0-3902-43c5-a7b9-3957e217126c · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.093662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.093662Z digest=sha256:462de9d6e6244a7384eeee4c4033427033c6751cb8aeae13b675a9d87c71f085

Observation 37a6dc63-29f5-41d1-9baf-1838eeb6e3e1 · outbound

This paper cites Palm-e: An embodied multimodal language model.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Palm-e: An embodied multimodal language model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.098304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.098304Z digest=sha256:f97dbcb32ae2cfef2f4a44c5c2190ae7b78e54ebd1a37ffbd60e44f0b2de8a66

Observation bbd3ecab-970e-41ff-a67d-1968169d066f · outbound

This paper cites Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.103053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.103053Z digest=sha256:805cc54187d21c882a503c6a5c02e88074ee5cb734e5e24f75de6564e99ccfd0

Observation 66b63c11-d1cc-4147-82b4-d16af2d16d1b · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.107903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.107903Z digest=sha256:12f02cab80a055d0eccb19ba6c987b5085eeb4bcc0451f8ab7a9d615ba5cd379

Observation 1ae7c703-13b9-4e3b-a1b6-f9b37008791b · outbound

This paper cites Helix: A vision-language-action model for generalist humanoid control.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Helix: A vision-language-action model for generalist humanoid control

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.113547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.113547Z digest=sha256:fa1ea6b2520402366b8ae2af10ed7534961cb21830562830f466bc1528bd26cb

Observation 34badb41-4d38-4ea4-ad74-7936abfe692d · outbound

This paper cites Foundation models in robotics: Applications, challenges, and the future.The International Journal of Robotics Research, page 02783649241281508, 2023.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Foundation models in robotics: Applications, challenges, and the future.The International Journal of Robotics Research, page 02783649241281508, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.118351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.118351Z digest=sha256:9108c720d028b6e34aeffedcb666f4a8edcbc505ff636e3b002e7f697cbeda32

Observation f5cabab8-3937-42e5-8f40-3e27b6781869 · outbound

This paper cites Cril: Continual robot imitation learning via generative and prediction model.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Cril: Continual robot imitation learning via generative and prediction model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.125156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.125156Z digest=sha256:98404a167b12a87099c932f5dc1a4facf4a3a00f3e655da360c3379cc65ef80a

Observation 127ace33-b85d-4412-8d4a-0393891ec53c · outbound

This paper cites Transferring hierarchical structures with dual meta imitation learning.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Transferring hierarchical structures with dual meta imitation learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.130146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.130146Z digest=sha256:569d69f6521f433d4c470f3fd17da1a802e0f9a8ba8c678d33f6a20f8c8751c7

Observation 9446cf73-ef99-4828-8a0b-a21c70511013 · outbound

This paper cites Iterative interactive modeling for knotting plastic bags.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Iterative interactive modeling for knotting plastic bags

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.135031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.135031Z digest=sha256:4f292d6c531d4a11f5e6e348dce400c2a18ecf3b033c1a24c04603c87d3b2f24

Observation 9cfe5a7f-32b7-4311-92b8-5082c6c98d24 · outbound

This paper cites FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.140114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.140114Z digest=sha256:af70808b13a427a7da578a59310201de69b6630f418d90c70941c12b3d5ede44

Observation 7bc351a4-0c81-4930-868d-ad444cef80fa · outbound

This paper cites Peract2: Benchmarking and learning for robotic bimanual manipulation tasks.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Peract2: Benchmarking and learning for robotic bimanual manipulation tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.148990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.148990Z digest=sha256:ea0b6c8a57d44ac4d749d114c6165cb8d70bfa252ee3c001fc4de41d3be796fd

Observation 4a794c05-14d3-468f-b720-cebc3611754b · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.154410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.154410Z digest=sha256:4363be8a9f51f8fc521348617f53a0208449ebdf0e4aabd9a89423bfdeb3e101

Observation 3784f67d-ee61-43dc-b36e-e58ca98ac549 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.160457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.160457Z digest=sha256:c86a8fc6e2bfe1338785398f22c1a0d1acbc9b6681b7332c6c968ba0b620d599

Observation 25c08f81-c3f0-459f-bb21-d5422ca2b2f8 · outbound

This paper cites Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.165223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.165223Z digest=sha256:da178ced01af5a6c4199abcca81628c9c4bac2d451224023249f0041987c4762

Observation ccee3d34-3d57-4f4f-8959-24fc8e79b7af · outbound

This paper cites Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation.The International Journal of Robotics Research, page 02783649241304789, 2023.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation.The International Journal of Robotics Research, page 02783649241304789, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.169851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.169851Z digest=sha256:d306ff163156ef3c0e5e30860d9476bd2b98bec56bd7012766ee704c1e7d7e37

Observation 90667bea-1a84-4668-95d3-4a67e3bc3843 · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.174266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.174266Z digest=sha256:29ab5a293f51c6fbe7c66c70905cb9327c0ed20fce6b6f2ed88f82c0cc684519

Observation bbb32158-d27e-4b4d-86df-2f3fd0827424 · outbound

This paper cites Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.179019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.179019Z digest=sha256:a3af741a659e88d91bb9253a13dcb38b1b4bc6855922733100b10d648c5b90b0

Observation a4570be2-5ab4-45b8-95b0-e48641531207 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.183764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.183764Z digest=sha256:776dcebc5a11d73fae89e95560195182110fc9ae9fd7e7f0a3926107a603006a

Observation 44974a0e-f713-441e-bc0a-d6345852dd51 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.188840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.188840Z digest=sha256:981944ce0302a15a1d4073850dafd95c044d27b0f6cec9034fe6c54b1a55eed5

Observation 12bfa25a-b0d4-4149-bc4b-29f0e6a50825 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.193785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.193785Z digest=sha256:0aa82f1bbfd7fa8b4fb2f7feb220cee9cbe90bcf20d8fcbbb7775afa178c2300

Observation 62ad3dcf-ac87-4f66-93eb-402d0ee45e2f · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.199126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.199126Z digest=sha256:6b39421def7ab944ca6d7deafdf6726b678b30c65d649f7931687fc628fe0170

Observation 4528cd50-4c89-417b-9b54-83855159888b · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.204775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.204775Z digest=sha256:cf37142adae0998034a2ec175e295001f0f522fed6148031da787701426a11ae

Observation b1a1af65-5c05-468a-97a2-41e65c9067c3 · outbound

This paper cites CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.210219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.210219Z digest=sha256:507529634c186648b7f16a159ba09cc4c5c01687e980a6976387b84ca27aa73c

Observation b3d386a0-5c77-4e66-86a3-d21109583148 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.216285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.216285Z digest=sha256:a1b0f1be6e5c1a7a51417a245422f715226ab1f8e935d68c46ba3b7d4ce2bd49

Observation ce8e387d-0f3a-4f19-af60-efa5883cef10 · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.221449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.221449Z digest=sha256:72e5fd9ac595b53bad3fe66f03e2270bdd600f7d6efbde24d65af908202c3cb8

Observation 58bede9a-0671-4b96-8878-712b6279513d · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.227125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.227125Z digest=sha256:ec997dae41eb16129f01756625c8e30757d6003cc0f5d496505d22a2db730187

Observation 19f38507-8370-4ef5-92ca-796f0a2e8502 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.231900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.231900Z digest=sha256:21a202000e24cf2aecd6449115e1c08a3eb4e9897b8107c138b913bc0d13c6b2

Observation 51b81117-1059-4855-a127-d0f2bf9b9184 · outbound

This paper cites Segment anything.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Segment anything

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.236696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.236696Z digest=sha256:bcbf9fc05ddc3cfc4221a75e93018e4df52209674fb7e6f50ca143d2616057bf

Observation d9f8815a-0023-4b5c-8fc3-037002ce252d · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.241005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.241005Z digest=sha256:6026c31100492810cecd9cc9bacc39b01e02980870b19075ac86ef018fbe2196

Observation 1d1d8ce3-4fa7-43eb-8964-f2492252831c · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.245859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.245859Z digest=sha256:4bbb88f227bcea2d11f4ca46f3a2098bddac043072a82021e19d80498e0c53da

Observation 62c8773d-0242-4c74-ac32-6ed1f9555e05 · outbound

This paper cites HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.250566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.250566Z digest=sha256:46c9f26e516567f49d9cfe8966336b3c15eb2fcff8979287b302436707afeb4a

Observation 0cf68b7d-6d66-4354-8bef-253c7b323f79 · outbound

This paper cites Code as policies: Language model programs for embodied control.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Code as policies: Language model programs for embodied control

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.256457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.256457Z digest=sha256:6e047b32e79248dc29c5e5a95cd52ae24205360ed58e9a4903114abcf53ec76e

Observation ade42fed-3a92-4aa6-9ef0-fc298af79248 · outbound

This paper cites Flow Matching for Generative Modeling.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Flow Matching for Generative Modeling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.261844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.261844Z digest=sha256:6c8d8bbe581002c0739b385bc06bf25aa267cc15aa7f3e7162e679145d059112

Observation ea34b8be-fa20-45e9-976b-5a5ca993a081 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.267226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.267226Z digest=sha256:e383a36135f6cbffa0e1ac9b1487b513e8889f1b527492c3760da226832d78b2

Observation 52d21b8f-5b49-4905-8601-b7eb7574d7e6 · outbound

This paper cites Visual instruction tuning.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Visual instruction tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.271777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.271777Z digest=sha256:8ea8769981632040923c51f60c730ef710cb6d95a7c2dfd5f21dcbb21f1db36a

Observation 5efa89d5-6566-49d4-bd5f-119da8ab5bb8 · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.276576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.276576Z digest=sha256:3421917173ee473098828448b2f5db78c8cad239be1f2ac2c5701ff6fb3c36f2

Observation 084827d1-0369-4d6e-9e27-ab208db2038e · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.282155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.282155Z digest=sha256:df23bc6e1520f738aa039f47fd49f3a22b2f9e90ecfb47818c5c87878b80133b

Observation 9e47f889-03a2-47de-b7ea-7d58862239d7 · outbound

This paper cites Structured World Models from Human Videos.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Structured World Models from Human Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.287296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.287296Z digest=sha256:c0f9f8e266ea728910009bf3c6af9869e1cce800dad9ace65a94473ca466b2dc

Observation 54ebc49d-9fa4-43e9-b1e1-27180a0e8fe3 · outbound

This paper cites RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.292201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.292201Z digest=sha256:24a835bedee66cd430ad22b14420785c5f4be1ceaa26d0d47d930cce6454ae22

Observation dba92367-874c-441b-b2c1-f972801ef107 · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.298511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.298511Z digest=sha256:ef5fbef2a900a4cef26150b56eb9bdfe4137d70cc53f31051dfff2fdb011805a

Observation 1fd7db5f-5287-417c-80ba-b875d3cc9c9f · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models DINOv2: Learning Robust Visual Features without Supervision

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.303883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.303883Z digest=sha256:196bfc112e51334e6c286f3cf682804b211d7e0910c63aa17109e454898570ce

Observation 8c4e211a-9b94-437d-ab62-bb3861ccf149 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.313763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.313763Z digest=sha256:665d1c2dcdb0eaaf862ec20a88c1a9e03dbc1d3d7bd42a28ec0561b3f92bf5f9

Observation 1e14c9f2-3112-46d7-99ce-e879ab1510c6 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.318235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.318235Z digest=sha256:c83bfc3d40c6089ba9b8d3c095bd16a96d42f7fe29d1bbd24011a0891792dc8d

Observation 11d6191d-ae5e-4a7f-9370-f7941eac36d7 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.323117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.323117Z digest=sha256:403ce05595e08b7a1520260061ce396189bb1afc2bdc782327b4226d22a0f4a0

Observation 09d1547a-b8d4-4fee-ac8d-ed8b48754385 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.328677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.328677Z digest=sha256:16feb8043cdeb71cec752295ce8f98e7945bef3c61496a3540c43d8e8c947e72

Observation 370a51f9-414d-4da0-b0c4-d64ae2220bbb · outbound

This paper cites THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.333639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.333639Z digest=sha256:72a5c49b2dcebeb5111f415a84ae9afd058456e8c051f899790f1247af7298ae

Observation 183f2236-2e46-4520-87be-7fef6f368d8a · outbound

This paper cites Open-vocabulary Mobile Manipulation in Unseen Dynamic Environments with 3D Semantic Maps.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Open-vocabulary Mobile Manipulation in Unseen Dynamic Environments with 3D Semantic Maps

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.338620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.338620Z digest=sha256:6effb95c6eedc8103681c3fccc5b1771a274e03bc391b74c1468ef0042f7111d

Observation e56934cd-e5c5-4973-ab06-af846be78952 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.343826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.343826Z digest=sha256:ad92288c0db4c63e025baaca96c60779b58959a7fec57312ab1442c24c37b690

Observation 59875891-dd11-48c4-897b-f2c7e59d664a · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.349176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.349176Z digest=sha256:721a63aef644af7fe4b1e041966fa4261460f752f3bcc5919892afc66d39f546

Observation c15eb20d-766d-48bd-bafc-58f858d803e3 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models SAM 2: Segment Anything in Images and Videos

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.353687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.353687Z digest=sha256:f7bf70627e4f0ee812f86a9621ce0f7ce9073b0c7038f53e99e8141ea1676d21

Observation 54e0ec3e-c0ef-4d4c-8f7a-77b6477186c6 · outbound

This paper cites Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.359020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.359020Z digest=sha256:8b2e765a418cfbef2e78e816e6be9a5eac64e8683cc7907a50f7e2c342115b26

Observation 9d8dd8cd-9f4f-4400-aa9d-dbdfbeaf9398 · outbound

This paper cites Learning to Act without Actions.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Learning to Act without Actions

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.480043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.480043Z digest=sha256:fc66481c8ecb4b8f58648fc13dec8e0be0f35960ac08d7ca7d1a9c8265db477f

Observation 4479e97e-078f-44d4-8deb-75ef88d4657f · outbound

This paper cites BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.484823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.484823Z digest=sha256:7bc9f0c7c16ce6addc0c3b4b63f7c63c65dcf430e898d20c6014e975155cbd7c

Observation 8dc55914-ea42-4a7d-8a3b-6e358567cd2c · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.489750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.489750Z digest=sha256:1573a7a5987e187e0eccd991500a3f0f4c1d8aa5d0346202f1375b84621c299a

Observation fc06fdbd-7fe4-4365-b8b2-ca056bab4c37 · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Progprompt: Generating situated robot task plans using large language models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.494695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.494695Z digest=sha256:76a0095a0f54f1129363cef80d198de8f9af2ecd1cfa028680bc2a4c81bfcfbe

Observation 9a1c103d-1ee5-48af-8ff8-9e139eea3969 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.501457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.501457Z digest=sha256:47a7a9fe954054a999620179f69ad3f22960113f060427885cd2792d8d4ef4ec

Observation 3ffd21c3-9ee9-48a5-9617-a1283b5a6467 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.507119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.507119Z digest=sha256:66c6ec047e49d1dab8c8f1149c4b44274ec859851969b1b7ad578942a2d64d43

Observation 97e66d6e-ecd0-4584-b8ca-637fa755131d · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.511441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.511441Z digest=sha256:e42d60aacf05446192505f0e2029ba45552f8228d4b31546350dcd8f55f91f24

Observation cb1b856c-bce7-4bf2-8e19-17ae716835eb · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.515569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.515569Z digest=sha256:60f5209f71df13d601e3bf0b6e70a8411a8c7e12a676078dc8feee5b6c8ef7aa

Observation 315a5d10-127f-4f33-9e51-dc51f21ae169 · outbound

This paper cites Llm^ 3: Large language model-based task and motion planning with motion failure reasoning.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Llm^ 3: Large language model-based task and motion planning with motion failure reasoning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.519672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.519672Z digest=sha256:8eec0dfe851fd97b64f0bf0e3b2c7295c3f17e3e8e3543d96ba935fe3fd95a3e

Observation c8f36ef7-0f0f-4ef1-903e-3b4793e5bd2e · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Chain-of-thought prompting elicits reasoning in large language models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.524313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.524313Z digest=sha256:20cf301dcf952a2746b50e1f9ad139c21223ad40a58ab86d6b4dfc657e6442bd

Observation fd3a2dd9-3a41-4008-996c-361cf17c92fa · outbound

This paper cites $\mathcal{D(R,O)}$ Grasp: A Unified Representation of Robot and Object Interaction for Cross-Embodiment Dexterous Grasping.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models $\mathcal{D(R,O)}$ Grasp: A Unified Representation of Robot and Object Interaction for Cross-Embodiment Dexterous Grasping

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.529499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.529499Z digest=sha256:a0ce012501084e7bf8ccd9fc7300508f08d4ba9aee2fb082c5a5115f704b89a6

Observation 6ebc802f-158b-469f-9730-d32e35db6c1a · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Any-point Trajectory Modeling for Policy Learning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.534917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.534917Z digest=sha256:2884b4692f692391ae51d0c92d95285dc8a8ad9eabab6bee1b1758b17bd2a1a8

Observation aa52847d-d72f-4257-a581-a7e6d0f8791a · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.540220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.540220Z digest=sha256:e90b9ca6dceb55746c4bdcb7700c97d572c031d17fbe1fe86a793a1bfd4f7f70

Observation 988e8b70-1b10-4ca3-aae5-44bd9776cdb3 · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.545942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.545942Z digest=sha256:81a61a5c7ba1b0de207b727ef824ceeedeaf6120a1ca94f0d5cfe6f710a61f82

Observation d93be186-6d36-47d8-a0cc-864385877482 · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.557930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.557930Z digest=sha256:3500b03f6fcd607d85da1ec2699bff46421909bb2b2ea22a59d957e4cbef07d8

Observation d7ae7b81-399b-4d01-916b-0c10a3d845d9 · outbound

This paper cites Sapien: A simulated part-based interactive environment.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Sapien: A simulated part-based interactive environment

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.566206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.566206Z digest=sha256:8fb7d001084add2cc783e2fed634b3878f104a5f0f92098ebdfa28ad74d405ad

Observation 9dd4d635-b4ad-4bc8-8f42-22e06d3cfd02 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.575264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.575264Z digest=sha256:51ab235f201deade3938435180ca937dc82abea2e9c91a93f4e177b4c5eee641

Observation a9605982-a5e8-4281-bd7d-412c877c068b · outbound

This paper cites Manifoundation model for general- purpose robotic manipulation of contact synthesis with arbitrary objects and robots.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Manifoundation model for general- purpose robotic manipulation of contact synthesis with arbitrary objects and robots

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:10:29.184845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:10:27.580101Z digest=sha256:2ab3707e8d91786d54c39c6d5c8bd77d5a49e58b246b42bdb08474e76beb946e

Observation 2eff18e7-954d-4c87-a07e-d8a0f04ac6db · outbound

This paper cites Robomm: All-in-one multimodal large model for robotic manipulation.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Robomm: All-in-one multimodal large model for robotic manipulation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.584573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.584573Z digest=sha256:b9dd039be0fe3c51e8b9f2eda7c1b32aace5b933e663d57473b8a5fc51462f04

Observation 64e90c23-6fcd-4d53-b330-be5927ca307a · outbound

This paper cites Qwen2.5 Technical Report.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Qwen2.5 Technical Report

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.589117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.589117Z digest=sha256:1818474c9020a44cab870ab49840c96b3bdcf6d5de50b55b40211184ccaf60c6

Observation eaa65cd4-029e-4e8e-a201-dd3da6cd5161 · outbound

This paper cites BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network Quantization.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network Quantization

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.594157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.594157Z digest=sha256:81473a75e6a14ded68b7a4e7f3e08d38c63fcb77e6e2a101d23f62d06d589711

Observation 8219064f-8bc8-4ea4-a728-f1485c32fdee · outbound

This paper cites Learning Interactive Real-World Simulators.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Learning Interactive Real-World Simulators

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.598959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.598959Z digest=sha256:9fd52b28fa219ca91f711035cea27ff3dd51576b26368ef31b5ce78b182a142f

Observation 47a7d6b6-d25e-40ab-a63e-2e5e1386a40f · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.603536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.603536Z digest=sha256:47027b1081c9d7af9aed3317095309b996e964f6e7b38eb0dd86eac628d975b5

Observation 9b59f4b3-ec40-4353-a1e8-0dacdbc317c4 · outbound

This paper cites Latent Action Pretraining from Videos.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.608248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.608248Z digest=sha256:c861212fb63880b21c45f1001335a65afb7dc77a38a499ebd62b3d1a1a16f7cd

Observation 8444b4ae-aabd-4e78-9011-0d699d113629 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.612740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.612740Z digest=sha256:f8860546534c6fd079696c5ab6a406eda8aafcf48d01f867b206adfee03cbbe0

Observation 2d95718c-b065-4b13-a942-bd12653ddf7a · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.617413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.617413Z digest=sha256:1fac5f07ee86abefbfd31be8a117bb0a27982fdf62b771731b62eaf76fae958a

Observation d7d5761f-c4cc-41e3-8a1d-4b9652b7e651 · outbound

This paper cites Sigmoid loss for language image pre-training.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Sigmoid loss for language image pre-training

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.622572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.622572Z digest=sha256:c6168ccbc8ed2db3c872bd445982a52727caccf46a995f47c47d188f84cf697a

Observation e21363c4-de58-43e7-addd-118d46879e4d · outbound

This paper cites VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.627121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.627121Z digest=sha256:68a8df74636e3302f83fd46cb341835f8fd023f38ef8abaf0e66498473b3abe2

Observation 700946ae-e552-4edd-a701-778f98350994 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.633325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.633325Z digest=sha256:00e79f1356f089c6cfa51575c93bcbdacc341bdde0c04a41aeeb612a8ca2aafd

Observation 49be3601-4f4f-4377-9ba8-78e6b763d4c4 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.638634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.638634Z digest=sha256:cbca400084ebb08c3eb2afb48af98e58df1f865a160a3b72ffe9bd6afb5ec4c4

Observation 43390db1-f8f4-4f41-8aab-4331000e72e6 · outbound

This paper cites ALOHA Unleashed: A Simple Recipe for Robot Dexterity.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models ALOHA Unleashed: A Simple Recipe for Robot Dexterity

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.643489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.643489Z digest=sha256:8dc133a4843b7f6d3d84e719f3f4d06ac85513df88062a1bb2438c3bde903f77

Observation 9ff03f4e-97e5-4870-9e1c-803d3c3de889 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.648196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.648196Z digest=sha256:fbbb2c8ca20f9ea33b0cf62b6e78cc1aed5d4d453e87702a27f13989ea6088d6

Pith citing papers

Observation 1b1409b6-a8a0-4b15-bccd-5823be6bceb9 · inbound

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:43:24.461538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T20:43:24.417901Z digest=sha256:51a2e61cb8655cb4f96e66ec7f3e39b27f9679533e078fa96819b928b5b47873

Observation 4d9825f4-f8e2-4510-b434-fc1732b5d445 · inbound

Galaxea Open-World Dataset and G0 Dual-System VLA Model cites this paper.

Galaxea Open-World Dataset and G0 Dual-System VLA Model VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:31:09.958850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:31:09.958850Z digest=sha256:ef402f655754857729d849f0b3930ff10c6160bf752e0617b35fac025a340ed5

Observation c3266e4c-70c4-4b32-a59f-e9a0947a83dd · inbound

Continually Evolving Skill Knowledge in Vision Language Action Model cites this paper.

Continually Evolving Skill Knowledge in Vision Language Action Model VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:04:09.142663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T06:02:31.638120Z digest=sha256:c21f42d13b1517c11a926f2b0eae4ae87a358a22788687f4cf23ec0b9e5d2116

Observation c19d9121-4ffe-4a89-83d4-3ddb452ead54 · inbound

Mixture of Horizons in Action Chunking cites this paper.

Mixture of Horizons in Action Chunking VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:32:26.811080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:32:26.811080Z digest=sha256:ef1fc3b69ae7f81d72eef3e844b11a28efa4f7832659709907b3b70ec957a32c

Observation a8ef4a7a-3f03-4953-a533-2a3faadffc76 · inbound

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts cites this paper.

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:26:10.231298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T09:11:21.715023Z digest=sha256:8aa685c0089abd8402da213e05f0546967df20db5f9b6e6a80e6d56b35382a3b

Observation 9fce7ef5-7bec-4bce-a585-0e7d327e8a39 · inbound

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation cites this paper.

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:25:52.929046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T02:25:23.710842Z digest=sha256:121bde034a197faff20ec15525dba1c32bffd8369cbacbcbdc2e1b77e342c17d

Observation d7a4b6d9-4f86-46a6-bb5b-8b5e954b8d57 · inbound

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation cites this paper.

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:46:04.519797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T04:46:03.020800Z digest=sha256:f2eab340f7936f2f7e84c0bd233f538d8c73aaaa21b24e530aad8095afe46f0e

Observation 44aea398-8255-41bc-b582-90bb43ed3b0e · inbound

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance cites this paper.

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:44:48.603453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T15:36:13.134340Z digest=sha256:5ce611e2819eaeeece9bff0cd79c38648f7f88b617ac8fd000f1c83c15c6cde4

Observation 32b0416f-6a3d-4e5d-ab44-9ef473cb0564 · inbound

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents cites this paper.

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.571669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T13:36:56.552395Z digest=sha256:f979c06774bc43a20cbfe649a065002c7657231456c6cfd64c8f7e56fdbf7af1

Observation 27885bd4-0d9a-4583-aaca-4507cc721d18 · inbound

S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation cites this paper.

S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T04:43:06.349780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-29T04:41:57.643116Z digest=sha256:3e7a03c82f39ce17ba0ff90a6fcdf85aa884ac1a879715fbb075080655fc9499

Observation 3541f0cc-245d-4df8-a6af-7f0f93129804 · inbound

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging cites this paper.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:ab247a79b83308724b564aeabf7677f64c6defe76de7bfb331612ca12c2132d3

Observation 9c481aa7-258e-4ac7-bea4-c374833899c7 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:35.190360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:35.190360Z digest=sha256:49cc03c1294c75fdfe3e31e8e0412377d00ea585a2a21c18179efe05a85dcfb4

Observation 05b03699-62f1-40f9-b3c7-5e0390014d80 · inbound

Transforming Remanufacturing Automation with Large Language Models: A Forward-Looking Analysis with Case Studies cites this paper.

Transforming Remanufacturing Automation with Large Language Models: A Forward-Looking Analysis with Case Studies VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:47.724159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:47.724159Z digest=sha256:5097ef26ffee61684957e942069bd73e908ea219f37d1be6096b99c497ca95c5

Observation e9e93aa3-4f76-4be8-883e-411a146f47b1 · inbound

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation cites this paper.

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T19:38:41.712173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:38:41.712173Z digest=sha256:dec717e86b00a504585e5f35250da550e1b9d5f5253a9954f8d0cf97cb539782