Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:10:27.648196Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 14 inbound Pith citation observations for arXiv:2506.17561.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:10:27.648196Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T19:38:41.712173Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
100 of 109 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 9ec5de35-6096-4d2f-8296-10f60ef2fd82 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d926090-8c7b-46aa-a059-5a883583aeea · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d03ec14-9f3e-48c2-b990-72f4cc939298 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Dexart: Benchmarking generalizable dexterous manipulation with articulated objects
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b7bbc82-118a-485e-9d98-338d4d1fac28 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Minivla: A better vla with a smaller footprint
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aa62eb3-b3cc-410f-809d-fd521ce12374 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RT-H: Action Hierarchies Using Language
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf2ae482-fef2-4bab-ad87-67c9bd34d633 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models PaliGemma: A versatile 3B VLM for transfer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35b47c6d-5c46-4310-b54f-4e5b814b9559 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ad6bc2-204a-4beb-8e66-71771cb81664 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d421add-9cf8-4ff3-8c02-89de5e8cd8b5 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models MIT press, 1982
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0f215ab-bfcb-460d-8122-9fb53d629c09 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae157066-4690-4f23-9fe7-7273fe336801 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c2fa5b2-242f-4fb2-9b04-f0fc4d1e03f9 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf6eb3d-16b3-43c7-a976-d831098c78f1 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Univla: Learning to act anywhere with task-centric latent actions
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d42b62d-b0fc-4115-a016-0d3b712ff8e8 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models MetaFold: Language-Guided Multi-Category Garment Folding Framework via Trajectory Generation and Foundation Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcbf8451-f69e-4dd3-85ef-88806e85856a · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ddca0e3-d259-40f1-a1f5-800e9a5cb0ce · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 720b0bd8-8bf7-44c7-a166-e0a64d9670ce · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D Manipulation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c7cd1f0-3902-43c5-a7b9-3957e217126c · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37a6dc63-29f5-41d1-9baf-1838eeb6e3e1 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Palm-e: An embodied multimodal language model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbd3ecab-970e-41ff-a67d-1968169d066f · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66b63c11-d1cc-4147-82b4-d16af2d16d1b · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Scaling rectified flow trans- formers for high-resolution image synthesis
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae7c703-13b9-4e3b-a1b6-f9b37008791b · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Helix: A vision-language-action model for generalist humanoid control
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34badb41-4d38-4ea4-ad74-7936abfe692d · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Foundation models in robotics: Applications, challenges, and the future.The International Journal of Robotics Research, page 02783649241281508, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5cabab8-3937-42e5-8f40-3e27b6781869 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Cril: Continual robot imitation learning via generative and prediction model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 127ace33-b85d-4412-8d4a-0393891ec53c · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Transferring hierarchical structures with dual meta imitation learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9446cf73-ef99-4828-8a0b-a21c70511013 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Iterative interactive modeling for knotting plastic bags
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfe5a7f-32b7-4311-92b8-5082c6c98d24 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bc351a4-0c81-4930-868d-ad444cef80fa · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Peract2: Benchmarking and learning for robotic bimanual manipulation tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a794c05-14d3-468f-b720-cebc3611754b · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3784f67d-ee61-43dc-b36e-e58ca98ac549 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25c08f81-c3f0-459f-bb21-d5422ca2b2f8 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccee3d34-3d57-4f4f-8959-24fc8e79b7af · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation.The International Journal of Robotics Research, page 02783649241304789, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90667bea-1a84-4668-95d3-4a67e3bc3843 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb32158-d27e-4b4d-86df-2f3fd0827424 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4570be2-5ab4-45b8-95b0-e48641531207 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44974a0e-f713-441e-bc0a-d6345852dd51 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12bfa25a-b0d4-4149-bc4b-29f0e6a50825 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ad3dcf-ac87-4f66-93eb-402d0ee45e2f · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4528cd50-4c89-417b-9b54-83855159888b · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a1af65-5c05-468a-97a2-41e65c9067c3 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d386a0-5c77-4e66-86a3-d21109583148 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce8e387d-0f3a-4f19-af60-efa5883cef10 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58bede9a-0671-4b96-8878-712b6279513d · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19f38507-8370-4ef5-92ca-796f0a2e8502 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51b81117-1059-4855-a127-d0f2bf9b9184 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Segment anything
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f8815a-0023-4b5c-8fc3-037002ce252d · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d1d8ce3-4fa7-43eb-8964-f2492252831c · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62c8773d-0242-4c74-ac32-6ed1f9555e05 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf68b7d-6d66-4354-8bef-253c7b323f79 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Code as policies: Language model programs for embodied control
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ade42fed-3a92-4aa6-9ef0-fc298af79248 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Flow Matching for Generative Modeling
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea34b8be-fa20-45e9-976b-5a5ca993a081 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d21b8f-5b49-4905-8601-b7eb7574d7e6 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Visual instruction tuning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5efa89d5-6566-49d4-bd5f-119da8ab5bb8 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Rectified Flow: A Marginal Preserving Approach to Optimal Transport
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 084827d1-0369-4d6e-9e27-ab208db2038e · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e47f889-03a2-47de-b7ea-7d58862239d7 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Structured World Models from Human Videos
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ebc49d-9fa4-43e9-b1e1-27180a0e8fe3 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dba92367-874c-441b-b2c1-f972801ef107 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fd7db5f-5287-417c-80ba-b875d3cc9c9f · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models DINOv2: Learning Robust Visual Features without Supervision
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c4e211a-9b94-437d-ab62-bb3861ccf149 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e14c9f2-3112-46d7-99ce-e879ab1510c6 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11d6191d-ae5e-4a7f-9370-f7941eac36d7 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d1547a-b8d4-4fee-ac8d-ed8b48754385 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 370a51f9-414d-4da0-b0c4-d64ae2220bbb · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183f2236-2e46-4520-87be-7fef6f368d8a · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Open-vocabulary Mobile Manipulation in Unseen Dynamic Environments with 3D Semantic Maps
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e56934cd-e5c5-4973-ab06-af846be78952 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59875891-dd11-48c4-897b-f2c7e59d664a · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Am-radio: Agglomerative vision foundation model reduce all domains into one
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15eb20d-766d-48bd-bafc-58f858d803e3 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models SAM 2: Segment Anything in Images and Videos
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e0ec3e-c0ef-4d4c-8f7a-77b6477186c6 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d8dd8cd-9f4f-4400-aa9d-dbdfbeaf9398 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Learning to Act without Actions
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4479e97e-078f-44d4-8deb-75ef88d4657f · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dc55914-ea42-4a7d-8a3b-6e358567cd2c · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc06fdbd-7fe4-4365-b8b2-ca056bab4c37 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Progprompt: Generating situated robot task plans using large language models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a1c103d-1ee5-48af-8ff8-9e139eea3969 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ffd21c3-9ee9-48a5-9617-a1283b5a6467 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e66d6e-ecd0-4584-b8ca-637fa755131d · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1b856c-bce7-4bf2-8e19-17ae716835eb · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 315a5d10-127f-4f33-9e51-dc51f21ae169 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Llm^ 3: Large language model-based task and motion planning with motion failure reasoning
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f36ef7-0f0f-4ef1-903e-3b4793e5bd2e · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Chain-of-thought prompting elicits reasoning in large language models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3a2dd9-3a41-4008-996c-361cf17c92fa · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models $\mathcal{D(R,O)}$ Grasp: A Unified Representation of Robot and Object Interaction for Cross-Embodiment Dexterous Grasping
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ebc802f-158b-469f-9730-d32e35db6c1a · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Any-point Trajectory Modeling for Policy Learning
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa52847d-d72f-4257-a581-a7e6d0f8791a · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988e8b70-1b10-4ca3-aae5-44bd9776cdb3 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d93be186-6d36-47d8-a0cc-864385877482 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ae7b81-399b-4d01-916b-0c10a3d845d9 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Sapien: A simulated part-based interactive environment
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dd4d635-b4ad-4bc8-8f42-22e06d3cfd02 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Florence-2: Advancing a unified representation for a variety of vision tasks
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9605982-a5e8-4281-bd7d-412c877c068b · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Manifoundation model for general- purpose robotic manipulation of contact synthesis with arbitrary objects and robots
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2eff18e7-954d-4c87-a07e-d8a0f04ac6db · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Robomm: All-in-one multimodal large model for robotic manipulation
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64e90c23-6fcd-4d53-b330-be5927ca307a · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Qwen2.5 Technical Report
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa65cd4-029e-4e8e-a201-dd3da6cd5161 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network Quantization
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8219064f-8bc8-4ea4-a728-f1485c32fdee · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Learning Interactive Real-World Simulators
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a7d6b6-d25e-40ab-a63e-2e5e1386a40f · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b59f4b3-ec40-4353-a1e8-0dacdbc317c4 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Latent Action Pretraining from Videos
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8444b4ae-aabd-4e78-9011-0d699d113629 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d95718c-b065-4b13-a942-bd12653ddf7a · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Robotic Control via Embodied Chain-of-Thought Reasoning
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d5761f-c4cc-41e3-8a1d-4b9652b7e651 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Sigmoid loss for language image pre-training
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e21363c4-de58-43e7-addd-118d46879e4d · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 700946ae-e552-4edd-a701-778f98350994 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49be3601-4f4f-4377-9ba8-78e6b763d4c4 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43390db1-f8f4-4f41-8aab-4331000e72e6 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models ALOHA Unleashed: A Simple Recipe for Robot Dexterity
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ff03f4e-97e5-4870-9e1c-803d3c3de889 · outbound
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b1409b6-a8a0-4b15-bccd-5823be6bceb9 · inbound
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4d9825f4-f8e2-4510-b434-fc1732b5d445 · inbound
Galaxea Open-World Dataset and G0 Dual-System VLA Model VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3266e4c-70c4-4b32-a59f-e9a0947a83dd · inbound
Continually Evolving Skill Knowledge in Vision Language Action Model VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c19d9121-4ffe-4a89-83d4-3ddb452ead54 · inbound
Mixture of Horizons in Action Chunking VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ef4a7a-3f03-4953-a533-2a3faadffc76 · inbound
VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9fce7ef5-7bec-4bce-a585-0e7d327e8a39 · inbound
Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d7a4b6d9-4f86-46a6-bb5b-8b5e954b8d57 · inbound
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 44aea398-8255-41bc-b582-90bb43ed3b0e · inbound
Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 32b0416f-6a3d-4e5d-ab44-9ef473cb0564 · inbound
What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 27885bd4-0d9a-4583-aaca-4507cc721d18 · inbound
S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3541f0cc-245d-4df8-a6af-7f0f93129804 · inbound
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c481aa7-258e-4ac7-bea4-c374833899c7 · inbound
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b03699-62f1-40f9-b3c7-5e0390014d80 · inbound
Transforming Remanufacturing Automation with Large Language Models: A Forward-Looking Analysis with Case Studies VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 134
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e93aa3-4f76-4be8-883e-411a146f47b1 · inbound
Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.