Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T21:37:50.617813Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 47 inbound Pith citation observations for arXiv:2412.14058.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T21:37:50.617813Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:55:52.186146Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T05:36:01.252344Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9f43182c-4218-4708-81e3-a183d7ba0b05 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6327640b-6520-4734-9dcb-076e52bd1caa · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d2b49e18-aab9-46ff-9f7e-53173b115b85 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots PaliGemma: A versatile 3B VLM for transfer
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bcb6b6ea-2692-4ab3-9e0e-7092a78ef483 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5b4e2902-16db-4ccf-a9b8-c08bcffd4191 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0205a320-063b-4fa4-a7bf-8f828a9ac790 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots RT-1: Robotics Transformer for Real-World Control at Scale
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8f70d6d8-e7a0-4264-823a-6d06bd4319d0 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9418149c-13e5-4fdc-84d4-6477d619c478 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bede542-7190-4ddd-980e-1361f1aaf294 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Diffusion policy: Visuomotor policy learning via action diffusion
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5c1194ee-339c-4ccb-8aac-2e6e19e9d6db · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1be9bf5b-b4c3-48c8-a1f7-7f4725d8ecee · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2bb8beeb-d451-4af2-9fbb-94ce424e233a · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Model-agnostic meta-learning for fast adaptation of deep networks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 84705487-e9db-481b-96a7-2f4d9c1ba5d4 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Gpt-3: Its nature, scope, limits, and consequences
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec2dab5e-509a-47da-9237-afa7bfcc9518 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Long short-term memory
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1bff4d0c-4c89-49f2-abc0-fa8959b304d5 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e78c6868-9ab3-415c-8280-acc279f32aec · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Perceiver: General perception with iterative attention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b6e94f44-696c-491b-9efe-f8a7f1ee7f22 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Bc-z: Zero-shot task generalization with robotic imitation learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cf67eb99-1617-40ce-a7c2-52496ef7c858 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots VIMA: General Robot Manipulation with Multimodal Prompts
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d55420aa-f5ab-40ef-8bfb-b696c5d40ff1 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 042a928e-46ca-4f70-822b-68c8bad0a779 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots OpenVLA: An Open-Source Vision-Language-Action Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df37c3b8-5d67-4673-8bb0-e408b0af1b93 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a307198-f89e-4159-b195-7ec5d128335b · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Vision-Language Foundation Models as Effective Robot Imitators
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b081c3b-5c2f-46de-a2e5-e7d32ea62e91 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Evaluating Real-World Robot Manipulation Policies in Simulation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ed40d69-2c67-4419-8d4b-30fa59529416 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Flow Matching for Generative Modeling
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 02893c26-0624-4d53-828b-b8e2cadf64bc · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3b7fd351-9604-4a32-b8fd-48b7d726c916 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Visual instruction tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3886197b-86db-422c-aa03-5c4ff445cc46 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Embodied intelligence: A synergy of morphology, action, perception and learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 290b5786-1f73-4067-a976-d11cdcb7f124 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1edfb71d-d626-4eaa-8ad5-5afc19a43340 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Recurrent neural networks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aac55734-6788-43db-aa99-6dd726fbc1bb · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 877634e0-c398-45ea-bd70-ae617c33e6f8 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Attention bottlenecks for multimodal fusion
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c37c1c17-84f1-47af-b705-f85104062e88 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots R3M: A Universal Visual Representation for Robot Manipulation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc30b2aa-16b7-48f4-800c-82aa1ba557b4 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9919b5a0-4dbd-4ad9-9e9f-be9110487b93 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b33b9288-af0f-45a5-a5e7-5f48c27ddbdb · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Real-world robot learning with masked visual pre-training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e91eae74-9107-415f-8f31-3356a00f644b · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots A Generalist Agent
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 345cdf66-8b1d-49ad-bb48-0028d7eba8e1 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 288a73e2-40a5-425a-a837-c0f7d44d0638 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Octo: An Open-Source Generalist Robot Policy
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 53be6ba4-0d61-4c3e-81dd-6bd5782e8c53 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation de5219c6-f0ea-4710-8d13-ef8de3f2de45 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Uform: Pocket-sized multimodal ai for content understanding and generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dca11815-5cef-4c2a-8bf5-6bd5aea0f969 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Attention is all you need
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 450d7ae3-4b84-4d3d-b226-755bdec2009b · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Moondream, tiny vision language model
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 11ef8a8c-0395-4273-b77e-63458688bd34 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Bridgedata v2: A dataset for robot learning at scale
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3680f861-18c4-4642-a5d4-6e6b6f7a5924 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 45b2bbee-8bf4-46c8-a56c-125019b862e6 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f274156-4873-4ef9-ab0f-4a87e6bc7137 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation affd3e59-1ee0-4735-8917-06ed7dde6ea3 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4b663d66-8493-4372-a97a-1ac02bb8f69b · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Latent Action Pretraining from Videos
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6b7cc466-0404-408a-9c6b-c4427ec8cb95 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f20cad6-497c-44f3-a411-8e205100dd2a · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Robotic Control via Embodied Chain-of-Thought Reasoning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 26e200d2-dd94-4409-b05e-db71f1772e7a · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a0045528-21f9-42a5-af08-787fab6777af · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots Sim-to-real transfer in deep reinforcement learning for robotics: a survey
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a3c3806-e811-4803-a045-c40be984e3f8 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots 3D-VLA: A 3D Vision-Language-Action Generative World Model
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7b90b889-cb7a-4be7-bbaf-b74925dd9442 · outbound
What Matters in Building Vision-Language-Action Models for Generalist Robots ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f02b858-b7d1-47a3-81c8-de4291621983 · inbound
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3e5c5f32-e6a1-4867-b353-6c4259945e00 · inbound
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2d9fe60-5df8-4a13-999d-af208700494e · inbound
UniVLA: Learning to Act Anywhere with Task-centric Latent Actions What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 14a5a2ee-4f77-46f6-a14f-3c287f392eb1 · inbound
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e5560d0c-f671-43d8-b56b-f9f4c0d786df · inbound
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aa552093-d4af-4d3e-82ba-3b401294b059 · inbound
GR-3 Technical Report What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c6f04bca-d4aa-40e5-bc1a-5929e0aaff62 · inbound
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 941c0693-7845-4c8e-900e-1492f4a0ccc8 · inbound
Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05320bbe-34ca-4a1e-a85c-b9294b827571 · inbound
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9002ed8-887a-49fa-9d4f-fdc60d2ae041 · inbound
FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6887cc6b-0618-4f86-a323-39a4b08a9c20 · inbound
LLaDA-VLA: Vision Language Diffusion Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad47c271-72d7-4eb1-b772-2be9468b1762 · inbound
RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43c0e31a-92c7-4be8-ac5a-6b9597484dea · inbound
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae57cef9-d09c-4f3a-a437-919db1abb4e4 · inbound
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfca82c3-030d-4500-ba8d-b02ce8fcf60f · inbound
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0823ed07-a098-43fd-bb65-84a784894e10 · inbound
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6a301265-94fd-406f-b8c2-9b42cef5380c · inbound
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17df520a-3ab1-4f7b-8b08-0b58539720cf · inbound
Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde4c1cf-9f0f-45c3-a83d-eca0a94d1a50 · inbound
PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ada60f4-cda7-444f-9bde-9c004195bb68 · inbound
Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87aaae01-72b2-4596-833b-9378e6f387a7 · inbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54a2240a-1abb-4040-8116-622980df07a2 · inbound
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76a9d1ea-a4be-41d3-9f12-e9bd4cee6c22 · inbound
A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27b6b2fc-8d27-4e7d-a3c6-33eed07672a0 · inbound
JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 32ca6a19-7b21-40a5-9197-1f9b20191251 · inbound
Bimanual Robot Manipulation via Multi-Agent In-Context Learning What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86bfb119-9e23-4e0e-8518-b0d0aa742974 · inbound
VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 909513c3-68d7-4d84-a971-79577968e705 · inbound
VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 737b6042-96ad-4d5f-bfa9-86e7651e7b7b · inbound
Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a725cd6-4940-4c71-a2f7-9ca3778c0476 · inbound
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c45d145-87a7-47b6-a5fe-19764a7312ed · inbound
RotVLA: Rotational Latent Action for Vision-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4008f89-b3fa-4021-b7c3-d0a3999d9938 · inbound
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9fdcabab-c52d-45b7-8230-af335289e687 · inbound
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ba6987e-b015-410e-a7b4-6cb70b030d51 · inbound
PhysBrain 1.0 Technical Report What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5fa77bc2-1fb2-489a-96f9-3d1cca4507bf · inbound
OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 33aced77-c7df-4dee-bb7d-8c91acfe6590 · inbound
VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76a1b0c9-b239-4667-8ca3-67eb07f8b574 · inbound
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ed99c49c-fc2a-4b33-86ad-b7653c49ee90 · inbound
TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dfd867b6-7c44-4bfd-888d-9ad0abaf0381 · inbound
3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aedf29e1-5680-4cfa-a61e-0f673e20da5b · inbound
LARA: Latent Action Representation Alignment for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1fe5ca7b-9480-4b7b-ad3c-be6f2500dcbc · inbound
LARA: Latent Action Representation Alignment for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8380a49b-2ac3-4eea-a7bd-13f20dfd7dfe · inbound
What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 92c15b20-0b1d-4bff-b1a3-7b61a365a99d · inbound
VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9891ac9f-4a8e-4538-86fe-69fe07102273 · inbound
HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0dde5dc2-ca81-465d-b62b-251bf7a85223 · inbound
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9dc5e2df-7a63-44a3-9226-42faad756336 · inbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2bb9c8-09c2-484a-9e7b-208f63fd66f2 · inbound
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c5727ca-8a7d-4120-8242-c4e4033704a9 · inbound
Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.