Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T21:37:16.250864Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 3 inbound Pith citation observations for arXiv:2602.19710.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T21:37:16.250864Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:11:17.634768Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T00:39:17.424396Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c7836cd2-72d2-4240-9beb-46ff0e3673c0 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Objectron: A large scale dataset of object-centric videos in the wild with pose annotations
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13787cf2-8a07-4f77-812b-a6d664dcbc45 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies On the representation degradation in vision- language-action models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc1b9b76-3585-4d13-ad50-241b1b6d04a8 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f14761d-9d52-456b-9e60-da1e6a74cf26 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies PaliGemma: A versatile 3B VLM for transfer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4356ceb-fefb-4b66-918b-6a8d9d2f8b4d · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50fa0fdd-f25a-4d27-9cdf-b650b250858a · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32de1248-beff-414b-9a66-056fb6ba3064 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies In9th Annual Conference on Robot Learning, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af5f446-d60d-4fdd-8490-c16efd2d18e7 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Omni3d: A large benchmark and model for 3d object detection in the wild
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3665d143-5566-4892-aacd-59f730cda58a · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies WorldVLA: Towards Autoregressive Action World Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e7fb85d-047b-4cca-a206-12cb859cff1b · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 695d346a-885e-4bbe-ab8f-f360d97c2d60 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Training Strategies for Efficient Embodied Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a825d5c6-361f-4032-84ec-2cb6bbbc3e42 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Language-Image Models with 3D Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f395ef63-03ef-440d-867d-58d13819b45d · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332ec85a-a121-4926-8f00-0b3157c5210a · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Vla-0: Building state-of-the-art vlas with zero modification.arXiv preprint arXiv:2510.13054, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9158f500-ae05-4265-8f33-ca70f58be589 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Seed1.5-VL Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6bc258a-8622-4d21-b897-b3622abba245 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Pow3r: Empow- ering unconstrained 3d reconstruction with camera and scene priors
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36e77b2e-5337-40a6-a07c-f381598307c7 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Don’t blind your vla: Aligning visual representations for ood generalization.arXiv preprint arXiv:2510.25616, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab137981-81d6-46f0-8282-1c87ec346ec8 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies OpenVLA: An Open-Source Vision-Language-Action Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608225c9-c9d3-437c-9ee3-100846ec7415 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies MolmoAct: Action Reasoning Models that can Reason in Space
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e19ce90-9856-4c74-8aa1-fab6df6129f2 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Spatial forcing: Implicit spatial representation align- ment for vision-language-action model.arXiv preprint arXiv:2510.12276, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff1c7ab5-0453-4562-ab5d-224f51eef893 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9490399-d3bd-4063-bc71-c22b33f310db · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Onetwovla: A unified vision-language-action model with adaptive reasoning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6855142e-b978-4985-9626-8d6b5cb7039e · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Flow Matching for Generative Modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6781868-2bad-440d-a316-42773c16ae41 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36, 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ce6743-9a5f-41dc-8a83-b2bb6d4a9225 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4041a377-8649-4145-baa2-65e6f5f33f11 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Rectified Flow: A Marginal Preserving Approach to Optimal Transport
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89b1aa89-e22d-4ce3-a247-bbb8a6ffa1cd · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0d59919-a6a7-49dc-8105-3f2c08e04de7 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50fa8582-1f65-4fa2-8cc4-bab0e4de3690 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Locateanything3d: Vision- language 3d detection with chain-of-sight.arXiv preprint arXiv:2511.20648, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b02e4c95-c70a-488d-b79e-fd43c711c706 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Spa- tiallm: Training large language models for structured in- door modeling.arXiv preprint arXiv:2506.07491, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58dc68ba-727b-4dc8-a9e4-3b0ed97a74f7 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies FAST: Efficient Action Tokenization for Vision-Language-Action Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e84639-41aa-45db-95b8-40ce6cf16bc0 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Eo-1: Interleaved vision- text-action pretraining for general robot control.arXiv preprint arXiv:2508.21112, 2025
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a9f7163-7f55-49c6-b515-1e2d44d64994 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7908e9-4133-462c-8987-902a9cf46eea · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Qwen3-vl: A frontier multimodal large lan- guage model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d727fc-7eb0-4d66-a788-1f2b3873e63e · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbb2ba4e-6064-464e-9287-655ed5bf425e · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Sun rgb-d: A rgb-d scene understanding benchmark suite
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ec095cf-0eec-4824-bdbb-0def141888c1 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Emma-x: An embodied multimodal action model with grounded chain of thought and look-ahead spatial rea- soning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1212fad9-76b7-4c03-8e71-0e9887b77dc2 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Gemini Robotics: Bringing AI into the Physical World
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99aa2b4a-aeeb-4118-bade-36b97fa5adf2 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Octo: An Open-Source Generalist Robot Policy
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 408a266b-fe64-4d6d-a97c-0ced98d15a1f · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9ff308-f506-4843-ada3-067a6187d278 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies N3d-vlm: Native 3d grounding enables accu- rate spatial reasoning in vision-language models.arXiv preprint arXiv:2512.16561, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3354e406-cad3-44e4-a168-aa6be5881022 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b496158-cb66-4d09-bcd6-0b6d266b69c2 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8597b77-d3c7-4985-b394-26b413f23c41 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Instructvla: Vision-language-action instruction tuning from understanding to manipulation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c0049f5-38d9-4087-ba75-7af585525d4e · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Robotic Control via Embodied Chain-of-Thought Reasoning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f4fbad-b8c3-4acc-9277-2b09dd81eb7c · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Sigmoid loss for language image pre- training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d1f9c5-29de-4ab9-b147-807980952787 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5045d1ff-c7a7-44e5-9a3c-0d2322629782 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7774f0c-7958-4a18-9b98-868e92ec8d31 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Cot-vla: Visual chain- of-thought reasoning for vision-language-action models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48770eb8-b9d6-4007-856f-b98651a35fdc · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Chatvla: Unified multimodal understanding and robot control with vision- language-action model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a1e215-21a2-49a5-982f-4a1c8d8ce4a7 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Rt-2: Vision-language- action models transfer web knowledge to robotic control
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cf70152-9102-4e65-b5fd-273de537d6d3 · outbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14adc3dd-c331-4945-844f-7b8981d089e0 · inbound
MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation db88cc04-c30a-4c28-864c-8cfed2513b62 · inbound
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation debb4867-1ec5-4d65-ac42-eeedc4ffc602 · inbound
GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.