Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:38:19.310675Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 10 inbound Pith citation observations for arXiv:2507.05116.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:38:19.310675Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:16:49.933343Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T05:36:01.255549Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e1cd1b59-a6c2-4e96-a21a-a9beda525bf7 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting PaliGemma: A versatile 3B VLM for transfer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 412e0ec6-68d9-4fa9-a3e1-3e63db8f54cc · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting RT-1: Robotics Transformer for Real-World Control at Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7463e2d-0405-4365-9883-6ac655ee6bca · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting PaLM-E: An Embodied Multimodal Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fb197cb-d664-4eda-9327-1b385cbe971f · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting OpenVLA: An Open-Source Vision-Language-Action Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e686f197-da58-4351-bf9a-d9d4175998c4 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80dbc449-ee99-4313-a9d9-87e0e3b600ad · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7e28558-4f5c-4ce8-8fc0-60800e645546 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfab95f5-b1b1-4426-857b-10a4dc6ec490 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting FAST: Efficient Action Tokenization for Vision-Language-Action Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 040cb837-d7d1-4b3f-930e-17b550d5d8d8 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37666799-45b7-4c9d-ab1b-dc77f887f718 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44e84d90-ef23-442e-bc9c-307ac520d915 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Exploring token pruning in vision state space models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 33ff373d-8cb1-47fe-86ab-5060eeca896a · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98b1c47-714e-44a6-af33-980ec31863d2 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5952d23-a5c3-459e-8b2a-932ede49858d · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bdefb09e-276c-4a49-a037-e06ccad9c3b2 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting 0 20k 40k 60k 80k 100k 120k 0.05 0.1 0.15 0.2 0.25 0.3 spatial object goal long Steps Action Loss Figure 5: Training Loss Across LIBERO Datasets Compare with OpenVLA-OFT Kim et al
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 429de2dc-cda3-4a75-82b0-5f082e2f6e08 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bcdd5043-7cc5-47d9-9973-bdd448afb883 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Otter: A vision-language-action model with text-aware visual feature extraction.arXiv preprint arXiv:2503.03734,
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6007afda-9712-48c5-9c39-f17709a3a19d · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9298607-3ab8-4e82-b19c-77faaef27559 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7baef221-06b4-4b86-8311-b766e3dd8be0 · outbound
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6df99e5-6f91-4b17-aa77-f34e0db2dec4 · inbound
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a1651af6-f46b-4280-9f52-a8537f756317 · inbound
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d7b1da5-4b15-4a25-9913-1dc72f49ae4b · inbound
Human Cognition in Machines: A Unified Perspective of World Models VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6a7c1953-01be-49dc-be94-3e098ffe693e · inbound
AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e788380-0f35-4f94-8021-b386b0d88566 · inbound
AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1936401-242f-4e89-94b2-bf03d872615a · inbound
PhyWorld: Physics-Faithful World Model for Video Generation VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 67461939-4baf-44ee-9823-4ca56e207122 · inbound
QuoVLA: Quotient Space for Vision-Language-Action Models VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 94f44ce0-77ed-42e4-926c-ddb43c4e056c · inbound
Flash-WAM: Modality-Aware Distillation for World Action Models VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0816f57c-57e0-4f93-ad67-b58a0d26215a · inbound
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0c8f6e06-90fd-4860-8a5f-b45a3efb5698 · inbound
Self-Evolving Embodied Agents via Skill-Harness Evolution VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.