Pith. sign in

Paper Citation Record · LEDGER

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

As of 19 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 10 inbound Pith citation observations for arXiv:2507.05116.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05116 v5

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:38:19.310675Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:16:49.933343Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T05:36:01.255549Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1cd1b59-a6c2-4e96-a21a-a9beda525bf7 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting PaliGemma: A versatile 3B VLM for transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:17.284320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:17.284320Z digest=sha256:ed26c1d38697859e79dd1b66ac1b8505c53d6ed11a5f936ec48e7f61b0b88db3

Observation 412e0ec6-68d9-4fa9-a3e1-3e63db8f54cc · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:17.655957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:17.655957Z digest=sha256:486776918d4ffdee2a3d15dc100cf851ed61dc078e8f640914cb4b84fa813015

Observation c7463e2d-0405-4365-9883-6ac655ee6bca · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting PaLM-E: An Embodied Multimodal Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:17.931214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:17.931214Z digest=sha256:83ab50faaf36f49ce5a4b2880c8c9300b0440b03ea0e12a87867a4dc3360eff4

Observation 0fb197cb-d664-4eda-9327-1b385cbe971f · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting OpenVLA: An Open-Source Vision-Language-Action Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:18.159423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:18.159423Z digest=sha256:f43e58c58809438eb2e9189936559881343b8389c62e943444b4e07674adf760

Observation e686f197-da58-4351-bf9a-d9d4175998c4 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:18.346384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:18.346384Z digest=sha256:dba2a84c81283bf106d5b41e2000d9feb7752c3e222159778fcd5078619da64d

Observation 80dbc449-ee99-4313-a9d9-87e0e3b600ad · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:18.477731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:18.477731Z digest=sha256:e89ff0448fe75d59a9762a5e61500e611341a4e49465980128c12735cd6219ad

Observation b7e28558-4f5c-4ce8-8fc0-60800e645546 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:18.585964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:18.585964Z digest=sha256:822faf6e2779727f6a1744b1f4c7fb87fa99e36a652d226573a599fc629138f2

Observation cfab95f5-b1b1-4426-857b-10a4dc6ec490 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:18.681530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:18.681530Z digest=sha256:caceeb3588a7c206028065fc4cef9b3cf2a8bd4958db51c6dc2c0a08242068f6

Observation 040cb837-d7d1-4b3f-930e-17b550d5d8d8 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:19.001984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:19.001984Z digest=sha256:7d7fd156f8685665d62ea25f355f1a240f32fbb43184cf9fb8284d467ee3276f

Observation 37666799-45b7-4c9d-ab1b-dc77f887f718 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:19.176024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:19.176024Z digest=sha256:41b42b0e5b91f4dc15956db26d885658fb18ae6621df0f855a65f165f47e9039

Observation 44e84d90-ef23-442e-bc9c-307ac520d915 · outbound

This paper cites Exploring token pruning in vision state space models.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Exploring token pruning in vision state space models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:19.575538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:38:19.293695Z digest=sha256:f8d2464c3d9d718069c3af458c28f668ab15ce4521e61b5f310c236136491aba

Observation 33ff373d-8cb1-47fe-86ab-5060eeca896a · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:19.296867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:19.296867Z digest=sha256:c545c3233e381c48c7566bc206a9fb571e3e8320b654885f0f28b3bae687ea3f

Observation b98b1c47-714e-44a6-af33-980ec31863d2 · outbound

This paper cites ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:19.300782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:19.300782Z digest=sha256:7ebcbb8008c65ad18ba5914d00299d923e5dbc140d6ae9647c118f15255249fe

Observation b5952d23-a5c3-459e-8b2a-932ede49858d · outbound

This paper cites an unresolved cited work.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:38:19.566737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:38:19.304066Z digest=sha256:c39f080c65d7b3056a05fb4b3ee9f1f3a7e35b0b031e8db3ef8db17590f0e33e

Observation bdefb09e-276c-4a49-a037-e06ccad9c3b2 · outbound

This paper cites 0 20k 40k 60k 80k 100k 120k 0.05 0.1 0.15 0.2 0.25 0.3 spatial object goal long Steps Action Loss Figure 5: Training Loss Across LIBERO Datasets Compare with OpenVLA-OFT Kim et al.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting 0 20k 40k 60k 80k 100k 120k 0.05 0.1 0.15 0.2 0.25 0.3 spatial object goal long Steps Action Loss Figure 5: Training Loss Across LIBERO Datasets Compare with OpenVLA-OFT Kim et al

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:19.558049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:38:19.307349Z digest=sha256:1de74919802e1ecbfe1466fa37bdca00d4fffcdd65ad9dbd8b5958650f0742eb

Observation 429de2dc-cda3-4a75-82b0-5f082e2f6e08 · outbound

This paper cites an unresolved cited work.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Unresolved cited work

Reference 20

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T19:38:19.548391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:38:19.310675Z digest=sha256:02dc531eaf58740c1d6735c2e4ed5ff3914e2b48dafd27006ec563062cba24f8

Observation bcdd5043-7cc5-47d9-9973-bdd448afb883 · outbound

This paper cites Otter: A vision-language-action model with text-aware visual feature extraction.arXiv preprint arXiv:2503.03734,.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Otter: A vision-language-action model with text-aware visual feature extraction.arXiv preprint arXiv:2503.03734,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:18.047728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:18.047728Z digest=sha256:248a09a80ad29fe8c6ba6588508f3a290e4f7c0ac259954b009d9429d69c53dc

Observation 6007afda-9712-48c5-9c39-f17709a3a19d · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:17.779201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:17.779201Z digest=sha256:47fe5be9146aa3df03ee3070943fec1c8963aade576bdbec680d69f93351559a

Observation b9298607-3ab8-4e82-b19c-77faaef27559 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:17.491803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:17.491803Z digest=sha256:8fbc12f061577a9be5e969b445f7aa5305873dcb9bb597efca29165e7f242834

Observation 7baef221-06b4-4b86-8311-b766e3dd8be0 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:18.829099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:18.829099Z digest=sha256:b2a8ecc6d1bcd6a9b94718156986c3d421533b85889f58ecef6315a156be8fe8

Pith citing papers

Observation e6df99e5-6f91-4b17-aa77-f34e0db2dec4 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:20:27.970763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:76fd1a87bb141aa4c5f20072c015b84eee94769485483b02438ad96ce43f3959

Observation a1651af6-f46b-4280-9f52-a8537f756317 · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:15.654858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:15.654858Z digest=sha256:425edb862b7660a80fa8b0c3466f4356be9a9a6cc008127b0086bdcadfc85d17

Observation 6d7b1da5-4b15-4a25-9913-1dc72f49ae4b · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:20:27.970763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:caedda89faf49f2541a544982dade5b688ceebe81117e7e216558e43c012349c

Observation 6a7c1953-01be-49dc-be94-3e098ffe693e · inbound

AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models cites this paper.

AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:20:27.970763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:11:52.984680Z digest=sha256:9fc9f0bfa2d13808556e0e3517e941f20c4ab2e72e8b10cb488b345d0eb419cb

Observation 8e788380-0f35-4f94-8021-b386b0d88566 · inbound

AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models cites this paper.

AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T15:56:08.326817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:56:08.326817Z digest=sha256:0c5124c6b7538d278179392139aca460a5654328196f687049762584382c5b7d

Observation d1936401-242f-4e89-94b2-bf03d872615a · inbound

PhyWorld: Physics-Faithful World Model for Video Generation cites this paper.

PhyWorld: Physics-Faithful World Model for Video Generation VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:20:27.970763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:f647fe4f2a3116274b8b1d0029fde3a1f6ddc3758210d53dffdbe1a3989b151b

Observation 67461939-4baf-44ee-9823-4ca56e207122 · inbound

QuoVLA: Quotient Space for Vision-Language-Action Models cites this paper.

QuoVLA: Quotient Space for Vision-Language-Action Models VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:20:27.970763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T12:09:12.124995Z digest=sha256:09349372a81745bfab2e3b130822e24af58b840b32b1264a18c3f90d1de0bbe6

Observation 94f44ce0-77ed-42e4-926c-ddb43c4e056c · inbound

Flash-WAM: Modality-Aware Distillation for World Action Models cites this paper.

Flash-WAM: Modality-Aware Distillation for World Action Models VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:20:27.970763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T07:00:46.566214Z digest=sha256:17db57e108e474780e5dfb5fa766a5e8fbdef6b4671238b0eed63b628e2c2f36

Observation 0816f57c-57e0-4f93-ad67-b58a0d26215a · inbound

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.257153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:2bc90e9581fecd5b8cba97ee894adb6074f6d8f8c07acca65e49d4908f710cec

Observation 0c8f6e06-90fd-4860-8a5f-b45a3efb5698 · inbound

Self-Evolving Embodied Agents via Skill-Harness Evolution cites this paper.

Self-Evolving Embodied Agents via Skill-Harness Evolution VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:49.933343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:49.933343Z digest=sha256:f039204172741e54b8b8ca7522ea8e89cd9e249affdd2b40cee9fd932836775f