Pith. sign in

Paper Citation Record · LEDGER

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

As of 12 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.09448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09448 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:19:03.483273Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d7485df-f8a9-4342-bac0-51a73e9298b0 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.219335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.219335Z digest=sha256:5ffb2bc279c6b1700620f09fb4a3e22552f57ac525179c83146be8ef0f66b3e8

Observation 56872ec0-c873-41ae-8df3-c23457c4d512 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.224395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.224395Z digest=sha256:0c5e369907e2614cdf85038e3def2c061c4e004c4df6f9b251c4b48c8ded46d0

Observation 8b03f082-cc58-4d17-b300-139368ba46b5 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.231232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.231232Z digest=sha256:600eb397d26372b0e103355c3dae7838e18c9c752be46326895bca4b96e3ec8a

Observation 0c9d7b80-789b-46dc-a6e1-bbc1bee4502e · outbound

This paper cites StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.240874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.240874Z digest=sha256:4fdd1a284314e0df55e4128379966ade69b03206c4cfc957988ae3380ad17d93

Observation c59441c2-fb7d-42db-abcf-fc0c42e565d8 · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Test-time training with self-supervision for generalization under distribution shifts,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:19:04.868097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:19:03.257251Z digest=sha256:42385ca8097a86f19d2cea666f5c0e1c1ddd8f0f9731b806e89252be943e781c

Observation cc23585c-deac-4580-b185-afe683077b5d · outbound

This paper cites TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:19:04.439639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:19:03.262124Z digest=sha256:6bdb14adf881d4860295507c5b789c3b53b3adee860f9adb18df15c89b4a8243

Observation a30b8e02-da1c-4cef-8b79-1a36e9be4b31 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.284744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.284744Z digest=sha256:220fb0c32a87245be5b5e8e112c27c581f0132258d06207e912588b3fc361957

Observation f9a924dd-2190-42fb-b07e-a996187ae127 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.294272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.294272Z digest=sha256:8a244202f9de62b363691ec9bbf340a07c3272817277516465925c8352203ce0

Observation 25be8609-38ff-48cd-855b-73c55161efe6 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Octo: An Open-Source Generalist Robot Policy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.299674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.299674Z digest=sha256:8c67d4978a1787a1c0878ac7b04da7503b67637854d7cbb5288a8458f66998d6

Observation 431a9520-447a-4ec7-9630-cf1a8639fa70 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.304665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.304665Z digest=sha256:6939694e5ab0c65e0ded4e7890dc6efdd685de6742453c667885ceb9f41374f6

Observation 6ffd3e06-e620-401f-a280-ba04036768ba · outbound

This paper cites FutureVLA: Joint visuomotor prediction for vision-language-action model,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction FutureVLA: Joint visuomotor prediction for vision-language-action model,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.309570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.309570Z digest=sha256:70cad3834c60e8a21495ec5526ab85de0940ef25fb4006bca1a2e357360adeee

Observation c2b2479d-ce4f-4399-a707-15d4549e6003 · outbound

This paper cites Causal World Modeling for Robot Control.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Causal World Modeling for Robot Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.318964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.318964Z digest=sha256:a3b2fddb8bfd70818de23f7b45308e8e93105f84a6fe7d37b29c1f7603f86ac7

Observation 507a4fbb-eaff-444e-879e-5d6a75d44da3 · outbound

This paper cites Tent: Fully Test-time Adaptation by Entropy Minimization.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Tent: Fully Test-time Adaptation by Entropy Minimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.325990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.325990Z digest=sha256:2a9040d520233c1ce57ce6fabde84cba06a929a80eea1d1d6d2a4a707f4cc00f

Observation 5a8deba6-e879-48b9-868e-e8218785dc6f · outbound

This paper cites Continual Test-Time Domain Adaptation.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Continual Test-Time Domain Adaptation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.331608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.331608Z digest=sha256:a9f4895b25671bd189ea3b6785688e5f095f301b715d499c8cef8c1f8103a44d

Observation 669c05a8-3646-4d7b-bdff-4ebb53931613 · outbound

This paper cites Efficient Test-Time Model Adaptation without Forgetting.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Efficient Test-Time Model Adaptation without Forgetting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.336650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.336650Z digest=sha256:839e4f4a8460157863bd9ebaae6b7f97b82ce55939de8a55ee7d322b61e6a667

Observation 8d192dd6-7ec8-41c5-b9c0-cbcd48f40c5b · outbound

This paper cites Test-time training on nearest neighbors for large language models,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Test-time training on nearest neighbors for large language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:19:04.821475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:19:03.348569Z digest=sha256:17988a21023152453dd42205bc76f3ef7508686b57728d6c71c35f38cf1ed989

Observation 0f38ac48-59ff-4e89-97da-d4612b37b7f6 · outbound

This paper cites Test-time prompt tuning for zero-shot generalization in vision-language models,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Test-time prompt tuning for zero-shot generalization in vision-language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:19:04.767806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:19:03.357886Z digest=sha256:27e65e9f639932f8ccdf305afa91e3329eff8d89c464d9e77d7eeb758eacaa9e

Observation 5db3f6fa-2134-466a-a388-6c8160cffae2 · outbound

This paper cites Test-Time Training for Visual Foresight Vision-Language-Action Models.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Test-Time Training for Visual Foresight Vision-Language-Action Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:19:04.126100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:19:03.366148Z digest=sha256:6343ae81422b5d6ecec55c22c6cef964f5696ba05f15e4f995b992052ac74966

Observation 61a32ede-4fde-4d5b-8d89-370c8697a2a5 · outbound

This paper cites On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.371771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.371771Z digest=sha256:e75db91430c019dcad3104ab62aa3dc33478a15098be48dc28817109da3afe74

Observation 3815bd1a-2863-4d54-9054-1e93fd68bfbd · outbound

This paper cites BridgeData V2: A Dataset for Robot Learning at Scale.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction BridgeData V2: A Dataset for Robot Learning at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.376838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.376838Z digest=sha256:5906ced4bd31df9f8e02a9acc79b926f1da49ecdbdbb6077618c40020b18446e

Observation ca3555e0-070e-4d5c-8d7c-3e949bd6e7b5 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.397660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.397660Z digest=sha256:04dcd65953ca9d31e478699bb20b52ba33ddbe088a2d4a43fcb5ed671278b130

Observation 617e83ed-c743-4141-9be3-5d16abf9e83e · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.402213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.402213Z digest=sha256:f62e705230cd71131dfe541edea485dffca45c5238a8b38acb46b438faffe175

Observation 949a4614-ee69-47d1-94fb-140d02b9eb01 · outbound

This paper cites Magma: A foundation model for multimodal AI agents,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Magma: A foundation model for multimodal AI agents,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:19:04.712163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:19:03.414747Z digest=sha256:df16f51db20c60f0f6e4d76b36ebd18c8ea643954c676243481d99e1eb08c3eb

Observation cdf2ca80-3f8d-448b-a9df-71d0c728e1e7 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.424992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.424992Z digest=sha256:dcb3ee0aec9d486db723324cda682f79693b12601fc5508099ce721e4807190c

Observation 62b9da78-3944-4c60-8a6e-594ceaeb8195 · outbound

This paper cites InstructVLA: Vision-language-action instruction tuning from understanding to manipulation,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction InstructVLA: Vision-language-action instruction tuning from understanding to manipulation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.432904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.432904Z digest=sha256:944302524cef3c389fc9cd34afc759e1ab9955cbfaf80f465d03155b23e72429

Observation 99c7043f-3511-4787-ab9b-83851fb181cc · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.440970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.440970Z digest=sha256:5ada3a898c547e5f71e4efa7285b8d264713621164f8ac4b08089e872238ef23

Observation f0e64b66-f9bc-4581-abcd-e1a046a8e5dd · outbound

This paper cites ThinkAct: Vision-language-action reasoning via reinforced visual latent planning,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction ThinkAct: Vision-language-action reasoning via reinforced visual latent planning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:19:04.627133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:19:03.446342Z digest=sha256:60cb6874917d1fe68c9d61e079df13fb2135d5f8471c0ff7c0f280d54df1ae76

Observation b6156ae5-abb0-4e46-8bc1-5b01b162e02c · outbound

This paper cites TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:19:04.594758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:19:03.459498Z digest=sha256:738549a1b503ae3552fdf63568bb162697cfc58fae7565a55930d383e393288f

Observation 77a40b28-65d4-491e-98f1-f2e9adcb9a84 · outbound

This paper cites Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.467276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.467276Z digest=sha256:9c4f393f4f407b2fa8f283488032282c19a3ff8f4d02042c9f9297fdf856d4cd

Observation bff9a389-2130-40cc-a8da-2115d00bb57a · outbound

This paper cites FAST: Efficient action tokenization for vision-language-action models,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction FAST: Efficient action tokenization for vision-language-action models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:19:04.574056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:19:03.475861Z digest=sha256:c43446747abff21c0b05e282fde3c8367edaf72761742f98050c9e014fb005bd

Observation eaeba563-a6aa-4136-9fd9-111e1369df7b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.483273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.483273Z digest=sha256:45094b1f1bf7aa6593e7446dfbc419d5c2d4ca22bce67e44ec87da87fdaac95d

Pith citing papers

No inbound Pith citation observations are available.