Pith. sign in

Paper Citation Record · LEDGER

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models

As of 14 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2603.05147.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.05147 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T18:48:54.233217Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2db258e-e7c0-4e4d-8348-5ca9f54cb7cf · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Robotic control via embodied chain-of-thought reasoning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:51.342213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:51.342213Z digest=sha256:a2b749f69834d0577e1d2dcfa472ea0ffaefad8ce403ea389a7b2ba14e5052f8

Observation 24964ee5-06f2-4113-bc0e-76975f88b8d7 · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:51.431250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:51.431250Z digest=sha256:6e390227e39c759be87830dd60077a7dd30e643077715d5e0347fc617e366193

Observation 3b1ef08a-15bf-4565-a708-c708f6643013 · outbound

This paper cites Fast ecot: Efficient embodied chain-of-thought via thoughts reuse,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Fast ecot: Efficient embodied chain-of-thought via thoughts reuse,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:51.539689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:51.539689Z digest=sha256:8778393adfff354446e0da680ac6a4d6a0eb06d79cd00b4d6f51c76ab19a466b

Observation 6f98cfc5-8a08-4985-a5e5-d3a24d0d8305 · outbound

This paper cites What matters in building vision–language– action models for generalist robots,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models What matters in building vision–language– action models for generalist robots,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:51.625938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:51.625938Z digest=sha256:8e024842971fdb3e070193ec29e09c823f45d40188ce53e64613b4b2c5915526

Observation a05419c4-9e03-4566-9c81-759097979ccc · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Bridgedata v2: A dataset for robot learning at scale,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:51.722517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:51.722517Z digest=sha256:0c169f1b2d6d0c18e450c59f8918da2c8ad87f6de7a5f9a3d10f1decbab3a71f

Observation 6d9a28d0-581e-44b9-9ee3-631407d90cc5 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:51.844145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:51.844145Z digest=sha256:8d6b7d9f43c6d3b4f49c529383bbcbc88d6bf3dd5c21f53c9d4d1ab216c38041

Observation 944a38e9-5f64-4e69-9ceb-bd0979ed8521 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:51.936566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:51.936566Z digest=sha256:154c9df2b1bfd58abac60fcf2d513774cdd0b57a1069c19a6f050b4eee34e194

Observation ae8431e6-7195-4066-9964-bcf1e8d8736c · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.069550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.069550Z digest=sha256:32eeb5d82a7bf3e7b980d07475cc4eb789df15390df6592f800715ca0c58f7c2

Observation 423453be-298b-4b5d-b5ba-3d088f9b4197 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.167088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.167088Z digest=sha256:c572bcb6a623ee20ba43d362f3c2f6dbf4becf976c7c2e3dba40e61424f6386f

Observation f0589c35-40dd-4b9d-8a10-16b87376a1de · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.236235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.236235Z digest=sha256:9af98e61dbdd406bd15b2071059ce4782fed9b692aee3b7d7145365cc6daf32a

Observation b7755610-7e70-4994-87d7-d40accde8d50 · outbound

This paper cites π ∗ 0.6: a vla that learns from experience,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models π ∗ 0.6: a vla that learns from experience,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.321488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.321488Z digest=sha256:35cb3c059b9dd1320e2d262f053a17d7361b9d645103ae8126d41700d7025de4

Observation 71402313-8f76-4efc-a5be-c6222c35c19a · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.392191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.392191Z digest=sha256:e77b2f837d3b31fe207d13471eb95ebef0e37648bc59275e880c21b01f8f7cec

Observation ba4bc401-fab4-48ee-8754-4ba42047d549 · outbound

This paper cites Smolvla: A vision-language-action model for affordable and efficient robotics,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Smolvla: A vision-language-action model for affordable and efficient robotics,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.509812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.509812Z digest=sha256:7e224ec5ba58c46e2536c2e3ba55cbce928d5c72f4f54f79d22469f228cf9ada

Observation 9d7e39c7-d0fa-47a4-8a04-aa18ad26340f · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision- language-action models for robotic manipulation,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Tinyvla: Towards fast, data-efficient vision- language-action models for robotic manipulation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.685186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.685186Z digest=sha256:deb587ff4f506387ce1d63b8c3a6c0516a9383ee122abf22e4594d86b4b1514f

Observation efb29793-8700-4c0d-b4e4-e7c8cee26164 · outbound

This paper cites Instructvla: Vision-language-action instruction tuning from understanding to manipulation,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Instructvla: Vision-language-action instruction tuning from understanding to manipulation,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.777610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.777610Z digest=sha256:7e9af3d076f465808436579d24114ec700a4853f9f4b42e0d9427a02efd9795f

Observation a0d047ee-566a-446d-b825-7676369d0279 · outbound

This paper cites Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.849158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.849158Z digest=sha256:46850b5fa4f02fab79f11189796120612f63cfbddab2f1e414d80688f4ebe6ac

Observation b709fa55-c3e7-4efd-ae11-97c23f2140f5 · outbound

This paper cites Onetwovla: A unified vision-language-action model with adaptive reasoning,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Onetwovla: A unified vision-language-action model with adaptive reasoning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.933599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.933599Z digest=sha256:51c5966c1fa6f103af0a3b88df316961837c88df39cf0e86ef65b76eddf66ac3

Observation 6c97d867-f5df-4a45-9579-1bea54c45a00 · outbound

This paper cites Safe: Multitask failure detection for vision-language- action models,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Safe: Multitask failure detection for vision-language- action models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:53.033786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:53.033786Z digest=sha256:a4c91e730f59532e8c07d0bcd31a441f1212d3619956bcc17e92a9a5a4c38914

Observation 0cb36279-bc22-4038-bd8a-8f48011192a7 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models SmolVLM: Redefining small and efficient multimodal models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:53.101905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:53.101905Z digest=sha256:2f458e53c1a85f3abe0ddff0f57d93997c32e483c1dc6d02b5ca4303e578a882

Observation 1d2a65e5-348e-4ee3-a952-5440288eca47 · outbound

This paper cites Flow matching for generative modeling,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Flow matching for generative modeling,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:53.309315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:53.309315Z digest=sha256:fbc4932b5f42ab1dab65ebcf08fa57c9da73b8714830fe2cee9e9838ae29aef9

Observation e3ab23f1-9163-4fa2-907d-6c58b8fe9beb · outbound

This paper cites The Llama 3 Herd of Models.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:53.460763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:53.460763Z digest=sha256:a7993f1fd3be70be4dce518b91c119bfbe718b80d664c29bbb98547e7e6f0f58

Observation 4e859834-cb50-4da4-b9bf-4cacc0b97dd6 · outbound

This paper cites Gaussian mixture models.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Gaussian mixture models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:53.619512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:53.619512Z digest=sha256:796fd7302a60078f14061e2e7952da9e2efb317f70ff27aa38516b0a61f11d26

Observation d3ad6f08-25c5-49e5-9cd6-ff58e53d7c47 · outbound

This paper cites The ma- halanobis distance,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models The ma- halanobis distance,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:53.716648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:53.716648Z digest=sha256:2d4e3c304fb82867819e09684abe06aa25497fe89519b2dd33c407bdc4a4d83f

Observation ad71fbf7-8820-4815-85f7-70ee195fca93 · outbound

This paper cites A well-conditioned estimator for large- dimensional covariance matrices,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models A well-conditioned estimator for large- dimensional covariance matrices,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:53.857715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:53.857715Z digest=sha256:f78ea048885d243b639e70c6ff48c76d03cd3b72252ee7713d836f014f112ff9

Observation 37e1eb46-4955-49ad-a21d-4cecec6ec22b · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learn- ing,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Libero: Benchmarking knowledge transfer for lifelong robot learn- ing,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:54.012491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:54.012491Z digest=sha256:15d0dbd303c3e8751f1b3d770146b026b338ac6c4eef384ebccf99fa546d9298

Observation 55b79511-27f2-49c5-92f6-f3e19302baa5 · outbound

This paper cites LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:54.118494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:54.118494Z digest=sha256:a3b683ea3fc0aec71e4d9a83125a6cd967e0351900f8596acab2643e056415aa

Observation 224d4ea4-6cbf-45bd-b00a-c6be90abfeb6 · outbound

This paper cites mixup: Beyond empirical risk minimization,.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models mixup: Beyond empirical risk minimization,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:54.233217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:54.233217Z digest=sha256:62f6792c2d52a29eb2f627aac17973ecdf397cc49f256d7ccb8de6a1e21d47c4

Observation 0aa0d5b0-414a-4e09-a3d6-7684c7023ea6 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T18:48:52.616387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:48:52.616387Z digest=sha256:3202a776802182ebc2ef57c9897ccee7acd7c02c3b5e28748861e1e383380ba7

Pith citing papers

No inbound Pith citation observations are available.