Pith. sign in

Paper Citation Record · LEDGER

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models

As of 19 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2608.11671.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11671 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:39:40.605712Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact4
  • verified fuzzy15
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f5c8d6a-811d-4c38-ab78-5befcc6692e4 · outbound

This paper cites Qwen3-VL Technical Report.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.159696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.159696Z digest=sha256:73f36157fe9a8825392079f1d8d4a21022717bffedc3fbdbfefed2d6346c8d55

Observation 67d75d5f-0648-4b58-b217-e545c44d0264 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models RT-H: Action Hierarchies Using Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.166058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.166058Z digest=sha256:6dc21834d04063eac8886833de30a94235c2ade5c4c29b80120017dc03f577e1

Observation 6dac024d-934a-4e63-9c28-17108811c050 · outbound

This paper cites Motus: A Unified Latent Action World Model.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Motus: A Unified Latent Action World Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.171426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.171426Z digest=sha256:1c4466b975193dd4740cc44d399e194f305e4887222be7cb334f87fb7bd7d095

Observation 3d188d7a-3dcb-41a8-9f41-95057c8b2db6 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.177152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.177152Z digest=sha256:fd362fe9c967d22564960310df45db0bb78d203dd31cf9cb270ca974a0269abd

Observation bc553b3e-ccae-4b9a-9985-c34c2ecf9c15 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.187843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.187843Z digest=sha256:f765202e94fabb179cbb5ec4ad5f6552e21f6d0fbe62a8b3c6373d8f8a0cec7e

Observation 0d1af0ab-83cc-4257-9d54-d35208203a99 · outbound

This paper cites See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations.arXiv preprint arXiv:2512.07582, 2025.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations.arXiv preprint arXiv:2512.07582, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.193129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.193129Z digest=sha256:a70ff02bde881ebf618b963e0b96e02431fd7d9a4122bd12337148471bc0ffe8

Observation 35e919bf-d023-467d-b3ed-a0256c2a742f · outbound

This paper cites Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.198497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.198497Z digest=sha256:ebb5265c3c6c1464525eeae6116641fddda96fba51ecf049a664a332868a9011

Observation 78dbd86f-6df9-4060-87f0-f9c36e3b9296 · outbound

This paper cites From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.204306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.204306Z digest=sha256:3ffb9492eeac366766d166ded654b38db284b7d9c04fbcd73b3875db6ab4fdca

Observation 92387e10-3135-4277-ae10-736ecc3fef1a · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.209987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.209987Z digest=sha256:3f48afe31f1a20024106e40142a69b3b8652862a09a8a7a3cb4a96249e8a4702

Observation 59b30ab7-0fa0-4397-a48c-9c111f559777 · outbound

This paper cites See what matters: Differentiable grid sample pruning for generalizable vision-language-action model.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models See what matters: Differentiable grid sample pruning for generalizable vision-language-action model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.138077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.215675Z digest=sha256:777ccea01148dfbf532554d89e46002313d9cedcefb80512f24d619c9d48bf4f

Observation e1295b98-35fa-4107-820e-4a881ae8b493 · outbound

This paper cites WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:39:41.579937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.220903Z digest=sha256:7814b110c408588f7c1d141fd477e3110a6e42a8218ba48a568b39a185471733

Observation f09b947f-1824-459f-b2ee-31d585b4c5ce · outbound

This paper cites Icrt: In-context imitation learning via next-token prediction.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Icrt: In-context imitation learning via next-token prediction

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.118206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.226605Z digest=sha256:4e36adef7eeb175f974b29f9c5274399d751add5e266184e844529af0b366fee

Observation 5ca88b4a-50bf-46f0-afc4-e195fcc9caa2 · outbound

This paper cites DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.231644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.231644Z digest=sha256:78e3557233f62d1c034c53c66c04f82b876e50141a2469f8a7bc23720c0226a6

Observation b74b918f-6ff3-4a17-bc9a-f7d4e5c201ab · outbound

This paper cites Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky, and Anirudha Majumdar.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky, and Anirudha Majumdar

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.237022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.237022Z digest=sha256:1ac2f3c6d34806c453c8f13837b75599bae1627fee87a9ad025b00cec151a71f

Observation bc03fbe7-39c9-4a07-a4cb-08641716db5b · outbound

This paper cites Motion dynamics learning for few-shot embodied adaptation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Motion dynamics learning for few-shot embodied adaptation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.098520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.241659Z digest=sha256:c6e699ce6257ffbfd4e29eb4a95267081f0017c230114ede632bca532dfdd535

Observation dc9ad96b-87e0-4072-bc9b-3499977baaba · outbound

This paper cites Thinkact: Vision-language- action reasoning via reinforced visual latent planning.Advances in Neural Information Processing Systems, 38: 82782–82802, 2026.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Thinkact: Vision-language- action reasoning via reinforced visual latent planning.Advances in Neural Information Processing Systems, 38: 82782–82802, 2026

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.079782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.246629Z digest=sha256:0e25b55ccf14161e9802c8cffe17577a1b69cdbe118354778dc5e776df7d577f

Observation 21ad72bd-a67f-4dfd-a627-4ab833b294fc · outbound

This paper cites Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.372922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.372922Z digest=sha256:0cbc9c667068f2b4f7983080f3a67feda95d5f7503062cf3d054c8588b4abed3

Observation 289bd257-aba7-4d12-a147-a9f0cf32edeb · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.378753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.378753Z digest=sha256:c620c8f16780bcc7c26daaf3ec6c3cf51f057764afae94912354b1f75a9462b6

Observation dd1e365e-fbf5-4751-a3d7-ed85140f453d · outbound

This paper cites Ra-vla: Retrieval- augmented vla for test-time adaptation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Ra-vla: Retrieval- augmented vla for test-time adaptation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.061087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.383828Z digest=sha256:bb03c5ee27d12fa7768ef50947b930e9f0056bbc26b81b9e7e096605c7becf19

Observation 767bc788-6f2e-467b-b80d-9b6644aeec2e · outbound

This paper cites Ra-vla: Retrieval- augmented vla for test-time adaptation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Ra-vla: Retrieval- augmented vla for test-time adaptation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.040916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.389054Z digest=sha256:f73a144e67175b0021c29bb8f2876f318be81dc0609c04ef5694498298e7aee4

Observation 66fbaafc-a41f-4c2b-b8aa-67c6d9b5983c · outbound

This paper cites RoboTTT: Context Scaling for Robot Policies.arXiv preprint, 2026.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models RoboTTT: Context Scaling for Robot Policies.arXiv preprint, 2026

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.019250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.394348Z digest=sha256:0285413fea06e508dbf45671ab0e38cadf6b952d1016cfc2ff0ab0917024ea11

Observation ebb4701a-c999-4648-a9d6-2ed844b4d514 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.399303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.399303Z digest=sha256:040c449233caa4e4f4f5639e82db92baa2e87bcacde80f14e472b6b85d2f0a6c

Observation b433c465-28c0-4c40-9039-47a72fc38fa3 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.405179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.405179Z digest=sha256:5b9f657e1f2fa6be44182fdc477b36be9525ce0ae701c191d1a9454a7af79a40

Observation 0faafd41-c780-4e71-95a5-70a99bf761fd · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models MolmoAct: Action Reasoning Models that can Reason in Space

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.410246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.410246Z digest=sha256:e34fc4e2b0e3b1d1a56a96b71952c0f02eae64fd02108f1c3346c0d25531e2d4

Observation af116d28-9cb8-4d60-be99-bb45c4d5b95f · outbound

This paper cites CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.416285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.416285Z digest=sha256:e20a4ea373ae87ffaf1e3d73d82592eab970063b6c19d0eab26282f4246ff40f

Observation 81313a33-46db-4747-aafb-d1ca6837f4af · outbound

This paper cites Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.421768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.421768Z digest=sha256:6b544c2429f05a00e4f2d6966ec573940a5df7c111c079fe5ca527293d4b14b5

Observation 6a9848be-5e94-4b18-97ae-ed3eaedd743d · outbound

This paper cites LA4VLA: Learning to Act without Seeing via Language-Action Pretraining.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.432225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.432225Z digest=sha256:2cfedbbd45ff631e3a8555e2c42d827c4cbd4a358e9a2b20204e4e34d6f4e873

Observation 4b40be1e-1cc9-498f-b0fc-8bb09341e4e2 · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.437052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.437052Z digest=sha256:92c4c0499eca063a18762818b4e743022be2660bc2fe1c6201a9880b37d3ab9f

Observation cea6bdc0-b60c-4b96-8748-1bd43bed470a · outbound

This paper cites LocoFormer: Generalist Locomotion via Long-Context Adaptation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models LocoFormer: Generalist Locomotion via Long-Context Adaptation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.000853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.442473Z digest=sha256:a622349d0f5db6c3991902b0c97c312b7cfc7b62d193f501622c7eeaa57864ef

Observation cb272567-7abf-4b62-a6ba-ad96b84fc08e · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.447818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.447818Z digest=sha256:4f3157f416fa43a63f15e8c7a0118c54ffaf3a90ae70598537c4e9008d69ef08

Observation 984c62fa-cf3d-49b1-b3d5-6a27c331b914 · outbound

This paper cites Behavior Prompting Policy: Demonstrations as Prompts for Manipulation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Behavior Prompting Policy: Demonstrations as Prompts for Manipulation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:39:41.258504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.452698Z digest=sha256:7d4f69c89397013bf88a6ae145bff773447acf424ace6c7c7c94cb1d80684946

Observation c8b74a6f-d11d-4d31-bba9-b692ebb2c443 · outbound

This paper cites Action-aware dynamic pruning for efficient vision-language-action manipulation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Action-aware dynamic pruning for efficient vision-language-action manipulation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.982044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.458321Z digest=sha256:fcac14fd0713a70ff6e2f5d20efa3948c727fd65c27b810793bf7cc5c1e75969

Observation 85f6ee72-9c7b-4f32-9dcb-9b5b11063e0c · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.464253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.464253Z digest=sha256:ed08e357e26776b675c5bd8027473a629fa1f6ae2e322477734704cb5f215aed

Observation 93bb0d3b-252d-4ab8-bff1-f5be796a5a06 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.470070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.470070Z digest=sha256:721af42f946193f7750d8f369b542fbe2ab7e4eec151e5730a92a78146c83a27

Observation 0b5fe059-9135-45c7-92f3-d01d315ab801 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.475421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.475421Z digest=sha256:50dfadee0a4076f596b6b7d6fa6569ff862706c6df4b3331de23ad7994d0e860

Observation 70f45c2d-7d60-4706-93b6-f443ff7cef90 · outbound

This paper cites Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.480220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.480220Z digest=sha256:0dd55432706faf5a96f5bb778f3069ac1aa9d2856a5f281d54db0300aa3d4536

Observation cc00cbe3-e0d0-45ba-b9f1-25cf18ce645f · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.485551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.485551Z digest=sha256:edaa369b58239824fa8c5ef234d73fc560bb22ff312f4e1fc42deb506bec2082

Observation a0b64ab9-d276-4788-96aa-e8b824aea016 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.491234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.491234Z digest=sha256:b8c835dedbe37d748caba6c9c096628c071fbcadeda7dd8f51d55bb88d3c424a

Observation 7b68db90-0d21-4b5d-8502-9ea65dd1f2be · outbound

This paper cites RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.496275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.496275Z digest=sha256:bdc423a1f67ce38ec34dd961b84a163e6582c59be175d5c8deb27cb1acbfa6c9

Observation f1865d81-e0a2-4a80-8d75-fcf68bc8ea9b · outbound

This paper cites StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.501945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.501945Z digest=sha256:9fb29a80a45aa90aa46009b348a1d0b4786cd46f81a6ec2e3a090fd57dd09eaf

Observation 19f7e96f-fdaa-4a2a-988f-5868b448f60d · outbound

This paper cites Test-Time Training with Self-Supervision for Generalization under Distribution Shifts.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Test-Time Training with Self-Supervision for Generalization under Distribution Shifts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.508228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.508228Z digest=sha256:94a0f39338904d98549a4b4e5a0d30ff7375e2a86d87a89995aa8d3ac6224ef4

Observation 41f76f0c-886c-4559-8c30-c280a22a1732 · outbound

This paper cites X-OP: Cross-Morphology Whole-Body Teleoperation via MPC Retargeting.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models X-OP: Cross-Morphology Whole-Body Teleoperation via MPC Retargeting

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:39:41.067131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.514334Z digest=sha256:6cdce98025a1d76c690942afceb78ae1ccb9fd1be6dd312510ec34b95f6ccff4

Observation 479b66dc-ecf4-4bec-b458-164d4580e070 · outbound

This paper cites From Foundation to Application: Improving VLA Models in Practice.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models From Foundation to Application: Improving VLA Models in Practice

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.520150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.520150Z digest=sha256:8af2f945a96ad5f929cf5f5ac34ad6364d476c5822d594fdc820b6be14757b71

Observation 70ca2af7-44b9-4716-b5d6-e4d845a1c7b2 · outbound

This paper cites A V A-VLA: Improving Vision-Language-Action Models with Active Visual Attention.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models A V A-VLA: Improving Vision-Language-Action Models with Active Visual Attention

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.964349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.525612Z digest=sha256:074e01048830db8b16de896fde6770615f97b0cb1d706633ffa2c8eacd177092

Observation afe0bddc-4a58-4ea8-9541-71736e7cf7f9 · outbound

This paper cites Towards efficient embodied reasoning: Mixture-of-depth compute allocation for vision- 17 language-action model.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Towards efficient embodied reasoning: Mixture-of-depth compute allocation for vision- 17 language-action model

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.947017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.530558Z digest=sha256:1ce79a94d812b333dd8583a3b8681597813fa7587eaf0f0f54150b523a478945

Observation 878b0789-b711-4571-8299-f330b4f86d39 · outbound

This paper cites Vla-cache: Efficient vision- language-action manipulation via adaptive token caching.Advances in Neural Information Processing Systems, 38:164448–164473, 2026.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Vla-cache: Efficient vision- language-action manipulation via adaptive token caching.Advances in Neural Information Processing Systems, 38:164448–164473, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.535540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.535540Z digest=sha256:452da50f9c7574b380eba2ea4b7c5812da88abafaa4974b3664fac7dbd314b1d

Observation 45094728-647b-4228-89c2-1813c683421f · outbound

This paper cites Affordance field intervention: Enabling vlas to escape memory traps in robotic manipulation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Affordance field intervention: Enabling vlas to escape memory traps in robotic manipulation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.918310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.541031Z digest=sha256:c69ec5d0b7bffe2686b8864a1f6e07b2028fb669f70d506f8fa7300884e65378

Observation ce9f1612-a1dc-4988-9b10-dde7eb44ebcf · outbound

This paper cites Latent Action Pretraining from Videos.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.545967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.545967Z digest=sha256:b544e8961734bdf729cbd67c4cff7229973ed24daf607480855fa32a756182da

Observation 60d41e58-ce4f-472d-984a-7d2b30795eae · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.551955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.551955Z digest=sha256:027b2e060e722a0f998877410391b203b095123a10d410722512fe95c8b1b931

Observation bc9939bf-886c-425e-a7e4-3c06e6301528 · outbound

This paper cites Hancock, Mingtong Zhang, Tenny Yin, Yixuan Huang, Dhruv Shah, Allen Z.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Hancock, Mingtong Zhang, Tenny Yin, Yixuan Huang, Dhruv Shah, Allen Z

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.557935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.557935Z digest=sha256:562f5e07fe27c1b06c03570528440d3936c8c6569216371e25961d9565d06987

Observation 4b74d118-0022-4d72-b5a7-8e64f05164a0 · outbound

This paper cites VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.562799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.562799Z digest=sha256:904c9276d2af9a14cbf9953e749c1883ad193f5ed7fe3e59e954f9ec3ae28568

Observation 17528a3f-09df-4915-b63c-4472aee2cf7d · outbound

This paper cites Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:39:40.751326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.568412Z digest=sha256:d8a9f9d33e94e02be410a1c8ccebed38e1f9c3abfe8dbf9228d5481890ba01dd

Observation c9c63866-d224-4bf7-840f-dddfeabfb829 · outbound

This paper cites TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.574154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.574154Z digest=sha256:b8a76f97345b1a28258ae55bbaac32b453694006ec2352293054d2b4cb4c217c

Observation 106b0c91-6b38-467a-8957-6146547487af · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.579241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.579241Z digest=sha256:58bf8ae319ca11365c4a6791c55ef663cce2bf4d2f3e258ec5795eed0d142f89

Observation d009941a-08dc-4f66-b5a6-f5cbe72df6a1 · outbound

This paper cites Retrieval-VLA: Training-Free In-Context Adaptation for Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Retrieval-VLA: Training-Free In-Context Adaptation for Vision-Language-Action Models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.901566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.584585Z digest=sha256:d4f8bfeadff5b840f94a2f46608e8e3c0e7c21d1fd8c6fcd252d36d29584fa6e

Observation af0c0057-cb24-4949-a60b-f467895629d2 · outbound

This paper cites Retrieval-vla: Training-free in-context adaptation for vision-language-action models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Retrieval-vla: Training-free in-context adaptation for vision-language-action models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.882434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.590092Z digest=sha256:34ab553f83013c27fa145e262ece88ce6f5040297f6cb04e1a853ac153d5bab6

Observation 0f32b0c2-d391-4553-90e7-afdd546692c2 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.595363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.595363Z digest=sha256:4f8588913802ab65b7c08cabf00a3638ec6e8b3ea49fe6398624e6a4b1e5148d

Observation d1be27be-a1ef-4583-917c-99e785ed12d8 · outbound

This paper cites ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.864856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:39:40.600800Z digest=sha256:b96d59b5c54a4825dc04e59a327adc84dd6e2faef0ecfd2e47a9a7f752b62a7f

Observation 0f1c1164-5e90-4490-af96-a4da7be0e23c · outbound

This paper cites LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.605712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.605712Z digest=sha256:4a635a33a0083548c0f0a525d857f9a8d306ab3dfc44c7b2c093365794ca4528

Pith citing papers

No inbound Pith citation observations are available.