Pith. sign in

Paper Citation Record · LEDGER

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning

As of 14 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.04591.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04591 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T16:45:53.967918Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 313a52c0-7026-44db-843f-5e2484a77d48 · outbound

This paper cites X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:a5bc93a5d8fd8a000286aaa8d4096efa7d52f392982a280232a68830cab679e7

Observation 8df5747b-704a-400a-bab5-479f771ccb04 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:f7c18c072200a67c4bdf80464507e68edf9d72d1b96dbe884781af6e6622cea4

Observation cc598254-cdbf-46ed-b8ad-d48bd0c583a2 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:ef325feaeb5e2ea6fdbd73bc4045e833c472d3ebbb26f6199fb8f1afe425bf33

Observation 039af18f-9f18-4ba4-9cdf-4cb93af55865 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:7f0fec3caa4b18e4997a7a4a61b37e1f9e70278bab7773c47509b160cf5c3255

Observation aa354b34-f166-45ec-b304-f551b0cd70f5 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:cc105e1b64bc0e410a718d3586b82ba4e7bec23bce648eb7898598799b61b952

Observation 9316d916-16b3-401c-8cb8-fba211344a84 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:63c916cd47cd2956c111d3f68edfdf780bcfe19a9acdc65f9cad7e7426131fe1

Observation f6f8eec5-a9f2-460a-8afc-78f68c53c823 · outbound

This paper cites Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:b5fed7701d9c0a45926652142df935ced4248fc4f458562d3296b78710618223

Observation 6dd8b39e-6c8a-498b-8380-123b3402a4f5 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning VIMA: General Robot Manipulation with Multimodal Prompts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:d0f616a0089d48f7d0cb2905db46579aed0418ea20ef2bd0a437dcf5e658eec8

Observation 423d678b-0d08-442f-92f9-9f007724e024 · outbound

This paper cites What matters in learning from offline human demonstrations for robot manipulation.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning What matters in learning from offline human demonstrations for robot manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:2da2f47fd60d84283203df1770b280a75687857469c00d3709a2376be8e86035

Observation 722f4f77-4444-4cf8-988c-d8ea7b837484 · outbound

This paper cites Imitation learning: A survey of learning methods.ACM Computing Surveys, 2017.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Imitation learning: A survey of learning methods.ACM Computing Surveys, 2017

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:aba9c46abd08044781f0e37650f4e5210674f83c95de649f210d25b980b32452

Observation f1984fc8-5717-441a-b2d1-6502b7b4dd71 · outbound

This paper cites Alvinn: An autonomous land vehicle in a neural network.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Alvinn: An autonomous land vehicle in a neural network

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:038d06585142d4d8d553d6e03a1c033c801322ab6de68e0133df036c8ff84e49

Observation 40437ddf-c081-4a0c-9de4-2b1cbb7a686b · outbound

This paper cites Overcoming exploration in reinforcement learning with demonstrations.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Overcoming exploration in reinforcement learning with demonstrations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:9ba9bb98d5653133e6b5826a1bb8f214be3d99747d5f69f1cc6d27e0cb2d32ca

Observation 00558d34-6413-4f1f-a4a5-6784a9d6b5c8 · outbound

This paper cites Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:af2041bc96da9edbdfe31b719b70700093e187bb37b2eda227c05ee9e1096626

Observation 601c7383-88e6-4e21-8849-ff184de03e1b · outbound

This paper cites Robonet: Large-scale multi-robot learning.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Robonet: Large-scale multi-robot learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:c18ced04f73a2223cf93321d55a3be1ce7af125c783610605303c328fcb6d21e

Observation 61052563-a315-46e1-8524-87a3121d5c2f · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Open x-embodiment: Robotic learning datasets and rt-x models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:f0f17d8c4f4c2dfad54f6b896c2f2e29837138a83c48dc2742308214208ca54f

Observation f45e6a8e-c63c-43b2-9fd3-587d396d0133 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning OpenVLA: An Open-Source Vision-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:e41deb46a86cc79c2a0f552f9c761e3e5b74095f829980e2e96df22e236a66fe

Observation 6c8ec250-f1d1-425d-bc37-4dc5884354d9 · outbound

This paper cites Emergence of human to robot transfer in vision-language-action models.arXiv preprint arXiv:2512.22414, 2025.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Emergence of human to robot transfer in vision-language-action models.arXiv preprint arXiv:2512.22414, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:89cf33c748c768841970aaa3bf62883a42aab57f61e9e505c7dbcf8c62365fa6

Observation c4ada3ac-a62c-4137-a6eb-d3940168a90a · outbound

This paper cites SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:fe4e34fb171b3ac71f65280320640bf2df1d7a475caed557d7cd81a62c4aaecb

Observation 82b64311-2b55-4304-b6dd-df3e40d4bc38 · outbound

This paper cites Curriculum learning.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Curriculum learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:0785feacfa8e12a18e220b2fb3cd32db0ec3f4a89bc5caa9f771b4788ead5770

Observation dfa72942-71bc-4465-8e70-1d06dec59403 · outbound

This paper cites Reverse curriculum generation for reinforcement learning.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Reverse curriculum generation for reinforcement learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:2c614422b73e22c4099da1a75154a5097991649ad2415c0e68a625bb5bedf052

Observation f9ac2e84-edfb-4fd1-acc4-36766321b5a5 · outbound

This paper cites Feudal networks for hierarchical reinforcement learning.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Feudal networks for hierarchical reinforcement learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:7afad9b875fe2dd2abc104a16ee4ca49819e192ef40f177e880edb2b63a7f14c

Observation 68a4bdfd-24bf-428a-b78d-455abda9cb17 · outbound

This paper cites D4rl: Datasets for deep data-driven reinforcement learning.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning D4rl: Datasets for deep data-driven reinforcement learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:3c783371ab78790ebd96556d6abdec385b8534b0dc474b0bb1330778ff572101

Observation 4c1c6e72-b91f-4e54-8fd9-312d1567b0f7 · outbound

This paper cites Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:385ed2dfdc4bac928b24de852029d3a0543bfc9cf0838d57af2408e55047e4d5

Observation c7718bf3-352e-41ff-a860-4d415c3e95b2 · outbound

This paper cites Self-Evolving Cognitive Framework via Causal World Modeling for Embodied Scientific Intelligence.

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning Self-Evolving Cognitive Framework via Causal World Modeling for Embodied Scientific Intelligence

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T16:45:53.967918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:45:53.967918Z digest=sha256:26d3801e5e27103a13d57175819ca50383bbabb20d2f24054998d9e27af06167

Pith citing papers

No inbound Pith citation observations are available.