Pith. sign in

Paper Citation Record · LEDGER

$\pi^{*}_{0.6}$: a VLA That Learns From Experience

As of 6 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 100 inbound Pith citation observations for arXiv:2511.14759.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.14759 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T10:34:59.134604Z

measured 187 of 187 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 100 of 184 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:46:48.385585Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact35
  • verified fuzzy45
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5dcd0a3f-fb95-4417-97d6-22b819deb5dc · outbound

This paper cites MIT press.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience MIT press

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.628520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:d5e161c9d5691cfa05e260f2a71919b1459fc03a2ef6b69b93de72fb1823a1d6

Observation f2239f95-3fb5-48ad-8321-dbd2b30598d7 · outbound

This paper cites Riedmiller.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Riedmiller

Reference 2

Resolution
verified exact
doi, observed 2026-05-12T10:34:59.202868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:0de6b387a3fc5b98699f89b155b12c23b6b9bbc1594bab9d0aaac086ef2e7778

Observation a8beca63-e9f0-4739-8ca2-675a1853e085 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:34:59.242880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:1e8ff48db3ee1debe3a678b5b7a17aa878bc960dbae80dee6ab0daf43bfa714f

Observation c674f213-37de-4d0d-bacb-bec5bad07779 · outbound

This paper cites Diffusion Guidance Is a Controllable Policy Improvement Operator.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.250729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:03a0e7dd421a34627df0c5cdabf3966bd87c31510f8ad2f77490d824e186e946

Observation e437fd5b-1f71-4cab-9bd9-9a33ad1c500c · outbound

This paper cites In9th Annual Conference on Robot Learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience In9th Annual Conference on Robot Learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.640156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:221b2fa5bc43dec2944400edab8deb7f229f80e01ad45975dc92e7b370ce1dcc

Observation fa478c9a-afc6-40cc-bcc5-c80fd89d0599 · outbound

This paper cites an unresolved cited work.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-12T10:34:59.529595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:2d8360d7a2905d1d07108e65140cf9d572e54260c391dbd68265a3f8dcfdeb69

Observation 12c10b72-88f9-4ac2-8cc5-ff4de9f5fbd4 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience A reduction of imitation learning and structured prediction to no-regret online learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.525394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:eb450a0f27f9c78fa487927db5d78bc68aeedc65600f881eaca65124a939a015

Observation 63c05c4a-f4c5-4ae9-9545-65a8bf51f63f · outbound

This paper cites Shiv: Reducing supervisor burden in dagger using support vectors for efficient learning from demonstrations in high dimensional state spaces.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Shiv: Reducing supervisor burden in dagger using support vectors for efficient learning from demonstrations in high dimensional state spaces

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.620826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:e9076011bd167a4f3f626da2303e0ecde0903ec8bd1a5cfc03f048b6acc5066b

Observation 33adf923-a9f3-4145-8b09-b9e428f2244d · outbound

This paper cites In: 2016 IEEE International Conference on Robotics and Automation, pp.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience In: 2016 IEEE International Conference on Robotics and Automation, pp

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:34:59.184958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:ea2b40de93c138d20fc9723bd4ddc958be3912be33b99bd8159418c711ebe05c

Observation f7eb68c1-41c6-44a2-aec1-1b761dbb3a5f · outbound

This paper cites Dra- gan, and Ken Goldberg.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Dra- gan, and Ken Goldberg

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.616102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:824819cd47fb7ab3b26b82ae13d00cedc2fe453e1a68a9d9d76e748fe097a405

Observation 10f7c1e5-0b4e-4858-a56c-9f0e7db0931d · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.563569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:3a31a124a7db65f584f8ecf0b5ef669cf204424ce507b2063a8ac14156f41f08

Observation a06ac923-93be-4b52-bcbf-00d503e5f779 · outbound

This paper cites RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.258097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:aee4f04159aeda41b4194a14db0f34168916f651c7038c0019c5e728ee49f302

Observation 30bf1ff6-c63f-49b3-9388-2e56e7ccbbee · outbound

This paper cites Hg-dagger: Inter- active imitation learning with human experts.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Hg-dagger: Inter- active imitation learning with human experts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.539271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:3f641d3db17ff163812e013da399738f1c391ad631723000f5759c3068361054

Observation d8e94625-a3ca-4068-9548-1bd8780f0867 · outbound

This paper cites End-to-end training of deep visuomotor policies.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience End-to-end training of deep visuomotor policies

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.544242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:38b2edfe34670f48738430b8f9225f41c70ef50edab97c5951f912238c7c80c3

Observation 0a795517-5ea7-4ca0-922b-cf519cd4f3a5 · outbound

This paper cites QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.265492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:594f3858cb332ccca1aa8de6519ff976a14139a85e0e6106f80ebe7af5bec9d7

Observation 245d0e33-8e3a-4f39-83f8-c0095b40de9c · outbound

This paper cites Iris: Implicit reinforcement without interaction at scale for learning control from offline robot manipulation data.ICRA.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Iris: Implicit reinforcement without interaction at scale for learning control from offline robot manipulation data.ICRA

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.553493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:60e4f0b22cb523f58b7e92d6d0e31125d91af7c64593a0d0ea6f1cf0d13295a0

Observation a4612fd2-75cd-4478-a963-e12d53ba35aa · outbound

This paper cites Ahmed Ahmed Rehaan Ahmad, and Chelsea Finn.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Ahmed Ahmed Rehaan Ahmad, and Chelsea Finn

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.558379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:7a8e7bfe86d64e0ab20b41fcb752d88a755087fab2138b7c66ba76766d1c6c0b

Observation 0b5ab1bd-e684-4e94-a076-009a709ce7c0 · outbound

This paper cites ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:34:59.197814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:f2f8243e679aa6eef8e1049ecde5b12a7747fbd231339f8fbb09dff347722a08

Observation feb8868e-dfd1-4400-be0a-29b092bf2ad3 · outbound

This paper cites Continuously improving mobile manipulation with autonomous real- world rl.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Continuously improving mobile manipulation with autonomous real- world rl

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.568147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:0e661716e6edf5949d867490645740a32e5982ad8e1d6c2230fd06f3d65fa4f6

Observation 6ce860c6-f134-4433-9a46-83bc854ad6ae · outbound

This paper cites Serl: A software suite for sample-efficient robotic reinforcement learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Serl: A software suite for sample-efficient robotic reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.573045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:3f423722fa4d79e9e8b2576e6f14d8676571af5fd351a943ba2a13964065f146

Observation 9aa860a1-91ee-41a7-947e-78bd7d88c623 · outbound

This paper cites Residual off-policy rl for finetuning behavior cloning policies.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Residual off-policy rl for finetuning behavior cloning policies

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.277473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:b1e39499c9d16a597e0729e2e0ba19b261d80a00e38b1bcf8bccab26ad08cfbd

Observation 08871e90-a466-4c7c-a63d-0d90562e2a38 · outbound

This paper cites 10610948.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience 10610948

Reference 22

Resolution
malformed identifier
arxiv_id, observed 2026-05-12T10:34:59.285013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:19efee08bceb874c7b63c54b0005caf9d7e07229617330a2bbba6cbd09e04dbc

Observation 7915e3fe-a197-4270-b62c-5412084809c3 · outbound

This paper cites What Matters for Batch Online Reinforcement Learning in Robotics?.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience What Matters for Batch Online Reinforcement Learning in Robotics?

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.291343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:ac031e31bcc70427610199dd608b2f4b2421a2fbec951479dfb5734148ab0fe9

Observation 8abd7861-ce6c-4201-870a-edd8d0987e0d · outbound

This paper cites Ren, Justin Lidard, Lars Lien Ankile, Anthony Simeonov, Pulkit Agrawal, Anirudha Majumdar, Ben- jamin Burchfiel, Hongkai Dai, and Max Simchowitz.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Ren, Justin Lidard, Lars Lien Ankile, Anthony Simeonov, Pulkit Agrawal, Anirudha Majumdar, Ben- jamin Burchfiel, Hongkai Dai, and Max Simchowitz

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.591296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:55bd40822c48ad4df69be91a237e1baa29528916a49aa07a3623f646b221c5e1

Observation c7f91144-73d1-4c16-b0c5-a6bb0338e430 · outbound

This paper cites Rl-100: Performant robotic manipulation with real-world reinforcement learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Rl-100: Performant robotic manipulation with real-world reinforcement learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.298836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:6227957d5b8064c2b3bf86d3eafa280a38af330cd8bee6dd9cead86c969095fd

Observation 78e5e364-1465-43f5-ab5e-307a1dc66238 · outbound

This paper cites Mt-opt: Continuous multi- task robotic reinforcement learning at scale.arXiv.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Mt-opt: Continuous multi- task robotic reinforcement learning at scale.arXiv

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.602663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:b80ebddacdf8dd6d84afdf4ac1e2b5415fd10e59bef7e3e6a445787d6a257da7

Observation 1c0dd07e-16d1-4e17-bc8e-5406a3eb8608 · outbound

This paper cites Zhao, Vikash Kumar, Aaron Rovinsky, Kelvin Xu, Thomas Devlin, and Sergey Levine.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Zhao, Vikash Kumar, Aaron Rovinsky, Kelvin Xu, Thomas Devlin, and Sergey Levine

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.606908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:086f85218efd439cd55c4b40de4f08a13eba3c5cfdbdc5f6e2a3a473e2f14c08

Observation 52fe8996-0e6c-4e63-9037-c81cb7567ed4 · outbound

This paper cites RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.306163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:6f21e43a09e2dfa91ed165e15cafb2e89aebb08be76bf440945050683f92ef9d

Observation 9d63d13e-2499-46c3-8d43-496b22386510 · outbound

This paper cites Pre-training for robots: Offline reinforcement learning enables learning new tasks from a handful of trials.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Pre-training for robots: Offline reinforcement learning enables learning new tasks from a handful of trials

Reference 29

Resolution
verified exact
doi, observed 2026-05-12T10:34:59.190623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:5b7288889a75ef1c0f76bfdf65a8733a2b1ab86415793b4f27ff55763a278341

Observation 33e7d781-4892-4a66-b4e3-24285b00a19c · outbound

This paper cites 10610948.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience 10610948

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:34:59.228044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:a7198cb6ea1e12c30cd2ecf0ea3785043df341e4f6a2ee8aa12a820eb016820a

Observation 5f6ecec2-6ead-4619-853d-ba97fdcad20f · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Interactive Post-Training for Vision-Language-Action Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.344448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:6c5722b6da6b69232804c48250223176fd9f9591528805ff9a7f0df3586fba9d

Observation 18c063c8-8a2b-4822-85b1-e981fe93cbcf · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:55:40.591564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:5a7f5322c0e5050c84b55db8872f27de4c47c28763f0afb76d6d0a6ee6474277

Observation a31667a7-a6a9-471c-a68e-17e40c2c4aee · outbound

This paper cites What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.325937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:3c1a07ce18bd0bbd22d056c8204b5f87732a44d839c0667b2c7d326d9bf3181b

Observation c86132ea-c38b-4ca5-9eb4-75fb2e0eb079 · outbound

This paper cites pi rl: Online rl fine-tuning for flow-based vision-language-action mod- els.arXiv preprint arXiv:2510.25889.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience pi rl: Online rl fine-tuning for flow-based vision-language-action mod- els.arXiv preprint arXiv:2510.25889

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.334967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:8141f71cc29a4c44adf76e6c1ab9feeba20175ba0931bb0d9833a5fef0e0f67b

Observation df616bc2-710d-412b-99d3-f631e51d532e · outbound

This paper cites SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:02:11.500343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:c9ec2842db48487770ed8a1183fbdd4b2587d2c6100eca9c5ea5eef844f70dca

Observation 96f1fb71-e18f-481e-972d-a1fafab88cf6 · outbound

This paper cites Improving Vision-Language-Action Model with Online Reinforcement Learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.347237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:667e967a1d31cd9819c5f47a1001dd8500c07187757dc3fb2ba99e28b1f72655

Observation 2c61e69c-06d0-4bc3-a4a7-24f95a767725 · outbound

This paper cites Self- improving vision-language-action models with data gen- eration via residual rl.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Self- improving vision-language-action models with data gen- eration via residual rl

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.647039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:02dbbe3e7ff0e1097058238b9ebf3efe54e0d6139c14fd19b51f7f9b3277680c

Observation 9cd1f206-e1e4-4e35-9094-886b02fa9b09 · outbound

This paper cites ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.353437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:381f927dd8f23ec74fe45e366a4d0d7be34b1c7878b8150dbb16c808d3998395

Observation 609b4a33-52c6-4b10-aecf-d739fcbb9438 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.359504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:206d2effe1ae3a3d818e78d9b3ef5c4935cdf34d8ea10c0b7a463a8f46468990

Observation b240e196-727f-4d3d-a44f-6d163d7e5038 · outbound

This paper cites Steering your generalists: Improving robotic foundation models via value guidance.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Steering your generalists: Improving robotic foundation models via value guidance

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.664038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:a571a314410319ac7b44934166fda1b18d8adf28057396cea9a04bf016c290c5

Observation 2b94fcba-d01d-45a2-92f5-c1861269572d · outbound

This paper cites Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.368442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:d8b6838bbb4be87e6b690baadfb9b992e7fe7c238655fe6ad19891f150da57a5

Observation 87644b33-10a7-4362-a00a-c8c7bc9e6212 · outbound

This paper cites Steering your diffusion policy with latent space reinforcement learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Steering your diffusion policy with latent space reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.459439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:69905266b978f743efe5cd63b14cd2059d57c127c41fb23dd7a5e40948f637a8

Observation 298d2bfe-c995-4e35-9711-a71c32863790 · outbound

This paper cites RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.374709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:2cf31e8728c3dff2f787d70abf52aa67fb061aae7c8bf7d0916d546bf48d3082

Observation f3d6fb73-8d81-4939-ade9-134f28c319f0 · outbound

This paper cites CO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement Learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience CO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.381175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:79628ad8d4fab33dad2fbe4fad017a488dfcd7f488697d6c6ab8ea7bda6183b0

Observation 1dd654af-0a47-4061-a394-7a1a3f3bf613 · outbound

This paper cites GRAPE: Generalizing Robot Policy via Preference Alignment.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.387695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:c60a1bf41bff8c0799742b292c86661897386e49f3342110c268de18606cc187

Observation 6a4a20d1-9c89-4675-9771-3952f454a83e · outbound

This paper cites A vision- language-action-critic model for robotic real-world rein- forcement learning.arXiv preprint arXiv:2509.15937.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience A vision- language-action-critic model for robotic real-world rein- forcement learning.arXiv preprint arXiv:2509.15937

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.393800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:9b4b91b991dcf04d6933e3d78bbf8acca55f726709440e5b34525fb7913844a3

Observation 788af721-76c2-4e37-aba2-3075430651d8 · outbound

This paper cites Self-improving embodied foundation models.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Self-improving embodied foundation models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.399105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:c0698df91010454f96b8bb1aa356f28f621660f0d6abf6c5e3aaec3b4cfbd945

Observation 7542ae43-ba7b-45cf-b733-1f13a4991a15 · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.408653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:8292e25f595ebf6df55edefdd6f49cc68e50794d085c6ee3bdcab5823dc667b4

Observation 334c2d12-73fa-4d3b-9df4-d5254ff8fd4c · outbound

This paper cites Reward-Conditioned Policies.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Reward-Conditioned Policies

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.416893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:84da72eaa88548c79d0b8a06f3e64ef0543f2a2981da2e5606da980e1a1895f0

Observation 60020006-bd82-428d-b465-b0999b87336a · outbound

This paper cites Decision transformer: Rein- forcement learning via sequence modeling.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Decision transformer: Rein- forcement learning via sequence modeling

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.492892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:4bf33e79e9723d97dc16f2d80bae8548a5e9e5375096eb17218a29b9bba1e9a5

Observation 74147d76-9be9-4914-bc10-e079fcedf03c · outbound

This paper cites When does return- conditioned supervised learning work for offline rein- forcement learning? InAdvances in Neural Information Processing Systems (NeurIPS) 35.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience When does return- conditioned supervised learning work for offline rein- forcement learning? InAdvances in Neural Information Processing Systems (NeurIPS) 35

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.497502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:4b4535b33bbc5dc9a64dfe9849c2c77ae3991dc90d36d637fe99c4e07836f75e

Observation 31d3abc4-1555-4f36-87cd-2aec7df1529b · outbound

This paper cites Rvs: What is essential for offline rl via supervised learning? InProceedings of the 10th International Conference on Learning Representations (ICLR).

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Rvs: What is essential for offline rl via supervised learning? InProceedings of the 10th International Conference on Learning Representations (ICLR)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.501395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:6f363f07284578da46d5f6ba4b6b6c54ca128294b64d4a88fdfca214b3444a17

Observation 1636963a-4305-42b9-a179-d85bb12eae5f · outbound

This paper cites Generalized decision transformer for offline hindsight information matching.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Generalized decision transformer for offline hindsight information matching

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.505991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:47ca72c71b0b58c35cebfd0daefa4c0082d12c0ebd4cf88441bfbe792e526708

Observation 8160b27b-73b1-420d-a612-54682522becb · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence mod- elling in offline rl.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Q-learning decision transformer: Leveraging dynamic programming for conditional sequence mod- elling in offline rl

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.511141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:42ccf878ed50a4f2e2ae5c3bca03f52b9e893411c67ab1db220d3654956d3d9b

Observation 9cd398b2-c4fa-4d0a-83f3-29bbc23202e7 · outbound

This paper cites Online decision transformer.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Online decision transformer

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.515926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:daa8e64e1d088ee2b7a9d06bd3a7fb56a6caa13c30f548961bb865a4d7ba3dd3

Observation 73edf8f2-ab91-4ee6-8d2e-1b58ed15f4a4 · outbound

This paper cites Advantage-conditioned diffusion: Offline rl via general- ization.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Advantage-conditioned diffusion: Offline rl via general- ization

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.520844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:0ddc888b707f97c60755bf5376d6c9cf59b6c9d05f3b519d3734c34b7126f332

Observation dc7a5f55-249a-4ef7-8a6e-7287f9caab4b · outbound

This paper cites Elastic decision transformer.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Elastic decision transformer

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.218576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:55235137d7d0e8c1dbda86b5b476ea7b96d876dd33095dc872ec7e5b10f974eb

Observation 1b47a4de-1ed4-47f8-88e5-30d2dfcae865 · outbound

This paper cites Concept2robot: Learning manipu- lation concepts from instructions and human demonstra- tions.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Concept2robot: Learning manipu- lation concepts from instructions and human demonstra- tions

Reference 58

Resolution
verified exact
doi, observed 2026-05-12T10:34:59.211955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:76f7890ae78f1020e7484b79b0ff42fe5b83993924ab94433301cf8b6c66dcb0

Observation 97208df7-d5c7-4075-b661-9248a6a0936d · outbound

This paper cites in-the- wild.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience in-the- wild

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.534382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:a06f172b74fe421eb600e2c0bea0c18bb659aab4059a6082c6cc54943cb1b80b

Observation 5868eb68-fd47-4081-96e7-3733a3261135 · outbound

This paper cites Learning language- conditioned robot behavior from offline data and crowd- sourced annotation.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Learning language- conditioned robot behavior from offline data and crowd- sourced annotation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.548981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:7df6b36be9e7f939cf08be9748fed80d63d060f922138271e0b1f2ff4f4a75ec

Observation 28565aa5-5f5d-4ca8-9c11-5a0c5e38f1fd · outbound

This paper cites Sontakke, Jesse Zhang, S ´ebastien M.R.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Sontakke, Jesse Zhang, S ´ebastien M.R

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.577261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:8c3e2fc0ef387e52222092a2a77a8827d4975fadb15119062cad1d4e3692d404

Observation 1079ff25-f206-45dc-998e-a346556f0e0e · outbound

This paper cites Language to rewards for robotic skill synthesis.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Language to rewards for robotic skill synthesis

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.581726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:afac36179db4c04178f595133426efc855be591567737ca2e0c8566ffbfc3ca4

Observation b25b6e64-0c71-4a0d-8e6e-9ce9d0041fa7 · outbound

This paper cites Lim, Jesse Thomason, Erdem Bıyık, and Jesse Zhang.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Lim, Jesse Thomason, Erdem Bıyık, and Jesse Zhang

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.586493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:99bfc1b9edd1ec73e863575293ebe2e4ec716dd78b329cd4edc32bfb2fb65938

Observation 1bf69143-533a-4a3e-8e5a-dc40e515d2e8 · outbound

This paper cites Video-language critic: Transferable reward functions for language-conditioned robotics.Transac- tions on Machine Learning Research, 2025:1–22.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Video-language critic: Transferable reward functions for language-conditioned robotics.Transac- tions on Machine Learning Research, 2025:1–22

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.598140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:6a037d3b6d45cd97c0f7a154384cf4621593d3a867b68caa56dfa7d0306e1177

Observation eb9a043b-8d9c-4cd4-9765-f92ee79e5b02 · outbound

This paper cites Liv: Language-image representations and rewards for robotic control.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Liv: Language-image representations and rewards for robotic control

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.611605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:6a29074f2a22eeb2bfba8b8122e9ce189985143231c65cc45f49c72fbe92d377

Observation 96b606e6-3bd1-4ba6-81f1-6a49dc31a0c9 · outbound

This paper cites Vision language models are in-context value learners.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Vision language models are in-context value learners

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.624912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:873a21f0447db45aaf9b6f99eb913be315131bbf9fbbf0dd1a95d6d9e52d0a16

Observation b251e6ef-ce49-4f58-8b11-b29d2320ae6d · outbound

This paper cites Proximal Policy Optimization Algorithms.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Proximal Policy Optimization Algorithms

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:34:59.235701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:836be858de09b5c3540620c1d1ca596c9d4674d184db8c1a154459fb014b1074

Observation 212203cb-f019-4248-b4e0-ccf0eb861e57 · outbound

This paper cites Maximum a posteriori policy optimisation.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Maximum a posteriori policy optimisation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.632562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:aeff6eba12b4133a9e096a90e6701f83f13ddab4d329b0ca4681f7ca1abda209

Observation 5e2c454d-e7f1-4db5-8325-b01196157eed · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:34:59.422852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:bef12948b309230b9e0f260d0106a025e9002e996c428d2b743fe7a1026b3a24

Observation 3fffcf59-8ba6-471a-a24b-116aa2c13322 · outbound

This paper cites an unresolved cited work.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Unresolved cited work

Reference 70

Resolution
verified exact
doi, observed 2026-05-12T10:34:59.207330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:3c34ff2bec49a873895e552ea4e89d4b65b8065716cbe89940c3ada39eac4d80

Observation 69649c11-9323-49c3-abe7-5afef49fd71a · outbound

This paper cites Rel- ative entropy policy search.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Rel- ative entropy policy search

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.643617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:70e44b8dd070d7d0d3027dea9b70114af5f12bfb139ea33df8edf11366fd59bb

Observation da2778fd-a054-454f-bc79-3c004ceba287 · outbound

This paper cites Exponentially weighted imitation learning for batched historical data.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Exponentially weighted imitation learning for batched historical data

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.650399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:a40e4395ea58f3feb4eac2feb33ac1321697160e7aac2c4d0991473396fda716

Observation de86c718-e261-440f-9ccd-ce84835fc8eb · outbound

This paper cites A distributional perspective on reinforcement learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience A distributional perspective on reinforcement learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.659808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:8441d051cefd2258407011fdf62f4dfe5c8384df95232be8828b2fc8cae2f4e1

Observation 5e52e2cc-5f9f-4b73-b839-ec2dd811298e · outbound

This paper cites Knowledge insulating vision-language-action models: Train fast, run fast, generalize better.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Knowledge insulating vision-language-action models: Train fast, run fast, generalize better

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.455997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:a54c8abc13e591db829badced2208e97da00066092cfec291f06bcb262ea1066

Observation 30a3883e-2db7-46e1-b7a4-8605c3f7c401 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.ICML.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.ICML

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.463116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:e0979aea1c07c2ca1457fdf0857982dd7681bad29a2a15583fc2aa075cf618fe

Observation c3734c80-8077-4bff-9667-e848fd0f809c · outbound

This paper cites Critic regularized regression.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Critic regularized regression

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.466649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:e45882d1061e7beab63f4d82c0266d2b7f5fd7f4baba5b09fa7546c23affe0c7

Observation a32ca1de-c36c-459f-a4a8-276481a9e877 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Offline reinforcement learning with implicit q-learning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.470485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:48ec5fc6c2aa974af59a36fc2125b910e66e8f46cc11015359e8ea1313fa02c9

Observation ccbfcb12-752d-4658-8525-31dfdc40743c · outbound

This paper cites FAST: Efficient action tok- enization for vision-language-action models.Robotics: Science and Systems.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience FAST: Efficient action tok- enization for vision-language-action models.Robotics: Science and Systems

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.474216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:fea97d1cf207a9b35edc3dc40a203fa406ac036b459dde53de13b3f4f39c1ef3

Observation ab3e6269-6ed5-4742-a243-a1e7972cffed · outbound

This paper cites an unresolved cited work.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-12T10:34:59.478607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:24fa8153273d6bbc3b10730c3e64c4ad11aa64276031a0d2f99ea5a1407a2d58

Observation 5dbc5840-8bd8-4a10-a4a5-cfd8ed472dab · outbound

This paper cites Flow Matching for Generative Modeling.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Flow Matching for Generative Modeling

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:34:59.428506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:1fe2df6c40a8e56bc51528202db72f4b373af95d3530aa178afeda3f92ef141d

Observation 3010726f-688c-4e34-91de-673ea3383a76 · outbound

This paper cites Understanding diffu- sion objectives as the elbo with simple data augmenta- tion.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Understanding diffu- sion objectives as the elbo with simple data augmenta- tion

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.488745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:3c60079617db7badde49550d1389fcc4391d3223132f76563dc52a1658b04996

Observation 01ef8a05-630b-4073-bc33-f59c021e03cc · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:34:59.433800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:d8c2bdcfe66525784e850443463ed5c77838b11a7c947f5eb6423ba6161c9cdb

Observation 44e99cfc-de44-4698-b672-c8c86153c0ec · outbound

This paper cites Flow matching policy gradients.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Flow matching policy gradients

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.483666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:06a245195ea72f20dbd4d0d64319988b4a7d7bfeb8ba3a47026bfb1539df5eba

Observation f3196016-432c-4f6f-bf90-8630c3dea41b · outbound

This paper cites Simple policy optimization.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Simple policy optimization

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:34:59.636641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:41d18d3896516567b43c2c38d7769e3ebe42b17c57c516f0e69df8c6ba133dab

Observation adb13b4b-71bd-4495-b751-9bdfe8dba9c1 · outbound

This paper cites The Ingredients of Real-World Robotic Reinforcement Learning.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience The Ingredients of Real-World Robotic Reinforcement Learning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.439726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:0d6f429c4e8e924250a721802b18eb3915762e3bacb2b77327299bf3e1911122

Observation 43626bdb-b668-4803-872b-89dfce2c51de · outbound

This paper cites Autonomous Reinforcement Learning: Formalism and Benchmarking.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Autonomous Reinforcement Learning: Formalism and Benchmarking

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.446657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:ef1f6f7f501af24333e4592b2dbe8f1181c6b949cafa9b0ca141ec83844d4b0d

Observation 3c453cca-20f2-4bff-873f-c86d8923f044 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:9f1b6c6400ed091f91dd9f65d024b74bfdf7d1bcad9733f5137e3ffa6de0afe6

Pith citing papers

Observation d656d3fc-cfa3-458f-b907-a720ca3ae3ff · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:25:54.273628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:1de84c58d7b5516a1f52893b98e0f75206142be0ab497efeeb8a67c23156239d

Observation b700710a-bf86-40f9-b41a-f96b56b5d410 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:08:01.701442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:a84e4cb5b79d3e4aa0270d21baa9c793b075557a6c6f881ced8f91bcfac99a89

Observation 069f214e-bc61-41d4-9ce6-5dce93bd9be7 · inbound

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance cites this paper.

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:07:47.840570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T11:06:23.555610Z digest=sha256:c87c151b1c8e3ea4bac90e7b23273b4b78ee7e51605152d57edfb790109f240b

Observation 492e2bcf-f89e-435b-ad68-3c7b5935f288 · inbound

Language Movement Primitives: Grounding Language Models in Robot Motion cites this paper.

Language Movement Primitives: Grounding Language Models in Robot Motion $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T05:17:28.099404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:17:28.099404Z digest=sha256:fca9d81e42d0bf764cdc232658c2fa2c25c16b8c16a21153e1e69e80f28948f3

Observation 456df98f-5d2b-4699-af17-2cc0c13daf57 · inbound

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy cites this paper.

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:47.154587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:56:47.154587Z digest=sha256:1741863881e19832dc7897ab3cd2253118af2cd585fe7089c8bd1692f63661ad

Observation 08c15000-9bea-496d-bf00-ab0ec2c4dfe4 · inbound

Action-to-Action Flow Matching cites this paper.

Action-to-Action Flow Matching $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:57:29.225522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T06:53:35.153155Z digest=sha256:03e305a4c1192cd2683b691d54b328292c9e565d92983ba3a95e0ab2d1762ef3

Observation bd027e24-351f-4214-a143-d25652b4c97a · inbound

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning cites this paper.

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T14:04:11.970546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T14:03:48.795572Z digest=sha256:140991a76bd51bc8fe137471ea520db7095d0315436f4e700f35fecea213e391

Observation 5399d154-392f-4eac-9870-fa665a2caff4 · inbound

Combined Constrained Sampling and Reinforcement Learning for Robotic Manipulation cites this paper.

Combined Constrained Sampling and Reinforcement Learning for Robotic Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T03:18:42.039197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:18:42.039197Z digest=sha256:ab52b1d5fb73b0367dc883f8808b87737201965487163cb39c99748416aab284

Observation 14a13e2d-bc73-4c3e-9803-cb92aa2eaca7 · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:37:13.860172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T03:36:09.272019Z digest=sha256:611a26f7fca637d36c3337f273814cef8c19d3dbab0db4863d7e255b61b281cd

Observation 30299ff5-a561-4b6e-865d-6883cb23ea8c · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T02:53:16.108040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:53:16.108040Z digest=sha256:aa07a1c189d0fa9e9d8a58be0e73caeb6501015bced2260997a8ec665318f544

Observation 71bb122c-598e-466f-81fd-07ce3b3ed833 · inbound

RISE: Self-Improving Robot Policy with Compositional World Model cites this paper.

RISE: Self-Improving Robot Policy with Compositional World Model $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T02:30:32.143070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T02:28:37.997148Z digest=sha256:518bfa99f35a5852764c7af2737b503533324779c0a64fc41efd7bcfbf453d6a

Observation 65c3e564-9f56-4902-ac84-9ac87c0dd074 · inbound

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models cites this paper.

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T23:48:57.949446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:48:57.949446Z digest=sha256:3f2511c044a70fdb494c93af79e2f3c3c6e04d82a506f85adfa70548f946205c

Observation 6ec06c83-7a1b-4938-b7b9-d086d9509fb2 · inbound

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training cites this paper.

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T23:47:55.179001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:47:55.179001Z digest=sha256:562ddd9ccd7fed58dfbb9fc9797b4fba5a92b57437b834b835cfca742794acd0

Observation fb9fa478-13d0-4ec9-b9ca-ca2f7c7e2e85 · inbound

VLANeXt: Recipes for Building Strong VLA Models cites this paper.

VLANeXt: Recipes for Building Strong VLA Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.863649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:d19062e185126c568fcf88d312ed725bb8c695379055bd7f633f70c023a26095

Observation 0dc9129e-996e-4fcf-9081-d6df4b8aefca · inbound

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics cites this paper.

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:43:07.989656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:43:07.989656Z digest=sha256:02f6ddd3ff8970b7ca694d00832d1d4bf665ed6d58e7dfa884a4c9c336c5f5a1

Observation fc8a87a6-2d43-4351-a88b-8e605df661b5 · inbound

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning cites this paper.

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:06:33.769825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:05:30.324309Z digest=sha256:cc5ca0f73760bb60de2e98dd925a5400431456bfbd974baa9fae5cc7ac0915fc

Observation 2e410113-341d-4f7c-84b7-7cda7ac62ed0 · inbound

RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design cites this paper.

RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T19:42:46.514018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:42:46.514018Z digest=sha256:485956dfb728dc835d2ab722f9ea287a532ca73dba39af8d8c5ea9dfcbede821

Observation cf568421-15ba-4304-824e-22bb63730f5c · inbound

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons cites this paper.

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:00:12.574856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T17:59:52.365630Z digest=sha256:42b7005c3606b493a0b72db7bcc421b841923324064cf07293eebd66a6339b0c

Observation 701392f9-6ab7-48d0-9f4f-c19492a94438 · inbound

CoFL: Continuous Flow Fields for Language-Conditioned Navigation cites this paper.

CoFL: Continuous Flow Fields for Language-Conditioned Navigation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:20:10.476539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T17:18:38.294448Z digest=sha256:5e0b2ba3e8f049b9b2199c16059ba32d3ff6a30344f62cdbf0ef79a181432284

Observation 3df4d643-291c-4e93-a3ba-19629da9680a · inbound

Optimization landscapes of variational quantum algorithms cites this paper.

Optimization landscapes of variational quantum algorithms $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-15T14:44:57.978537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:44:57.978537Z digest=sha256:23e757c6a61a5b43e65c383e390c274dd83dc9f8eb7496fa0b174bc2c62fa1c9

Observation bcb4c41c-a3cd-4f7d-9ba7-bd8c7bbb1db9 · inbound

Safe-Night VLA: Seeing the Unseen via Thermal-Perceptive Vision-Language-Action Models for Safety-Critical Manipulation cites this paper.

Safe-Night VLA: Seeing the Unseen via Thermal-Perceptive Vision-Language-Action Models for Safety-Critical Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T18:44:43.510752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:44:43.510752Z digest=sha256:a34426af159a64a7336f03eb6b5faaee8550a299cf8d31e64deb057e004a5f59

Observation ac17d9de-4eea-4839-91d7-f606ebe49032 · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:647fb2558c8318e46fcc4daac4d80c9a3f199795547e4deb15529f6915c7063b

Observation 106e493e-314a-41a6-86f7-e14ec4c7fae0 · inbound

You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector cites this paper.

You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:55:24.225788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T09:53:45.151347Z digest=sha256:7d69c0b555da75227b34ce933de0416f6517905973d32f92308151058b925c5c

Observation e425baf7-6393-4ef6-8402-519b01ff1be0 · inbound

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation cites this paper.

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:49:54.496767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T09:49:03.333757Z digest=sha256:aebdb726987bf05a35ab1046ac066801efdbdd8b259eca80ffb9ebf3742db092

Observation beedb161-72f4-4344-9bc6-d192029c325e · inbound

MemoAct: Atkinson-Shiffrin-Inspired Hierarchical Memory-Augmented Policy for Robotic Manipulation cites this paper.

MemoAct: Atkinson-Shiffrin-Inspired Hierarchical Memory-Augmented Policy for Robotic Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T05:46:48.385585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:46:48.385585Z digest=sha256:0d15ce471c5cd77c29f81ba9995493c5cc57333555cd0384b095ccdee98f7d3a

Observation cdd44350-9566-4b35-944f-ee1a87cc6dd6 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T08:05:15.387111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T08:02:13.188363Z digest=sha256:2bc4c4a61c72768a20df50f35c75129b3b9780db2681fc0451627c4e8d892b14

Observation 9254c69b-ab9d-40f5-b1eb-c3e4d4f0f9e3 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T10:50:01.094248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T10:48:56.280105Z digest=sha256:009aa445cb4478f9a1f3200c9d141adc79738c01a5eebbb546e103c6e72ba3d6

Observation ba0b2f20-8946-4f9b-b859-3408db88f193 · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:a010daf3888f554b7fae4d9c242876bc829f48948c6f75aac26ab521e28f6997

Observation 341318fb-f226-41d3-9cae-558f220d78dc · inbound

Open-Loop Planning, Closed-Loop Verification: Speculative Verification for VLA cites this paper.

Open-Loop Planning, Closed-Loop Verification: Speculative Verification for VLA $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:48:11.455634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T19:45:17.852646Z digest=sha256:bc13214cec74b995dff27b7fa7fdcd5338961f959607d9d18e49bbead2fe74a6

Observation 27d0a391-d7f8-447b-98d2-0124304f98fe · inbound

ARM: Advantage Reward Modeling for Long-Horizon Manipulation cites this paper.

ARM: Advantage Reward Modeling for Long-Horizon Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:28:09.926570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T19:25:46.693189Z digest=sha256:a9b63098caa00015320099f218ca346dbb19ff418b0ddbf51640cd65404cd552

Observation 099112c4-1561-49ef-aabf-015a9e4faecd · inbound

Redefining End-of-Life: Intelligent Automation for Electronics Remanufacturing Systems cites this paper.

Redefining End-of-Life: Intelligent Automation for Electronics Remanufacturing Systems $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 134

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:23:09.520029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T19:21:34.849729Z digest=sha256:b71a738b986c863b411095aeb0fb9c27cea31ecce101bbd499a4e4fd10f56578

Observation 1da0bcd2-2a79-413e-9c2c-9f66411ef23b · inbound

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model cites this paper.

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:58:08.855870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T18:54:07.081457Z digest=sha256:65650ffd7c386f04ab3a90b909e4b64da214641400365bc432c53ba75701f7a0

Observation d6d35575-9365-4cbc-bc1a-8c7c767b4f5a · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:00:48.378858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:21:45.365156Z digest=sha256:f84756a8d0ddef3b0fd4d3f3d95aa5300673937ba6f8bf373df6fb340a60c9a3

Observation f2cb6e96-cc98-47ee-825b-4ea282c16191 · inbound

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment cites this paper.

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:55:52.412697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:27:12.286456Z digest=sha256:10316f89eb724e75bb560c11af16a3b0720bd74ff6f97982127e1e1ece0c5291

Observation 96cc6f07-83ad-4ed4-912b-311ad1d1a9c0 · inbound

Action Images: End-to-End Policy Learning via Multiview Video Generation cites this paper.

Action Images: End-to-End Policy Learning via Multiview Video Generation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:50:53.514156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:51:05.206602Z digest=sha256:154e0b398509ad22033068d5ed456c0f4a8c5031d19d05cc2a1dbb3035c9a73a

Observation 28bb7d2c-38e4-4ca0-8c27-16dead3fbd92 · inbound

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning cites this paper.

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:25:59.082501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:12:08.970164Z digest=sha256:d6b7e9c1ac04e4bc4e02838b5911200ce15e8566396164c5bae456c8c1175b5b

Observation aa38d91c-7794-46a2-a4ea-787c90f77b9c · inbound

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence cites this paper.

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T00:03:53.609175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:03:53.609175Z digest=sha256:a21cba043dd8c7c994f909abc8b08d64285718ec362394058bf9c5fec71cc36b

Observation d040512f-ed95-4966-8cc8-6cd42c87e1fa · inbound

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis cites this paper.

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:11:01.142615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:16:31.378588Z digest=sha256:b6c7997e1becf111b37e588310175d49e0fec5969b4fc227694aeb63443e1f6e

Observation 60535b7f-9c12-4054-9027-6c34510a347d · inbound

ScoRe-Flow: Complete Distributional Control via Score-Based Reinforcement Learning for Flow Matching cites this paper.

ScoRe-Flow: Complete Distributional Control via Score-Based Reinforcement Learning for Flow Matching $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:26:00.670288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:02:36.261596Z digest=sha256:01c1e386b7dbe22f184c3b5fa55b507237e75dea20dea7c948997e7a853fa91e

Observation 849a9f66-883f-42de-b293-8ff9b5fac114 · inbound

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities cites this paper.

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:45:21.702816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:42:34.409651Z digest=sha256:ef252dbb9bc2b3f7588684f4692fa1da3d29d42b11a27082f79f4f8afe699fbb

Observation 11bbe7f9-32f6-437f-ba42-ed36faef8f79 · inbound

OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL cites this paper.

OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:28:39.271234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T05:16:57.241755Z digest=sha256:e62f222b6440e908ad506455e256da330a319245a5db791a1e41cf5401328289

Observation d2ef1d48-631d-4ff8-8db6-427455325208 · inbound

SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model cites this paper.

SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:10.899312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:34:29.624029Z digest=sha256:f7a5b94d926da0f065dee13b75809b204478ef9303d393f8b31dd955c3229143

Observation 79dc5373-be84-473e-ba4a-68b154ed655c · inbound

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models cites this paper.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:04.018345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:584f6e6f54722551f8b5f24d2d947b1ffa324e0f1e8d50514b8b7343f6256f88

Observation b5e9749d-9299-4baf-9de1-9b59876d33bd · inbound

FASTER: Value-Guided Sampling for Fast RL cites this paper.

FASTER: Value-Guided Sampling for Fast RL $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.263461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:47:36.475845Z digest=sha256:cb9d152041714a497513e3c70dd0bcee1eeb894521e32974ecb54920e7f23928

Observation 1aa59379-c0d1-4e27-bee2-2f0ee93b78ee · inbound

RL Token: Bootstrapping Online RL with Vision-Language-Action Models cites this paper.

RL Token: Bootstrapping Online RL with Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:09.184643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T11:56:34.978806Z digest=sha256:992f946391cb7da1ff464b1f7451e6373384403f43e67ff8b0bbc1834899fff3

Observation d75b9113-b2cd-43fa-b10f-fbe96e1cfa2e · inbound

Cooptimizing Safety and Performance Using Safety Value-Constrained Model Predictive Control cites this paper.

Cooptimizing Safety and Performance Using Safety Value-Constrained Model Predictive Control $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:21:11.452892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T05:52:11.090003Z digest=sha256:b81a4fcf3b93251b8452d2b2b7d05c66960c25dc9f1df0873223e233de91e803

Observation 11f572ef-8fe3-4103-837a-9962c0a0b575 · inbound

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors cites this paper.

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:46:11.935381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T02:30:30.319810Z digest=sha256:28602314f40326d47e6b24e46914a8da33e287a287353405525ff88a789ead38

Observation 52f3a094-dec9-40ca-a545-fc9f7b17afdd · inbound

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors cites this paper.

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:35:33.445649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T08:32:26.160039Z digest=sha256:7759891f61604918bd1021fea3087195cfa9591343734b57acf33f3d5043c919

Observation 7db32639-0052-4126-a62f-dc40a36e8002 · inbound

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations cites this paper.

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:51:29.124042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T08:56:32.164424Z digest=sha256:1f38db62a7a98e688aeb82b562a58f3e4ad4c7f732a0ea70740d27efb8b0a431

Observation f33c21b4-a7c7-4353-875f-222d94b944c1 · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:31:30.071764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T05:47:17.494531Z digest=sha256:d359e2f57615253d9bbaf58b751e103d461b22c4b0dbf5dbdecde98732e2afe7

Observation 7bd5b30e-bf2e-4c10-b421-b84ef469c9f8 · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:17:04.478386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:00:26.352130Z digest=sha256:2cb2546a626fc5b88d02ff02b5214e273ebbf51427dacfe8d18bfa71b242792c

Observation 821bdf88-6d1b-478d-8eab-c3df931350d6 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:10.922526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T19:31:38.069592Z digest=sha256:b05b89bf335e387eb32f9e1c316dbe880fa5bfa41603e20c6008c1543096d41b

Observation 31042599-b993-42b1-9b4f-31a07958c008 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.263986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:1f236ede1707a6ac217281971016eafe6ec5cd05e6f71fd5ff45091abc60d8fc

Observation 0cd918c7-08e8-44e9-a758-020f720c804a · inbound

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation cites this paper.

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:35:38.643323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T18:21:00.089755Z digest=sha256:534c8a05eee0771ee80fb3d4eb0344b84a5d5ca307e928476ef88a150ad470f1

Observation 366acde0-6660-495b-91cb-b42c6ed4a338 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:51:30.983227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T16:06:10.595164Z digest=sha256:3d9f3da7767190d308a29558525bf2dc6025984c75f8b672d137e59e082d682d

Observation fa180a3b-b5db-44ba-8a1c-f25384a8d997 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:35.153838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T19:02:19.756092Z digest=sha256:f8eb46a36ab7ca49b1fc390f3da1f8846bf71ba40ac7a96753694a66f158521d

Observation e24d812c-f698-4b26-8779-0c2e81e705f0 · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:41:08.137796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T11:20:19.749415Z digest=sha256:398431024733edd5d67551f6e04006e492603a727ce6b643afe28c4dd6587f80

Observation e1727c4a-b0e3-47a2-b37f-63bcd996ca70 · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:24.464065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T05:05:29.984355Z digest=sha256:82ebbe293633910e433688421577c8ae92853367b8d0a46554889e3a21db1335

Observation 9de942b0-daca-4732-a02a-971cd3a69ced · inbound

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models cites this paper.

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:57.016888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:31:30.530595Z digest=sha256:92f7bdd045ae77a5071eb05228420c58f3ce74572eb520f677f0f094d2ffa3ee

Observation 653d5306-3366-43d4-8e0e-62ea18775bfb · inbound

How to Utilize Failure Demo Data?: Effective Data Selection for Imitation Learning Using Distribution Differences in Attention Mechanism cites this paper.

How to Utilize Failure Demo Data?: Effective Data Selection for Imitation Learning Using Distribution Differences in Attention Mechanism $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:55.326380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:08:26.682345Z digest=sha256:97a839ce270702a06eea1bb0468c99902851da3e83a2a54b9e672d10ab311dd8

Observation d404f7cb-cff8-4653-b92d-35a61034809b · inbound

How to Utilize Failure Demo Data?: Effective Data Selection for Imitation Learning Using Distribution Differences in Attention Mechanism cites this paper.

How to Utilize Failure Demo Data?: Effective Data Selection for Imitation Learning Using Distribution Differences in Attention Mechanism $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.285237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T08:19:30.310013Z digest=sha256:4bf874889b7ed4d481a71e788de9b0888e4f4e6d030e269941a70b438927a58d

Observation cea9e3d8-b191-4f83-ac5b-e72cab9015ed · inbound

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation cites this paper.

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:41:20.081536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:40:22.192349Z digest=sha256:d71759f30b1fe09a96d7ed161da9f9c638c0e35e2df7f66f1fcaa7b0405e6a83

Observation 642b95dc-67fa-4f8e-8b92-201795b07d38 · inbound

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models cites this paper.

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:25.887746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:39:32.715650Z digest=sha256:e60159f2704940158e2843f2f650302095f182311645d76592c7449d60ce39d8

Observation cee6d915-3da8-49d6-96cb-a32a0428e05d · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:25.616019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:08:43.222818Z digest=sha256:44137680c7304d5fcee94f3b6a54491777fabaeff1ca22e341aa5f672b5aba40

Observation 3ee5fb14-c9e3-499d-8773-69be5697373f · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T14:26:01.108089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:26:01.108089Z digest=sha256:ec83e341f6f50f56bf90df1a48a5a1febba1c6dced59efea1bea114bf1ffca71

Observation 243ca8ad-4a03-4ec7-b345-de498c102a9b · inbound

Offline Policy Evaluation for Manipulation Policies via Discounted Liveness Formulation cites this paper.

Offline Policy Evaluation for Manipulation Policies via Discounted Liveness Formulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T02:17:06.392865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:14:57.505609Z digest=sha256:6ffb601e8a27c5f78c331bc5f6b96edfa706010a675140dda357fa46abfb9833

Observation e4264a32-1d3d-49d1-9a89-c8c7440005cb · inbound

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning cites this paper.

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T04:22:14.539966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T04:22:07.953637Z digest=sha256:94474de9c48cdb42a1a7a8e4e2e01a89972722136870a10d5a0880f14a721220

Observation 477f222f-4cf3-4775-8aec-db52d2a03170 · inbound

Reinforcing VLAs in Task-Agnostic World Models cites this paper.

Reinforcing VLAs in Task-Agnostic World Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T04:22:14.806687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T04:17:51.349213Z digest=sha256:1b17755e4974f10b6525d32236594b6db98358c4271f898ec8b6c780b341a5b9

Observation d6057664-e984-408a-8b65-ccf9339154e9 · inbound

Reinforcing VLAs in Task-Agnostic World Models cites this paper.

Reinforcing VLAs in Task-Agnostic World Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:14:03.120777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T08:13:40.975340Z digest=sha256:d764f6901968b277cad25158e4d68a68db5c8309ef48dbccedea756266495e0b

Observation 0d479dcf-08d8-4105-a675-33e93d0b1b76 · inbound

Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic cites this paper.

Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:38:01.112967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T21:33:43.740593Z digest=sha256:5fc02ec5af4798226eee9a7bfbafcec8744a77833096126b47c6c184689440a2

Observation 708bd060-aef4-4677-b570-80664becf849 · inbound

Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic cites this paper.

Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:05:01.918390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T05:03:14.358407Z digest=sha256:a393be34ca0192f4f64d680ee28d3274a26422186c80dc9204eafd18cbd09cbb

Observation 0528ad19-a077-4bd0-8fb8-1e9a2e37f3c6 · inbound

RotVLA: Rotational Latent Action for Vision-Language-Action Model cites this paper.

RotVLA: Rotational Latent Action for Vision-Language-Action Model $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-14T17:49:23.121112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T17:48:06.734816Z digest=sha256:fbd48efa51cd42bb3fffd692036ac1cb319d2be2abc61af6982d5a48c114e57c

Observation af909c11-c285-409d-b62a-a28fbf9940b2 · inbound

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model cites this paper.

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:25:46.144820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:39:15.662180Z digest=sha256:d103e36e70763b8d20bd5c1dd7a3829ad090377aa19497a4bdd76298ecdab4d6

Observation 11c89708-b46b-4666-8511-5819eda6a17f · inbound

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention cites this paper.

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:09:42.194160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T03:09:26.858170Z digest=sha256:0a63d94d38cdde87aa83ce2eb72b4062b1045cafba9dd6e9edc82e962b617544

Observation 547d1aa1-9442-4d13-adb5-101ebdf8b4ec · inbound

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention cites this paper.

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:34:05.577883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T08:31:57.077346Z digest=sha256:1fbe235e4c75ce2cd96a6ddd2497308b178bbae33f7c998cf592b8d48dc0c356

Observation bbf4e9fb-371b-497c-bc86-212e66694060 · inbound

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment cites this paper.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.733969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:4e4bb44cd0278e0fb302b1891f0b3713fc4ad378ad098dd5f88d2cb22f5edafc

Observation a8ad6fdb-2863-43b8-8ae3-94d9e214331b · inbound

ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics cites this paper.

ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:58:11.298331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T09:55:17.865828Z digest=sha256:7e302b36a318518d43eef2456095214b299107b46bc67aea3d77b0b35f238ea7

Observation 1bbed8e8-9328-4cbb-9cdf-cc0c8222db22 · inbound

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models cites this paper.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.916497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:0eb558506cd9713378417750c1f03319f9081b63a3d9f012e08bdedaefb0671e

Observation 403221ff-756d-48c9-a339-941450a8772b · inbound

Beyond Action Residuals: Real-World Robot Policy Steering via Bottleneck Latent Reinforcement Learning cites this paper.

Beyond Action Residuals: Real-World Robot Policy Steering via Bottleneck Latent Reinforcement Learning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:13:03.351681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:12:45.553510Z digest=sha256:dd348583022adb7f6971ac8cfe351f284cab85fbe1b0a73c3af1585b12f53216

Observation 91aab0bf-fc79-4861-be4e-4ae4888e955c · inbound

VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models cites this paper.

VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:59:36.440417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:58:35.695638Z digest=sha256:a204c1660819e4db17508ea3a86435959ac7a26c6874d328393a9c729df7138a

Observation 4dbb3016-900d-4394-90f8-4d375e570ff3 · inbound

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis cites this paper.

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:14:42.831553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T07:12:02.612292Z digest=sha256:0dae4e252af9167cb520041dcdd9d10232d362d8ecdb826b986f7bb342ccfb78

Observation 74fd6aeb-5451-4804-992f-b9c84461e9ca · inbound

ParkingWorld: End-to-End Autonomous Parking Reinforcement Learning from Corrective Experience in 3DGS Simulation cites this paper.

ParkingWorld: End-to-End Autonomous Parking Reinforcement Learning from Corrective Experience in 3DGS Simulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T16:35:50.720005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T00:26:22.777878Z digest=sha256:7955579a65aeed15568e1e3d719bfaaa9b2528eca978013c882e5eee280c4a65

Observation 29701e97-9c2c-43b4-9f76-3528a17411ff · inbound

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models cites this paper.

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.942734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:10:08.682307Z digest=sha256:79761d8d74c24c6f258dd2c3386d8b7549b3d25d4343731471ec4a1bfaf59c3b

Observation 36e776f9-fe90-4494-9c13-671cb8343c63 · inbound

TapSampling: Inference-Time Sampling with a Task-Progress-Understanding Verifier for Robotic Manipulation cites this paper.

TapSampling: Inference-Time Sampling with a Task-Progress-Understanding Verifier for Robotic Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:04:00.522441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T21:55:45.761276Z digest=sha256:1389c642717ac61b80450afd972a70fb675b12fa408c2c0ff70b8dd5e2dca7bb

Observation 7d89163c-bb81-4dd7-b6c4-3e9bb2facfbe · inbound

Trust Region Q Adjoint Matching cites this paper.

Trust Region Q Adjoint Matching $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:13:52.794109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:04:04.109255Z digest=sha256:b256d18db01ec6281efceb70785dab304a5de7b5934fa51b476a11c22c3b63c1

Observation 0e1d34ea-87b6-4f86-b9e3-ff665c722e1f · inbound

PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology cites this paper.

PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T14:13:31.047658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T06:45:09.301480Z digest=sha256:5204695dbe5e7ea36fcc92ed44f027562d5c0cd5121e99a85f7a3a6dd36a3925

Observation 12a7c66d-2bfb-47dd-b1a9-86fc9af5fb00 · inbound

MARS Policy: Multimodality Only When It Matters cites this paper.

MARS Policy: Multimodality Only When It Matters $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T06:43:10.272087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T06:42:33.949547Z digest=sha256:6a2f83df5915866a517fca8a01039c6155238e7c5601a3a8a8acf304d7d183a0

Observation e894690e-f8fd-47ba-8873-8d707a76205a · inbound

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models cites this paper.

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:23:12.765963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T07:18:59.266263Z digest=sha256:99143f73d83050efa68efb51a8d6bd330e911a29a09516ce6b62bcd228ceabae

Observation 14378c67-19f3-4148-9481-a6f631bf846f · inbound

Feat2Go: Visual Feature-Grounded Value Estimation for Embodied Reinforcement Learning cites this paper.

Feat2Go: Visual Feature-Grounded Value Estimation for Embodied Reinforcement Learning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.165638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T22:41:42.394450Z digest=sha256:0646ce8b0f0fdefc9f90407f260e6bb3515eabbffa727435c6b048403763f9e9

Observation c7c2f300-52f1-4ac5-99ad-46e28abb972f · inbound

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring cites this paper.

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.564956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T22:39:15.211339Z digest=sha256:6f6cbd1a73c21a0b59ef6a5a9fdd3881a59adf957c24394e024499e596569a4a

Observation f25ab638-847b-491d-a10f-6304913606e0 · inbound

DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation cites this paper.

DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:56:10.403716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T21:58:48.611343Z digest=sha256:262a8afef0492fcba80702fa3728024c2f007e01d9d573d3b3723601744cc207

Observation e04488e6-68e4-4b8f-bb07-81620419a62e · inbound

Set-Supervised Diffusion Policy: Learning Action-Chunking Diffusion through Corrections cites this paper.

Set-Supervised Diffusion Policy: Learning Action-Chunking Diffusion through Corrections $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:26:21.980669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T14:25:37.987916Z digest=sha256:27e6bd86da3dcc867644e40e881a2b439c928592eee42734a058373cdbe838f5

Observation b8fd35c3-3510-4838-bb18-a42bc7998cad · inbound

Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation cites this paper.

Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:56:34.450797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T09:37:58.434897Z digest=sha256:535299649d768f31e0f1aee4a45517acabc310af50d63eb87dc792aa72b66b68

Observation 1b030e9c-c707-4ca2-9d50-7611f9179f14 · inbound

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization cites this paper.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.372377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:fd573a5ebf67cb8f3bc40693f06bace60918997ee82890e5689e20bd1ebe6024

Observation d554ed0a-01b6-471c-93d0-a43f932aa4e7 · inbound

What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos? cites this paper.

What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos? $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:58.988013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T01:04:26.608953Z digest=sha256:55996da4c0e790a3293de9cbeece855c5cb860901a94faf1eaac1d6207b5bc4b

Observation 651a7c5f-a301-4467-86d2-997c9a4ff715 · inbound

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation cites this paper.

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T19:07:18.115064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:40:00.330510Z digest=sha256:c65dc6ca11e1cb094ad53ca300c31374279728ec9236edf060cc3db8a91d2219

Observation 3486b7c9-061a-4cc8-b3b9-e2968784ef6a · inbound

PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation cites this paper.

PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-27T18:51:07.725191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:8bbd9d650b5ca0bd1c7c4ecb42188933f114aaac551af5752e54865bbc8bdc7a

Observation 7febade0-ee7d-4ffe-89c9-bbe104161c40 · inbound

Scaling by Diversified Experience for Vision-Language-Action Models cites this paper.

Scaling by Diversified Experience for Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.754454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T17:12:50.216492Z digest=sha256:13ea7de5598a43882222815d772df62c5f16bd5d6dd608ee210a7ca956f075c0

Observation dd797f3e-e316-42e9-84b0-789fbdbef3d1 · inbound

TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation cites this paper.

TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:31.335128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T16:30:35.385161Z digest=sha256:d581a5729af75eafbea96984db843876d649261d49ec6a966b8ec99110003670

Observation db3cb8b3-808f-41af-bd49-de6b64e2e1f9 · inbound

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience cites this paper.

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.747483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T16:47:22.504175Z digest=sha256:c8de45fc3d6d2bdea3a4109b6a32a4ea20af0427b3450b766d4bdbcade631ab1