Pith. sign in

Paper Citation Record · LEDGER

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies

As of 20 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2608.02958.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02958 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:02:41.650877Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0e327eb-cc37-41fc-ac9f-ad43b5fb7db0 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.415362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.415362Z digest=sha256:eb4760cb607c69dca5949df748d356b8d212f278869058123079ef0fbe062e8e

Observation 4bdb01c7-c19c-49aa-8284-2675c899ca72 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.421479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.421479Z digest=sha256:b98be0dac775c0003075915985acd01b9878a22a2c2c255ae434b7fe0de0313f

Observation a6f130c9-fd8b-43f6-b4fc-e6d31fc55b12 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.427337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.427337Z digest=sha256:961d26993e291e35ee1a2c0581fa647d5ff79fdbaed79c5de98a6bb733500c19

Observation f2044ad1-7e77-4b81-a8ac-c3c54c9fdfa2 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.432578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.432578Z digest=sha256:198ecf75582f443a8370b3d4fe42f82a94d93cf990d97b5bd172eb867783ec3f

Observation c5494d4b-18ed-4c80-a060-212a70fe5ec2 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.437785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.437785Z digest=sha256:e5de55b3c41c4ecd8319ab4a48729ecc4497866033790e4186c8cfc099b98a4e

Observation 190a658c-ae30-41eb-9636-7f2fd317abcf · outbound

This paper cites GigaBrain-0.5M ∗: A VLA that learns from world model-based reinforcement learning,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies GigaBrain-0.5M ∗: A VLA that learns from world model-based reinforcement learning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.442971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.442971Z digest=sha256:9e0752bd7126bed675e4e3d66443e8e0e2d74447767a37d6b66d9172b7175cc1

Observation 0fe84d86-5341-478a-ad6b-ec8f4a0381f2 · outbound

This paper cites LeRobot: State-of-the-art machine learning for real-world robotics in PyTorch,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies LeRobot: State-of-the-art machine learning for real-world robotics in PyTorch,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.698854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.447649Z digest=sha256:9e14a21e50812384380690d62b0187af19b69d748b7b51b44eb3ed65ba12ac30

Observation cf8cc2e8-5f26-4ed8-b694-d0f93b49890b · outbound

This paper cites DINOv3.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies DINOv3

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.453078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.453078Z digest=sha256:e4d7b24c7cb639965e469efe74312dcf6161a550b4685b24203aa304e165bb3a

Observation 4fa4a75d-27f3-4c63-9114-223b557c1e9d · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies RT-1: Robotics Transformer for Real-World Control at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.457922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.457922Z digest=sha256:91079b5ed033167ed243c8043dd4fb6c4624a335fa75754e91aefeb801c8bc08

Observation 94c88663-bab4-4128-9f0f-ef5be6899801 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.462863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.462863Z digest=sha256:ea15b64ee54136f8d65cbf2ecd8eeec98b99e85cbad099e1afaf8f98436d206f

Observation 0edee28e-0d3e-471f-8f6d-99308dc3d347 · outbound

This paper cites OpenPI: An open-source implementation of theπ 0 vision-language-action model,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies OpenPI: An open-source implementation of theπ 0 vision-language-action model,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.683287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.467599Z digest=sha256:1b62dc874f38af13044ed92f37d7bdb563c221ea610670ec5260d870fd64789a

Observation 38447433-88ab-41eb-8e8a-928dedf3afd4 · outbound

This paper cites Flow matching for generative modeling,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Flow matching for generative modeling,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.472878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.472878Z digest=sha256:2214b9e66082dad471ed813d2421d91f6ac88491ae6c929b43dcfb3b8f5e2b98

Observation be3c095b-6159-4b72-bbf3-e2a3846b4de1 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.656926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.477834Z digest=sha256:0253501d9cf69b23baa8898817767efae1e7148831060b929a6e734bdc214477

Observation ac4a517a-a463-4a62-ae3e-2f16db055145 · outbound

This paper cites Learning fine-grained bimanual manipulation with low-cost hardware,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Learning fine-grained bimanual manipulation with low-cost hardware,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.482362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.482362Z digest=sha256:37dc690e7842c8a756ba4564ff79bce23b2ef66e6d2dc383dbc3b9a3281d2ae4

Observation 2526b958-28da-4b52-877e-9471b3ff58fb · outbound

This paper cites Xiaomi-Robotics-0: An open-sourced vision- language-action model with real-time execution,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Xiaomi-Robotics-0: An open-sourced vision- language-action model with real-time execution,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.486801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.486801Z digest=sha256:a129c35f5cfef3f85120754c99881063401b5c6eeea70e47012b90af6631a2cf

Observation 4c587d16-b1b5-4a7e-8262-35ad00ff8500 · outbound

This paper cites Human-level control through deep reinforcement learning,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Human-level control through deep reinforcement learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.491888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.491888Z digest=sha256:88bae33447d126cf549bf67b1182c763750a2d611308a5d59f9647d87d1d96a7

Observation d9682bcd-ab6e-4fac-bd90-88541120cb38 · outbound

This paper cites Continuous control with deep reinforcement learning.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Continuous control with deep reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.496940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.496940Z digest=sha256:ed9f705feb5bd1ecf54099c5e094765bf21c97b098f8087b26a7ac830aac7ab7

Observation d12c4145-0e4d-4519-ba34-4be611a74e51 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.623620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.501647Z digest=sha256:2722ce4a52cce7e6efedfd359894f67aa5862ba0303aa72d303730d998f23bf3

Observation 05f1db3b-a716-4f3f-8a38-7b3b1cb1ffff · outbound

This paper cites Q-Transformer: Scalable offline reinforcement learning via autoregressive Q-Functions,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Q-Transformer: Scalable offline reinforcement learning via autoregressive Q-Functions,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.609498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.506037Z digest=sha256:6d0fabbbcd0b55d2289c7931f25dbdfe40cd769ddc2cc49e0fb88f7e7f17500f

Observation bca37e54-10b7-41b6-bb0f-a8bc0544a626 · outbound

This paper cites Offline reinforcement learning with implicit Q-Learning,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Offline reinforcement learning with implicit Q-Learning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.594889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.510778Z digest=sha256:c032345c04f4da89705070037e525afdc7340b2efbbad0fac6ad1fc33b2bf325

Observation 55cee351-4331-4ed6-9459-d2b387895242 · outbound

This paper cites Conservative Q-Learning for offline reinforcement learning,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Conservative Q-Learning for offline reinforcement learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.580342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.514974Z digest=sha256:3b0e92f26e65dea0ad422d81a6837c89ed63623004f1af955cc6ce79d8c0952d

Observation 347c65e5-ea50-4851-8e64-3b1ffe7c2ea7 · outbound

This paper cites VIP: Towards universal visual reward and representation via value-implicit pre-training,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies VIP: Towards universal visual reward and representation via value-implicit pre-training,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.565820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.519302Z digest=sha256:4499752c203895b91949e1224bfe811da0d9cf0a23a95ca2eebf0ff97812ffec

Observation fc3462e8-1e6c-4a63-8dc3-52e4950688b6 · outbound

This paper cites LIV: Language-image representations and rewards for robotic control,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies LIV: Language-image representations and rewards for robotic control,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.551612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.524128Z digest=sha256:d905fe7e78c2b107aa14f948d0b81984f4caa6b09e7108cf6d62599a4649df02

Observation b16b8d36-e70a-4112-bdd4-d778613395ff · outbound

This paper cites Contrastive learning as goal-conditioned reinforcement learning,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Contrastive learning as goal-conditioned reinforcement learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.537740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.528700Z digest=sha256:84adcc35b9cf963b848b17caaf630378f9c3ab90923cded11a01653116ed403e

Observation ad680993-cb28-44dc-9e86-70d19154b112 · outbound

This paper cites GoFAR: Offline goal- conditioned reinforcement learning via state-occupancy matching,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies GoFAR: Offline goal- conditioned reinforcement learning via state-occupancy matching,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.522202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.534071Z digest=sha256:9c4625dabbd3ef4fc117ae9c5d9320a751623b815b2771f6922471f28ff6cffd

Observation c02acba3-526e-475e-b87b-4c162cd52012 · outbound

This paper cites Diffusion Policy Policy Optimization.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Diffusion Policy Policy Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.539164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.539164Z digest=sha256:7140732d83f03b9556dd60643b3e57d43f38d3ad0e35897a5d9181d9995987fa

Observation c413955e-7141-4bfe-8f89-3cdea09e62de · outbound

This paper cites VLAC: A vision-language-action-critic model for robotic real-world reinforcement learning,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies VLAC: A vision-language-action-critic model for robotic real-world reinforcement learning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.544101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.544101Z digest=sha256:b3fe3796ac52ceb03eae598833928fff67eaf2e591e2084af6c2ef0698b38727

Observation 22522702-9309-4bb1-9acc-1a0fc399b256 · outbound

This paper cites SAFE: Multitask failure detection for vision-language- action models,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies SAFE: Multitask failure detection for vision-language- action models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.548852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.548852Z digest=sha256:814d30a5bcbeaa3775b1e9a52e1557c06c9e68936113ac8c3b9becce55b1cd2e

Observation 65428792-2496-40f1-9f1c-1ffa74e6bf32 · outbound

This paper cites AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.553251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.553251Z digest=sha256:88988166ecc16b423a25aaa3f2e7278cec350902dd3978c37b81da2562f5af06

Observation 3f8d0f19-fa02-4d52-828c-45d1482d5488 · outbound

This paper cites I-FailSense: General robotic failure detection with vision-language models,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies I-FailSense: General robotic failure detection with vision-language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.558114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.558114Z digest=sha256:64520dc2afbae9d77af12ffdadd42a592f07ac787b0356259cf75ad948cd9564

Observation 8289811b-fdef-41ff-affe-6afbfc7166bf · outbound

This paper cites Score the steps, not just the goal: VLM-based subgoal evaluation for long-horizon manipulation,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Score the steps, not just the goal: VLM-based subgoal evaluation for long-horizon manipulation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.563786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.563786Z digest=sha256:5fa2830c4f15020e88def943a9ca66b2757fac211d10d56c5a396c0408b54179

Observation 3a095653-1d88-491b-bbaa-b4452fae635c · outbound

This paper cites RoboCLIP: One Demonstration is Enough to Learn Robot Policies.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies RoboCLIP: One Demonstration is Enough to Learn Robot Policies

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.568317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.568317Z digest=sha256:ae484948643cafc6fb076789853728af410043cd4469c91b3152d3f03284d40b

Observation 81ba3366-5daa-4143-b4ad-8f142ef64257 · outbound

This paper cites Eureka: Human-level reward design via coding large language models,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Eureka: Human-level reward design via coding large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.507851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.573081Z digest=sha256:4656b5b40b765e324005712423e8171c1ca30d7308e7d0705c732ecfb42ef123

Observation d37e56e8-90bf-4972-8fb0-9738bf696814 · outbound

This paper cites JHU-ISI gesture and skill assessment working set (JIGSAWS): A surgical activity dataset for human motion modeling,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies JHU-ISI gesture and skill assessment working set (JIGSAWS): A surgical activity dataset for human motion modeling,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.492478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.578714Z digest=sha256:52d68efbae83b878e5d5ce7d5ce5e3974042a8d79698f08c26baef1967ac85b9

Observation 6598b3e5-267f-4211-a448-c469afd72baf · outbound

This paper cites EndoNet: A deep architecture for recognition tasks on laparoscopic videos,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies EndoNet: A deep architecture for recognition tasks on laparoscopic videos,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.476569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.583014Z digest=sha256:effb3f2cad2371845463387679aae2b2583e82828b777b1952a89c80cb7c9a3c

Observation a611b781-5a24-4f2d-88ce-c215b0f87aba · outbound

This paper cites MS-TCN: Multi-stage temporal convolutional network for action segmentation,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies MS-TCN: Multi-stage temporal convolutional network for action segmentation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.460133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.588456Z digest=sha256:ba966b56c4cf94235154ee6c68623e37cb6c81220024c14e2042e6f9810112e8

Observation fcd55768-862d-4696-b916-1da56be9d3e7 · outbound

This paper cites ASFormer: Transformer for action segmentation,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies ASFormer: Transformer for action segmentation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.442471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.593049Z digest=sha256:4257241b4e0181e374bda5bdbcf0fc1fbb4031c186c382140a25e8f515d36e36

Observation 590cca07-9941-4f4c-9828-ec86c1fdceb2 · outbound

This paper cites The THUMOS challenge on action recognition for videos “in the wild.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies The THUMOS challenge on action recognition for videos “in the wild

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.427607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.597504Z digest=sha256:9d35a12614f195595114aec0019e211e7860e98a855c920151971c605b92b620

Observation ffd4ba50-0acf-4444-9620-757526ca789e · outbound

This paper cites ActivityNet: A large-scale video benchmark for human activity under- standing,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies ActivityNet: A large-scale video benchmark for human activity under- standing,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.410480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.602421Z digest=sha256:e24dc791655ff00ed3c962aef676362113443166779a4249f365c94dc1ded2a0

Observation defdb187-7a36-42a7-ad83-d31e400ab0fb · outbound

This paper cites Scaling egocentric vision: The EPIC-KITCHENS dataset,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Scaling egocentric vision: The EPIC-KITCHENS dataset,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.392963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.606917Z digest=sha256:100e5be18a808303d683da8052d04b934491cdf798985bed334a10dc61a6d230

Observation 4389f750-1222-45d5-bd87-865c7e4dd82b · outbound

This paper cites ActionFormer: Localizing moments of actions with transformers,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies ActionFormer: Localizing moments of actions with transformers,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.373937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.611388Z digest=sha256:bd1b251f03b76e542d77034151eae657a095341b6e13d176ea650cb77feb77b2

Observation 6b5a9b44-287b-4d51-b70b-6f2908687eba · outbound

This paper cites Let’s verify step by step,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Let’s verify step by step,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.615797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.615797Z digest=sha256:233ec01d328505e4a262272fd853860b0480e2d9bcb46fd0e75825f71116c98b

Observation fefe5a37-e52d-4331-ace0-616a3a092fa6 · outbound

This paper cites R3M: A universal visual representation for robot manipulation,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies R3M: A universal visual representation for robot manipulation,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.620209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.620209Z digest=sha256:ff67021ccb5a9bad262596041025b86fc9e5f8a0c353f9f0f0963dcf75041ef9

Observation 720ce021-625e-40e8-a0ff-8f81bb2e8cf1 · outbound

This paper cites Value prediction network,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Value prediction network,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.337585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.624683Z digest=sha256:a091bbee6a0c5c4739b63862d1a35a33702fc2e6cea54685654f9fd705c8a25b

Observation 162dcab4-6c66-4f8c-b8d4-591017289f75 · outbound

This paper cites Hi Robot: Open-ended instruction following with hierarchical vision-language-action models,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Hi Robot: Open-ended instruction following with hierarchical vision-language-action models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.320148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.629081Z digest=sha256:27e3bd63cefc02812deea7befaf726c4c5ad0aedc5ca2c1e49f8dfd29583dfe1

Observation 09de69db-ef78-412b-af36-2eeaff55e989 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:41.633989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:41.633989Z digest=sha256:9c5eafe0b454e1d24fbcc0f9a50f884b6469912530d38d3cec2f2587500870ff

Observation e21eac49-62d1-421f-ab97-bc72cb148efe · outbound

This paper cites Inner monologue: Embod- ied reasoning through planning with language models,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Inner monologue: Embod- ied reasoning through planning with language models,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.303592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.638484Z digest=sha256:a73ed95bab388d51e393b8620ce9ab2b8ea90181e98e56fac22729e7baff4dae

Observation 7a40187b-8cd6-41ef-bcf6-ab63ff94a761 · outbound

This paper cites Interactive language: Talking to robots in real time,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies Interactive language: Talking to robots in real time,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.287712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.642742Z digest=sha256:91d1ed7baad1bb2a7d7fb7ef6a765a08ba12ec51b7db6807a1ea815427905800

Observation 64aec2f3-cf86-4567-9900-1a457c730b57 · outbound

This paper cites What matters in language- conditioned robotic imitation learning over unstructured data,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies What matters in language- conditioned robotic imitation learning over unstructured data,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.272114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.646714Z digest=sha256:bb97ccbe8972dd89602285f6a62f0f55f7cd3a79ed4d6e134bf635e73308eaa5

Observation 400348e5-c209-434d-be2e-52a6acbba034 · outbound

This paper cites HG-DAgger: Interactive imitation learning with human experts,.

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies HG-DAgger: Interactive imitation learning with human experts,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:42.256203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:02:41.650877Z digest=sha256:2ab1b930056b89e7b800136f51c2282729d4755bed1f3785379c6507a5298436

Pith citing papers

No inbound Pith citation observations are available.