Pith. sign in

Paper Citation Record · LEDGER

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 0 inbound Pith citation observations for arXiv:2603.28730.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.28730 v2

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T16:10:12.689957Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

76 of 76 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved76
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62a3df6c-1b59-466b-8c5c-f437165af3be · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:9eaa4881f35cb717e7fa0538f9820e47105c5e4afc9f943230c4a51fe8150a50

Observation 9598803c-ec5a-4079-addf-f7da5a91ce3c · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:f204febaa57afe9e0876f71283f33f42ce95e95f6df5a967e50e9f85bd2d182d

Observation afa849d1-9faa-4e20-893f-9ab9872b4bf5 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:8d8c5fc6377c3ffc71c0788e0c0d31cc8846b717889dbe90c6854e0dcdb2599a

Observation 9ff64ef4-43cb-4187-8a91-b9730e3e8d1b · outbound

This paper cites Vision-Language Models as a Source of Rewards.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Vision-Language Models as a Source of Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:19b54ffd40a384098c0311da27648cf3da89910e1d49528c6f22977cd0e9d6d1

Observation 4f345bdf-4225-48aa-9cb4-a2ab1daeb8ba · outbound

This paper cites J., Platt, R., van de Meent, J.-W., and Wong, L.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning J., Platt, R., van de Meent, J.-W., and Wong, L

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:543b1dab8c25bab93df0a699a4dd720ee61d201cac76cf173c484ea208b37978

Observation 0025c0cf-90cb-4a0e-9c54-a1b7d4601509 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:a90c80bd3ff3ee0f233b4de242069b0beec0fbcc0e490d026564eba9cea28fdc

Observation 49283563-5e29-423a-bb3b-c6dfb84b6485 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:d980e3a3ace911aade99e62dceb2f2eb67f1a5057c402e88598002d738d12553

Observation dc63545e-e5f0-43f7-a984-90ee3e098ed3 · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:ef7734c508944e80d4cf822c42154fe4f8fa61577cd5eff2065a647ef372b144

Observation ab07615c-fe48-4386-9ebc-e6962892d043 · outbound

This paper cites Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:e8b97676e36261ddb21a2a373222e63505199a958613ef3aaf07165c25ee9062

Observation d3a58e38-4f87-412b-aec0-44aa1c89bf6e · outbound

This paper cites J., Ren, Z., Ratliff, L.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning J., Ren, Z., Ratliff, L

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:47df99881fc160abaf512644b32524c95b3354bcf2388d301de28fd25e7dad69

Observation bc7eda4d-f748-434c-985a-f2d8f32aa19a · outbound

This paper cites ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:b5741df769618f1e45f8223e368e345b33f86a6cc94afc933836293a138f63a6

Observation 195496e1-79b7-473b-beae-810a615dee70 · outbound

This paper cites Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:1aad31685c0e03e60f0d83994c75bddda0cc3cb315c498c6829f5131d52918ab

Observation c249368a-f755-42bb-9682-bb3984fde106 · outbound

This paper cites Vision-Language Models as Success Detectors.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Vision-Language Models as Success Detectors

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:d69bbc79ecacde54d608b534e3ac9caba3bdb593bdf70746b7ffcc2c34cc20f0

Observation b8cfdf0c-b4b8-45e2-824f-3d59261c22db · outbound

This paper cites Gemini 3 Pro Model Card.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Gemini 3 Pro Model Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:bb2e292a8a2081aff706d6ddbcfeb61bbc8e09f8bda21c670784260bb82cb944

Observation 5ec04593-5322-4cc8-aa79-1ea6192af76a · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:c4d1e335ded51b8cc5f0d34d0d15241709ba722c66574b6f914434c6bb3ab4d1

Observation c498d802-e594-43b6-bad9-edf44e82e018 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:a3826e9a8e9e61364cb59b8a8d9df933e79575d617825dd67e28ddf29f5b5c82

Observation ac860e48-7b16-4106-9f8a-1ea66bde5243 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:8194b849d6897b208e30fecfc476ab865d8f6f5e49b9e9ebb627dd1a58bf5e2f

Observation a8a96bbb-80b6-49fb-ba57-6ba6f5a515d8 · outbound

This paper cites GPT-4o System Card.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:61f955747889d432e06a734eb93a45192f1c0b6a483b58c5f13427addb655842

Observation ba0b2f20-8946-4f9b-b859-3408db88f193 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:2b7694f71f1592a02d58f7c762931950b857b85a6031f452837c03ba51711d8c

Observation 7c10364a-a134-4915-bea2-4dc4370cfd46 · outbound

This paper cites and Berg-Kirkpatrick, T.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning and Berg-Kirkpatrick, T

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:3f53a9fe2413f528f0ebc2aa839d66a79dc2ba196b942a6ac5a0f629f3791c19

Observation 48594a0f-4176-4eaa-83a8-5351f2880570 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning OpenVLA: An Open-Source Vision-Language-Action Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:5482900117bdc484fdcf52a6de33b1fa16b507f9ee75536b58f48ef887fc4098

Observation 19de2312-eb00-4f82-9d85-e4152773d7cb · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:887d3de1ce5e7db3407c3e7d5693bb0f545b7c513079ba3c76953e6f0e621f72

Observation 88914eb7-203c-4aac-817e-05a1b2e84b48 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:fa2d87af64e65284da4f8442a77404446ef23362b638f93f0ab2c76cc671869e

Observation bcc89693-6875-479c-b6e3-e677596372cb · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:9b15d437b9613747f5aacfed1073bfed4c0d792f19bf349f7fa2fa559cf2c733

Observation f37521e2-0fbc-45b0-9a93-87fd0b3e23d7 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:e9bf68e1ec0ae2f0f7520ad69cf9d6c142f51baa09f851c7eee9515b5388b0ff

Observation 34e70e25-17e8-4609-9dcb-7ec131393ba1 · outbound

This paper cites Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:c2f6e40ec8d6d8b6038df2efd8a062d73e23aceaeae77cbf5fe6b0ca4c14b8af

Observation af931e7b-966f-49d5-ae32-94e19f6246f7 · outbound

This paper cites Let's Verify Step by Step.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Let's Verify Step by Step

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:a7fe0e228144f9bdf39d812f4f2ad8d8c24d5ff3f26e7754726064cd8cb73fd8

Observation 342f070f-c0a6-4305-b008-e204e5b415bb · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:c2e066ca9613f752008b4d9d8aefeaeee695b7881d2b4f4c87ee72e5d6957dd2

Observation 7baa9a60-b7cb-47c7-ae76-730587a39393 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:54fe28d30fe8a4c16947390f7a24633c8254acc5627c90a52c87db360bbefce3

Observation 68425c62-3fa6-4362-aa6f-8e035412073a · outbound

This paper cites L., Berg, J., Sharma, A., Schaal, S., Finn, C., Gupta, A., and Levine, S.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning L., Berg, J., Sharma, A., Schaal, S., Finn, C., Gupta, A., and Levine, S

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:31467385f766f0163f714a49d7cfda1b0c220a89de951030c82907c0bf376deb

Observation c5b4e255-8cd4-44c7-b2db-b736049f7868 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:23b2a85e6f4d612a50526c18bf032324784685847a5c6c586a862bfb20a9c801

Observation 51f198e4-d4b5-4957-a463-d0b23162bab7 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:89071d1779848a0ed631c105f3d611d3c4f3e6b5fca0c5ebd6a707e2c8c3b81b

Observation 0e9d3458-f541-44cc-bb1f-2ae14dbf8d0c · outbound

This paper cites Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:17471687415d36db5a7fa1e4a3353748e2cc5394484d73961a3592f399aeb8ef

Observation 81a86104-4936-48e4-a241-7236a7661ccb · outbound

This paper cites J., Hejna, J., Fu, C., Shah, D., Liang, J., Xu, Z., Kirmani, S., Xu, P., Driess, D., Xiao, T., et al.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning J., Hejna, J., Fu, C., Shah, D., Liang, J., Xu, Z., Kirmani, S., Xu, P., Driess, D., Xiao, T., et al

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:e663d38b29ea3536918f6005a48978665fbc9e52401975513b7c4866c6d8b026

Observation 6d0827d1-adc2-4eab-bcd0-d3b0fc0e9b57 · outbound

This paper cites Vision Language Models are In-Context Value Learners.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Vision Language Models are In-Context Value Learners

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:a90f96846f30cf051d7dcd7b0600f564df7b4efd6e3827a39174fd3a181b2a10

Observation 1af7843f-5822-4b11-b749-04224eecdc08 · outbound

This paper cites J., Kumar, V ., Zhang, A., Bastani, O., and Jayaraman, D.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning J., Kumar, V ., Zhang, A., Bastani, O., and Jayaraman, D

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:e91f8e5bda682307408023d5300ca0141d9cfe59ff4bd2b492e5bd149cd1353d

Observation 00521669-6dec-42f6-945a-8f2473994a8e · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:96ea2948725514ec0b2534e92b657c621987682b9737bfb4764d16624193dece

Observation 5880f5f0-3f52-4486-a046-4f97ab4df8ca · outbound

This paper cites Continuously Improving Mobile Manipulation with Autonomous Real-World RL.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Continuously Improving Mobile Manipulation with Autonomous Real-World RL

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:55349c43edb0b0acb045245cc97a07999b53940bb4721749781f32a8509bc9f2

Observation 5de3b609-1808-437b-a62d-46500b7d2574 · outbound

This paper cites Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:a91d5ce2d82c0e88dcea13ffd351d2f286a63b2092edfe529d9ab96a727d2336

Observation 451c0277-5997-4e89-aeab-2e8b5d7e67de · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:2271d25b6a4a8f2c3beeefb83ba99dbc4029c93269e5a6065534bd11075ab6b6

Observation 64cf7f14-e689-4011-b9e2-ba1b2b6d5818 · outbound

This paper cites Y ., Harada, D., and Russell, S.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Y ., Harada, D., and Russell, S

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:4603135c62548a3ff42ba0c0fb750cd05909924482ee4117492b6e4d449b9dca

Observation 7eb8b812-ad1d-4715-ae8e-702283ff209b · outbound

This paper cites L., Sanketi, P., Vuong, Q., Xiao, T., Sadigh, D., Finn, C., and Levine, S.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning L., Sanketi, P., Vuong, Q., Xiao, T., Sadigh, D., Finn, C., and Levine, S

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:7fe6b9565a60c70a5d2109040d12d14661809087e3a759fa6d3061ee4f4f9518

Observation 08295e29-a059-4c83-962f-f8079be655bc · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:799a894655d813a8be5e953250106c25a41aed952188f7175d5445d9f4d1d12f

Observation c1316c06-19b0-47ea-afef-97cd8674b049 · outbound

This paper cites Spacethinker dataset.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Spacethinker dataset

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:c83873b22fd2ef6bca0cf908fc8527c19f580525902f38137d40b4f8f41bda35

Observation dc570533-9b5b-43b7-a841-9e82520f575a · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:3eaf9034ce9fb92dadbd25d8b677d3704097e1df7a6c267bf71877e6ba83fafd

Observation 326ddb2c-398e-47b5-8d62-9041f74f5f33 · outbound

This paper cites Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:86d63eeb619d9c33ef77914acbe369071e9065ea629612c3220e41ac022b6be8

Observation c52db252-dd66-4867-92a9-4d7c494d345f · outbound

This paper cites J., et al.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning J., et al

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:deab585d0ce1b165a233fe448701d3353bad3de995c7c822e3d01c6f43580086

Observation 6c65d2ae-070b-4d36-8ce4-7cf974c484fd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:bc62ed8a4907b6bbbad0a19761dcf2ef84cb2bdb91cc98478dee7a80104a8e52

Observation f568d937-7887-47af-bb6d-07833bd943e1 · outbound

This paper cites VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:7d4473de37f20efe7cb0152fd46893852d3dc6569e1bde02ab4b66893e75f947

Observation b108c7c0-ade0-40f9-857e-b866973cab5a · outbound

This paper cites OpenAI GPT-5 System Card.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning OpenAI GPT-5 System Card

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:a4c5d858f2718c2d61774ff29b8a2e588507091c8d37c12b9c1702e04ec7ebf6

Observation 001e821a-ec99-4940-a6d2-5d6b75be96af · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:18ded407edf440932a6789212661ea2f6e1ba82f222ff46648682f726f36e2e9

Observation 71b8a253-e382-444a-82bd-f86446f8b43c · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:653a3ae70feb0e2bdb6dd320f46cdd29dd28c0a9860d5ea70f604230cdfacca0

Observation 91189df2-b21b-4b71-b119-1d036bc5e6ac · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:2e55939d733026bed1e425bcb0132b3038921ddc989c1b9f9a20b89f39f85e3a

Observation ae592448-af17-411a-9e61-6674e22f573c · outbound

This paper cites Real-World Offline Reinforcement Learning from Vision Language Model Feedback.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Real-World Offline Reinforcement Learning from Vision Language Model Feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:d2e3b4c01e9e65c3ed6ace9d565cd9a7245cfee0120646144d15c2ec9fc973da

Observation 8e582b40-d12d-421b-8bb9-ac4cc3893c17 · outbound

This paper cites Code as Reward: Empowering Reinforcement Learning with VLMs.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:1770f792a84bafc126ba618e549ed620a28b354101e759f678115db73c65e733

Observation 6a8c27e3-8764-4a95-9a01-36ee6c1cc348 · outbound

This paper cites Steering Your Diffusion Policy with Latent Space Reinforcement Learning.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Steering Your Diffusion Policy with Latent Space Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:77da29cabd0231b9f2ca10fdd3cf6420fb788c0044b4c9d36032e0fc96488bc7

Observation 7db9bcd0-e23d-48d7-b94b-bf583e8e8325 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:6aff3d966c1af2f67aadb459b53bc23bed06aba0301546ae41f574b978e53b7a

Observation b0f84f00-831b-4459-8c57-1980f1a76a1b · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:4bb06a7c5e6a537f66b3ff4587718c9e03ca1183670ef794fb642aac139324f4

Observation 3031dd73-9c45-4839-aa60-fa3855aed234 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:4a5f063596bdd2c2ceac33764efd39c54f01fc015b1a589fee344662abe5ffad

Observation 033617b1-8153-48ca-9491-c069e6a8cd91 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:9f79052e37879ee496d9959486fdd36fc8c716b5c0a6199879eda81588d658e5

Observation 00749b07-771a-41d1-bdf4-5fc232c54668 · outbound

This paper cites SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:4bb397306433817d1609d725712907bd7fb6f68791e48c7284bd60143a54fc1a

Observation 2d00360e-f8df-449a-ac97-4fe3418eca0d · outbound

This paper cites Rank2Reward: Learning Shaped Reward Functions from Passive Video.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Rank2Reward: Learning Shaped Reward Functions from Passive Video

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:2c1b1a8b7e2ee7523ffa719c9f841d55d9de6661f54644e14b78e951ac073cde

Observation 223277ad-c3a3-4dce-9c58-a67a02a08077 · outbound

This paper cites Adapt2Reward: Adapting Video-Language Models to Generalizable Robotic Rewards via Failure Prompts.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Adapt2Reward: Adapting Video-Language Models to Generalizable Robotic Rewards via Failure Prompts

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:1ef8174cf725f067f7e087fe3514c6453d163173ce7717faaee087ebdca2eb6e

Observation 04a77363-684b-4a62-bb6f-769c80e96f6f · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:e967cc90aef3913cf2cd3aec6378d0a2e2127de5b1a267fe15c16ad02ad42968

Observation 95281469-68af-4221-a3e6-435307e72a7b · outbound

This paper cites From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:a648790fac8b56b274694b9ea2af9ae5e098c2ee468e39eaf19897fa9c63eb7e

Observation 391014b3-5fa1-4d9b-9f1e-4a2936501141 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:56052e4fe294d3dd45ed31af6b174d242a6ea3d4152d9e590724dd38296059a7

Observation 61759c34-9b6a-41c9-83d3-d3a92d7212db · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:89b438ea49c7e0cc3c66b7e7ee3b7484c70238193269b144be6ad195fa53bfaf

Observation 26aefde6-bd9e-49d4-b247-8d3e204fef35 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:55a485dfc36e7b0d9a59f9a05ffe9defb8e4d12385fa87ae90a8e8372eeef583

Observation a2a8eb6e-c0ef-45a2-9704-62907eda66af · outbound

This paper cites A., Lim, J.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning A., Lim, J

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:3d4be34e3e1450c60cf6ebce988c2d7a0be96dd66ed1a6864c7498b4e401438f

Observation 1d8c31d4-55fa-42aa-a26e-d307f82e0179 · outbound

This paper cites A., Lim, J.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning A., Lim, J

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:b7199c7b8df623e9051147fd678daa8404534b6f83a4e1794f24d4ee361ffdc0

Observation 90f2bd49-a0c0-4c77-93ee-3e95ab17149e · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:217100fae0b6339517b85630c87ddf931d8dab2c996f67f4550142c9f5bb5661

Observation 904c8a16-d0a8-4b19-8bae-3e0284a62a0e · outbound

This paper cites GRAPE: Generalizing Robot Policy via Preference Alignment.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:5e985290bdf5c7c15f0affcc5ffbc0b82b8dba92a629ffcf5d3b217224560391

Observation f83b1ae3-17e5-4c8f-a1e6-5a4e5a693fd2 · outbound

This paper cites close drawer.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning close drawer

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:2c45fa779243ed4585d2a3ef8aac4538acd2755d61fdeca2dc9f5f3e79b703b7

Observation b32a9b70-1144-4ada-b084-601e50e53e42 · outbound

This paper cites To prevent optimistic extrapolation and reward hacking, training data must include authentic non-expert behaviors spanning varying degrees of task completion.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning To prevent optimistic extrapolation and reward hacking, training data must include authentic non-expert behaviors spanning varying degrees of task completion

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:14cf155bd66b5b9ca2804301d857dd5cb988bef942a32a5d7b35f41ea79dda0a

Observation 00fb0e76-b612-419d-9415-4527fe0e1fd3 · outbound

This paper cites an unresolved cited work.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:b5e8f4983ccf3f871681fc8962ac465c9d50405e2d497fadf3ad7994a3ee900c

Observation 3a13f084-b3b9-4dfe-871e-f589391090c4 · outbound

This paper cites near the handle.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning near the handle

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:d9855e4160e15113c607754ed9c750b0347ce1590d6f4cf422edba5699d2748b

Pith citing papers

No inbound Pith citation observations are available.