Pith. sign in

Paper Citation Record · LEDGER

Code as Reward: Empowering Reinforcement Learning with VLMs

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2402.04764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04764 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:24:02.379856Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5ef0f259-b734-480b-83f0-32aa46ffb8de · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T12:24:02.379856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:24:02.379856Z digest=sha256:c71ab02371a1468036cd62d9c93bc58ebaaf32bdc5a5a0e7eacb9120fea83b76

Observation 8e582b40-d12d-421b-8bb9-ac4cc3893c17 · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:01d9b3afe29c11f9bfe879a7f0baca6ef04117a4fb4885400a63c78eb6e46c46

Observation 71bedf21-e11d-47d6-abf7-0821cf3432ae · inbound

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence cites this paper.

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:44.083013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T03:57:19.028507Z digest=sha256:2d2781372a6a92d55741a06ab836b04cfae3d72c03fd48e4742ded42a40a0af6

Observation 6e83b4c1-aca7-4359-b8df-6992451c2968 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.529491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:68c4a5c75e469acc08dfb9f6628c4159c6755de29fd89ca37a6820819caa3dd3

Observation d7d7e84b-5cbe-4546-be36-ed692400209a · inbound

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents cites this paper.

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:05:28.982000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T05:59:13.631078Z digest=sha256:b49ab7fc2e0d3fae7f9a7825d5a79d1163260f79a12109b5eb038b4f5a40cd4e

Observation fd35c397-5693-4c02-86b2-8a0975e6cef4 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 245

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.270452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.270452Z digest=sha256:cbe6bfc2724038b44b2f4dece190d462f34b72071eaa83ffddf2649fe87ad03a