Pith. sign in

Paper Citation Record · LEDGER

Code as Reward: Empowering Reinforcement Learning with VLMs

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2402.04764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04764 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:24:02.379856Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5ef0f259-b734-480b-83f0-32aa46ffb8de · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T12:24:02.379856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:24:02.379856Z digest=sha256:c71ab02371a1468036cd62d9c93bc58ebaaf32bdc5a5a0e7eacb9120fea83b76

Observation 8e582b40-d12d-421b-8bb9-ac4cc3893c17 · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:2774d1fe4f2e4a03fc0ba283c27a25b5f70f6dd4b950f2ee00545976d5a55d6c

Observation 71bedf21-e11d-47d6-abf7-0821cf3432ae · inbound

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence cites this paper.

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:44.083013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T03:57:19.028507Z digest=sha256:63c518c0d30b9ebe84d57c8fbafd77b2e73545336c85e1a2445344f997bcf968

Observation 6e83b4c1-aca7-4359-b8df-6992451c2968 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.529491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:fc3ab44ec658284a36f8b7ac6b99c061ee2b364b6cc13919dc5e0c50322e958b

Observation d7d7e84b-5cbe-4546-be36-ed692400209a · inbound

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents cites this paper.

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:05:28.982000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T05:59:13.631078Z digest=sha256:e6583a6279cec553f0eb87598b98bcfacec282e8127f964f2f7448292ae7071c

Observation fd35c397-5693-4c02-86b2-8a0975e6cef4 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Code as Reward: Empowering Reinforcement Learning with VLMs

Reference 245

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.270452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.270452Z digest=sha256:801e2a26cb4054ff8f2f40027e3543c93e6a86b39ae96580aa5fd2e68147a092