Pith. sign in

Paper Citation Record · LEDGER

Process-Supervised Reinforcement Learning for Code Generation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2502.01715.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01715 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:31:54.286136Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T12:44:40.230064Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 18e4674b-e365-46e7-8ec7-e59fadbb27af · inbound

Improving LLM-Generated Code Quality with GRPO cites this paper.

Improving LLM-Generated Code Quality with GRPO Process-Supervised Reinforcement Learning for Code Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:54.286136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:54.286136Z digest=sha256:e6b53b79e783c7ff26402b3dab89f63cbde3b7c009166d408e3be64eb2c9de0d

Observation aaffbf05-6329-479d-bf10-f10098639611 · inbound

Reinforcement Learning in hyperbolic space for multi-step reasoning cites this paper.

Reinforcement Learning in hyperbolic space for multi-step reasoning Process-Supervised Reinforcement Learning for Code Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.819837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.819837Z digest=sha256:73c9ba7db1dca20254be4ad6050219a820a244e4243fbd09b2c9b9dc04f86973

Observation 1ebf6c85-9541-44f4-ba5d-45ae9cf8dc64 · inbound

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs cites this paper.

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs Process-Supervised Reinforcement Learning for Code Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:38.089096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:27:38.089096Z digest=sha256:3a4e9eabb64fc59191af7142bd6df2341bdebc0eea89e94839594e658ee15acb

Observation 862d71d8-3201-45e2-a169-e6611ec11551 · inbound

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation cites this paper.

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation Process-Supervised Reinforcement Learning for Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:35.177823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:19:35.177823Z digest=sha256:d04f5ab1a7d5057e0ed9ce29b582fb4281551cab7c42c92bf532bc3aedec3a82

Observation e0374b31-8d81-4623-8e1e-6d44d37a3bfb · inbound

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning cites this paper.

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning Process-Supervised Reinforcement Learning for Code Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:38:18.788744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T21:36:14.478007Z digest=sha256:12e9326983e35586e0102ccec1cc2acecb5474b2a8b2086ba2f112c45605be9d

Observation 6a3feea1-6dfe-49be-aa00-6b9d94158540 · inbound

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs cites this paper.

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Process-Supervised Reinforcement Learning for Code Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:45:59.984342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:58:10.013475Z digest=sha256:b174ac5866d99d43715b23749ce8706c836c10233e7f0676a4715921a09f1a3e

Observation cd0ae5ae-5a2c-43c2-abbf-20c4a72c161e · inbound

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning cites this paper.

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning Process-Supervised Reinforcement Learning for Code Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:01:00.233468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:51:19.555272Z digest=sha256:225c26ef18185ebf56fbd2249d4f8d61197ec48b51bad13afcd2e73c3eda27cf

Observation 1d03fe18-37d5-47bd-8db7-482b1e88f803 · inbound

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling cites this paper.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Process-Supervised Reinforcement Learning for Code Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.288936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:290c2220034ea633e4155fed2f95b566537e31da6ed52e38388d0ed43cb703b5

Observation fae0bdb1-6d17-4048-8adb-07a787463b99 · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Process-Supervised Reinforcement Learning for Code Generation

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:40.231715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:38f3a214c3151c268256c7edb57278091e8502864c954ff9856a5c49925b02c5