Pith. sign in

Paper Citation Record · LEDGER

Improving LLM-Generated Code Quality with GRPO

As of 8 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 2 inbound Pith citation observations for arXiv:2506.02211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02211 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:31:54.358241Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:14:04.101113Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T17:26:04.783434Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7d8571f-bed6-4320-94eb-c67e61bc8125 · outbound

This paper cites Program Synthesis with Large Language Models.

Improving LLM-Generated Code Quality with GRPO Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.092772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.092772Z digest=sha256:c585d07c53ada7c8aac6a4f5905c8b800825427a8a7495b62812b651ecdde149

Observation 27ef6c89-a7dc-41ec-84a4-cefc545ee0fd · outbound

This paper cites StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback.

Improving LLM-Generated Code Quality with GRPO StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.249493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.249493Z digest=sha256:a7d4236056c428630dc5d210654b7dc92a5c253b42446451f7d91225b211e6cd

Observation 42e8fd50-6085-47b2-9879-b8b77ad352ed · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

Improving LLM-Generated Code Quality with GRPO RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.421299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.421299Z digest=sha256:e0c06e3d1acbf86f589c0743715aba771af81e419370eec734fcf6e284ffa724

Observation e9bfec68-2cfd-40d0-a16a-60fa32f7e1e5 · outbound

This paper cites 2 OLMo 2 Furious.

Improving LLM-Generated Code Quality with GRPO 2 OLMo 2 Furious

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.700853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.700853Z digest=sha256:ffd06a6e300171b1511db6b64380e9531ce631676f0240d88f320aa17975605a

Observation 4e323f52-8cda-464e-850c-c18facdd882c · outbound

This paper cites Qwen2.5 Technical Report.

Improving LLM-Generated Code Quality with GRPO Qwen2.5 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.801015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.801015Z digest=sha256:79927d751a013c0a4f5ed9f9502a19abe74dd2eba97db6a19f1066909724fca5

Observation db786141-1b76-41c9-9c68-35409a456dcc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Improving LLM-Generated Code Quality with GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.979789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.979789Z digest=sha256:d8493b075a7757d3e23fbacb516267e10104730db954b870a5fc32847f364fcd

Observation 18e4674b-e365-46e7-8ec7-e59fadbb27af · outbound

This paper cites Process-Supervised Reinforcement Learning for Code Generation.

Improving LLM-Generated Code Quality with GRPO Process-Supervised Reinforcement Learning for Code Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:54.286136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:54.286136Z digest=sha256:bfc04304ab30462118dd36ec38724d772bde3893fc1229826a54ad71d5e16d54

Observation 378849d3-2eeb-4410-9d09-7f42e5337343 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Improving LLM-Generated Code Quality with GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:54.358241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:54.358241Z digest=sha256:4049b82524f49629dae316bb02a4982c499b9971f4d4a28c102983731658f108

Observation 16b8eb3b-5ed2-474b-9cc0-dc6b5efd9bd6 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

Improving LLM-Generated Code Quality with GRPO Measuring Coding Challenge Competence With APPS

Reference 1977

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.495008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.495008Z digest=sha256:e1bf6af8d81beea79613123a5e444e07b3c7bf5fe9d8cf1c9773489b3a6c6072

Observation f03c37c0-91bf-4042-876b-76f091be84d8 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

Improving LLM-Generated Code Quality with GRPO CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.592291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.592291Z digest=sha256:e88de6340541e3d898f81eea4e0875c4c97b53c3d88d00cc6a8a4828778c676e

Observation 6ca98cb9-9959-472b-be3d-e676d2059e9d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Improving LLM-Generated Code Quality with GRPO Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.908094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.908094Z digest=sha256:cba1263c034130efd806a931934dcafd2dfcbf8d74c686a5b3da9b10ae1b4f87

Observation 7aa5ae01-ac28-44cb-aa26-bbf9606695d6 · outbound

This paper cites Iterative Self-Training for Code Generation via Reinforced Re-Ranking.

Improving LLM-Generated Code Quality with GRPO Iterative Self-Training for Code Generation via Reinforced Re-Ranking

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:31:54.623757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:31:54.067709Z digest=sha256:a8e2f8896294e926b1c048dadd0f905011b02d7599e21ba55aa23d469def1304

Observation 6bbf3403-b156-47ef-9a44-b9cca6cd9fd6 · outbound

This paper cites The Llama 3 Herd of Models.

Improving LLM-Generated Code Quality with GRPO The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:54.201555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:54.201555Z digest=sha256:b14d446ad02c9a16a8665c5cd4bda641056815918e1070207181512ae52159e0

Observation 70f3149e-36fb-4255-b48f-bdcf2e5f911e · outbound

This paper cites Process Supervision-Guided Policy Optimization for Code Generation.

Improving LLM-Generated Code Quality with GRPO Process Supervision-Guided Policy Optimization for Code Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.155905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.155905Z digest=sha256:f52d249d6d7c1fe268ed0241dd166864ac2d518c1b219eea1e72313145adf39a

Pith citing papers

Observation c4a76982-bab2-4716-b3a1-911f1a06ad74 · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models Improving LLM-Generated Code Quality with GRPO

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:04.101113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:04.101113Z digest=sha256:1da1def9c3f1a6848cc90556f11c5966eeff6b370885adb65b9551dbd7d2b7d8

Observation 053b9dab-34da-4aa9-834e-b03b938dc901 · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code Improving LLM-Generated Code Quality with GRPO

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:26:04.786091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:99c6bcf20f492fd533211b9e3291ddf5c8eedf164c87c5d6f9fb16848cf71e90