Pith. sign in

Paper Citation Record · LEDGER

StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2402.01391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01391 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:39:09.585165Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 493ee7ce-f59d-4a86-87b1-2c8bae5e5a44 · inbound

A Survey on Large Language Models for Code Generation cites this paper.

A Survey on Large Language Models for Code Generation StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:18:06.739448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T20:18:06.304134Z digest=sha256:47d66031e7dc3e71162ae84155f7fe766be78d129b2defd33bb1f2072664962d

Observation a6fe8069-4af6-42a9-b8a4-475274dab7e2 · inbound

MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic Testing cites this paper.

MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic Testing StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:18:31.317120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-23T22:17:02.681066Z digest=sha256:605d78dbed94727e6ae9cfe269d540468c4da48493bd8122bf07b77dcaba4ed4

Observation dae15ffe-6ac2-4676-832a-1328af38d70a · inbound

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs cites this paper.

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:05:15.099853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:05:15.099853Z digest=sha256:4ff20bbc5c78be7e1454864aba0ea336751ac6054152e0a111e606aef11a4edd

Observation 6bbc4a7e-b8cc-4b9e-86d0-5d060bdf148d · inbound

Preference Optimization for Reasoning with Pseudo Feedback cites this paper.

Preference Optimization for Reasoning with Pseudo Feedback StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.560207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.560207Z digest=sha256:0db39131af4a441895486956252da82307fad7033dd4d91b5f1f11115a3e3342

Observation 7f1aaa35-7be3-47ad-a19d-409ca0333f57 · inbound

Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning cites this paper.

Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:11.364925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:11.364925Z digest=sha256:0e030e6aaa55176cc0264184389da934cb6ef113028489eaf21b4384c2fb1be8

Observation 9d034c76-dece-451c-96ae-c69ff824e8a2 · inbound

Distilling Desired Comments for Enhanced Code Review with Large Language Models cites this paper.

Distilling Desired Comments for Enhanced Code Review with Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:28:51.280416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:28:51.280416Z digest=sha256:44362a69ffe98712ebb34f00a6f54acbd45c0afda30c6037209d22cca696b0a0

Observation 7a48100f-4cdd-435d-8ebb-f657497a38e3 · inbound

ACECODER: Acing Coder RL via Automated Test-Case Synthesis cites this paper.

ACECODER: Acing Coder RL via Automated Test-Case Synthesis StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:53:49.686606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:53:49.686606Z digest=sha256:b9a9b2471ff07cfe0cc7ec03dc66c0c3b5d9792d4b013cbdda55e2e840768607

Observation 1c172a57-6eb3-497a-9cd2-7e8bdbb84290 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 164

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:23.489044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:b516af652feb8dc13315caab8e6cfee08fd9e81e6f3ddae9da3b0ed695358285

Observation 3403b63f-cad3-4b34-a036-c1401269f0ff · inbound

Themisto: Jupyter-Based Runtime Benchmark cites this paper.

Themisto: Jupyter-Based Runtime Benchmark StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:39:09.585165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:39:09.585165Z digest=sha256:c46e57126eb0fb07b4f18027d34221497793974ffac63f5ad9f02689e7f9a959

Observation ea30a76c-e77e-40e4-b4e9-06aaaf99fb45 · inbound

Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs cites this paper.

Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:34:01.054883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:34:01.054883Z digest=sha256:8d6d0ef08addca2371841eb7fd3dd47e6208d22e6df3b557abfe6f6ec68b4c0f

Observation 44bd3ffe-f9ed-4004-ad79-63854f7f6591 · inbound

CRPE: Expanding The Reasoning Capability of Large Language Model for Code Generation cites this paper.

CRPE: Expanding The Reasoning Capability of Large Language Model for Code Generation StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:22:19.042890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:22:19.042890Z digest=sha256:7ed9f0f77d9ee728ec201e9aa2859b6166db6aa272d5f89315cb15450f83d351

Observation 59474258-9f83-4571-bd6e-e0f7e94a8363 · inbound

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution cites this paper.

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:03.712435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:03.712435Z digest=sha256:495c883cdc5bbe47a3ad6f1144eb2ae5ac9c3b357c5f1e709694875e9065e1ad

Observation 6bb10acb-174c-40f0-81d2-ae3bcbc23691 · inbound

Training Language Models to Generate Quality Code with Program Analysis Feedback cites this paper.

Training Language Models to Generate Quality Code with Program Analysis Feedback StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:04:58.439514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:04:58.439514Z digest=sha256:e5ab8b6f8b9179458c63b008aa1b4699e6499063e60a0cee7c10d636a243b90b

Observation 457bd8db-44ed-4218-8753-a5053c9b1667 · inbound

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization cites this paper.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:24.562628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:24.562628Z digest=sha256:39d7ce183f45d82c37c10f4404787ed9c2afa82fd7118c4a2bb2d8299381c0b1

Observation 4887bd46-fc39-4dd5-b312-36bb30a7e0e8 · inbound

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review cites this paper.

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:28.236569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:28.236569Z digest=sha256:8a175b8f7d2377540e3cb963465faf68bb039f10c1954e9011f41e2026215bbd

Observation 27ef6c89-a7dc-41ec-84a4-cefc545ee0fd · inbound

Improving LLM-Generated Code Quality with GRPO cites this paper.

Improving LLM-Generated Code Quality with GRPO StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.249493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.249493Z digest=sha256:ae5f9525673207f89f0903f1d79005395ab68c699a3fca01d6b9cfb3c82ebeb7

Observation 14c18af4-3176-4b4b-a3dc-f9fdb4b4e445 · inbound

D-LiFT: Improving LLM-based Decompiler Backend via Code Quality-driven Fine-tuning cites this paper.

D-LiFT: Improving LLM-based Decompiler Backend via Code Quality-driven Fine-tuning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:41:56.947388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:41:56.947388Z digest=sha256:65dfa384127b3cc1c67102a8f76277f3f4a4e0255db7dfdac6a8a8994090c005

Observation 7c073181-d46a-45d0-908e-9845c553223b · inbound

SysTemp: A Multi-Agent System for Template-Based Generation of SysML v2 cites this paper.

SysTemp: A Multi-Agent System for Template-Based Generation of SysML v2 StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:10.403693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:10.403693Z digest=sha256:415e5ef405944b795f36965ce198c5bdf22723e97a91bc1345a9710f10bda5d0

Observation 7a511d87-afd3-4702-83c6-9d560eabbca4 · inbound

ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle cites this paper.

ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:47.513069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:47.513069Z digest=sha256:792816a449ab08d772b67ac3eaa81e47ec709a04ed683e3efb52e8896291739c

Observation e11cf801-ef24-4569-aa48-584de858c0cc · inbound

ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge cites this paper.

ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:32:01.539614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T03:29:23.464348Z digest=sha256:3fd2a56bf0ab46b844935fce859219d3d61fcaea4fdf3fcf571077997247a75d

Observation 1182d1ba-ef9a-4b20-a296-7ae0aea84ed5 · inbound

Repair-R1: Better Test Before Repair cites this paper.

Repair-R1: Better Test Before Repair StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:19:19.040967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:19:19.040967Z digest=sha256:217bd43458f55d5eb539ca90aea62c053649b2c7de06711c75511d2bdf544b2a

Observation 983fec0b-a38e-41d8-aba5-5a554a3e9642 · inbound

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention cites this paper.

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:56:50.471275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T20:54:30.449792Z digest=sha256:4451755f5b15fe154cf56e087aff1f9ac8a52715f3473885bafd155df9885ef6

Observation bbe70dc4-8e78-4d4e-88a8-b072a15b3a7e · inbound

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models cites this paper.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.749999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.749999Z digest=sha256:60838a0545844f3c762934f9c2d7ad240ca94be04ea80172e73f8a43b82f7059

Observation 26cfeec9-3740-40b3-b940-d4cd569333d8 · inbound

Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data cites this paper.

Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T09:02:26.930703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:02:26.930703Z digest=sha256:e1b1716fc6eaff499990002aea4729117ed6961488d24717a44dd025b98424d6

Observation b21c6ace-e1f4-4553-8d50-6b21cb945071 · inbound

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping cites this paper.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.826796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:4590bd5f503b8de94c2c0d2e1c83a25138311c85d81947b10f4b661038c32038

Observation 9d74a1e8-b880-4c2e-bf55-8761bab3834c · inbound

Towards Enabling An Artificial Self-Construction Software Life-cycle via Autopoietic Architectures cites this paper.

Towards Enabling An Artificial Self-Construction Software Life-cycle via Autopoietic Architectures StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:45:23.468601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T12:43:14.903173Z digest=sha256:c7c7da025fecfb9cec3ed82599ea094ec1697e06bbdae0e28a286f12e03f7832

Observation ae70bde9-1c87-4c90-aff3-4b911f425f8f · inbound

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems cites this paper.

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.660929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T07:02:02.992871Z digest=sha256:e301d683adde6bccab045c463a30a91e0dca25029de3ccea87869e173f95d20b

Observation 52d9851c-dd6b-438e-b8a0-86ca5b61a069 · inbound

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning cites this paper.

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:19:46.690201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T00:18:18.340391Z digest=sha256:efef5839799a2155a45cab9afd3fd31d888c55e2512ee6b0f23aa8f5a26105fb

Observation 81432182-5677-4745-bd85-ee10950b3dc0 · inbound

BoostLoRA: Growing Effective Rank by Boosting Adapters cites this paper.

BoostLoRA: Growing Effective Rank by Boosting Adapters StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.696129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T08:00:21.048408Z digest=sha256:1a337a7f7b9662b48d037596aa087e3aa095029349b51d7fce63fbc79c2b7553

Observation c8fd815c-4d19-4ff3-9422-04b899c04fca · inbound

Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning cites this paper.

Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:25.357376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T19:21:30.956682Z digest=sha256:28bcd296e28434e6d946448d4330d6b787587eab4353dc6f4e29ab518d7e24b3

Observation 2836c5b3-c6a5-4726-9aa4-7c5e83278da3 · inbound

Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation cites this paper.

Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:08:21.267584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T14:03:45.869373Z digest=sha256:5ad1c6101e5bb7368d52b39ba49fd77321691a0b9d0b8792a27143c827234c21

Observation f96a4253-1842-4991-b3d9-ad54b92d3e15 · inbound

Distilling Game Code World Model Generation into Lightweight Large Language Models cites this paper.

Distilling Game Code World Model Generation into Lightweight Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:04:44.656018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T13:58:37.956333Z digest=sha256:b775e6cc844a0442a8e5b31a5ca15a69671062dddc192154286415b690e9c715

Observation c3ffb77b-e6c1-4d7a-89fc-d70816408c83 · inbound

Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning cites this paper.

Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:03:24.427771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T11:54:07.211895Z digest=sha256:14e67faf1dde3a68c992b7760f88bdd7cdabd85e86784177c8a3be6fbd15765f

Observation 819bc428-11f8-4236-a84a-3f8b2dd030b1 · inbound

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback cites this paper.

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:53:32.268045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T06:11:06.960389Z digest=sha256:148e8a930daa60eba87079ad522a9fdacf90cdb204d5c9142a503d2bc0f34b7f

Observation 1e40299e-9fa3-4d54-a36f-a848fbcf333f · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:10:55.862963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:7700080d50fbc0271b17d41f89c8a037f5a3de86c3389f119c259cb370147afb

Observation 8b70a479-b2a7-40a8-b68c-7f5172e8d7ab · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:00:19.817888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:6cbf68dd82bc0bdbca447128b3a1196edd15702540b1164857c95fee6d702fd7

Observation a8185407-e9a5-4fd2-906a-d6a1d7386175 · inbound

MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning cites this paper.

MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:35.169574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T09:31:56.712360Z digest=sha256:12968f3427c9c8cc32a400d8e0bd2450ab16f5d75e8ba8bdb154469f460be825

Observation e5986654-96d5-413b-8126-212f1c8c847e · inbound

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study cites this paper.

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:33.953898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T17:07:21.486960Z digest=sha256:52c72975e44cdcd9a8961c8e54d608877209d0ad897926f049074f713675ce23

Observation 396d5770-455f-4181-9e4a-5c07690b25d9 · inbound

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning cites this paper.

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:53.759973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:53.759973Z digest=sha256:2fdcadc9baee229cf8cd52d6fa593209cab7b6f078d26d800b5b20434c42d886

Observation ab486f05-3228-4ee3-9670-56dd3dd2c8af · inbound

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning cites this paper.

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:42.958248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:27:42.958248Z digest=sha256:42d19c01f3ef3bef53b6b61bff49d4492fa7866f9e8bb0a817fe2de1d69c3237