Pith. sign in

Paper Citation Record · LEDGER

RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2410.02089.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.02089 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:17.249805Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 97c36278-9d4c-4032-bfe0-f980622569b1 · inbound

ProSec: Fortifying Code LLMs with Proactive Security Alignment cites this paper.

ProSec: Fortifying Code LLMs with Proactive Security Alignment RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:10:55.447728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:10:55.447728Z digest=sha256:eb0729e6a062146a2e7c436ef7f75099fd6a42b0ee5fb392d3ac7b737b90bb78

Observation a9639fc7-815e-4410-a095-b720063636ae · inbound

AlphaPO: Reward Shape Matters for LLM Alignment cites this paper.

AlphaPO: Reward Shape Matters for LLM Alignment RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:51:07.835200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:51:07.835200Z digest=sha256:1322c2e418cfe900916e244a50f5718c96c301d75f715455f3a2d8e23e346df3

Observation 91e1c7b4-1331-4b0d-8ee6-10cdad3b9da5 · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:20:59.399537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:914ce66bb3a37f686b402613f2b608fb107ea120e7f0b05c889c3eb7326c7f95

Observation 689841cb-aab3-40e6-af6c-213939b0fd53 · inbound

ACECODER: Acing Coder RL via Automated Test-Case Synthesis cites this paper.

ACECODER: Acing Coder RL via Automated Test-Case Synthesis RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:53:49.690579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:53:49.690579Z digest=sha256:f382ab793c7ff216b4ca19c15a7a83fd2b3bbdc9cc7a7f60ecaed258d2bbe4c5

Observation 4f86d814-6f5b-4ad7-a836-828d61c88d09 · inbound

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution cites this paper.

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T10:27:56.341144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-15T10:27:56.185943Z digest=sha256:5acedc80ad11d95b61153f1714fe4f7f888784c7353f5517387e711f9934a043

Observation 404dd99f-12dc-46b2-bc31-ea15c2f6cafe · inbound

What I cannot execute, I do not understand: Training and Evaluating LLMs on Program Execution Traces cites this paper.

What I cannot execute, I do not understand: Training and Evaluating LLMs on Program Execution Traces RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:14:54.193857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:14:54.193857Z digest=sha256:9880604737c98f1666f5cd7776128e65e16db0ff9934c74013d0b9dae9fc0aa7

Observation dfd6e967-1c5b-4d72-be45-fe8404b61f48 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 206

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:23.564056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:a2bcf508a49f7a285f2b3d8bc112c024f157bbc1bad3c9c43c6094d244f5114f

Observation 43bae483-6515-488d-813f-dedbbcfb33ce · inbound

Gemma 3 Technical Report cites this paper.

Gemma 3 Technical Report RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:22:12.170943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T22:18:55.976503Z digest=sha256:76b0e0e3d6b1d2dfff33cdb85dac37b6243627e1e12e3c16e14859a184d631cf

Observation aed3e921-9a1d-4c03-862b-3803fd15b4a2 · inbound

d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning cites this paper.

d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:17.249805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:17.249805Z digest=sha256:c1a584f9f7cc33700ee8779d6e2040d8c3ba7bc4982b9c711d016a32fd80c1c7

Observation cd65e293-366e-4696-91f7-50f144370466 · inbound

Themisto: Jupyter-Based Runtime Benchmark cites this paper.

Themisto: Jupyter-Based Runtime Benchmark RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T12:39:09.591657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:39:09.591657Z digest=sha256:3c1fb12fff73ae77e85a733b1b39ed2d2549fe5bf63fb06771a9a933069f1f5c

Observation c56626ec-d521-44db-b591-9947b9a4ba79 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T19:32:00.981090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:e7223a6d05aba7b5cfe70bb3b2e429a922972e0889ef57a72ddf56ab9eae9b67

Observation 450fdcf1-762c-496f-9437-fde06d88c6da · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 166

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:20.515021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:20.515021Z digest=sha256:62434e87cd2c07acb59586fedac187327bc75ef730cbedc471926d5e28f8bced

Observation 219f6c5b-4ee5-462c-9e17-3cd06d722d74 · inbound

Aligning Constraint Generation with Design Intent in Parametric CAD cites this paper.

Aligning Constraint Generation with Design Intent in Parametric CAD RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:12.210307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:12.210307Z digest=sha256:4085a01cd4376eb2d1e072c68aab0b3d1e9396ce73500538e3dc807e391aa40f

Observation b7b590a6-db9c-4d15-9e3e-03b5f9ac56fe · inbound

The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models cites this paper.

The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T00:45:04.892077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:45:04.892077Z digest=sha256:fd2e5e444f37ac3673fd2820d904c73767bdb20f15e566dcc53321f2421a5860

Observation 41cd47ac-0822-4c76-8982-68592cfe2bbc · inbound

Reinforcing General Reasoning without Verifiers cites this paper.

Reinforcing General Reasoning without Verifiers RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.613928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.613928Z digest=sha256:389e38e76b11ffa93ff27a964ead576cd558bdadea4571bb611c86539cd340d4

Observation 484deb15-1a48-4b08-a34d-273bc18125e1 · inbound

Training Language Models to Generate Quality Code with Program Analysis Feedback cites this paper.

Training Language Models to Generate Quality Code with Program Analysis Feedback RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:04:58.623446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:04:58.623446Z digest=sha256:8fdc777856d2b0866ebd440d0751619aa455a459c1401b13f1ce8c0389e0eec4

Observation 5badb4da-5bed-4308-8c0b-fd622e9b4717 · inbound

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization cites this paper.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:24.815311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:24.815311Z digest=sha256:00ad3d99547ffbf7106c2d31b1b1bb70b181a4b1169f23c7e52d4b08d855053d

Observation 1aff602b-b69d-462f-8792-7d66865a4e5d · inbound

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training cites this paper.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.658504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.658504Z digest=sha256:fd7202297a2571b71d7d9307759b3cdef53a9a297cf0c3255307515351f75ca6

Observation 42e8fd50-6085-47b2-9879-b8b77ad352ed · inbound

Improving LLM-Generated Code Quality with GRPO cites this paper.

Improving LLM-Generated Code Quality with GRPO RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.421299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.421299Z digest=sha256:805b898a3941facf0482cb784ea8972efe9a43151582cac5c8456c812e8252a4

Observation 43b36156-2531-46f0-b9dd-afe892cd943c · inbound

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models cites this paper.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.562765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.562765Z digest=sha256:1f03eef25c5257d59261ad40acec1d2376de75842d471616fc2074998af97d00

Observation 3b9a162a-60ac-446a-8ddc-30c8f0d508da · inbound

Reinforcement learning fine-tuning of language model for instruction following and math reasoning cites this paper.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.805270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.805270Z digest=sha256:3841999759e0e327184562ddccb8c72b02eeac3a545a792a961bfe33ca7dc1c9

Observation 9b939a93-5116-451a-b01f-0489329ea5b0 · inbound

CodeGrad: Integrating Multi-Step Verification with Gradient-Based LLM Refinement cites this paper.

CodeGrad: Integrating Multi-Step Verification with Gradient-Based LLM Refinement RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:41.213688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:12:41.213688Z digest=sha256:b886e020836fdcc41bb8bdd551073a2d4a401cba00e28d6b9ca09e4c7445e069

Observation ea83a352-454f-448a-810f-8d2cc060d73a · inbound

VERIRL: Boosting the LLM-based Verilog Code Generation via Reinforcement Learning cites this paper.

VERIRL: Boosting the LLM-based Verilog Code Generation via Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:28:23.056538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:28:23.056538Z digest=sha256:3a518a866ca9381fb313614faa531a6c669b3296993cdfe2e09837fc87a0e0c1

Observation 97650b8c-9389-4fae-9456-c96ce892e39f · inbound

Short window attention enables long-term memorization cites this paper.

Short window attention enables long-term memorization RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:11:21.813433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T12:10:42.646127Z digest=sha256:69d17d6f9ac3e70c679694eb4de1a07e8e969c4e77926f4bddd3e4bf955521ab

Observation 9651214a-fb95-4701-9ae9-a09d438c2181 · inbound

KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta cites this paper.

KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T13:45:23.514060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:45:23.514060Z digest=sha256:051e7bd2ed6c49325b355e7fb94048121e19d404923623a68cfa438a7329075b

Observation ba509e48-0377-4279-bbfb-229312551200 · inbound

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation cites this paper.

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:31.452883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:19:31.452883Z digest=sha256:983600b14b8c1031dc95c99335e97f3ea0b825972a849c007e88dba2789a75c8

Observation 6f6b0ecb-48a8-4d63-a242-e6c291cdc467 · inbound

NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable Architectures cites this paper.

NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable Architectures RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T10:38:11.981408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:38:11.981408Z digest=sha256:66c85074307afe09a449694b653ec6bbdfd2395b6c3b32738829698b25654a6f

Observation b38a1b66-17ac-4be1-844c-5a99239250b9 · inbound

Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints cites this paper.

Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:10:47.874733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T20:12:53.062214Z digest=sha256:10904b8bf8289174d7aef825cd7fc21f2151811d42553c6e3927942465b60130

Observation f5e72726-7d9c-4e30-9322-09cbe9cfdc61 · inbound

An Iterative Test-and-Repair Framework for Competitive Code Generation cites this paper.

An Iterative Test-and-Repair Framework for Competitive Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:48.523439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:44:32.950977Z digest=sha256:73c58cc52a3866d82d5b313371dd98d622aef50f8df961180a28d63223c7ef52

Observation 98fd10ca-0dbe-4f9b-bc8d-ac1fcde25d4e · inbound

An Iterative Test-and-Repair Framework for Competitive Code Generation cites this paper.

An Iterative Test-and-Repair Framework for Competitive Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T09:22:44.557413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:22:44.557413Z digest=sha256:0c9c534c890654681fc55758fa6f2c3dbba785dcc86ada2b41b1b7d063b84121

Observation b6d7d84f-326b-455e-8077-d709f2e0fac7 · inbound

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning cites this paper.

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:30:53.249113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:45:40.915428Z digest=sha256:d103cc7623bad8a2893f10de5ddbb679b419c332917c7ad3142a94db1cfcebf6

Observation 1027fdd9-2fe5-4d0a-8b08-6c6c3423f5e0 · inbound

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning cites this paper.

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:58.374374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:42:57.596073Z digest=sha256:7cf1f85e0e8b150186f5e09426df0c833e0d6d4c1e43134f0555f8bdff4e9514

Observation d4a371a1-8bf5-49fd-8ab3-d17134a5225b · inbound

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents cites this paper.

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:25:35.836993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T12:22:13.551709Z digest=sha256:9a3f32362c95871e83ae10f02544202c821db5fbc2c9c1c7fd2e048dc4e23dfc

Observation 5335f5e0-a2ef-4c60-aca3-0809b19ee7ab · inbound

CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora cites this paper.

CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:29:25.217935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T04:59:44.880102Z digest=sha256:335e818f11c3c2b6105a3b2998f5ce365ac31962bb56a18533529c3b8ab9dbb4

Observation de5616c2-4498-479a-80ef-866762313403 · inbound

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents cites this paper.

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:26.080342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T01:12:37.970638Z digest=sha256:3008d8432c34a2edbf458a40a599dce17454808dabe3c11997f49f4a35d757dc

Observation 1e6959cd-4a2f-4088-9d8f-8fa9aee28c31 · inbound

BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models cites this paper.

BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:11:18.979788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T03:09:40.321500Z digest=sha256:7f43c16342cb6e5cbd73e84b7b58a4638e28725b2a7a180e3b2fb6ee14a02b7f

Observation 71e6a669-defc-4543-bd0c-571fbabf2cb5 · inbound

BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models cites this paper.

BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:07:22.367744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T06:03:32.270553Z digest=sha256:9af2eda96ea2a01412ae56234ae63eb725a17d8e9af12284e1299b38ca6e06c1

Observation 61dd3dc0-56a7-474c-93e4-1a540b0357bf · inbound

Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards cites this paper.

Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:43:27.601951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T01:41:31.631505Z digest=sha256:68f9141532d748392418d8dcb0e7366724942f5c928f8da52453de44da7599b1

Observation 706ce023-1d4b-400f-9605-f8ed0fb8259a · inbound

Self-Supervised On-Policy Distillation for Reasoning Language Models cites this paper.

Self-Supervised On-Policy Distillation for Reasoning Language Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:43:22.222344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T14:42:55.368104Z digest=sha256:61d49a879e2cac4cda07d9d032c135316463fbbae66853d51c2db08252458817

Observation cbbfb1d5-2ad0-4a09-851a-d943557595ac · inbound

HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL cites this paper.

HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:28:18.834772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T13:28:03.965754Z digest=sha256:0ddf1beb97cef7f2dcae22fe7c43db4ab5fed269de91421841dff3826e000815

Observation 4b8c36f2-35bd-48e7-b0fe-ebdaa2ae94e0 · inbound

Code as Agent Harness cites this paper.

Code as Agent Harness RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 104

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:58:14.301565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T10:54:54.558241Z digest=sha256:44fcfaf642af1a88b67b6b4b13e136efdde42eee33c06196169b412718842fd0

Observation e60c5963-68fb-4953-90eb-5fc7254f05f9 · inbound

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models cites this paper.

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:13:58.422663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T05:11:35.785227Z digest=sha256:d94b13854535691082f2cacc5005f37a3f4b2b0b95544775d9edcd6e92d96bc0

Observation 59c618b0-de54-4a45-99fb-e3b65c89d6f4 · inbound

Self-Policy Distillation via Capability-Selective Subspace Projection cites this paper.

Self-Policy Distillation via Capability-Selective Subspace Projection RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T05:34:40.386519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T05:31:42.803465Z digest=sha256:d363b6a8f9827340d522f01c54b583c98a8efa31a7a078f28157e15768127316

Observation 86301b5d-5ca7-45a6-93ff-f7022016e5a4 · inbound

Learning the Error Patterns of Language Models cites this paper.

Learning the Error Patterns of Language Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:03:29.670798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:56:04.187007Z digest=sha256:e1d5582b7d2eb59e6a030df165ad625a2ce0fb9022fa3e02bbbf12989c4b02e4

Observation cf5dea75-95b2-4735-be05-ef2454e75c9d · inbound

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs cites this paper.

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:46:33.093345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T09:40:35.258818Z digest=sha256:777faa780b3f8747f7a5c300a04b5d7be76664d14014d546e679235a0270aae3

Observation af5d5b50-2658-4b97-b35b-aa63b36d48df · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 298

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:38.270687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:bc0409b2aaa929b09cb8be14dd2aee22b286fb6e5a655a9118b400139e04a0ff

Observation 0db63301-0fce-4dd1-a679-b74a2b8679f3 · inbound

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs cites this paper.

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-06-30T01:34:09.489245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T01:29:42.919461Z digest=sha256:c78e3120b0a8313832f496ec640968b215769ef82a791ecbf70161436ebc9480

Observation 94718f18-42ee-4043-b317-adf8860d58f3 · inbound

DecompRL: Solving Harder Problems by Learning Modular Code Generation cites this paper.

DecompRL: Solving Harder Problems by Learning Modular Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.791771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-03T16:30:34.793328Z digest=sha256:c0cb891cfde39df8f73e26ea2d4177efe01f020fb83e7d5e21ccb631945e1924

Observation 16ed8299-fb9d-4606-9c31-af71c070ade7 · inbound

Beyond the Need for Speed: Energy-Aware Code Generation via Simulation-Guided Reinforcement Learning cites this paper.

Beyond the Need for Speed: Energy-Aware Code Generation via Simulation-Guided Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T17:00:48.664985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:00:48.664985Z digest=sha256:2cc04de158bf4643ecb679a9d319c83c3c4ffa842b7dfc366f21c6be82755b8e

Observation dfd47062-a949-4b64-969c-22bb38ce2cf8 · inbound

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards cites this paper.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:3287ad1f016c7e750549c9d49e2c322149d1f74b75d2c6aa32aeeab85b88c875

Observation ae5926b7-33cb-4f8d-93b0-fb5d51422def · inbound

Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation cites this paper.

Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T13:39:15.458226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:39:15.458226Z digest=sha256:2fa2d67faafbdba5b9a4a85a2ed6b719fca06dd24eb33074dd6f55f6c907dc7f

Observation 873b42e5-cbfb-4569-9dae-a9a5bf5f35a1 · inbound

CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation cites this paper.

CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T09:12:22.385849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:12:22.385849Z digest=sha256:79f5425af51f9d3821a3a5c58ff55f118cc0f2c1d77b7801c45f3544eff28d26

Observation 66add268-a6a3-4546-9455-f27ca8322428 · inbound

Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation cites this paper.

Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-15T15:35:43.868636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:35:43.868636Z digest=sha256:ce64bc40fb137c62e7f629d28d5fce0e3e54d1bd7816f47811e7b3ea7fbc9e6b

Observation 8597521d-8618-40b5-9d1a-cfa2c439d041 · inbound

Training Large Language Models for Self-Explanation Faithfulness cites this paper.

Training Large Language Models for Self-Explanation Faithfulness RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-01T08:36:30.457193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:36:30.457193Z digest=sha256:1f3af83d86f677ebe43beaec88b06b2811f80d57ef0bad3abd12f620a53e0fa2

Observation bba3d182-65cb-4f0c-a6a5-429435763b84 · inbound

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills cites this paper.

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T04:28:45.552110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:28:45.552110Z digest=sha256:96c621aaa09098f516fc26ec19f6a5b4b0254bb54663a865c0291ac944a73249

Observation 9989336a-ed70-49c6-9097-e3ccdfc965fc · inbound

RLPF: Reinforcement Learning from Performance Feedback for Code Generation cites this paper.

RLPF: Reinforcement Learning from Performance Feedback for Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T10:53:08.880515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:53:08.880515Z digest=sha256:3eeadeeab02e5eb627a8fe9813de65e0f7fa27e9fd9692ba1403f7bffc62f6ea

Observation 4b0e5296-46c0-4996-9f78-40f3af77103c · inbound

LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation cites this paper.

LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:38.260775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:38.260775Z digest=sha256:6204f8705f02e5fca722a23c4ce668b4644aa1094ef08cd84e618127a9b8571e

Observation 6b017c0d-f03d-4ccb-8932-104c8caa8a47 · inbound

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning cites this paper.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.219331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.219331Z digest=sha256:a1a3b6cdc047add3d0ccf4ac0f24fbf428f6a5c0e4d6fa3a5db05595d93bd361

Observation de34f7b8-461e-4d9d-9e54-8a1d419f59c9 · inbound

ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation cites this paper.

ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:44:40.247401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:44:40.247401Z digest=sha256:71446ec1020532a48c798921be115883df32e1c6e9cfea8d1347cc3799603905