Pith. sign in

Paper Citation Record · LEDGER

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

As of 25 July 2026, this Paper Citation Record lists 31 of 31 outbound references and 16 inbound Pith citation observations for arXiv:2510.18471.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.18471 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T05:16:28.008746Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-25T06:30:59.84592+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T14:30:20.959431Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact24
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cb50c0e3-7c6a-4401-9a99-147dea01ef50 · outbound

This paper cites OpenCodeReasoning: Advancing Data Distillation for Competitive Coding.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.609566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:a1a13bf401a21535c2293b34493ef25acce3bbd13ded1d04b31203d7f18459c4

Observation 133023d3-99be-4eae-9140-1fb305078c90 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Evaluating Large Language Models Trained on Code

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.602730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:b5c1767b29c671a52f534d387aeafcf42c7fd477e43da32658e8464452a17448

Observation 35c381a8-345a-4eea-b64f-bc280f5d75ff · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.624448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:a296b78e175141152890e5e9d2a2fb74be9ec644e7480204d995244334230156

Observation 9f1ccf0c-4e9f-499f-85d0-45291f2e4ac4 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Process Reinforcement through Implicit Rewards

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.579753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:aefb7ea7f8d40bdfbce20baa581b20b3172d13d3f3b6b75c8233480c08461aef

Observation 19022fa2-d24f-4201-97be-b9d8355ac07a · outbound

This paper cites CodeScore: Evaluating Code Generation by Learning Code Execution.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment CodeScore: Evaluating Code Generation by Learning Code Execution

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:20:54.620410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:e733d64b06f4c7b54c8c2af9081a553aa5174a2a23025c7b73cf41b5df1a69d9

Observation f3ba0fca-58f2-4a83-98fe-1e1ceb21b23b · outbound

This paper cites A Survey on Code Generation with LLM-based Agents.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment A Survey on Code Generation with LLM-based Agents

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:02:21.366092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:46d5db210098109a3b54b990fc2b9bbae78cd71a9e2b19a495672e64cbb84215

Observation 33cd7339-6bfe-4775-ac0e-7e589667d67e · outbound

This paper cites ReCode: Reinforcing Code Generation with Reasoning-Process Rewards.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment ReCode: Reinforcing Code Generation with Reasoning-Process Rewards

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T05:20:54.545492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:2b08c0389fb50b859cfac79ebf79eda44f7613161451b2dd648919fd082299bb

Observation a92c6490-3bd4-4fd2-9249-c8876bf892ad · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment MiniLLM: On-Policy Distillation of Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.589881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:77aec189780a2a7b3739a7ee6ad73010a864e19e743f1644ccf7c9c0c0146772

Observation 42fdce16-fd98-40e5-8748-dc2aaa83dbe6 · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment The False Promise of Imitating Proprietary LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:54:31.902059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:b199c495e6a8529fe65b79f1520e8c7c56c0842beb01c7d59c7ac4198aeff77f

Observation 86d251eb-e19e-42fc-aca2-e7b41d675a63 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.606098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:6bb34e7c0d3fb3b9d54526fa727660d482795aff88779a01fb4f0c8594e13920

Observation 4509fc09-955f-49be-9835-cd4b484f275c · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Skywork Open Reasoner 1 Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.576707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:b4da97a0a94400351939fc909c75bfd87c174154b537e37d9d416a87be4afb12

Observation c92de9ca-b71f-443b-a17b-025fe33c0ba0 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Measuring Coding Challenge Competence With APPS

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.613065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:a5cd7587c719277f01e7807f2cd39846c2ff2b046f0fcd21418ae091646295c8

Observation bc0ebe0d-b6de-4509-8776-3e848d148cc1 · outbound

This paper cites Designing and interpreting probes with control tasks.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Designing and interpreting probes with control tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:55.141010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:aa9008e7c6ee314c8cbb0f9ca950ab6eb9a86966e5fc11ae9061ca70f8028299

Observation feda402c-46a4-464d-9774-51e523e764a1 · outbound

This paper cites Qwen2.5-Coder Technical Report.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Qwen2.5-Coder Technical Report

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.522277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:aab7714defa5bc66b8126c44a7ff9ac2bcfef3ac2083dde1c32e2206ea729030

Observation 428da1f5-5f6a-42a0-9113-5b1de683fc1d · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.599756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:53cd19cb60c91f6cf2075e2673fe1500c5f9bab630e99063ea12289c25484711

Observation fae20ebe-62cb-41d4-8060-b9a257df0488 · outbound

This paper cites SEED: customize large language models with sample- efficient adaptation for code generation.CoRR, abs/2403.00046, 2024a.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment SEED: customize large language models with sample- efficient adaptation for code generation.CoRR, abs/2403.00046, 2024a

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.555648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:5127770630b29c04f4fdb2a46e8e67dd640a7e7d59f94d90b101594e7508eff7

Observation 883330c4-a3b1-4746-b00f-3c2b18ea02a0 · outbound

This paper cites How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.586746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:b4d06ddd6e3d38d65a07a3344d77d15559fce7e81bfc5d6865b6b1b381db1cc0

Observation 498c88f4-de7e-41d7-b0fb-2772d6d3bf35 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment TACO: Topics in Algorithmic COde generation dataset

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.527865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:50b8c4d3e9114103caeda04151d3a654ec39e30fe4a2521e24f45374f21f8024

Observation 67f68427-9c13-4e8f-8a61-4f441c4caee8 · outbound

This paper cites GPT-4 Technical Report.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment GPT-4 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.616387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:f6d8f135d060bae3c72935ff5b75977302e69af0e07ef8c2a3c772ed0024060d

Observation d75a2d4a-d82e-47e6-924a-1b12a506e19b · outbound

This paper cites Code Llama: Open Foundation Models for Code.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Code Llama: Open Foundation Models for Code

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.631628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:adb4f52ec867d66963caa46aa2bc2e09568115668eb6c10f8dcb654d823e097b

Observation f338b61f-d491-4256-adb8-294aac569a5d · outbound

This paper cites Proximal Policy Optimization Algorithms.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Proximal Policy Optimization Algorithms

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.627938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:1bf3c206c540d06c4b293c83f391869a0e1a264c644815d4c5d77bcca25729dd

Observation 1ef83e66-5b18-4934-a089-3cab40d05de3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.559062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:03fe45b7f002b96ca85b92951789536f0c3ff84eb8ccf1f772b57568fcb8589f

Observation 7b23b830-ed05-4bf7-99dc-b4222ef13cee · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment HybridFlow: A Flexible and Efficient RLHF Framework

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.549662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:e51b7092289af970c33a0880f3a2532bfbc335cfc61802b4ac99e6c70a9578d7

Observation bf918166-6fc3-4071-8ec0-d28d694f7339 · outbound

This paper cites CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.583092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:28fb98a5464ae4049dcb0acb865c32f0b2c3484fa5feebac9e23b6c341540ee1

Observation 7bd90871-6fa5-4941-b7e0-192da2733377 · outbound

This paper cites CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.573655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:6dd9d18aa90ac735970ae9f48f8b2571eae66b6f52e18ede72c1c2aad856cd40

Observation c8a3f9db-b5ba-4130-b6c0-afbef25450ab · outbound

This paper cites Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:20:54.562712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:fb739efa883f7d6eebd0d5c15583eaba90836b28229c0ddd3daf18bd629360bd

Observation c5509687-8d24-4de0-8b8a-917d97791d98 · outbound

This paper cites LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.593415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:e444357882bc852e2980fea472ce1608d760b6f6413c99d46a15d403437898d5

Observation d6eb99f1-bcfd-41d2-bf6a-f13127b68e17 · outbound

This paper cites A Survey on Knowledge Distillation of Large Language Models.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment A Survey on Knowledge Distillation of Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.566348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:ed8eff40662d7fe5e9a37850cba92f2404017a54e71c1da109af5efe4a294574

Observation 835bd351-177b-4e26-abc7-1012ea600259 · outbound

This paper cites Here's a step-by-step approach to achieve this:1.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Here's a step-by-step approach to achieve this:1

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:55.133531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:314c21792ac2095b670ef4be3c9ecce294a0f8d1a73932f5b5d1e04fc0c2e319

Observation fb4308cd-a36e-4705-927c-258161822413 · outbound

This paper cites Initialize a Dictionary to Track Blocks: Use a dictionary to map the top-left corner of each 2x2 block to the count of black cells in that block.2.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Initialize a Dictionary to Track Blocks: Use a dictionary to map the top-left corner of each 2x2 block to the count of black cells in that block.2

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:55.136519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:c92604f75d36b4bffb763b84a0876744ce71bb00aea9e40157e564e80c52f685

Observation 815fb496-5bf3-4a86-b4e5-5b659f91ef37 · outbound

This paper cites an unresolved cited work.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Unresolved cited work

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:55.138529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:c2dd82aa6498714f6aa2fd856e9792742092f00082ec239aa2c16a9bdbeb0f7e

Pith citing papers

Observation d84a2454-7244-45eb-b109-e7f1f8a552cd · inbound

Think Anywhere in Code Generation cites this paper.

Think Anywhere in Code Generation CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:18:25.612924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-13T23:16:42.782431Z digest=sha256:5d2719d0632eb0ce08265c756e7e691028e7f890fb152583ef3a0d2f3e3a8206

Observation 6a76c5aa-6361-4f47-8d5e-46524980184e · inbound

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning cites this paper.

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:38:18.780983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-13T21:36:14.478007Z digest=sha256:9e2e04fb9d4a38110543fb03374f7bc7adef1d8d467e0a5f18e83e3f8fa0691f

Observation 830188b8-49f9-4487-9787-53a440b74aeb · inbound

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy cites this paper.

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:53:11.753782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-13T19:51:08.304741Z digest=sha256:2710530ce440018a3dba6c852d67655b6a654d164507e7de2675d5303b19739f

Observation b9f4dcbc-5880-46cb-8691-fba541979f9c · inbound

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models cites this paper.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:01:49.478225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:e45a45452b94c2f04fa366a81387d821f665fad1b341525ca538c58e77934f70

Observation dfa64b51-5155-4abd-b316-84f35d5c18d6 · inbound

Schedule-and-Calibrate: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs cites this paper.

Schedule-and-Calibrate: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:26:10.469171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-08T09:11:20.733960Z digest=sha256:c38f7a8a193ebb29b04883858eefaf5af7b06260c3a450a818caf3bbdfb7d766

Observation bb4b95a2-b255-4026-908e-42ee84ae03c4 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:19:43.108008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:575839f0a4137aba5c7df944f998c7841b31b3e4aa95b2d47c5d988283e90059

Observation 7c6a563f-8d8c-4ba8-ad5f-9cff54d5c9dc · inbound

Code as Agent Harness cites this paper.

Code as Agent Harness CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:58:14.219194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-20T10:54:54.558241Z digest=sha256:720e56af7f6dfa12463ccf839c41deab1e1c2fa0377b36b7e03305d7340ec346

Observation a33f4fe7-b7ed-4d90-9ab7-71c8191ca18e · inbound

Distilling Game Code World Model Generation into Lightweight Large Language Models cites this paper.

Distilling Game Code World Model Generation into Lightweight Large Language Models CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T14:04:44.660858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-06-30T13:58:37.956333Z digest=sha256:5f4465654059c06f4c258b56db0126077a65303afc1f667ace9e2179d3bba3ee

Observation 737d4a5e-111d-4cbb-b10a-0c5d1446197f · inbound

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback cites this paper.

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T14:53:32.279726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-06-29T06:11:06.960389Z digest=sha256:78ec223ed2989eadff75794c0c3ae9b4e505477aa7c0e40d28ef650479421aec

Observation 45b56bcc-7699-4426-ac98-e248ca395c1b · inbound

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents cites this paper.

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:16:57.774024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=arxiv_source observed=2026-06-28T02:11:11.638029Z digest=sha256:1ad2acb1d742b8e56db0df0d9fa9cd957f68a7c0851fff95c07851389185a3be

Observation aa580c9a-1676-4f23-8d9d-52405be4a5f2 · inbound

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning cites this paper.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.933247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:4f14436a17acb60f36fa188102e960d6616bfa26713dfd977547034de34f9beb

Observation 54136184-f7a3-452a-b109-c60542fdb2ef · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-27T13:10:55.917848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:05bb5a31b17cd0681596fc1c80b04b157f67caeb994dd71309c25a3d1baf68f3

Observation 1fa12586-08a0-4488-af79-31c964d593d0 · inbound

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training cites this paper.

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T12:58:08.768615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-06-27T08:35:16.435272Z digest=sha256:b5eab4e9cec5c49ba45ab563a80ec4169d955e21ef4e9c681f2df17fffb0ec0b

Observation e2839e89-9d81-4a5b-9f33-b3947090af37 · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:00:19.806171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:657b498d4842155a3f960f1c471dc436506a0cd7ce0c07e389a96f8a5a6ad27e

Observation a9389d13-50cf-41b5-ad78-a794dfa3a7fa · inbound

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study cites this paper.

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T04:19:33.972007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-06-26T17:07:21.486960Z digest=sha256:343d83c365e45d6f67fd2d74a123499236db5b6f6e0139648ede6d9a68d009e8

Observation 401e2898-b4b6-464c-8117-91fcc19a8df3 · inbound

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment cites this paper.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:a23cb4e0479907d43a26527a6182f74a32a601d0baf30bd28b872294be8247a3