{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:DCDLMEBPG7YCYCHSVU3AGALMXG","short_pith_number":"pith:DCDLMEBP","schema_version":"1.0","canonical_sha256":"1886b6102f37f02c08f2ad3603016cb9a753d3acd909077ceb072dfad811c1a1","source":{"kind":"arxiv","id":"2412.20367","version":5},"attestation_state":"computed","paper":{"title":"Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.SE","authors_text":"Guangwu Qian, Hengyuan Xu, Junqiao Wang, Keqin Li, Kuan Lu, Kunyu Wu, Lewei He, Menghao Huo, Qiuwu Chen, Tang Jingqun, Tianyu Shi, Xinhang Yuan, Xin Yi, Xinyuan Song, Yangfan He, Yuchen Li, Yuyang Song, Zeng Zhang, Zhongwei Wan, Zihao Zhang, Zijun Wang","submitted_at":"2024-12-29T06:15:41Z","abstract_excerpt":"With the rapid evolution of large language models (LLM), reinforcement learning (RL) has emerged as a pivotal technique for code generation and optimization in various domains. This paper presents a systematic survey of the application of RL in code optimization and generation, highlighting its role in enhancing compiler optimization, resource allocation, and the development of frameworks and tools. Subsequent sections first delve into the intricate processes of compiler optimization, where RL algorithms are leveraged to improve efficiency and resource utilization. The discussion then progress"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2412.20367","kind":"arxiv","version":5},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.SE","submitted_at":"2024-12-29T06:15:41Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"45dbd0347b12c6c4d221f70ab7a1e7ec2de407d5d4802db7e0b54e71109527bd","abstract_canon_sha256":"cf79cfacd027adb85a2633983ef030747cb71dc8860aeeef80f1f60e5ed56c6a"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:49:49.505015Z","signature_b64":"GV3hwEHk/s0yn3UU/NlszzlNwIlZ8NHAHHZa5CnDorJ0hOpwECRuGTXEIT5XPFmif6aPAVwp+FyGYh9+nnYUCA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"1886b6102f37f02c08f2ad3603016cb9a753d3acd909077ceb072dfad811c1a1","last_reissued_at":"2026-07-05T11:49:49.504525Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:49:49.504525Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.SE","authors_text":"Guangwu Qian, Hengyuan Xu, Junqiao Wang, Keqin Li, Kuan Lu, Kunyu Wu, Lewei He, Menghao Huo, Qiuwu Chen, Tang Jingqun, Tianyu Shi, Xinhang Yuan, Xin Yi, Xinyuan Song, Yangfan He, Yuchen Li, Yuyang Song, Zeng Zhang, Zhongwei Wan, Zihao Zhang, Zijun Wang","submitted_at":"2024-12-29T06:15:41Z","abstract_excerpt":"With the rapid evolution of large language models (LLM), reinforcement learning (RL) has emerged as a pivotal technique for code generation and optimization in various domains. This paper presents a systematic survey of the application of RL in code optimization and generation, highlighting its role in enhancing compiler optimization, resource allocation, and the development of frameworks and tools. Subsequent sections first delve into the intricate processes of compiler optimization, where RL algorithms are leveraged to improve efficiency and resource utilization. The discussion then progress"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2412.20367","kind":"arxiv","version":5},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2412.20367/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2412.20367","created_at":"2026-07-05T11:49:49.504582+00:00"},{"alias_kind":"arxiv_version","alias_value":"2412.20367v5","created_at":"2026-07-05T11:49:49.504582+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2412.20367","created_at":"2026-07-05T11:49:49.504582+00:00"},{"alias_kind":"pith_short_12","alias_value":"DCDLMEBPG7YC","created_at":"2026-07-05T11:49:49.504582+00:00"},{"alias_kind":"pith_short_16","alias_value":"DCDLMEBPG7YCYCHS","created_at":"2026-07-05T11:49:49.504582+00:00"},{"alias_kind":"pith_short_8","alias_value":"DCDLMEBP","created_at":"2026-07-05T11:49:49.504582+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":12,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.05677","citing_title":"From Conversation to Contribution: Characterizing Coding Agent in Open-Source Software","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2606.21943","citing_title":"Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning","ref_index":216,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28707","citing_title":"BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards","ref_index":104,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27859","citing_title":"Rethinking Agentic Reinforcement Learning In Large Language Models","ref_index":90,"is_internal_anchor":false},{"citing_arxiv_id":"2509.02547","citing_title":"The Landscape of Agentic Reinforcement Learning for LLMs: A Survey","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2512.14018","citing_title":"PerfCoder: Large Language Models for Interpretable Code Performance Optimization","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2602.16548","citing_title":"RIDER: 3D RNA Inverse Design with Reinforcement Learning-Guided Diffusion","ref_index":51,"is_internal_anchor":false},{"citing_arxiv_id":"2603.19880","citing_title":"What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27859","citing_title":"Rethinking Agentic Reinforcement Learning In Large Language Models","ref_index":90,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27859","citing_title":"Rethinking Agentic Reinforcement Learning In Large Language Models","ref_index":90,"is_internal_anchor":false},{"citing_arxiv_id":"2604.09813","citing_title":"Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18027","citing_title":"CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora","ref_index":73,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/DCDLMEBPG7YCYCHSVU3AGALMXG","json":"https://pith.science/pith/DCDLMEBPG7YCYCHSVU3AGALMXG.json","graph_json":"https://pith.science/api/pith-number/DCDLMEBPG7YCYCHSVU3AGALMXG/graph.json","events_json":"https://pith.science/api/pith-number/DCDLMEBPG7YCYCHSVU3AGALMXG/events.json","paper":"https://pith.science/paper/DCDLMEBP"},"agent_actions":{"view_html":"https://pith.science/pith/DCDLMEBPG7YCYCHSVU3AGALMXG","download_json":"https://pith.science/pith/DCDLMEBPG7YCYCHSVU3AGALMXG.json","view_paper":"https://pith.science/paper/DCDLMEBP","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2412.20367&json=true","fetch_graph":"https://pith.science/api/pith-number/DCDLMEBPG7YCYCHSVU3AGALMXG/graph.json","fetch_events":"https://pith.science/api/pith-number/DCDLMEBPG7YCYCHSVU3AGALMXG/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/DCDLMEBPG7YCYCHSVU3AGALMXG/action/timestamp_anchor","attest_storage":"https://pith.science/pith/DCDLMEBPG7YCYCHSVU3AGALMXG/action/storage_attestation","attest_author":"https://pith.science/pith/DCDLMEBPG7YCYCHSVU3AGALMXG/action/author_attestation","sign_citation":"https://pith.science/pith/DCDLMEBPG7YCYCHSVU3AGALMXG/action/citation_signature","submit_replication":"https://pith.science/pith/DCDLMEBPG7YCYCHSVU3AGALMXG/action/replication_record"}},"created_at":"2026-07-05T11:49:49.504582+00:00","updated_at":"2026-07-05T11:49:49.504582+00:00"}