{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:NRK4ZO5D7HG4UBBR2CPTGSRBUG","short_pith_number":"pith:NRK4ZO5D","schema_version":"1.0","canonical_sha256":"6c55ccbba3f9cdca0431d09f334a21a1bf6882fbbcbb4175d94dd2485d34f8ac","source":{"kind":"arxiv","id":"2402.01391","version":2},"attestation_state":"computed","paper":{"title":"StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.SE","authors_text":"Caishuang Huang, Enyu Zhou, Haoxiang Jia, Junjie Shan, Limao Xiong, Qi Zhang, Rui Zheng, Shihan Dou, Tao Gui, Tao Ji, Wei Shen, Xiaoran Fan, Xiao Wang, Xuanjing Huang, Yan Liu, Yuhao Zhou, Zhiheng Xi","submitted_at":"2024-02-02T13:14:31Z","abstract_excerpt":"The advancement of large language models (LLMs) has significantly propelled the field of code generation. Previous work integrated reinforcement learning (RL) with compiler feedback for exploring the output space of LLMs to enhance code generation quality. However, the lengthy code generated by LLMs in response to complex human requirements makes RL exploration a challenge. Also, since the unit tests may not cover the complicated code, optimizing LLMs by using these unexecuted code snippets is ineffective. To tackle these challenges, we introduce StepCoder, a novel RL framework for code genera"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2402.01391","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.SE","submitted_at":"2024-02-02T13:14:31Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"d7ac05414d54ea4dea5bb4888e488ffc74fcbd02848ec9a0beff37a69ea5150f","abstract_canon_sha256":"8dcd4482672a89e54c21b776991cdca14ee4c7a044e9c12b35949da38728c640"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:41:04.179424Z","signature_b64":"3imAUqjNP2AZmI8ElaP885nW9RgdUp3x/YJEcEew70GGm9pfAk0KVZLqnV5W/FqRYA0v0XDvbFBkV/nmez3xCw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"6c55ccbba3f9cdca0431d09f334a21a1bf6882fbbcbb4175d94dd2485d34f8ac","last_reissued_at":"2026-07-05T07:41:04.178919Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:41:04.178919Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.SE","authors_text":"Caishuang Huang, Enyu Zhou, Haoxiang Jia, Junjie Shan, Limao Xiong, Qi Zhang, Rui Zheng, Shihan Dou, Tao Gui, Tao Ji, Wei Shen, Xiaoran Fan, Xiao Wang, Xuanjing Huang, Yan Liu, Yuhao Zhou, Zhiheng Xi","submitted_at":"2024-02-02T13:14:31Z","abstract_excerpt":"The advancement of large language models (LLMs) has significantly propelled the field of code generation. Previous work integrated reinforcement learning (RL) with compiler feedback for exploring the output space of LLMs to enhance code generation quality. However, the lengthy code generated by LLMs in response to complex human requirements makes RL exploration a challenge. Also, since the unit tests may not cover the complicated code, optimizing LLMs by using these unexecuted code snippets is ineffective. To tackle these challenges, we introduce StepCoder, a novel RL framework for code genera"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2402.01391","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.01391/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.01391","created_at":"2026-07-05T07:41:04.178975+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.01391v2","created_at":"2026-07-05T07:41:04.178975+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.01391","created_at":"2026-07-05T07:41:04.178975+00:00"},{"alias_kind":"pith_short_12","alias_value":"NRK4ZO5D7HG4","created_at":"2026-07-05T07:41:04.178975+00:00"},{"alias_kind":"pith_short_16","alias_value":"NRK4ZO5D7HG4UBBR","created_at":"2026-07-05T07:41:04.178975+00:00"},{"alias_kind":"pith_short_8","alias_value":"NRK4ZO5D","created_at":"2026-07-05T07:41:04.178975+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":19,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.20881","citing_title":"When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2606.20641","citing_title":"MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.24375","citing_title":"Distilling Game Code World Model Generation into Lightweight Large Language Models","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30478","citing_title":"Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2605.28409","citing_title":"Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2606.11052","citing_title":"Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2606.17682","citing_title":"From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2408.15815","citing_title":"MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic Testing","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17174","citing_title":"Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2507.21990","citing_title":"ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2508.16771","citing_title":"EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2406.00515","citing_title":"A Survey on Large Language Models for Code Generation","ref_index":71,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27308","citing_title":"BoostLoRA: Growing Effective Rank by Boosting Adapters","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2503.09567","citing_title":"Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models","ref_index":164,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00433","citing_title":"Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2604.11297","citing_title":"The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13934","citing_title":"Towards Enabling An Artificial Self-Construction Software Life-cycle via Autopoietic Architectures","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2604.16804","citing_title":"AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2604.20398","citing_title":"WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning","ref_index":9,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/NRK4ZO5D7HG4UBBR2CPTGSRBUG","json":"https://pith.science/pith/NRK4ZO5D7HG4UBBR2CPTGSRBUG.json","graph_json":"https://pith.science/api/pith-number/NRK4ZO5D7HG4UBBR2CPTGSRBUG/graph.json","events_json":"https://pith.science/api/pith-number/NRK4ZO5D7HG4UBBR2CPTGSRBUG/events.json","paper":"https://pith.science/paper/NRK4ZO5D"},"agent_actions":{"view_html":"https://pith.science/pith/NRK4ZO5D7HG4UBBR2CPTGSRBUG","download_json":"https://pith.science/pith/NRK4ZO5D7HG4UBBR2CPTGSRBUG.json","view_paper":"https://pith.science/paper/NRK4ZO5D","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.01391&json=true","fetch_graph":"https://pith.science/api/pith-number/NRK4ZO5D7HG4UBBR2CPTGSRBUG/graph.json","fetch_events":"https://pith.science/api/pith-number/NRK4ZO5D7HG4UBBR2CPTGSRBUG/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/NRK4ZO5D7HG4UBBR2CPTGSRBUG/action/timestamp_anchor","attest_storage":"https://pith.science/pith/NRK4ZO5D7HG4UBBR2CPTGSRBUG/action/storage_attestation","attest_author":"https://pith.science/pith/NRK4ZO5D7HG4UBBR2CPTGSRBUG/action/author_attestation","sign_citation":"https://pith.science/pith/NRK4ZO5D7HG4UBBR2CPTGSRBUG/action/citation_signature","submit_replication":"https://pith.science/pith/NRK4ZO5D7HG4UBBR2CPTGSRBUG/action/replication_record"}},"created_at":"2026-07-05T07:41:04.178975+00:00","updated_at":"2026-07-05T07:41:04.178975+00:00"}