Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:00:53.226387Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2608.07147.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:00:53.226387Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 96dd344a-c9a8-4a6c-8fe7-66aa61c3c76e · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Program Synthesis with Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b30ba429-aa4e-4f0c-a4a1-b7431964c91f · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A general theo- retical paradigm to understand learning from hu- man preferences
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c0caab54-2410-4fa9-88d2-a451dd47fc33 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A course in metric geometry, volume 33
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eafb808e-3f7e-4071-88ff-7ec9de1adfa6 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2414384-1702-4a07-87fb-580d91a55674 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training CodeT: Code Generation with Generated Tests
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604af2c4-39c3-4a80-8fb9-b61d552f1bdd · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee1b012a-c8f5-48a4-aed1-cd8f86d8ec9e · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Evaluating Large Language Models Trained on Code
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ded73128-d4c1-4245-9883-16835dc4b281 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Teaching large language models to self-debug
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3e686dfd-e70d-4e72-aaf0-8c0b698f8046 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c318a7f3-178f-4734-a570-a17178156205 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Deep reinforcement learning from human preferences
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a03c4c37-e90d-40b1-8b97-6a433e41c913 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Training Verifiers to Solve Math Word Problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52f58160-bd33-4651-bd7c-c33a04f1ecc4 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Process Reinforcement through Implicit Rewards
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0e354f9-bf1b-46bf-aa42-b7357d9270d6 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Memp: Exploring agent procedural memory
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bb008a4b-e7ba-45fb-beae-70e58aea9c77 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Group-in-group policy optimization for llm agent training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8a10f563-8b5a-423c-bcb6-b2ca68db2106 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training InCoder: A Generative Model for Code Infilling and Synthesis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86773d89-d8d2-452f-b7bf-0ed4afeaf70c · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b2aedc4-624e-4d1f-be0d-6ab78e6847fc · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Skywork Open Reasoner 1 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a60960-3f80-4c5f-bfe0-ab41e5cb9842 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Measuring Coding Challenge Competence With APPS
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34727b05-4c6b-43c7-a075-5848c0928136 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Cogagent: A visual language model for gui agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 715dce2b-b515-4371-ad60-dc1057a9ebc3 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Qwen2.5-Coder Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07eab6a8-4ab8-441a-8f55-1c5f7d143ec6 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72af3f08-9c54-4441-821f-77df178fe084 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Cure: Code-aware neural machine translation for auto- matic program repair
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5d23cf88-c9c3-4622-8cf1-58faf4a36a0c · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Self-planning code generation with large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation be636341-f7aa-4b45-a333-a47237e80770 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c30cfec5-0f4d-4ac4-a2e8-f83ee8545226 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-bench: Can language models re- solve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1c03404b-347f-446e-9ace-e665d587bd5e · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Defects4j: A database of existing faults to enable controlled testing studies for java programs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 39751694-7240-46a2-aa05-80913ae7ff9c · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ds- 1000: A natural and reliable benchmark for data science code generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c4df06f0-ca4c-4fbc-a9a2-8dc4f6508833 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Coderl: Mastering code generation through pretrained mod- els and deep reinforcement learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b7090b3d-68ee-4154-8aa9-abadb60c460f · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A systematic study of automated program repair: Fixing 55 out of 105 bugs for $8 each
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a4423c6f-381a-4223-ac9f-5611ff0c601e · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Meta-Harness: End-to-End Optimization of Model Harnesses
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e42778f4-00d3-49f6-8d85-2a101745a892 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training StarCoder: may the source be with you!
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2d316f3-4a1e-459c-a37a-b2ad8723f5ba · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Let’s verify step by step
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d98740db-f9be-47a6-a44b-9c8a68c49c18 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Automatic patch gen- eration by learning correct code
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 38cc9b54-22aa-4d9e-b27b-be5d3f894c4a · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9e86fcb-867e-407f-ac33-e219f224351b · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Gromov–wasserstein distances and the metric approach to object matching
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ffec4f66-7adb-44dd-abd4-e37dd26e4222 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training An analysis of approxima- tions for maximizing submodular set functions—i
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c8e1f20b-cf6b-4a58-b9b5-88a3e13e7256 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Training language models to follow instructions with human feedback
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59b2863d-88c9-43e5-8ba9-a930c36d89e1 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e2f8ecb-b654-4df3-9321-66e6bf130d60 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01e04f36-e2b7-454d-9f54-8ab530e9e8b7 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Can Language Models Solve Olympiad Programming?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34648ae4-8a93-47f0-aee9-adfb3921f37d · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Reflexion: Language agents with verbal reinforcement learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4ee0b14-3de4-4172-b51c-9c847a547bbc · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 221ec3e4-32fb-4fef-86d9-8561cf0a0d4e · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ex- ecverify: White-box rl with verifiable stepwise re- wards for code execution reasoning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 47c4081e-8bdc-4fff-bf26-1a51e6c534a5 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ex- ecutable code actions elicit better llm agents
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0901c2c9-f0e5-41af-add5-75d7ad709401 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Open- hands: An open platform for ai software developers as generalist agents
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 588315d6-711c-47f0-a712-50ff196738c3 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Lightweight Self-Knowledge Distillation with Multi-source Information Fusion
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e38ea274-93e9-4d8f-9c39-8e34500c770f · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Multi-label self knowledge distillation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8e70c78f-71f1-41ae-9751-2a01ef3524bd · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4d604580-d058-465c-a1bb-13163f739bdd · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ojbench: A competi- tion level code benchmark for large language models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b43e4fac-d08d-4843-b78e-462087c6991e · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca742a1c-a1c6-4dee-81a6-3d791f4b3904 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824– 24837, 2022
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3bbe2263-4e0c-4dc1-8f32-39249168a903 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-rl: Advancing llm reasoning via reinforce- ment learning on open software evolution.Advances inNeuralInformationProcessing Systems, 38:78500– 78525, 2026
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b5487387-c251-4b8f-a5bf-fcec858ac467 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Agentless: Demystifying LLM-based Software Engineering Agents
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3508a87c-76e5-4c2d-8369-6cd95c41bad6 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f320710-abfb-4221-abeb-35ea8925fee4 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Icpc-eval: Probing the frontiers of llm reasoning with competitive pro- gramming contests.Advancesin Neural Information Processing Systems, 38, 2026
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fac97e41-cf4b-4f13-a262-43bc38f2b12b · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-agent: Agent-computer inter- faces enable automated software engineering
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ea1ef2f3-8f9b-46eb-8f3d-373ca0326653 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Webshop: Towards scalable real-world web interaction with grounded language agents
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cfdc3d24-d520-4968-9b00-34595f0ed208 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training React: Syn- ergizing reasoning and acting in language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 50916849-3599-4b50-beee-4b8a7bb28f7a · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Dapo: An open- source llm reinforcement learning system at scale
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1aeb31ca-10f6-442b-b360-d62a0cecf979 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Group Sequence Policy Optimization
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f4b7440-a570-47d8-b43f-e20b6eeeca5a · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A syntax-guided edit decoder for neural pro- gram repair
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5d6bbe64-8f09-4f69-80c5-bd33ead92af8 · outbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training GAGPO: Generalized Advantage Grouped Policy Optimization
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.