Pith. sign in

Paper Citation Record · LEDGER

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training

As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2608.07147.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07147 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:00:53.226387Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy29
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96dd344a-c9a8-4a6c-8fe7-66aa61c3c76e · outbound

This paper cites Program Synthesis with Large Language Models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.706852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.706852Z digest=sha256:dd5cb6ac73491ead761f24beacc387ee281721b519133e8c150a5f9c93c607ba

Observation b30ba429-aa4e-4f0c-a4a1-b7431964c91f · outbound

This paper cites A general theo- retical paradigm to understand learning from hu- man preferences.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A general theo- retical paradigm to understand learning from hu- man preferences

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.891649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.715509Z digest=sha256:0221849fc19e6375f242b71c59679f3165f19a70f4fab93987635473c49eae83

Observation c0caab54-2410-4fa9-88d2-a451dd47fc33 · outbound

This paper cites A course in metric geometry, volume 33.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A course in metric geometry, volume 33

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.868084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.722903Z digest=sha256:488e0ea953d37ba617f30fe2c37adde32f3a733f87cf32ba58773eebb319aa76

Observation eafb808e-3f7e-4071-88ff-7ec9de1adfa6 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.729442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.729442Z digest=sha256:d889af5ca802090392cfaea396c11769e588fa9eb892d45fc12980cc01eb8234

Observation a2414384-1702-4a07-87fb-580d91a55674 · outbound

This paper cites CodeT: Code Generation with Generated Tests.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training CodeT: Code Generation with Generated Tests

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.742000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.742000Z digest=sha256:49bd8e8d62982ad9d69abd924c0a67ca406a43dbb427ddc70f7d1f4764d681d6

Observation 604af2c4-39c3-4a80-8fb9-b61d552f1bdd · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.751101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.751101Z digest=sha256:9939e8853dbed09aeb5eda375eb3810ec2783e12394152857df364cdffe28245

Observation ee1b012a-c8f5-48a4-aed1-cd8f86d8ec9e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.758668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.758668Z digest=sha256:e49119905741d02e5f76a869498f1096e89c73e27b7668e307834a56d8dc4ad4

Observation ded73128-d4c1-4245-9883-16835dc4b281 · outbound

This paper cites Teaching large language models to self-debug.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Teaching large language models to self-debug

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.849027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.766615Z digest=sha256:8212d8f3218b6f74f745937c1635bebb79d004d5c2608220bdd781546406daa3

Observation 3e686dfd-e70d-4e72-aaf0-8c0b698f8046 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.774989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.774989Z digest=sha256:4e2c4a452471ee9b489201f975a6c0f10391d1f5cb4fde6ee59ace75a483ff66

Observation c318a7f3-178f-4734-a570-a17178156205 · outbound

This paper cites Deep reinforcement learning from human preferences.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Deep reinforcement learning from human preferences

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.824505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.780959Z digest=sha256:5195ee2222e7c484e2990d3ec1d856bd805d638f8a005cecf99d45ba9b4a779c

Observation a03c4c37-e90d-40b1-8b97-6a433e41c913 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.788428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.788428Z digest=sha256:853917bc843775f1a1c6701bca6a06177d9033388af6e544ff6c1ed768b9ce4e

Observation 52f58160-bd33-4651-bd7c-c33a04f1ecc4 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Process Reinforcement through Implicit Rewards

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.797382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.797382Z digest=sha256:8bb9743b2d94e3d8acf5c3fd772b2fa6f4fc9a9548696dcd0f950927495ccd19

Observation c0e354f9-bf1b-46bf-aa42-b7357d9270d6 · outbound

This paper cites Memp: Exploring agent procedural memory.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Memp: Exploring agent procedural memory

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.802083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.803557Z digest=sha256:bceded3e2321aa20f76bf6c81f7598c7cf6e870742d452bc91d0113aec03f85f

Observation bb008a4b-e7ba-45fb-beae-70e58aea9c77 · outbound

This paper cites Group-in-group policy optimization for llm agent training.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Group-in-group policy optimization for llm agent training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.781914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.810978Z digest=sha256:330e2522ba1bf5e4ac1e8573a39c7d6e0c42d3ff5f34485403ce2aa0f4d95c10

Observation 8a10f563-8b5a-423c-bcb6-b2ca68db2106 · outbound

This paper cites InCoder: A Generative Model for Code Infilling and Synthesis.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training InCoder: A Generative Model for Code Infilling and Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.821269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.821269Z digest=sha256:753db523dea34fcfd78742df819bf99671991cefbda940b0cfef26e2a935418c

Observation 86773d89-d8d2-452f-b7bf-0ed4afeaf70c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.827996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.827996Z digest=sha256:680163aea81ec86047f40272d9acf75afe03c0608025dee12f5fae9bd0cc11cd

Observation 3b2aedc4-624e-4d1f-be0d-6ab78e6847fc · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Skywork Open Reasoner 1 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.837073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.837073Z digest=sha256:27588c5910f9e81bc0f6a6401f11097862fffc2983780bf426d4a2a6a4ab62bf

Observation 97a60960-3f80-4c5f-bfe0-ab41e5cb9842 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Measuring Coding Challenge Competence With APPS

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.844229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.844229Z digest=sha256:e1eca57d6d08474fe32d0992d0b2cf4a19eabf11c3773d97763f7b4a6b0bf0ef

Observation 34727b05-4c6b-43c7-a075-5848c0928136 · outbound

This paper cites Cogagent: A visual language model for gui agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Cogagent: A visual language model for gui agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.852500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.852500Z digest=sha256:8d460e346794e27bb492d7ee44b12e139541cedca237420c453b37e7d860fef5

Observation 715dce2b-b515-4371-ad60-dc1057a9ebc3 · outbound

This paper cites Qwen2.5-Coder Technical Report.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Qwen2.5-Coder Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.859229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.859229Z digest=sha256:476f9352b702beebb874ed0f8986ac7fd3c6841cbd240ad18e8bc6651d9c2387

Observation 07eab6a8-4ab8-441a-8f55-1c5f7d143ec6 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.868170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.868170Z digest=sha256:6798f54608c8a893f618989a84b8af24548662a2e01a07f88a66d8933f4c410d

Observation 72af3f08-9c54-4441-821f-77df178fe084 · outbound

This paper cites Cure: Code-aware neural machine translation for auto- matic program repair.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Cure: Code-aware neural machine translation for auto- matic program repair

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.745463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.874648Z digest=sha256:c6d74a33514ada4189669be7c7885fcc0bda8c474a4d85536e07c88ebb41a4fc

Observation 5d23cf88-c9c3-4622-8cf1-58faf4a36a0c · outbound

This paper cites Self-planning code generation with large language models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Self-planning code generation with large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.719715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.887948Z digest=sha256:9a847892bcdeb555a5e8af2564b09708ef00aecd4c894c09662729f30fe9d409

Observation be636341-f7aa-4b45-a333-a47237e80770 · outbound

This paper cites CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.897000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.897000Z digest=sha256:a34cfa21241018c8434fb927c72a6ea4ca83272f75388b283352fe013dcfd1d6

Observation c30cfec5-0f4d-4ac4-a2e8-f83ee8545226 · outbound

This paper cites Swe-bench: Can language models re- solve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-bench: Can language models re- solve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.697644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.905481Z digest=sha256:c379dfc05e2b0b14354345a0944fcaae7ad16fcb3e620f5b7ed08bea0b2da8f2

Observation 1c03404b-347f-446e-9ace-e665d587bd5e · outbound

This paper cites Defects4j: A database of existing faults to enable controlled testing studies for java programs.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Defects4j: A database of existing faults to enable controlled testing studies for java programs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.677349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.917446Z digest=sha256:31f5931719d87c58fc863ac9e7e64b7ef15f54c43841b33dd306a648a182e512

Observation 39751694-7240-46a2-aa05-80913ae7ff9c · outbound

This paper cites Ds- 1000: A natural and reliable benchmark for data science code generation.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ds- 1000: A natural and reliable benchmark for data science code generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.655863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.924876Z digest=sha256:679d1a641a2787f16318e53c1e077d6b58b9a7beef050ca2aefd486828041689

Observation c4df06f0-ca4c-4fbc-a9a2-8dc4f6508833 · outbound

This paper cites Coderl: Mastering code generation through pretrained mod- els and deep reinforcement learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Coderl: Mastering code generation through pretrained mod- els and deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.635931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.935617Z digest=sha256:58058a5922c76a4095de49a3fde26cde5773f76af804b9eda18bed0031e7e160

Observation b7090b3d-68ee-4154-8aa9-abadb60c460f · outbound

This paper cites A systematic study of automated program repair: Fixing 55 out of 105 bugs for $8 each.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A systematic study of automated program repair: Fixing 55 out of 105 bugs for $8 each

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.613194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.945542Z digest=sha256:b289fdb9d0c616665b277ffea47a26158428f48c0ebb4f48f0b6261e424af2b5

Observation a4423c6f-381a-4223-ac9f-5611ff0c601e · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.951785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.951785Z digest=sha256:b37996fef8f54e809a7ce8accc342c3ef716150c569fa96d2b9251e29d4b0ed7

Observation e42778f4-00d3-49f6-8d85-2a101745a892 · outbound

This paper cites StarCoder: may the source be with you!.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training StarCoder: may the source be with you!

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.959772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.959772Z digest=sha256:724a660dea44d398f0a2630d4fcdfeb39d66f5b0e4ed8c752542d589bf91961a

Observation a2d316f3-4a1e-459c-a37a-b2ad8723f5ba · outbound

This paper cites Let’s verify step by step.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Let’s verify step by step

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.592579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.968533Z digest=sha256:7c694ec9345ba656f9205c539ef6cfd2217e95b99106ee3cb28b3c85886f098d

Observation d98740db-f9be-47a6-a44b-9c8a68c49c18 · outbound

This paper cites Automatic patch gen- eration by learning correct code.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Automatic patch gen- eration by learning correct code

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.573032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.976470Z digest=sha256:904dce834175f4ed2fce62ec93c6bb660e330647da700ed00445a833d4ef38e4

Observation 38cc9b54-22aa-4d9e-b27b-be5d3f894c4a · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.982894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.982894Z digest=sha256:144a185a82b2439f3e71386bddaa70731d72b7decfcbe375018f82b055185fd6

Observation d9e86fcb-867e-407f-ac33-e219f224351b · outbound

This paper cites Gromov–wasserstein distances and the metric approach to object matching.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Gromov–wasserstein distances and the metric approach to object matching

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.553305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:52.993757Z digest=sha256:561b2a5eaa9d93cdb5c72f70936959b2170832be8e853d84c62250a16d7b89f2

Observation ffec4f66-7adb-44dd-abd4-e37dd26e4222 · outbound

This paper cites An analysis of approxima- tions for maximizing submodular set functions—i.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training An analysis of approxima- tions for maximizing submodular set functions—i

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.529276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.000836Z digest=sha256:5a6223d45e06c6416a8b1e693698e70a27e60b37d2a1d49b04e6106e790bba35

Observation c8e1f20b-cf6b-4a58-b9b5-88a3e13e7256 · outbound

This paper cites Training language models to follow instructions with human feedback.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Training language models to follow instructions with human feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.010831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.010831Z digest=sha256:648e74e2adec11395e29a88efd5c51a8bc83f1213abf38b655c600b076ceabb1

Observation 59b2863d-88c9-43e5-8ba9-a930c36d89e1 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.018708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.018708Z digest=sha256:c1070a6d5a245674e6a381ac70695ca502fb8fe7627271ab085875404f15e58f

Observation 6e2f8ecb-b654-4df3-9321-66e6bf130d60 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.028899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.028899Z digest=sha256:b6c7d84d5a27bae350e310e840776896a03bab327fe2ce1f76d9a1dd36a9380e

Observation 01e04f36-e2b7-454d-9f54-8ab530e9e8b7 · outbound

This paper cites Can Language Models Solve Olympiad Programming?.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Can Language Models Solve Olympiad Programming?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.037166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.037166Z digest=sha256:badedf43e3337cd888f8e9bae27f6c01919c2f2789b5f051a55fcfb59740e43b

Observation 34648ae4-8a93-47f0-aee9-adfb3921f37d · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Reflexion: Language agents with verbal reinforcement learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.045807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.045807Z digest=sha256:0966f54392c3c81c2d0a0acd115c7063a14dc6b1cabb9c6f198439d28fd7bdc3

Observation b4ee0b14-3de4-4172-b51c-9c847a547bbc · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.055761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.055761Z digest=sha256:785e6b7e825bfeca4a45e70c1a6ae1ed6564d87ed83f99e2f6ccfdbdcc7ed786

Observation 221ec3e4-32fb-4fef-86d9-8561cf0a0d4e · outbound

This paper cites Ex- ecverify: White-box rl with verifiable stepwise re- wards for code execution reasoning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ex- ecverify: White-box rl with verifiable stepwise re- wards for code execution reasoning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.494879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.064958Z digest=sha256:e747dd8c37439e0cebc49067f3d8c65078f33e635f2934eea984699209f25a84

Observation 47c4081e-8bdc-4fff-bf26-1a51e6c534a5 · outbound

This paper cites Ex- ecutable code actions elicit better llm agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ex- ecutable code actions elicit better llm agents

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.475860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.072789Z digest=sha256:31ad4112bd512a4cb1c091cb94f7d2e13058124447d9efd5d0eb2efe1930de63

Observation 0901c2c9-f0e5-41af-add5-75d7ad709401 · outbound

This paper cites Open- hands: An open platform for ai software developers as generalist agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Open- hands: An open platform for ai software developers as generalist agents

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.451400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.083659Z digest=sha256:992b610abf406bb1b3d2beec7b25f26f2dc3d8089ded0aef6c97d150992fe4b9

Observation 588315d6-711c-47f0-a712-50ff196738c3 · outbound

This paper cites Lightweight Self-Knowledge Distillation with Multi-source Information Fusion.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Lightweight Self-Knowledge Distillation with Multi-source Information Fusion

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:00:53.632208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.093830Z digest=sha256:5e6f552f3a7b0573ed031183ad5a205e8adfbd57357fed25e9c95e76e5bc7fb4

Observation e38ea274-93e9-4d8f-9c39-8e34500c770f · outbound

This paper cites Multi-label self knowledge distillation.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Multi-label self knowledge distillation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.429162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.103235Z digest=sha256:376a26f0f18b4a5387c2e094481d3d67d195861d5654ee2ae92cb9fc767e23a5

Observation 8e70c78f-71f1-41ae-9751-2a01ef3524bd · outbound

This paper cites Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:00:53.586243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.117008Z digest=sha256:e310b43ac2e4adc937de97af981669dcb3d30967bbc5f1e8db1034c8eb562d5a

Observation 4d604580-d058-465c-a1bb-13163f739bdd · outbound

This paper cites Ojbench: A competi- tion level code benchmark for large language models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ojbench: A competi- tion level code benchmark for large language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.127123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.127123Z digest=sha256:5725fadeb6980aefef3c2545d535281a4d94284df447004403fd6718e7e431b7

Observation b43e4fac-d08d-4843-b78e-462087c6991e · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.133261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.133261Z digest=sha256:7775a4059e5aa97469d36ad0d6bb68c63f3d3361888e43bf258b96dd1aa51028

Observation ca742a1c-a1c6-4dee-81a6-3d791f4b3904 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824– 24837, 2022.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824– 24837, 2022

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.408999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.141550Z digest=sha256:1292f9f699e773d87ad0d26427ffaa5e1e32f7c6c5fb8a31a32e442cadcf6b86

Observation 3bbe2263-4e0c-4dc1-8f32-39249168a903 · outbound

This paper cites Swe-rl: Advancing llm reasoning via reinforce- ment learning on open software evolution.Advances inNeuralInformationProcessing Systems, 38:78500– 78525, 2026.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-rl: Advancing llm reasoning via reinforce- ment learning on open software evolution.Advances inNeuralInformationProcessing Systems, 38:78500– 78525, 2026

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.387694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.148353Z digest=sha256:28943048eede0a552661d72e569db75710b91406dab8cbda25ab807939116ead

Observation b5487387-c251-4b8f-a5bf-fcec858ac467 · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Agentless: Demystifying LLM-based Software Engineering Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.155623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.155623Z digest=sha256:7c50e6982f5092834252e87bb369992ee9b1d4d1a2040e817551c3b4264da04f

Observation 3508a87c-76e5-4c2d-8369-6cd95c41bad6 · outbound

This paper cites LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.161982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.161982Z digest=sha256:8e0ad1528f94cc045de2f954812c2ee1b7523b2a3d16579729cae6698d15456f

Observation 2f320710-abfb-4221-abeb-35ea8925fee4 · outbound

This paper cites Icpc-eval: Probing the frontiers of llm reasoning with competitive pro- gramming contests.Advancesin Neural Information Processing Systems, 38, 2026.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Icpc-eval: Probing the frontiers of llm reasoning with competitive pro- gramming contests.Advancesin Neural Information Processing Systems, 38, 2026

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.363007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.169064Z digest=sha256:53b7f483a4b33080e0629d8500b7a4c815dcc09104b9b2f76fb2b39da2c55e05

Observation fac97e41-cf4b-4f13-a262-43bc38f2b12b · outbound

This paper cites Swe-agent: Agent-computer inter- faces enable automated software engineering.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-agent: Agent-computer inter- faces enable automated software engineering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.335861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.177569Z digest=sha256:02d2cc1728ec7aa1eab420dd7ff518573a737452240dafa72af3f3937015da19

Observation ea1ef2f3-8f9b-46eb-8f3d-373ca0326653 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.314717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.184920Z digest=sha256:b5da94ac139107e43917fa3c12c9d5e03f304fdcb956f26b3806320cd6a54db3

Observation cfdc3d24-d520-4968-9b00-34595f0ed208 · outbound

This paper cites React: Syn- ergizing reasoning and acting in language models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training React: Syn- ergizing reasoning and acting in language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.290530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.193005Z digest=sha256:65738d81d23c730d8a08e8ec48cde6e0c31fb0f60d2ec5e2a493493068372114

Observation 50916849-3599-4b50-beee-4b8a7bb28f7a · outbound

This paper cites Dapo: An open- source llm reinforcement learning system at scale.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Dapo: An open- source llm reinforcement learning system at scale

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.272093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.199443Z digest=sha256:83f8f2c1c961c7b9f3e2acc8cb6f471a7ce0e00882d974f298561ec516a8048e

Observation 1aeb31ca-10f6-442b-b360-d62a0cecf979 · outbound

This paper cites Group Sequence Policy Optimization.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Group Sequence Policy Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.210565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.210565Z digest=sha256:09741876c3c6c21fd6c5849530f99dd2127f93c5ee369c914148f5f9fcec057b

Observation 1f4b7440-a570-47d8-b43f-e20b6eeeca5a · outbound

This paper cites A syntax-guided edit decoder for neural pro- gram repair.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A syntax-guided edit decoder for neural pro- gram repair

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.249562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:00:53.218339Z digest=sha256:689367cd3c19b0f45763cb5d980338adbcc7657472ed278dcb11faf247110ae0

Observation 5d6bbe64-8f09-4f69-80c5-bd33ead92af8 · outbound

This paper cites GAGPO: Generalized Advantage Grouped Policy Optimization.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training GAGPO: Generalized Advantage Grouped Policy Optimization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.226387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.226387Z digest=sha256:76104e2907a5b0395ed67a76d68790d102b31ddaf4ca98d4013514da17ca19ed

Pith citing papers

No inbound Pith citation observations are available.