Pith. sign in

Paper Citation Record · LEDGER

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training

As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2608.07147.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07147 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:00:53.226387Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy29
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96dd344a-c9a8-4a6c-8fe7-66aa61c3c76e · outbound

This paper cites Program Synthesis with Large Language Models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.706852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.706852Z digest=sha256:dd5cb6ac73491ead761f24beacc387ee281721b519133e8c150a5f9c93c607ba

Observation b30ba429-aa4e-4f0c-a4a1-b7431964c91f · outbound

This paper cites A general theo- retical paradigm to understand learning from hu- man preferences.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A general theo- retical paradigm to understand learning from hu- man preferences

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.891649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.715509Z digest=sha256:46606f6d75bab83aa5380156d8dae9ffffeea534f6d1b1fa8bf2a7eb7f28581b

Observation c0caab54-2410-4fa9-88d2-a451dd47fc33 · outbound

This paper cites A course in metric geometry, volume 33.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A course in metric geometry, volume 33

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.868084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.722903Z digest=sha256:d056a79d8bc3c5c23eed8a9b0893c3fb7ebf12c5a3e4471f2ec1c11afe389d58

Observation eafb808e-3f7e-4071-88ff-7ec9de1adfa6 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.729442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.729442Z digest=sha256:d889af5ca802090392cfaea396c11769e588fa9eb892d45fc12980cc01eb8234

Observation a2414384-1702-4a07-87fb-580d91a55674 · outbound

This paper cites CodeT: Code Generation with Generated Tests.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training CodeT: Code Generation with Generated Tests

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.742000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.742000Z digest=sha256:49bd8e8d62982ad9d69abd924c0a67ca406a43dbb427ddc70f7d1f4764d681d6

Observation 604af2c4-39c3-4a80-8fb9-b61d552f1bdd · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.751101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.751101Z digest=sha256:9939e8853dbed09aeb5eda375eb3810ec2783e12394152857df364cdffe28245

Observation ee1b012a-c8f5-48a4-aed1-cd8f86d8ec9e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.758668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.758668Z digest=sha256:e49119905741d02e5f76a869498f1096e89c73e27b7668e307834a56d8dc4ad4

Observation ded73128-d4c1-4245-9883-16835dc4b281 · outbound

This paper cites Teaching large language models to self-debug.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Teaching large language models to self-debug

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.849027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.766615Z digest=sha256:082f27bccf304ea1728314360974394847ca952c8446eec087b131a03aa744ab

Observation 3e686dfd-e70d-4e72-aaf0-8c0b698f8046 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.774989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.774989Z digest=sha256:4e2c4a452471ee9b489201f975a6c0f10391d1f5cb4fde6ee59ace75a483ff66

Observation c318a7f3-178f-4734-a570-a17178156205 · outbound

This paper cites Deep reinforcement learning from human preferences.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Deep reinforcement learning from human preferences

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.824505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.780959Z digest=sha256:e887751c0846df65b26cbbfc35116f2479798b30f40f91511728d446640baa61

Observation a03c4c37-e90d-40b1-8b97-6a433e41c913 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.788428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.788428Z digest=sha256:853917bc843775f1a1c6701bca6a06177d9033388af6e544ff6c1ed768b9ce4e

Observation 52f58160-bd33-4651-bd7c-c33a04f1ecc4 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Process Reinforcement through Implicit Rewards

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.797382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.797382Z digest=sha256:8bb9743b2d94e3d8acf5c3fd772b2fa6f4fc9a9548696dcd0f950927495ccd19

Observation c0e354f9-bf1b-46bf-aa42-b7357d9270d6 · outbound

This paper cites Memp: Exploring agent procedural memory.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Memp: Exploring agent procedural memory

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.802083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.803557Z digest=sha256:1d4076d6a78b69d61f1d4ce778d8dca3f3410b8fac7892adc39121dad4df9e0b

Observation bb008a4b-e7ba-45fb-beae-70e58aea9c77 · outbound

This paper cites Group-in-group policy optimization for llm agent training.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Group-in-group policy optimization for llm agent training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.781914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.810978Z digest=sha256:dba2347d2bc6400ccd733536399ce92f6a4342631dac4de9d3e799145c61fbab

Observation 8a10f563-8b5a-423c-bcb6-b2ca68db2106 · outbound

This paper cites InCoder: A Generative Model for Code Infilling and Synthesis.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training InCoder: A Generative Model for Code Infilling and Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.821269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.821269Z digest=sha256:753db523dea34fcfd78742df819bf99671991cefbda940b0cfef26e2a935418c

Observation 86773d89-d8d2-452f-b7bf-0ed4afeaf70c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.827996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.827996Z digest=sha256:74772ad7a4513ca173bc5864883c9835981c3432fc41fd59feb512c447186ef1

Observation 3b2aedc4-624e-4d1f-be0d-6ab78e6847fc · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Skywork Open Reasoner 1 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.837073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.837073Z digest=sha256:27588c5910f9e81bc0f6a6401f11097862fffc2983780bf426d4a2a6a4ab62bf

Observation 97a60960-3f80-4c5f-bfe0-ab41e5cb9842 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Measuring Coding Challenge Competence With APPS

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.844229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.844229Z digest=sha256:e1eca57d6d08474fe32d0992d0b2cf4a19eabf11c3773d97763f7b4a6b0bf0ef

Observation 34727b05-4c6b-43c7-a075-5848c0928136 · outbound

This paper cites Cogagent: A visual language model for gui agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Cogagent: A visual language model for gui agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.852500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.852500Z digest=sha256:8d460e346794e27bb492d7ee44b12e139541cedca237420c453b37e7d860fef5

Observation 715dce2b-b515-4371-ad60-dc1057a9ebc3 · outbound

This paper cites Qwen2.5-Coder Technical Report.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Qwen2.5-Coder Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.859229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.859229Z digest=sha256:476f9352b702beebb874ed0f8986ac7fd3c6841cbd240ad18e8bc6651d9c2387

Observation 07eab6a8-4ab8-441a-8f55-1c5f7d143ec6 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.868170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.868170Z digest=sha256:6798f54608c8a893f618989a84b8af24548662a2e01a07f88a66d8933f4c410d

Observation 72af3f08-9c54-4441-821f-77df178fe084 · outbound

This paper cites Cure: Code-aware neural machine translation for auto- matic program repair.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Cure: Code-aware neural machine translation for auto- matic program repair

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.745463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.874648Z digest=sha256:cd3533d703d60b81c5fdf2a2fee3c7ded12befde1fd895e7841a5bb5e2296894

Observation 5d23cf88-c9c3-4622-8cf1-58faf4a36a0c · outbound

This paper cites Self-planning code generation with large language models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Self-planning code generation with large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.719715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.887948Z digest=sha256:26f46ec1fd31a9b1da2c37a135beb4785dd1d0d76bfd54eb7219e92bde4105be

Observation be636341-f7aa-4b45-a333-a47237e80770 · outbound

This paper cites CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.897000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.897000Z digest=sha256:a34cfa21241018c8434fb927c72a6ea4ca83272f75388b283352fe013dcfd1d6

Observation c30cfec5-0f4d-4ac4-a2e8-f83ee8545226 · outbound

This paper cites Swe-bench: Can language models re- solve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-bench: Can language models re- solve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.697644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.905481Z digest=sha256:50619355dd6eebcd1af84895125fb625189b1cf7eebc86aa4e20f8ad1608bda1

Observation 1c03404b-347f-446e-9ace-e665d587bd5e · outbound

This paper cites Defects4j: A database of existing faults to enable controlled testing studies for java programs.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Defects4j: A database of existing faults to enable controlled testing studies for java programs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.677349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.917446Z digest=sha256:dd7c844394d85bdfc8130a9f35c44e7c91e1ae246e2c77e66c653b286e574e32

Observation 39751694-7240-46a2-aa05-80913ae7ff9c · outbound

This paper cites Ds- 1000: A natural and reliable benchmark for data science code generation.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ds- 1000: A natural and reliable benchmark for data science code generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.655863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.924876Z digest=sha256:4fb852a46475d110e0c56236ac3230e872b65db83df80b7a93c7d2713bd18b6f

Observation c4df06f0-ca4c-4fbc-a9a2-8dc4f6508833 · outbound

This paper cites Coderl: Mastering code generation through pretrained mod- els and deep reinforcement learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Coderl: Mastering code generation through pretrained mod- els and deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.635931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.935617Z digest=sha256:f8948c122c994d8536737a170323a59b118458a3f0cb4872c9aecb2f9f6e788c

Observation b7090b3d-68ee-4154-8aa9-abadb60c460f · outbound

This paper cites A systematic study of automated program repair: Fixing 55 out of 105 bugs for $8 each.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A systematic study of automated program repair: Fixing 55 out of 105 bugs for $8 each

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.613194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.945542Z digest=sha256:2da86d9cbd00dde9d3ba0a0bc1c03a7251b3285c83ed8b66131d9c3bb9e17d53

Observation a4423c6f-381a-4223-ac9f-5611ff0c601e · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.951785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.951785Z digest=sha256:b37996fef8f54e809a7ce8accc342c3ef716150c569fa96d2b9251e29d4b0ed7

Observation e42778f4-00d3-49f6-8d85-2a101745a892 · outbound

This paper cites StarCoder: may the source be with you!.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training StarCoder: may the source be with you!

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.959772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.959772Z digest=sha256:724a660dea44d398f0a2630d4fcdfeb39d66f5b0e4ed8c752542d589bf91961a

Observation a2d316f3-4a1e-459c-a37a-b2ad8723f5ba · outbound

This paper cites Let’s verify step by step.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Let’s verify step by step

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.592579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.968533Z digest=sha256:143932b60aa80692540f37cffbbcb704f680f4eb8763be62b6c8eb3e4c54c097

Observation d98740db-f9be-47a6-a44b-9c8a68c49c18 · outbound

This paper cites Automatic patch gen- eration by learning correct code.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Automatic patch gen- eration by learning correct code

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.573032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.976470Z digest=sha256:4d0a61e6ac92efd126565cb446ea458f9e32242e03b2fe340bcedc89d4ede6aa

Observation 38cc9b54-22aa-4d9e-b27b-be5d3f894c4a · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.982894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.982894Z digest=sha256:144a185a82b2439f3e71386bddaa70731d72b7decfcbe375018f82b055185fd6

Observation d9e86fcb-867e-407f-ac33-e219f224351b · outbound

This paper cites Gromov–wasserstein distances and the metric approach to object matching.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Gromov–wasserstein distances and the metric approach to object matching

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.553305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.993757Z digest=sha256:5a966e8bdf721b81343a43c86f69d51c79dc4ad2241d18e5f1b3d7d1d92a1508

Observation ffec4f66-7adb-44dd-abd4-e37dd26e4222 · outbound

This paper cites An analysis of approxima- tions for maximizing submodular set functions—i.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training An analysis of approxima- tions for maximizing submodular set functions—i

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.529276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.000836Z digest=sha256:bb1eac4d9560d0546905f5e56596957eb8701a7ab64ceb8d883a24be3ec4072d

Observation c8e1f20b-cf6b-4a58-b9b5-88a3e13e7256 · outbound

This paper cites Training language models to follow instructions with human feedback.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Training language models to follow instructions with human feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.010831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.010831Z digest=sha256:648e74e2adec11395e29a88efd5c51a8bc83f1213abf38b655c600b076ceabb1

Observation 59b2863d-88c9-43e5-8ba9-a930c36d89e1 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.018708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.018708Z digest=sha256:c1070a6d5a245674e6a381ac70695ca502fb8fe7627271ab085875404f15e58f

Observation 6e2f8ecb-b654-4df3-9321-66e6bf130d60 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.028899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.028899Z digest=sha256:b6c7d84d5a27bae350e310e840776896a03bab327fe2ce1f76d9a1dd36a9380e

Observation 01e04f36-e2b7-454d-9f54-8ab530e9e8b7 · outbound

This paper cites Can Language Models Solve Olympiad Programming?.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Can Language Models Solve Olympiad Programming?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.037166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.037166Z digest=sha256:badedf43e3337cd888f8e9bae27f6c01919c2f2789b5f051a55fcfb59740e43b

Observation 34648ae4-8a93-47f0-aee9-adfb3921f37d · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Reflexion: Language agents with verbal reinforcement learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.045807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.045807Z digest=sha256:0966f54392c3c81c2d0a0acd115c7063a14dc6b1cabb9c6f198439d28fd7bdc3

Observation b4ee0b14-3de4-4172-b51c-9c847a547bbc · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.055761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.055761Z digest=sha256:785e6b7e825bfeca4a45e70c1a6ae1ed6564d87ed83f99e2f6ccfdbdcc7ed786

Observation 221ec3e4-32fb-4fef-86d9-8561cf0a0d4e · outbound

This paper cites Ex- ecverify: White-box rl with verifiable stepwise re- wards for code execution reasoning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ex- ecverify: White-box rl with verifiable stepwise re- wards for code execution reasoning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.494879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.064958Z digest=sha256:30017bc42a637d33f74874fa5563ee5187dc6941e6c2f0eec1492e7b4d95ddfe

Observation 47c4081e-8bdc-4fff-bf26-1a51e6c534a5 · outbound

This paper cites Ex- ecutable code actions elicit better llm agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ex- ecutable code actions elicit better llm agents

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.475860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.072789Z digest=sha256:5601f85f3473c0537f416e6c2b1754326f695f7d7c489f11822ff5b12a8bc149

Observation 0901c2c9-f0e5-41af-add5-75d7ad709401 · outbound

This paper cites Open- hands: An open platform for ai software developers as generalist agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Open- hands: An open platform for ai software developers as generalist agents

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.451400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.083659Z digest=sha256:19f60f354212d7efaac39bae3a09d36eade28816e2354289830e740241b6e729

Observation 588315d6-711c-47f0-a712-50ff196738c3 · outbound

This paper cites Lightweight Self-Knowledge Distillation with Multi-source Information Fusion.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Lightweight Self-Knowledge Distillation with Multi-source Information Fusion

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:00:53.632208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.093830Z digest=sha256:56fba7c5528c7ec9c8675b7066f8decdec5521a4a315b5b148fb2524f08af1be

Observation e38ea274-93e9-4d8f-9c39-8e34500c770f · outbound

This paper cites Multi-label self knowledge distillation.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Multi-label self knowledge distillation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.429162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.103235Z digest=sha256:b8148b7dc61c6d2bb01b6ecbfc9eb0306b4499cf968e1925180d38bba2e95ffe

Observation 8e70c78f-71f1-41ae-9751-2a01ef3524bd · outbound

This paper cites Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:00:53.586243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.117008Z digest=sha256:d7fce6912c126fac25c718fd5497d723b20629d0e9b4f8f2da09a93d52733757

Observation 4d604580-d058-465c-a1bb-13163f739bdd · outbound

This paper cites Ojbench: A competi- tion level code benchmark for large language models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ojbench: A competi- tion level code benchmark for large language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.127123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.127123Z digest=sha256:5725fadeb6980aefef3c2545d535281a4d94284df447004403fd6718e7e431b7

Observation b43e4fac-d08d-4843-b78e-462087c6991e · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.133261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.133261Z digest=sha256:7775a4059e5aa97469d36ad0d6bb68c63f3d3361888e43bf258b96dd1aa51028

Observation ca742a1c-a1c6-4dee-81a6-3d791f4b3904 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824– 24837, 2022.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824– 24837, 2022

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.408999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.141550Z digest=sha256:c30be5576cf18715d0869d33a1b197e3c23e0791f938ac845c369626f4457d3f

Observation 3bbe2263-4e0c-4dc1-8f32-39249168a903 · outbound

This paper cites Swe-rl: Advancing llm reasoning via reinforce- ment learning on open software evolution.Advances inNeuralInformationProcessing Systems, 38:78500– 78525, 2026.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-rl: Advancing llm reasoning via reinforce- ment learning on open software evolution.Advances inNeuralInformationProcessing Systems, 38:78500– 78525, 2026

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.387694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.148353Z digest=sha256:6e7dafcf75e09dd4ce631628aee01dc14a9ac01fcae1f28f0c9fb7625b27c4d4

Observation b5487387-c251-4b8f-a5bf-fcec858ac467 · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Agentless: Demystifying LLM-based Software Engineering Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.155623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.155623Z digest=sha256:7c50e6982f5092834252e87bb369992ee9b1d4d1a2040e817551c3b4264da04f

Observation 3508a87c-76e5-4c2d-8369-6cd95c41bad6 · outbound

This paper cites LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.161982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.161982Z digest=sha256:8e0ad1528f94cc045de2f954812c2ee1b7523b2a3d16579729cae6698d15456f

Observation 2f320710-abfb-4221-abeb-35ea8925fee4 · outbound

This paper cites Icpc-eval: Probing the frontiers of llm reasoning with competitive pro- gramming contests.Advancesin Neural Information Processing Systems, 38, 2026.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Icpc-eval: Probing the frontiers of llm reasoning with competitive pro- gramming contests.Advancesin Neural Information Processing Systems, 38, 2026

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.363007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.169064Z digest=sha256:46e4891f2d50306c9f59366383cc32df2c90da7ef9fd461744290bdc26d0d78e

Observation fac97e41-cf4b-4f13-a262-43bc38f2b12b · outbound

This paper cites Swe-agent: Agent-computer inter- faces enable automated software engineering.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-agent: Agent-computer inter- faces enable automated software engineering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.335861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.177569Z digest=sha256:68b24fcd39b042f3d9cfdc7552c0d099a7ac8fb5a4dec8d959ba0696d8be785a

Observation ea1ef2f3-8f9b-46eb-8f3d-373ca0326653 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.314717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.184920Z digest=sha256:d753c5e419b91330018fc8082fd5906b73558b70b90bd13de5d42e078e421ee3

Observation cfdc3d24-d520-4968-9b00-34595f0ed208 · outbound

This paper cites React: Syn- ergizing reasoning and acting in language models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training React: Syn- ergizing reasoning and acting in language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.290530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.193005Z digest=sha256:d60bca6c8a3d52386762927665e035d7b61fb7843cb346f72b7c433f5db003c5

Observation 50916849-3599-4b50-beee-4b8a7bb28f7a · outbound

This paper cites Dapo: An open- source llm reinforcement learning system at scale.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Dapo: An open- source llm reinforcement learning system at scale

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.272093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.199443Z digest=sha256:2d18edca39b50c6107a7e39ebe81c8cbc1008ffa24ec86fc1ef88171d1a30e48

Observation 1aeb31ca-10f6-442b-b360-d62a0cecf979 · outbound

This paper cites Group Sequence Policy Optimization.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Group Sequence Policy Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.210565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.210565Z digest=sha256:09741876c3c6c21fd6c5849530f99dd2127f93c5ee369c914148f5f9fcec057b

Observation 1f4b7440-a570-47d8-b43f-e20b6eeeca5a · outbound

This paper cites A syntax-guided edit decoder for neural pro- gram repair.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A syntax-guided edit decoder for neural pro- gram repair

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.249562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.218339Z digest=sha256:a8ff7176e22f8cb8596e8835ecf2a02b07ca09e2e217ebe57a15cb49c9d553fc

Observation 5d6bbe64-8f09-4f69-80c5-bd33ead92af8 · outbound

This paper cites GAGPO: Generalized Advantage Grouped Policy Optimization.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training GAGPO: Generalized Advantage Grouped Policy Optimization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.226387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.226387Z digest=sha256:76104e2907a5b0395ed67a76d68790d102b31ddaf4ca98d4013514da17ca19ed

Pith citing papers

No inbound Pith citation observations are available.