Pith. sign in

Paper Citation Record · LEDGER

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training

As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2608.07147.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07147 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:00:53.226387Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy29
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96dd344a-c9a8-4a6c-8fe7-66aa61c3c76e · outbound

This paper cites Program Synthesis with Large Language Models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.706852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.706852Z digest=sha256:514d6fc2928628cae55a737b416b0fbb7f7c18014790be47b4bd13577a730381

Observation b30ba429-aa4e-4f0c-a4a1-b7431964c91f · outbound

This paper cites A general theo- retical paradigm to understand learning from hu- man preferences.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A general theo- retical paradigm to understand learning from hu- man preferences

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.891649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.715509Z digest=sha256:54b75bed434a43d30fb2d10d2659b182fb91ce8788e06e34260e3b0009931d4b

Observation c0caab54-2410-4fa9-88d2-a451dd47fc33 · outbound

This paper cites A course in metric geometry, volume 33.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A course in metric geometry, volume 33

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.868084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.722903Z digest=sha256:1e481c08d2bacf47b8ca290af328c194bd8712ff98e3ec9de16bf9fccc0121ee

Observation eafb808e-3f7e-4071-88ff-7ec9de1adfa6 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.729442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.729442Z digest=sha256:82c24256484b638711e43746d6b97a4372ebb016440b0ed6c39d3c362d219c59

Observation a2414384-1702-4a07-87fb-580d91a55674 · outbound

This paper cites CodeT: Code Generation with Generated Tests.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training CodeT: Code Generation with Generated Tests

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.742000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.742000Z digest=sha256:471da28a0f6ecdfc8f3f5085e9ab46bd5c0f29fd7338b19edc8ab0fdfff99ede

Observation 604af2c4-39c3-4a80-8fb9-b61d552f1bdd · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.751101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.751101Z digest=sha256:6991c75a052077dca8a6b166b9321652a5ef1c88e9c494c090d96031e01d315b

Observation ee1b012a-c8f5-48a4-aed1-cd8f86d8ec9e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.758668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.758668Z digest=sha256:a9eddf3a26c7e70dd48254dfe94828cf89e1f73c2aabb936744db747b08cce8c

Observation ded73128-d4c1-4245-9883-16835dc4b281 · outbound

This paper cites Teaching large language models to self-debug.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Teaching large language models to self-debug

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.849027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.766615Z digest=sha256:628f60db12078bb9df9edeffaf1049fa6fb0fc0579701180a8e6ae91ca9bc7f3

Observation 3e686dfd-e70d-4e72-aaf0-8c0b698f8046 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.774989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.774989Z digest=sha256:6c94a6522705ae8b8571e93eb89368e120989322a68258e40515021d32d8cb76

Observation c318a7f3-178f-4734-a570-a17178156205 · outbound

This paper cites Deep reinforcement learning from human preferences.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Deep reinforcement learning from human preferences

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.824505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.780959Z digest=sha256:29b459d8c846bec6d53d314410a80e19529ad86d2e30f5624ba217ce7a79c5aa

Observation a03c4c37-e90d-40b1-8b97-6a433e41c913 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.788428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.788428Z digest=sha256:c110cfffa42a542668efe2fbfd1f77eb8ce4332e68e48ea1f480f63dbad8fb38

Observation 52f58160-bd33-4651-bd7c-c33a04f1ecc4 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Process Reinforcement through Implicit Rewards

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.797382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.797382Z digest=sha256:680d2b49523dc8bc4bd1c1673e178f889ff1a8210c1e657d2dbae807752a52e3

Observation c0e354f9-bf1b-46bf-aa42-b7357d9270d6 · outbound

This paper cites Memp: Exploring agent procedural memory.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Memp: Exploring agent procedural memory

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.802083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.803557Z digest=sha256:7f998952968d0083b0213ee7c033dcf02b849081aa9ecd4a2e0c1f4119ce80dc

Observation bb008a4b-e7ba-45fb-beae-70e58aea9c77 · outbound

This paper cites Group-in-group policy optimization for llm agent training.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Group-in-group policy optimization for llm agent training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.781914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.810978Z digest=sha256:eca40ea8d1310c82f0c26762ede290e107b93d3705a539bca025cca56d3efe39

Observation 8a10f563-8b5a-423c-bcb6-b2ca68db2106 · outbound

This paper cites InCoder: A Generative Model for Code Infilling and Synthesis.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training InCoder: A Generative Model for Code Infilling and Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.821269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.821269Z digest=sha256:1d556231ad177e19c52263db2c4f7294bb0ff9e5ee2b348c8a14cd2322375b01

Observation 86773d89-d8d2-452f-b7bf-0ed4afeaf70c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.827996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.827996Z digest=sha256:a312e898bd862842aab6ab0cd48ad7d1189665d8626cfcdd172368049f52f3ae

Observation 3b2aedc4-624e-4d1f-be0d-6ab78e6847fc · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Skywork Open Reasoner 1 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.837073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.837073Z digest=sha256:446255d6d8ef7bcec2bc9a3ede01a4abf98f4567df0035220759af8c6aab27ed

Observation 97a60960-3f80-4c5f-bfe0-ab41e5cb9842 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Measuring Coding Challenge Competence With APPS

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.844229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.844229Z digest=sha256:3f65ea18d8b516314fdd19fbf27a679a084b558bd7670b98f6f8d8a9a40b9dea

Observation 34727b05-4c6b-43c7-a075-5848c0928136 · outbound

This paper cites Cogagent: A visual language model for gui agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Cogagent: A visual language model for gui agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.852500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.852500Z digest=sha256:8431c06d461b305d90f2bc2edaa4ebbf97ba3b55b4414138661fca2dbb72d6b9

Observation 715dce2b-b515-4371-ad60-dc1057a9ebc3 · outbound

This paper cites Qwen2.5-Coder Technical Report.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Qwen2.5-Coder Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.859229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.859229Z digest=sha256:a6c9b724be5764b7e63c783321e3e3d8af75bb8bbd2e41ee5ec11b6f60c1c552

Observation 07eab6a8-4ab8-441a-8f55-1c5f7d143ec6 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.868170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.868170Z digest=sha256:69a0f79b30c1afd288af87f5e3507a295e5639d2a33d6f0e50f485e76e10f65f

Observation 72af3f08-9c54-4441-821f-77df178fe084 · outbound

This paper cites Cure: Code-aware neural machine translation for auto- matic program repair.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Cure: Code-aware neural machine translation for auto- matic program repair

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.745463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.874648Z digest=sha256:90281c1cd39964a9be47a06256967640d16922673418401bfb4d39d0c8b7b1b6

Observation 5d23cf88-c9c3-4622-8cf1-58faf4a36a0c · outbound

This paper cites Self-planning code generation with large language models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Self-planning code generation with large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.719715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.887948Z digest=sha256:3ac0267c834387826afa4cd76408340ffa2522513665d3dd34de936ca3d22a74

Observation be636341-f7aa-4b45-a333-a47237e80770 · outbound

This paper cites CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.897000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.897000Z digest=sha256:d0287c0f34310146d8521149c4633de135c2e10dbbfd7dc221717cb991cbdb67

Observation c30cfec5-0f4d-4ac4-a2e8-f83ee8545226 · outbound

This paper cites Swe-bench: Can language models re- solve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-bench: Can language models re- solve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.697644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.905481Z digest=sha256:b766fd68755ff080895be15ae8192a4807e17e2a0325cb6de1948f610daea587

Observation 1c03404b-347f-446e-9ace-e665d587bd5e · outbound

This paper cites Defects4j: A database of existing faults to enable controlled testing studies for java programs.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Defects4j: A database of existing faults to enable controlled testing studies for java programs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.677349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.917446Z digest=sha256:2ec9b4e9bdeef4f84757ce80191aebaf084b7ae16201f3e298735f27fb1a1ce5

Observation 39751694-7240-46a2-aa05-80913ae7ff9c · outbound

This paper cites Ds- 1000: A natural and reliable benchmark for data science code generation.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ds- 1000: A natural and reliable benchmark for data science code generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.655863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.924876Z digest=sha256:bd96e82921af54c511e7e7cf94973d7e5764799704a6d2bdcff28c210b277e25

Observation c4df06f0-ca4c-4fbc-a9a2-8dc4f6508833 · outbound

This paper cites Coderl: Mastering code generation through pretrained mod- els and deep reinforcement learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Coderl: Mastering code generation through pretrained mod- els and deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.635931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.935617Z digest=sha256:c4464d7e21a6bf47489f7f1c85a0f8a55a8265d85e1d6a49c909cbd3ab542bc4

Observation b7090b3d-68ee-4154-8aa9-abadb60c460f · outbound

This paper cites A systematic study of automated program repair: Fixing 55 out of 105 bugs for $8 each.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A systematic study of automated program repair: Fixing 55 out of 105 bugs for $8 each

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.613194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.945542Z digest=sha256:bf8712feb52e7948b566bf2d24c28c0e6956afb4f1e6c0578f574726f0b87583

Observation a4423c6f-381a-4223-ac9f-5611ff0c601e · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.951785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.951785Z digest=sha256:b3035595d4cf5473893c8d7bb1a2657c0666ab64a1e55d476b4ce5e228ba205d

Observation e42778f4-00d3-49f6-8d85-2a101745a892 · outbound

This paper cites StarCoder: may the source be with you!.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training StarCoder: may the source be with you!

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.959772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.959772Z digest=sha256:8aa97ad812f659ccccc540eb4935881d08f4744719f1a089d1f99cf44664c199

Observation a2d316f3-4a1e-459c-a37a-b2ad8723f5ba · outbound

This paper cites Let’s verify step by step.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Let’s verify step by step

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.592579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.968533Z digest=sha256:602e8465fa136d40aa0a0e5a428aa2e6e2a8051692c768f93ca8546eeac58b90

Observation d98740db-f9be-47a6-a44b-9c8a68c49c18 · outbound

This paper cites Automatic patch gen- eration by learning correct code.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Automatic patch gen- eration by learning correct code

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.573032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.976470Z digest=sha256:71aaabc4ac8a4e9f1823a3d593fd7e70fadca7a8dc47e7eab648f7122a4c70e8

Observation 38cc9b54-22aa-4d9e-b27b-be5d3f894c4a · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.982894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.982894Z digest=sha256:b837b9fe3709a2c61f2072b97005553de10957767cfacd4bad3dc39f635883e3

Observation d9e86fcb-867e-407f-ac33-e219f224351b · outbound

This paper cites Gromov–wasserstein distances and the metric approach to object matching.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Gromov–wasserstein distances and the metric approach to object matching

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.553305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:52.993757Z digest=sha256:c924a9ec8e26070975f5deff664e23f314a59369e8915de6e559e7ed4a89388f

Observation ffec4f66-7adb-44dd-abd4-e37dd26e4222 · outbound

This paper cites An analysis of approxima- tions for maximizing submodular set functions—i.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training An analysis of approxima- tions for maximizing submodular set functions—i

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.529276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.000836Z digest=sha256:ea3fa5980b2aa01b6b67e25282ac4c67ebeef5dcd0b29534e86c833d1277b7ec

Observation c8e1f20b-cf6b-4a58-b9b5-88a3e13e7256 · outbound

This paper cites Training language models to follow instructions with human feedback.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Training language models to follow instructions with human feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.010831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.010831Z digest=sha256:d8cabb7f8f6f32f147022ecbba29fc485241064e1b084219763235440e975ace

Observation 59b2863d-88c9-43e5-8ba9-a930c36d89e1 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.018708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.018708Z digest=sha256:b40d54c2578052b4f35ebf337f1537e2ea6d84e685bab8bd9a472fad247df232

Observation 6e2f8ecb-b654-4df3-9321-66e6bf130d60 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.028899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.028899Z digest=sha256:0824fd8bb555594f1fa10d142bf2f2b9a11144c7e34655c7add075c03067eb85

Observation 01e04f36-e2b7-454d-9f54-8ab530e9e8b7 · outbound

This paper cites Can Language Models Solve Olympiad Programming?.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Can Language Models Solve Olympiad Programming?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.037166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.037166Z digest=sha256:48aac1a2d90aa8994bf7bc9b27c61f10548e6dfdebf3fc58fe7ac4a409ff404d

Observation 34648ae4-8a93-47f0-aee9-adfb3921f37d · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Reflexion: Language agents with verbal reinforcement learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.045807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.045807Z digest=sha256:b3371d8f5b7284388ab65f952a1bac9e12e93f2a70e893f11f1633bee7b815bb

Observation b4ee0b14-3de4-4172-b51c-9c847a547bbc · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.055761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.055761Z digest=sha256:6c998ecb176854295fe6c20c6c3f7a242a1f9530a9bfda91c8cd75f6f851bf8f

Observation 221ec3e4-32fb-4fef-86d9-8561cf0a0d4e · outbound

This paper cites Ex- ecverify: White-box rl with verifiable stepwise re- wards for code execution reasoning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ex- ecverify: White-box rl with verifiable stepwise re- wards for code execution reasoning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.494879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.064958Z digest=sha256:265ef090fd4d7ea6a736b42a48e3be8499c40b3d10136a3d20cafd714d84e43f

Observation 47c4081e-8bdc-4fff-bf26-1a51e6c534a5 · outbound

This paper cites Ex- ecutable code actions elicit better llm agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ex- ecutable code actions elicit better llm agents

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.475860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.072789Z digest=sha256:718ee378a9a7f5223a04d591ad245a91b33234458c1cdb8efd8389b20996027d

Observation 0901c2c9-f0e5-41af-add5-75d7ad709401 · outbound

This paper cites Open- hands: An open platform for ai software developers as generalist agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Open- hands: An open platform for ai software developers as generalist agents

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.451400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.083659Z digest=sha256:7a8444220a4115e5501de4664adcacaa0cf270b907db31124abe44a9debff764

Observation 588315d6-711c-47f0-a712-50ff196738c3 · outbound

This paper cites Lightweight Self-Knowledge Distillation with Multi-source Information Fusion.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Lightweight Self-Knowledge Distillation with Multi-source Information Fusion

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:00:53.632208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.093830Z digest=sha256:c4e70ef49ffa10b3a8f6d788399d678099ecebd62a2946feea12e7d35feb5e78

Observation e38ea274-93e9-4d8f-9c39-8e34500c770f · outbound

This paper cites Multi-label self knowledge distillation.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Multi-label self knowledge distillation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.429162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.103235Z digest=sha256:f5aeee5b2703129299b8a3bc35f2738e83b80606d58e193162f7ccbb3863cbef

Observation 8e70c78f-71f1-41ae-9751-2a01ef3524bd · outbound

This paper cites Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:00:53.586243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.117008Z digest=sha256:d006b113df5daf72c739605c8738ea066f32013b3980952d28dfe481c739b105

Observation 4d604580-d058-465c-a1bb-13163f739bdd · outbound

This paper cites Ojbench: A competi- tion level code benchmark for large language models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Ojbench: A competi- tion level code benchmark for large language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.127123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.127123Z digest=sha256:86fe57d85cd6baa6616f104eb68e04b99e7b8d6a10ecd11a286ceca21cd837a4

Observation b43e4fac-d08d-4843-b78e-462087c6991e · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.133261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.133261Z digest=sha256:0a34231c0660c2f570914866e75db3ba01f028c9b1d57d8ed5d5c68d0b144455

Observation ca742a1c-a1c6-4dee-81a6-3d791f4b3904 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824– 24837, 2022.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824– 24837, 2022

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.408999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.141550Z digest=sha256:7325f4034712183de75ccaf254123009e5c27b12b81abfcab5052ffd523a811c

Observation 3bbe2263-4e0c-4dc1-8f32-39249168a903 · outbound

This paper cites Swe-rl: Advancing llm reasoning via reinforce- ment learning on open software evolution.Advances inNeuralInformationProcessing Systems, 38:78500– 78525, 2026.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-rl: Advancing llm reasoning via reinforce- ment learning on open software evolution.Advances inNeuralInformationProcessing Systems, 38:78500– 78525, 2026

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.387694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.148353Z digest=sha256:333ad19f0c09dd1b643033fb023d934a50ead3cbadd814ddb611f9b82db218b0

Observation b5487387-c251-4b8f-a5bf-fcec858ac467 · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Agentless: Demystifying LLM-based Software Engineering Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.155623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.155623Z digest=sha256:8fb26217b1731a82e931394471e4be721bf613b29d6e5f14a896e57a3280151b

Observation 3508a87c-76e5-4c2d-8369-6cd95c41bad6 · outbound

This paper cites LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.161982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.161982Z digest=sha256:17e3bf538518734c910f42105676c0805b81a9eaaf2a12ff2932f7d18a837855

Observation 2f320710-abfb-4221-abeb-35ea8925fee4 · outbound

This paper cites Icpc-eval: Probing the frontiers of llm reasoning with competitive pro- gramming contests.Advancesin Neural Information Processing Systems, 38, 2026.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Icpc-eval: Probing the frontiers of llm reasoning with competitive pro- gramming contests.Advancesin Neural Information Processing Systems, 38, 2026

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.363007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.169064Z digest=sha256:76b4b2e48779c31c8487ac47d25706c0c777794383749bb6834640112bacea14

Observation fac97e41-cf4b-4f13-a262-43bc38f2b12b · outbound

This paper cites Swe-agent: Agent-computer inter- faces enable automated software engineering.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Swe-agent: Agent-computer inter- faces enable automated software engineering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.335861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.177569Z digest=sha256:5b117b89a6dca6ef4babf58716e6f03490743286742d880bd9f2622843eecc65

Observation ea1ef2f3-8f9b-46eb-8f3d-373ca0326653 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.314717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.184920Z digest=sha256:d028c5deb203066c6eecadff5e8cc31b04b6ffcad73a73a28f051ec31e66fb34

Observation cfdc3d24-d520-4968-9b00-34595f0ed208 · outbound

This paper cites React: Syn- ergizing reasoning and acting in language models.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training React: Syn- ergizing reasoning and acting in language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.290530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.193005Z digest=sha256:cb534614d1fb369ae1be71e588eacf3b474d1da6de3a42270338614c1144a7e3

Observation 50916849-3599-4b50-beee-4b8a7bb28f7a · outbound

This paper cites Dapo: An open- source llm reinforcement learning system at scale.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Dapo: An open- source llm reinforcement learning system at scale

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.272093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.199443Z digest=sha256:ae380286c763299a865c9e65deb5e39700017dab42f155f483c708b4231c7f4a

Observation 1aeb31ca-10f6-442b-b360-d62a0cecf979 · outbound

This paper cites Group Sequence Policy Optimization.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Group Sequence Policy Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.210565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.210565Z digest=sha256:9bcaaafbaf407855f87d273011fd6b1ec21ef6b7fbad858d2830f2505513a3b6

Observation 1f4b7440-a570-47d8-b43f-e20b6eeeca5a · outbound

This paper cites A syntax-guided edit decoder for neural pro- gram repair.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training A syntax-guided edit decoder for neural pro- gram repair

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:00:54.249562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:00:53.218339Z digest=sha256:8f0e967fa94765c3902a8eaac96edd2033b931e12e917597ebd4dfe2c9742200

Observation 5d6bbe64-8f09-4f69-80c5-bd33ead92af8 · outbound

This paper cites GAGPO: Generalized Advantage Grouped Policy Optimization.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training GAGPO: Generalized Advantage Grouped Policy Optimization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:53.226387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:53.226387Z digest=sha256:31a19316738ab18c35b9225fba95f18b93d08cd9596d56e7ff77bf3d87fb49fc

Pith citing papers

No inbound Pith citation observations are available.