Pith. sign in

Paper Citation Record · LEDGER

Differentiable Evolutionary Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2512.13399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.13399 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact21
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02d0721f-e551-4110-9805-1acf73e1e890 · outbound

This paper cites Concrete Problems in AI Safety.

Differentiable Evolutionary Reinforcement Learning Concrete Problems in AI Safety

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.742178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:b66305b1a0d2efb86860f2c4f28fc595ca42f7ba1387dd6a040bc80521371cfb

Observation e5d5fe68-ef3b-40f3-bda5-7b98e841a8c9 · outbound

This paper cites Learning to learn by gradient descent by gradient descent.

Differentiable Evolutionary Reinforcement Learning Learning to learn by gradient descent by gradient descent

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.718469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:8f3551cf4c4c6a778a3d0ea3f9fbee5958887e087d53cc9f1e974cea824ebc6c

Observation ae785887-de13-494b-b684-03be65f54ed3 · outbound

This paper cites under review.

Differentiable Evolutionary Reinforcement Learning under review

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T22:43:38.780793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:5bf282905c1668b62d0f463d829e525cec453c5c5ce58f49cdafa8e6e4e58033

Observation 7e28bf47-f990-47bc-8c8b-9558c4b3591b · outbound

This paper cites Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, and Quoc V.

Differentiable Evolutionary Reinforcement Learning Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, and Quoc V

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T22:43:38.786382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:e54bee3ec22f2676e4143607fdf25398c72e8e649cb6abfc069ed81e7f7279aa

Observation 69b1c518-69f0-4530-9e6e-f0f74d4be013 · outbound

This paper cites Xingwu Chen, Tianle Li, and Difan Zou.

Differentiable Evolutionary Reinforcement Learning Xingwu Chen, Tianle Li, and Difan Zou

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T22:43:38.782676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:8a8c5f127757d2659daadcae7e513c2fcc5b5fbc28b41ef242acce2ab14d5bf2

Observation ddef9326-458f-4dbf-b92a-4d74043a8a36 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Differentiable Evolutionary Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.676556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:d44ea21c09d771e8f7b360b8692288ed7e685dc140ebf8a8f62e6950bde74063

Observation 9d366ba9-077b-4570-8acd-a65603b4e8cb · outbound

This paper cites A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems.

Differentiable Evolutionary Reinforcement Learning A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.726446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:f160cf16ad37a04748d770a38ca1f68ec4e7366c1ce302bf7e2fdc15b08fc81c

Observation 77d0c7bd-4d6f-45a5-99c4-4b131f0726ac · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Differentiable Evolutionary Reinforcement Learning Group-in-Group Policy Optimization for LLM Agent Training

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.745778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:cc8ad36ad642ee21c08648a787ddfaa67ee14b3dbc909a60dfa0adcd2a39a177

Observation 11ce5c06-1edb-43c9-bd2c-3a02c38fd0ea · outbound

This paper cites A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence.

Differentiable Evolutionary Reinforcement Learning A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.734181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:123d7520c140976227f3e76cf85f0bf4b3b1379c3bc8633075b4c1bccdd9d583

Observation d49a2ebe-7c23-49c9-8545-a5733f40a76f · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Differentiable Evolutionary Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.752791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:8fef848fbba6819ba54b28cb00c14114fb441a464308f3897bd87805378ce1b4

Observation b64778b3-9b74-4f4c-b318-f75442a5e210 · outbound

This paper cites Population Based Training of Neural Networks.

Differentiable Evolutionary Reinforcement Learning Population Based Training of Neural Networks

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T22:38:37.424143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:06499a81e2c4eb92ac3ab10182133f41b516a8ec2c06cd79b5f5715deebcc8c2

Observation 08fc1195-f40c-453e-a9fb-c5cb81c61012 · outbound

This paper cites Wurman, Peter Stone, and Craig Sher- stan.

Differentiable Evolutionary Reinforcement Learning Wurman, Peter Stone, and Craig Sher- stan

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:38:37.672326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:974640f0faee7ba1b7a2d3148ac6bdd3a3bd64ec7cba43a5f286d0204e34bf7d

Observation 5c674f03-8bbb-46bf-b136-246e8b664e66 · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

Differentiable Evolutionary Reinforcement Learning AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.693784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:90228b6cf697235f9f428f77e5fbaffb1e06ffc7ff1083ae4c50424080b5dd8b

Observation ba872648-a8ef-4272-a4ab-2978d8bd1187 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Differentiable Evolutionary Reinforcement Learning ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.730345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:c9872a1cd574b38f4fe854476b12ee2980094ab8178eb579393c82b847a864c4

Observation 07155f3e-0191-48e8-bd72-0133ca926e9e · outbound

This paper cites doi: 10.1038/s41586-023-06924-6.

Differentiable Evolutionary Reinforcement Learning doi: 10.1038/s41586-023-06924-6

Reference 15

Resolution
verified exact
doi, observed 2026-05-16T22:38:37.428031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:8d4234f143ea181c73ca03da77c61a99da1894f5845d46df48ebabf1efa2b506

Observation c63261d8-6209-4c04-86b5-e20007f8028b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Differentiable Evolutionary Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.680496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:c0aec87046f3af04e17f1abc4c1f567f70073bf142f4a5e86a494a1a3eb63351

Observation be81840f-ea11-4071-9bd2-3f56d03f9ab4 · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

Differentiable Evolutionary Reinforcement Learning DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:28.487879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:2bd104d11b020aa7d646df49cbab12f153886f3e4bec5c29582b2ee81ade41be

Observation b0b1d0be-cdb3-443d-89cf-8d1894fea4dd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Differentiable Evolutionary Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.756424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:7744fc6cc371551f1703c32580e7390d494903a14e2f4f5e31f7d1ce11e94a0b

Observation a0131b75-0c1e-445a-9292-c4c2df1c1743 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Differentiable Evolutionary Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.749379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:b460f53cacc35370c3febf71853dcf80cc207f67b8c5c4522451fca3df03ce9e

Observation e40a3bff-d5a5-4a11-925a-abe9a37bca9a · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Differentiable Evolutionary Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.710448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:6b5b27169b06debd53267b8b7f27f7e244e6db8be70f00fc9dd072a7caca14cb

Observation 9d77cbbd-6597-417e-b19a-59dbf6b28ae3 · outbound

This paper cites Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning.

Differentiable Evolutionary Reinforcement Learning Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.722482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:ce4e7f832e44bd323fe10d1088872cbe863bba5c777500dd5e5c558d6b230cfd

Observation 6877c6ab-f718-438a-950a-ff54f58c54fa · outbound

This paper cites arXiv preprint arXiv:2510.04204 (2025).

Differentiable Evolutionary Reinforcement Learning arXiv preprint arXiv:2510.04204 (2025)

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:38:37.738299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:aba316620612b117f4a7888eaf21dc6e8cf8ebc402539659ba3fdf99231ee729

Observation d1b91593-7745-4878-bc2c-2d09e2cb6153 · outbound

This paper cites ScienceWorld: Is your Agent Smarter than a 5th Grader?.

Differentiable Evolutionary Reinforcement Learning ScienceWorld: Is your Agent Smarter than a 5th Grader?

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:38:37.714865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:3b2598c8220a986122975efb0722cf9f8be803e56922360dd123260ca672b715

Observation 2e67659e-2962-4cb4-b066-412c13cbcf6b · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Differentiable Evolutionary Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T22:38:37.697896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:f374db12ff1ca3cb826f5048614ac3f7d9d507532d14b9911d676861f534ad88

Observation 8c4c22c5-3eba-4d9f-b069-19726470732d · outbound

This paper cites Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny.

Differentiable Evolutionary Reinforcement Learning Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:15:55.820429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:b87540ee7d5a5119401621e90d20cee13011364fedfae86481d7f43ba4c161fb

Observation 593fab5c-e79d-490c-ac22-edad954cf38f · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Differentiable Evolutionary Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.668024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:ae4db1edfea33c695b3265f2df221297d985e4fbac5a297b980ec96577b87539

Observation 572b5b97-c897-463a-9d00-920974e3c058 · outbound

This paper cites Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents.

Differentiable Evolutionary Reinforcement Learning Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.702055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:130f9a31124a1e1617356c457c847b66c991a1e5f0562645f44f674f44338c3f

Observation 78e53603-8078-4053-9f3f-5712016ef43b · outbound

This paper cites RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents.

Differentiable Evolutionary Reinforcement Learning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:38:37.689797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:00b59388e93142c2db70afcc38989f5fd30f70ff04efaf3ef7ad5345a5c6c93d

Observation 568bd4b0-c1c1-4645-bcc2-ce4612a1e737 · outbound

This paper cites meta-gradient.

Differentiable Evolutionary Reinforcement Learning meta-gradient

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T22:43:38.784432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:27769535f988145f23d3192df494fdc5bc03043c56f7bb057b33da777bbf49a3

Pith citing papers

No inbound Pith citation observations are available.