Pith. sign in

Paper Citation Record · LEDGER

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2606.03108.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.03108 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T10:30:42.057301Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:32:38.114748Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact30
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7ab1db5-7afa-4111-bc66-8c630582546d · outbound

This paper cites 2026 , howpublished =.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning 2026 , howpublished =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:60ed3bad86652acdaea2eca3b3e7bfe92dbcd03a3d65c7b9689d8c65970d0a43

Observation 9aa53133-e85e-46b9-af69-3793cccce76b · outbound

This paper cites Darwin G.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Darwin G

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:ff47202ebd5f870c7bdf7ff15a7cf9588101756a13b0cd445759b4ff4387aa83

Observation 14836864-8dd6-4b2c-8f94-d1051a907ffa · outbound

This paper cites an unresolved cited work.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:b9f50f7d2e22debabae4ea8162d588536a14524b287ff5bbe6060833b796ff5c

Observation 9991e52e-6ec9-4eaf-83e5-09ce5f10699e · outbound

This paper cites an unresolved cited work.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:b195e7bdaac2f1ea177e646136c3b21d3c0d6dce75a314073586c54b01848e1d

Observation c0905de0-a5fc-4580-9d15-bbb4ae61df5f · outbound

This paper cites an unresolved cited work.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:c3107e2879da37baf891895a09a9a37fb3263dfc6e8c5eea93225ba0bf74e25a

Observation 4fd1cc0d-8646-496c-9302-9055feee126c · outbound

This paper cites an unresolved cited work.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:2600c3f891476167be82a982783a67bc67f5e617ec627d37ab2c9cc0d2610e96

Observation c80d199f-6e67-43ff-bece-84f5dbc9f20c · outbound

This paper cites Advances in Neural Information Processing Systems (.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Advances in Neural Information Processing Systems (

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:6d7f6e3986fd189549010aa3a10b8732df569eedb95e579e8d5d4eff556a0c54

Observation 8bd2eec7-180a-44ef-bd43-7e7037dad973 · outbound

This paper cites Proceedings of the Thirty-Second.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Proceedings of the Thirty-Second

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:71fd4f3df53a20ac144236f73d8da74edd81b46373e3e39db8d953432ba719f6

Observation 13fbd11e-5637-4764-990b-f94334b337ef · outbound

This paper cites 2026 , eprint =.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning 2026 , eprint =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:aab1a13aadd9febd118b424c4d60c70ec57d40051b9c1bef42265e4cb951a9f8

Observation 9f438de2-b65b-4608-b838-150e6ccc60d8 · outbound

This paper cites Bellemare.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Bellemare

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:1c2aa5194ff4bad0eda396cc0338b60c99ded912df64a57c60509c9838660e6d

Observation 4229a3b1-0471-4e0b-b1b2-0def05f79229 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.149115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:98a847c0f4199eb8c43a8fb7a8d62d427c02203c0f47e18a3ed64b6f15d1b02b

Observation eb5fcfef-e46c-468a-882d-7fe468754555 · outbound

This paper cites an unresolved cited work.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:1e11885362ba6eb7141ad7933b6d994d9512e7038c63bc42a4640aa553d36b7f

Observation 4f7bcaba-cd9d-469a-b9c1-8cade883d28d · outbound

This paper cites SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.172057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:a6e4f3ee3549cd75db9924c035b3cf20bc91b9265cdcd10c9466169d17cf876a

Observation 86c96cc9-33c8-4d74-9883-c7cccd025f6f · outbound

This paper cites an unresolved cited work.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.076490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:87721390577617dbe56537de3a2b399fd8d25b066c43fc360cc18bd3d7790016

Observation eabaf675-ae68-45a3-9704-648abfd971f8 · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.163648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:e8b259f3448e52885d1cbbb42204ffe479f860407b5e57d3b58c5cd1463a6268

Observation df0da249-623f-44ca-83e5-5bec76037bda · outbound

This paper cites arXiv preprint arXiv:2505.09655 , year=.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning arXiv preprint arXiv:2505.09655 , year=

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.064402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:0bdab69bec4cc83d2060002950406083dbfc0f394a67f2555764356f92dda53b

Observation 18e13c95-ede7-4f47-8d4e-394a87fb00bf · outbound

This paper cites How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.113814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:59d40aa5fb067a19f677addfba25a9565486876f995253ee834419cab8b9e823

Observation 201448c2-adc1-4eb4-b78d-a5d8f0035b45 · outbound

This paper cites doi: 10.1038/s41586-025-09422-z.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning doi: 10.1038/s41586-025-09422-z

Reference 44

Resolution
verified exact
doi, observed 2026-06-28T10:31:59.820179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:6af0d925c5ad847cea87d5d520aa07f894b6d8341309621ef75fe329193e2a67

Observation fd40d41b-5de1-4d40-8788-68d2d7c5f240 · outbound

This paper cites an unresolved cited work.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:66a066e66e7359ae337f29e24411f3b9152e7492a5e01914730cdcbdc914cc5b

Observation 9cc9bac4-8ee7-46e7-a7d2-dc24d144180f · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.137790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:35fe1264c7e06387013fd33901538abd5b77faa1eb7a0fe201fb72e0e5757e09

Observation bd8e608a-d9f9-4550-a120-4841c52efd17 · outbound

This paper cites On the Emergence of Implicit Curriculum in RLVR Learning Dynamics.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning On the Emergence of Implicit Curriculum in RLVR Learning Dynamics

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.141647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:6c3ed778edba36e588ab0722be7f5f3d64c073d7e297605db710f1f726957cc0

Observation fe3d4da6-9167-4967-aa75-8db0d155d2ba · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.157015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:d1e757ee1da6f9e7fa5d3de1876594a46fbdc1b085d6581355d38cfab6abec43

Observation 48063652-dd50-4ce4-9492-e02ce4bd88fb · outbound

This paper cites GEAR: Genetic AutoResearch for Agentic Code Evolution.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning GEAR: Genetic AutoResearch for Agentic Code Evolution

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:56:29.093104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:f952d6818c96c45dbc951350c9e3fc1ff0592264405853f5b44bb370224f3769

Observation 351c121d-b70c-4392-88a0-b6f0978d7e68 · outbound

This paper cites an unresolved cited work.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:dda0571703af154e56688d00c997c56f015749bdb97484777b61d43ed861123d

Observation 865f7c78-f573-4cbe-ac54-c39a44b538b8 · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.160351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:d5e3891a8919576626dfb7a181accac2a20286b0d2ab84ebde871c0523ba5de7

Observation f6d1447c-6761-4a27-8216-37b9b8595390 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning TACO: Topics in Algorithmic COde generation dataset

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.160398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:e992fc046f148686db6e9ef9231e350789fcfd3de87a093a359ac91495137bd7

Observation cea56aa0-0857-42f2-8911-aa873dc0d96b · outbound

This paper cites Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.149394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:348b8534f5e33f5de9fd44e4db74e91e09064ed6e5bd4d7c65de341c1dcdcb2f

Observation 16819099-b622-421a-ab33-4351d30918c8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.117233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:7b7951ed6059943c9546d23947c24ba7d24ebc44c908ea41e9e9229ef0b13385

Observation faee7cf6-a254-4a18-9814-36bb86ca4520 · outbound

This paper cites GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.176018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:5ad16a8c74c8dd7ba23a31732ca21c5f4fa883534fcf87ac65f922e9782f1aa2

Observation a6c679db-c435-4b5e-8c89-c3d8fc66ee43 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.163873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:0ce96650a806c1c6cbbb318c04600638290af7942e4258a153e66274894e3405

Observation e080dda4-992e-4013-ab71-81215dd3dd21 · outbound

This paper cites Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.168678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:3da3df06e99b4a31a3aa286a6eed88fed445bf10ecaac9d229416bbf2f9dba71

Observation bb38685a-7a15-4bf2-84fb-57702fa7e1d8 · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.084474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:21fba4aaeb4e51cca419e2647ec8de3a835284ede5539882c3baf0cfd71385d0

Observation ffa8521e-96be-4abd-a6a9-bb6e505ca90b · outbound

This paper cites Bilevel Autoresearch: Meta-Autoresearching Itself.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Bilevel Autoresearch: Meta-Autoresearching Itself

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.153238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:faecb8dbe1512ac592443b8e44a5b58a666fcaf4b60fec200764961321db1919

Observation 715a77f0-b58f-4318-8eaa-30ad14ac826c · outbound

This paper cites an unresolved cited work.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T10:30:42.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:22549ba62d7acec6c9e84e03b71e93e31e76652ccd465b9efdab9daecc4d7ad4

Observation fcf8ba75-fec6-4f91-b745-c32721ea6b61 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.156793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:a934223b6b8e305cdffa240a04af7005913fd0f41a69cddd6777e215bd77b47c

Observation a3c8f0e7-376a-488a-89ae-ae90e6cf71a0 · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning A Survey on Self-Evolution of Large Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.129913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:cfcfa39e22a0020cb54d852f7bf4353cab1be5ebee16b22ca0e856a2d521bf1e

Observation b6466a17-bf3a-442d-9222-54bb04b8c50d · outbound

This paper cites RAGEN-2: Reasoning Collapse in Agentic RL.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning RAGEN-2: Reasoning Collapse in Agentic RL

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.105430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:f8a8ac4a6e2fcceb07c0e90786c9ec9c1320b6c4c2a2a03fbc6f615338147de2

Observation 14f83664-ee98-4aa8-bae0-4d0bc5d1a688 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.172159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:c4dce1686a00322349a70bcacf6bd8783240adf5da343af61e3c5cd637a10bc9

Observation 3780e028-fdfd-4c3b-b32f-8729e728679b · outbound

This paper cites an unresolved cited work.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.125496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:f5cbebf78d08c5937788c0ff7df2c46812cb06d6b8890d68b84b9affa248900e

Observation 78483e8e-add2-4202-86b9-6f75ba1c7e46 · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.137487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:9fe082da980c3f9a635e320443428b1dce3cf0e52136c802ef23594ac71da3f3

Observation 06fe170b-1ac8-4284-bc72-4b7db3f2abdd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.113543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:ad70f04766dc1213edc2496c70e50905f08727e635eb6b270a64d60b7e6b2d24

Observation 5bcca6bf-293b-4752-b7da-98a8dde3362c · outbound

This paper cites Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.109678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:0f6d492f2396e6a193ab6eb8c5f4781f2a424e86ebd4d2d3fbbc6e748ac21253

Observation 1f8654e5-305c-46fe-9781-a682b495bb4e · outbound

This paper cites RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.080522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:e7f3d60dea4466b10582c587f06dc2a6c87416b80487ce9837fe712cfc0efb65

Observation 148e3c70-6848-4df5-ba8e-a552e8a8b7fb · outbound

This paper cites Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.145323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:91cdc77dcd16ce3037d63504bace2dc2b4bfe452fd1efbfc7da148b49318814a

Observation e2fc6487-0ff5-4d4e-a8ac-2b3ef55e6256 · outbound

This paper cites Group Sequence Policy Optimization.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Group Sequence Policy Optimization

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.145569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:47886b0621da5d27f664fbb513e972d1c89d234f2db45ef2e716e6ae0040cf31

Pith citing papers

Observation 7ddcb646-167c-4ac7-8c40-a26e10b1623a · inbound

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity cites this paper.

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T04:32:38.114748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T04:32:38.114748Z digest=sha256:1778ff4772792ceae6a265271046521143ba1b5c1ada0dd8d2893a376513d649