Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T10:30:42.057301Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2606.03108.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T10:30:42.057301Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T04:32:38.114748Z
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a7ab1db5-7afa-4111-bc66-8c630582546d · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning 2026 , howpublished =
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aa53133-e85e-46b9-af69-3793cccce76b · outbound
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14836864-8dd6-4b2c-8f94-d1051a907ffa · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9991e52e-6ec9-4eaf-83e5-09ce5f10699e · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0905de0-a5fc-4580-9d15-bbb4ae61df5f · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd1cc0d-8646-496c-9302-9055feee126c · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c80d199f-6e67-43ff-bece-84f5dbc9f20c · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Advances in Neural Information Processing Systems (
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bd2eec7-180a-44ef-bd43-7e7037dad973 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Proceedings of the Thirty-Second
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13fbd11e-5637-4764-990b-f94334b337ef · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning 2026 , eprint =
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f438de2-b65b-4608-b838-150e6ccc60d8 · outbound
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4229a3b1-0471-4e0b-b1b2-0def05f79229 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eb5fcfef-e46c-468a-882d-7fe468754555 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f7bcaba-cd9d-469a-b9c1-8cade883d28d · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 86c96cc9-33c8-4d74-9883-c7cccd025f6f · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eabaf675-ae68-45a3-9704-648abfd971f8 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation df0da249-623f-44ca-83e5-5bec76037bda · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning arXiv preprint arXiv:2505.09655 , year=
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 18e13c95-ede7-4f47-8d4e-394a87fb00bf · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 201448c2-adc1-4eb4-b78d-a5d8f0035b45 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning doi: 10.1038/s41586-025-09422-z
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fd40d41b-5de1-4d40-8788-68d2d7c5f240 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cc9bac4-8ee7-46e7-a7d2-dc24d144180f · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bd8e608a-d9f9-4550-a120-4841c52efd17 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning On the Emergence of Implicit Curriculum in RLVR Learning Dynamics
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fe3d4da6-9167-4967-aa75-8db0d155d2ba · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 48063652-dd50-4ce4-9492-e02ce4bd88fb · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning GEAR: Genetic AutoResearch for Agentic Code Evolution
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 351c121d-b70c-4392-88a0-b6f0978d7e68 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 865f7c78-f573-4cbe-ac54-c39a44b538b8 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Meta-Harness: End-to-End Optimization of Model Harnesses
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f6d1447c-6761-4a27-8216-37b9b8595390 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning TACO: Topics in Algorithmic COde generation dataset
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cea56aa0-0857-42f2-8911-aa873dc0d96b · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 16819099-b622-421a-ab33-4351d30918c8 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation faee7cf6-a254-4a18-9814-36bb86ca4520 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a6c679db-c435-4b5e-8c89-c3d8fc66ee43 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e080dda4-992e-4013-ab71-81215dd3dd21 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bb38685a-7a15-4bf2-84fb-57702fa7e1d8 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning AlphaEvolve: A coding agent for scientific and algorithmic discovery
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ffa8521e-96be-4abd-a6a9-bb6e505ca90b · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Bilevel Autoresearch: Meta-Autoresearching Itself
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 715a77f0-b58f-4318-8eaa-30ad14ac826c · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcf8ba75-fec6-4f91-b745-c32721ea6b61 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a3c8f0e7-376a-488a-89ae-ae90e6cf71a0 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning A Survey on Self-Evolution of Large Language Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b6466a17-bf3a-442d-9222-54bb04b8c50d · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning RAGEN-2: Reasoning Collapse in Agentic RL
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 14f83664-ee98-4aa8-bae0-4d0bc5d1a688 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3780e028-fdfd-4c3b-b32f-8729e728679b · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 78483e8e-add2-4202-86b9-6f75ba1c7e46 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 06fe170b-1ac8-4284-bc72-4b7db3f2abdd · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5bcca6bf-293b-4752-b7da-98a8dde3362c · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1f8654e5-305c-46fe-9781-a682b495bb4e · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 148e3c70-6848-4df5-ba8e-a552e8a8b7fb · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e2fc6487-0ff5-4d4e-a8ac-2b3ef55e6256 · outbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Group Sequence Policy Optimization
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7ddcb646-167c-4ac7-8c40-a26e10b1623a · inbound
Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.