Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 75 inbound Pith citation observations for arXiv:2508.05004.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T18:53:02.690293Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T18:17:33.697242Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 7ffd83e3-7cfb-4c48-9a3d-49d730adcb66 · inbound
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 57a2e0cc-3889-4def-a8ff-f498295de558 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 171
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 796cf684-1dd7-49c3-b966-fc4b361d7b33 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 209
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation edf8c7ba-09fc-4dea-9987-691406fe9548 · inbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57362e68-6424-4f58-8c70-6cc0d954ae4e · inbound
Discovering New Theorems via LLMs with In-Context Proof Learning in Lean R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 474f9f6f-50e2-4e5d-8757-d82d4147407c · inbound
Discovering New Theorems via LLMs with In-Context Proof Learning in Lean R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 932533c3-59a2-49ea-953d-ae24d7618eb6 · inbound
Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a79947-b1bb-4788-9a65-e5b761eee3d8 · inbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741c5291-7bb4-4b1e-a06d-918a2738ab69 · inbound
Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 044c5688-07a5-4fa4-8a99-bebeebdfb7e4 · inbound
Toward Training Superintelligent Software Agents through Self-Play SWE-RL R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation db63c098-f3df-4d55-b2d6-650361f6e571 · inbound
Toward Training Superintelligent Software Agents through Self-Play SWE-RL R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f181ed9d-8a1a-46cf-81c3-667feceb3c80 · inbound
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb056ad-3db9-455a-9ba4-619a857ff0e8 · inbound
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae2e50ca-c6aa-4988-bb67-c7def6526d50 · inbound
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation de9d955b-a5f3-4887-838a-7fde5a6f3931 · inbound
World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b99aa7a5-79d3-4d97-9221-2dfc48fc9928 · inbound
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c7cc9655-90f8-4692-a39b-64e4152b662a · inbound
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c144cfb-1d2a-4d57-8ae8-1ac2eff2f40a · inbound
ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision? R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 84f7baea-ea9a-4fb2-b508-457c9e3f8b59 · inbound
$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b37f44fe-de8e-4400-b5e9-1abd1a9f6aff · inbound
Steerable Instruction Following Coding Data Synthesis with Actor-Parametric Schema Co-Evolution R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 54e23510-0c45-4faf-b93e-b18794bb4407 · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fb71f621-e6c8-4b55-ae8b-dba0204d34ac · inbound
Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9f99459-8eb9-416b-800c-218163730cde · inbound
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a77039c5-391b-46f0-961b-c9bd3e6d4e0a · inbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 139dcdf1-28f7-49d8-b460-4252fe0181bc · inbound
Evaluation-driven Scaling for Scientific Discovery R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16ef9c3e-12c6-4542-bc13-dc746c5cc050 · inbound
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7fbb6323-6cbd-42f8-8566-5047837d1244 · inbound
Scaling Self-Play with Self-Guidance R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fdc49f2b-a5bf-455b-a401-5188d88780ef · inbound
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 960ef2ab-9b7b-4104-8d51-5c881c005fff · inbound
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b24d2392-11a6-4c58-b721-c98fc46a8d8b · inbound
SPARK: Self-Play with Asymmetric Reward from Knowledge Graphs R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 604e6e14-bd70-4c97-a5a5-01705bd5de45 · inbound
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7c4eb9bb-3fb9-4251-9a74-43b6899b805c · inbound
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3120205c-3e8b-4338-9a4a-72f079c2ab90 · inbound
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6c9dd457-c504-4db6-a2f5-5f4fe398d035 · inbound
SEIF: Self-Evolving Reinforcement Learning for Instruction Following R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f74eb4b1-c84d-4c0e-a2e8-47da77a058c5 · inbound
SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b6393cb0-1633-4ca8-b4b4-4b2b0f8da891 · inbound
SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 00663b6f-c180-4836-998f-9ced3017c357 · inbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eaaf08fe-e3dc-4698-9ce6-01f8d9783417 · inbound
Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 920ac4a3-4216-42b5-b86e-8fb209623b34 · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 00aac77a-2ba8-455f-9d51-171fd7248d4f · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 063ab454-8f56-48ee-af88-5b7bf9ce10d5 · inbound
Query-Conditioned Test-Time Self-Training for Large Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d87134a8-2484-40c5-9d8b-c6b62e397dac · inbound
Query-Conditioned Test-Time Self-Training for Large Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b43fa50d-6427-4a74-86e8-abc49f43c51c · inbound
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 824f0c7d-fdea-4f21-9a3d-5d062c17a7e9 · inbound
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d83bc4c-5dc6-41b2-bdc2-46c538c91a52 · inbound
PREPING: Building Agent Memory without Tasks R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e35b5a32-7a9b-4e6d-a465-a6fe29d55074 · inbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 216acf16-d54c-4aba-8334-00ae124539ab · inbound
Video-Zero: Self-Evolution Video Understanding R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95b23f76-a377-4462-9e73-9fcce8fc7a87 · inbound
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4beb94bf-65ed-49a4-9ccb-b64b56aa9530 · inbound
D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 669dc081-7464-44e6-ac44-502f597fbf8d · inbound
SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fbd7d8fb-b6e0-49cc-bcdb-2fc0dc63228a · inbound
RISE: Reliable Improvement in Self-Evolving Vision-Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ebaf286-a748-40bf-bb15-4ff6f62b0c77 · inbound
RISE: Reliable Improvement in Self-Evolving Vision-Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6b148deb-c6b4-41e5-9cfe-28af6f4d40dd · inbound
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6560bc44-3af0-4966-a5f8-ff2982c7974b · inbound
EVE-Agent: Evidence-Verifiable Self-Evolving Agents R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0a71aea-cdbb-47a8-b932-9d331bf073d4 · inbound
SEAL: Synergistic Co-Evolution of Agents and Learning Environments R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d3de43ee-4577-4c37-8e21-dff00f932eff · inbound
EvoRepair: Enhancing Vulnerability Repair Agents Through Experience-Based Self-Evolution R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a8cc42b5-a8b8-445a-be2a-df5579d060f6 · inbound
Trust Region On-Policy Distillation R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d1a69c57-a417-4519-9948-de4129734026 · inbound
BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9cc9bac4-8ee7-46e7-a7d2-dc24d144180f · inbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 695c6a50-058e-4f03-8d15-ee2a1f6d02fa · inbound
Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3b74462-1ca5-4b8f-9ce6-de9ed6b75ccb · inbound
Rethinking Continual Experience Internalization for Self-Evolving LLM Agents R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c90d88c-e40e-4fcb-9e07-285b7087ff05 · inbound
Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1095b9c6-65b2-46ca-8d40-2537da836808 · inbound
SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2cefef41-87fc-4502-8d1b-c1d94c72d276 · inbound
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fb9e5769-20cf-4bb0-9e43-29e2eb5fca8d · inbound
Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a11baa40-fae8-46b3-9765-17eafbc1ebf7 · inbound
Procedural Memory Distillation: Online Reflection for Self-Improving Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 807d88c5-6a1c-4eb1-9b2d-bbf29a9b926c · inbound
H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ea526f4-4baf-4b47-8ca6-1109d0908eb5 · inbound
Anchored Self-Play for Code Repair R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10135b71-3735-4e28-ad15-dad414f60591 · inbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e883d4-5917-4f10-98a4-694c69c58a7f · inbound
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3dc2e27f-1376-472d-9505-122daa81c3d4 · inbound
Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ebb69f22-ad95-466f-9db2-ab927f9b2f85 · inbound
From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fc197c8c-d538-4e0f-b32b-15e76bae5ea2 · inbound
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0f755cd-4170-43dd-ac7a-4442881dbc90 · inbound
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2273c5e5-9c33-4e6b-8e1b-f6694a665b93 · inbound
Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.