Pith. sign in

Paper Citation Record · LEDGER

R-Zero: Self-Evolving Reasoning LLM from Zero Data

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 75 inbound Pith citation observations for arXiv:2508.05004.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.05004 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 75 of 75 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:53:02.690293Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T18:17:33.697242Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7ffd83e3-7cfb-4c48-9a3d-49d730adcb66 · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:21:42.452862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:db46e3386efb4ba82bccb413debae532437825e631239e1e9421acd06f9e7f38

Observation 57a2e0cc-3889-4def-a8ff-f498295de558 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 171

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.062708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:83fc32dcebf310f74d3a108cca8e0cb3514bac9e5090dd55d78dff0e0b505667

Observation 796cf684-1dd7-49c3-b966-fc4b361d7b33 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 209

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:05:31.510709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:576f6b8bd442ff425b6087fcc71145262a386502d17a333bd3335c2cce4a5fe9

Observation edf8c7ba-09fc-4dea-9987-691406fe9548 · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.690293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.690293Z digest=sha256:525f9218a4d92b22ba4254266ad310b62f8b87a7730c20b3e3338172dfb8d448

Observation 57362e68-6424-4f58-8c70-6cc0d954ae4e · inbound

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean cites this paper.

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:11:35.906950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T16:09:39.975762Z digest=sha256:8abf8ba2d92020ab93260f53ddf96e4cb6016d8d074fdbf2146d1ed8dff523ee

Observation 474f9f6f-50e2-4e5d-8757-d82d4147407c · inbound

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean cites this paper.

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T16:40:59.805198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:40:59.805198Z digest=sha256:51b51563906d8ade73ac555e334e9f3f36065f03caa61d8fb4b6a524ed1f6289

Observation 932533c3-59a2-49ea-953d-ae24d7618eb6 · inbound

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL cites this paper.

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:30.568947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:44:30.568947Z digest=sha256:74ad799c80b6128428d7e73a99b415646284687fb299c346361909614797766d

Observation 91a79947-b1bb-4788-9a65-e5b761eee3d8 · inbound

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation cites this paper.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.613740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.613740Z digest=sha256:ee30af1bfe6f88116c9864bbf8e2e42a408c49f6eae79369d889d3b1b0ae70ec

Observation 741c5291-7bb4-4b1e-a06d-918a2738ab69 · inbound

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis cites this paper.

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:00:25.510812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T22:59:48.260121Z digest=sha256:5b56461e9cf44e7312490dc811547586cc4e3a36c086734b698dfdf1ac4203a4

Observation 044c5688-07a5-4fa4-8a99-bebeebdfb7e4 · inbound

Toward Training Superintelligent Software Agents through Self-Play SWE-RL cites this paper.

Toward Training Superintelligent Software Agents through Self-Play SWE-RL R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:10:20.187627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T16:07:48.570995Z digest=sha256:80c005a33e4c1850147d59bf0dd9c2a4fd38cd9877d47af49d46f48eac7b39a4

Observation db63c098-f3df-4d55-b2d6-650361f6e571 · inbound

Toward Training Superintelligent Software Agents through Self-Play SWE-RL cites this paper.

Toward Training Superintelligent Software Agents through Self-Play SWE-RL R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T15:02:11.420780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:02:11.420780Z digest=sha256:356a562c00077916c9f2a2b1b09ef5a8fb05756e1f8b6567136151dbdb6be3d9

Observation f181ed9d-8a1a-46cf-81c3-667feceb3c80 · inbound

Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability cites this paper.

Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T07:59:09.641194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:59:09.641194Z digest=sha256:48305986d90a7661f8b40d336f6f874d5ce58911d2c233e6f15b76908d2de8a9

Observation 5fb056ad-3db9-455a-9ba4-619a857ff0e8 · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:18.752065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:18.752065Z digest=sha256:95f79ca796762ad5f66159ed671de52c01d1f99ac189ecfe4c867255992c0c37

Observation ae2e50ca-c6aa-4988-bb67-c7def6526d50 · inbound

SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents cites this paper.

SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:16:27.869604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T11:14:47.943242Z digest=sha256:812497e84d8a844330d2440f864a02941c5ddc20f3baecb1540f72bdba39cdf5

Observation de9d955b-a5f3-4887-838a-7fde5a6f3931 · inbound

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry cites this paper.

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T14:05:26.303000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:05:26.303000Z digest=sha256:163af29258b11c71db5e287e259de8680629ebaac79fbed08905c59921759dcd

Observation b99aa7a5-79d3-4d97-9221-2dfc48fc9928 · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T19:23:10.901441Z digest=sha256:cf776be640cb145b90df849e05d83c0edc5569e23fefd3c1e4a1ab6bfac8c320

Observation c7cc9655-90f8-4692-a39b-64e4152b662a · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T16:52:27.075292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:52:27.075292Z digest=sha256:5b129b3260dcb4b6da8fd53f6c194efe08082438cd3b1c62785c1263bea4e7d6

Observation 4c144cfb-1d2a-4d57-8ae8-1ac2eff2f40a · inbound

ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision? cites this paper.

ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision? R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:36:23.127141Z digest=sha256:07c9e80958ba02cafd1b240dcb0607c03beca98a38db2754009b038823e5f88e

Observation 84f7baea-ea9a-4fb2-b508-457c9e3f8b59 · inbound

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data cites this paper.

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:43:09.581054Z digest=sha256:12e79a0fc5be6869789f648835459565fa6c8345f5606340522dec3e5f173e12

Observation b37f44fe-de8e-4400-b5e9-1abd1a9f6aff · inbound

Steerable Instruction Following Coding Data Synthesis with Actor-Parametric Schema Co-Evolution cites this paper.

Steerable Instruction Following Coding Data Synthesis with Actor-Parametric Schema Co-Evolution R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.355541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T18:00:39.404711Z digest=sha256:6a9bb456263f01d7af555779c8ed206b78c63126def012550eeb44266da3e6f5

Observation 54e23510-0c45-4faf-b93e-b18794bb4407 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:d6fa5bc29d51d298ba96e1948550ddbbae06b87187dd4d208174ce00f838f829

Observation fb71f621-e6c8-4b55-ae8b-dba0204d34ac · inbound

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration cites this paper.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:17ecf415ac0fc0502b57b4ca3b420579b7595176e5c1c48e80f75d88feedc2c9

Observation b9f99459-8eb9-416b-800c-218163730cde · inbound

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations cites this paper.

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T05:23:45.814042Z digest=sha256:44e29bfcabadf5518b21111cd198c6413e511cd3319468bdac3af037b7c64345

Observation a77039c5-391b-46f0-961b-c9bd3e6d4e0a · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:2c99d6b5e1377d5880b9a39251d835ed54d58dd3df9440d8e2427df002349565

Observation 139dcdf1-28f7-49d8-b460-4252fe0181bc · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:c8883e8f5480932db565497db686aeb8dfea44c4374eac25df640b6ce21210c5

Observation 16ef9c3e-12c6-4542-bc13-dc746c5cc050 · inbound

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text cites this paper.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:74bfc81401ca2caf41dbe5cbe5594300d070fb34b6e53ec9c292ee926c713fac

Observation 7fbb6323-6cbd-42f8-8566-5047837d1244 · inbound

Scaling Self-Play with Self-Guidance cites this paper.

Scaling Self-Play with Self-Guidance R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T01:31:06.090698Z digest=sha256:3e61f3054462e53db8c3ee197a5ebafb055f87952ed863dd6a6df40257ea2b81

Observation fdc49f2b-a5bf-455b-a401-5188d88780ef · inbound

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks cites this paper.

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:14:07.017420Z digest=sha256:c64f48b05b1ef14d902d818aabde6504cd2f3b051894713ecc173c3c6b38d847

Observation 960ef2ab-9b7b-4104-8d51-5c881c005fff · inbound

Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora cites this paper.

Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:13:30.648190Z digest=sha256:385d8146723f4fa590f8b60a3ac009285eaab82d108789fb47b0aa532156c9ec

Observation b24d2392-11a6-4c58-b721-c98fc46a8d8b · inbound

SPARK: Self-Play with Asymmetric Reward from Knowledge Graphs cites this paper.

SPARK: Self-Play with Asymmetric Reward from Knowledge Graphs R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T12:12:33.179209Z digest=sha256:870e19d1789e5ad91fdfe69796209575f183fdbbe2e3d967da3a0bdc5cd28f79

Observation 604e6e14-bd70-4c97-a5a5-01705bd5de45 · inbound

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key cites this paper.

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T09:35:47.501360Z digest=sha256:02d2d5b5594f340d1a7d0c247978a2b1ccc1a874cb5d40e01f185710bee84191

Observation 7c4eb9bb-3fb9-4251-9a74-43b6899b805c · inbound

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key cites this paper.

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T03:16:59.195706Z digest=sha256:dbce8d2104d7f55d41545850f1076b939951742ed20c93530a96df796203b3c7

Observation 3120205c-3e8b-4338-9a4a-72f079c2ab90 · inbound

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key cites this paper.

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:39:10.263076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T22:36:13.781114Z digest=sha256:8832935cb5b918e26ea383aa52661319714dc6da4887bf07d1633a52a4d8afbc

Observation 6c9dd457-c504-4db6-a2f5-5f4fe398d035 · inbound

SEIF: Self-Evolving Reinforcement Learning for Instruction Following cites this paper.

SEIF: Self-Evolving Reinforcement Learning for Instruction Following R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:51:09.514927Z digest=sha256:3cdb20871f6f5724c9a9f4e91d700992fe12ee1d32b5f6fe1362aa7bfc20b011

Observation f74eb4b1-c84d-4c0e-a2e8-47da77a058c5 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:21:44.087943Z digest=sha256:e5599b25be00f8b983192e35a84bf072a2b7eb538b293aa702dd02c5050cff6b

Observation b6393cb0-1633-4ca8-b4b4-4b2b0f8da891 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.661454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:30:42.766390Z digest=sha256:b042ace9c99125289f6b9c50bf8d46e24b2261a6f4e767aeb605adb2a819d3af

Observation 00663b6f-c180-4836-998f-9ced3017c357 · inbound

G-Zero: Self-Play for Open-Ended Generation from Zero Data cites this paper.

G-Zero: Self-Play for Open-Ended Generation from Zero Data R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:0a746ecc738f16244849dbf2bcef9712e21e0e5b4eb5880ce5ce2d848983927d

Observation eaaf08fe-e3dc-4698-9ce6-01f8d9783417 · inbound

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning cites this paper.

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:19:49.761472Z digest=sha256:a3f4ff3dec2879b0f5ca10ecdaa4d6b27d5e0cab8b261cf4231571bfb34dc8d8

Observation 920ac4a3-4216-42b5-b86e-8fb209623b34 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:c11c923eee34ca18255ce7f676b23e4b95aa831f883166de4d5ce7b8b9012a1b

Observation 00aac77a-2ba8-455f-9d51-171fd7248d4f · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:58.803288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:31dad19e6f5a6c612f17f173acf9e7b55e2964b86df06da4765eba7df8176c4f

Observation 063ab454-8f56-48ee-af88-5b7bf9ce10d5 · inbound

Query-Conditioned Test-Time Self-Training for Large Language Models cites this paper.

Query-Conditioned Test-Time Self-Training for Large Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:07:54.177291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:04:51.248797Z digest=sha256:e8547fe5541c1b594a789f947e9405cf68a7a07d22153e4116fcf35aaee37721

Observation d87134a8-2484-40c5-9d8b-c6b62e397dac · inbound

Query-Conditioned Test-Time Self-Training for Large Language Models cites this paper.

Query-Conditioned Test-Time Self-Training for Large Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:55:04.920434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T05:53:42.235498Z digest=sha256:7bc10bb374ccbe467237b6d2d2198e4eefae35b5529aabfe4bcf67098436c6eb

Observation b43fa50d-6427-4a74-86e8-abc49f43c51c · inbound

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data cites this paper.

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T17:54:50.325820Z digest=sha256:3a97ebe3c8b0548340ab1d7e8b0d7a6f51a693d3fb38d908d1185fbf0258c9c6

Observation 824f0c7d-fdea-4f21-9a3d-5d062c17a7e9 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:32:52.558475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:8a107e0b1eca1fd88fca618000a4282e8c4a9230f9e97bd5480de461e13260ee

Observation 0d83bc4c-5dc6-41b2-bdc2-46c538c91a52 · inbound

PREPING: Building Agent Memory without Tasks cites this paper.

PREPING: Building Agent Memory without Tasks R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:15:06.068820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T06:14:14.385586Z digest=sha256:6767ff830826ef38767dc4d3cebb27139cc611ac41e97432f021ac392c915134

Observation e35b5a32-7a9b-4e6d-a465-a6fe29d55074 · inbound

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis cites this paper.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.695285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:8561d8ea6ac51abc27807b05edaa1cf2b5fce5ab79a1b53d1bf26b8cf4035879

Observation 216acf16-d54c-4aba-8334-00ae124539ab · inbound

Video-Zero: Self-Evolution Video Understanding cites this paper.

Video-Zero: Self-Evolution Video Understanding R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.459746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T21:32:16.939563Z digest=sha256:8bb2205cbae8da3fbf1c8b66df7248e352bbc530a8542146f882b965caba1393

Observation 95b23f76-a377-4462-9e73-9fcce8fc7a87 · inbound

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play cites this paper.

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T21:42:48.064691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T21:37:56.570173Z digest=sha256:5481887e378abc9728ae74df33ca367cfc584bc3608d346ccb1dff922ef20d08

Observation 4beb94bf-65ed-49a4-9ccb-b64b56aa9530 · inbound

D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning cites this paper.

D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T20:22:45.321883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T20:21:48.657926Z digest=sha256:40cbf039f37b79c080528e727409930a363e5e8c6465e529de0d722126d5f3d0

Observation 669dc081-7464-44e6-ac44-502f597fbf8d · inbound

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation cites this paper.

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:24:08.702410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T11:21:30.867480Z digest=sha256:e30840ebda2d1bf45170eecc863fa2a71b2796b361560afab76423c996326a13

Observation fbd7d8fb-b6e0-49cc-bcdb-2fc0dc63228a · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:39:40.526852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T05:38:26.590720Z digest=sha256:2437283457e64be5ef8340f3ff377932971bba76598761c89e0c33384186e15a

Observation 8ebaf286-a748-40bf-bb15-4ff6f62b0c77 · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:24:57.539081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T17:19:09.596950Z digest=sha256:1b57ab1a24a72ac4832bca2315d7bf5bb0a62987e906f3d0027091925cb2b5b9

Observation 6b148deb-c6b4-41e5-9cfe-28af6f4d40dd · inbound

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models cites this paper.

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:21:12.892617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T07:19:30.508843Z digest=sha256:88325bf1eef0cd472011d3df585b63f73dc3ec27f8b244e769c5662b2ccce9b6

Observation 6560bc44-3af0-4966-a5f8-ff2982c7974b · inbound

EVE-Agent: Evidence-Verifiable Self-Evolving Agents cites this paper.

EVE-Agent: Evidence-Verifiable Self-Evolving Agents R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:45:23.884847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T05:44:42.464591Z digest=sha256:44020324e9644d8c7b4f55a0b5c7fa0a44c6d330566fa0b2e50261ad24b9891d

Observation e0a71aea-cdbb-47a8-b932-9d331bf073d4 · inbound

SEAL: Synergistic Co-Evolution of Agents and Learning Environments cites this paper.

SEAL: Synergistic Co-Evolution of Agents and Learning Environments R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:44:40.924437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T13:38:10.466713Z digest=sha256:6343e7e01f1cdc30f51ac8deab3620f5aac7fbd9c542c022e131698f02d2d891

Observation d3de43ee-4577-4c37-8e21-dff00f932eff · inbound

EvoRepair: Enhancing Vulnerability Repair Agents Through Experience-Based Self-Evolution cites this paper.

EvoRepair: Enhancing Vulnerability Repair Agents Through Experience-Based Self-Evolution R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:43:31.661411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T06:16:51.531156Z digest=sha256:753aae567de0a9f9cbfec85223c9e9e8d141350fd9bc9494fa83614e1fd9a423

Observation a8cc42b5-a8b8-445a-be2a-df5579d060f6 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.592288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:4d3dbbe05948e0862bf830d69e9334067cb436b538934d6cc9f94dfd42f07113

Observation d1a69c57-a417-4519-9948-de4129734026 · inbound

BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution cites this paper.

BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:36:14.649137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T16:49:50.653594Z digest=sha256:79a2f8368e86665d011362e0abee804afb195ef388cf46f67be043ef419235b6

Observation 9cc9bac4-8ee7-46e7-a7d2-dc24d144180f · inbound

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning cites this paper.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.137790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:961a365cfe8b8b4960bfaddef45b45f5a3d47ff04b4a92654deecdda2341da67

Observation 695c6a50-058e-4f03-8d15-ee2a1f6d02fa · inbound

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification cites this paper.

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:26:27.250046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T10:53:00.223228Z digest=sha256:57758b3f636eb7e4d9d03ceda4e6ec3dd79e6546dbbce69ea80de7b95715bcfe

Observation b3b74462-1ca5-4b8f-9ce6-de9ed6b75ccb · inbound

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents cites this paper.

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T07:56:47.775685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T06:29:11.398007Z digest=sha256:0e9a2366cbf054d1fbcfdb891d341135fb7a7a17a5f733a6587041d658a10ffc

Observation 2c90d88c-e40e-4fcb-9e07-285b7087ff05 · inbound

Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems cites this paper.

Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:06:59.488782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T01:33:39.607334Z digest=sha256:2055fbecd6fa9e11578b34ec3336fdab41f3148cf70ff0dc002128cb6ee6575d

Observation 1095b9c6-65b2-46ca-8d40-2537da836808 · inbound

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents cites this paper.

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:58:33.224115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T06:49:12.070481Z digest=sha256:e31b8ffb4428a9f0ec891300f781fa7f7928517bd496e8d7c49cbec95ffc3657

Observation 2cefef41-87fc-4502-8d1b-c1d94c72d276 · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 108

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:58:58.300001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:04ce56c800adc14e434f1c5421d8787c7c1a37277db9fd6609afd3679a04f3b7

Observation fb9e5769-20cf-4bb0-9e43-29e2eb5fca8d · inbound

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model cites this paper.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.309663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:2539cfcdf9c4690d1a5e49ee44e3bad56697be32f2ea77116ca66fe659048c0c

Observation a11baa40-fae8-46b3-9765-17eafbc1ebf7 · inbound

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models cites this paper.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.024959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:1207d24ab42119d9fb3245b615ecc30da448ae32d9dabb05e5d9db1a50e29be3

Observation 807d88c5-6a1c-4eb1-9b2d-bbf29a9b926c · inbound

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation cites this paper.

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T09:29:14.105502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T09:29:14.105502Z digest=sha256:08cb9743deb5a177099ac97cbac47d27dfd65756401f9b57c0875268eaf88ca8

Observation 4ea526f4-4baf-4b47-8ca6-1109d0908eb5 · inbound

Anchored Self-Play for Code Repair cites this paper.

Anchored Self-Play for Code Repair R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T01:53:44.674517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:53:44.674517Z digest=sha256:51925c60262a742ba0bb12ffbd13417be806364416bd88ed3796a631d5a3bc5b

Observation 10135b71-3735-4e28-ad15-dad414f60591 · inbound

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL cites this paper.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:151049853d216d74b7c06d1416998da793ed7850829c42025a98b613536582d6

Observation d1e883d4-5917-4f10-98a4-694c69c58a7f · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.692092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:2b89736055ba037c26f00eeca1c4f0edefc9671f795690dca6162e82e7b4be49

Observation 3dc2e27f-1376-472d-9505-122daa81c3d4 · inbound

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning cites this paper.

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:35:53.849536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T02:29:12.366018Z digest=sha256:ac0d42d2a6783e09c8e89957a9605dbfa840bc53dbb490f861a7e7fa04351c18

Observation ebb69f22-ad95-466f-9db2-ab927f9b2f85 · inbound

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier cites this paper.

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:17:33.698445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T18:16:31.176239Z digest=sha256:d2650514e71979e84a9572b66b1f16422663cc55c7c7707e0b300c46865e1d20

Observation fc197c8c-d538-4e0f-b32b-15e76bae5ea2 · inbound

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning cites this paper.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.890871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.890871Z digest=sha256:052bcfbabf6e087b2f9a0c4dda366ab596671b69321d928e68e4abeddf659bd4

Observation d0f755cd-4170-43dd-ac7a-4442881dbc90 · inbound

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills cites this paper.

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T04:28:45.732564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:28:45.732564Z digest=sha256:cd13ade0eb154fcabf97ed3ea2d733431f20cf558a807156efa20a8fe5dae87a

Observation 2273c5e5-9c33-4e6b-8e1b-f6694a665b93 · inbound

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering cites this paper.

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T01:39:47.105486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:39:47.105486Z digest=sha256:e9800efa862ffc4bc09c836d481dd30579b9d9816deeb7c4f78cab12b487283b