Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T03:39:40.780801Z
Paper Citation Record · LEDGER
As of 2 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2605.09959.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T03:39:40.780801Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-30T18:33:28.493350Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T03:45:55.708592Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a0bca19f-eea2-4e01-97e0-3a4d26c32514 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9fbd0733-d719-4076-b473-b0e906b14996 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 31459089-a7f6-43eb-9146-0701db6d5c45 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c16415e0-110b-4953-8524-66b384c9247a · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Serl: Self-play reinforcement learning for large language models with limited data
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 831c6896-fa2f-4a9c-89d4-62112c5abc83 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c0b57059-355f-42d3-afb8-5e498ca61b36 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data A survey on LLM-as-a-judge.The Innovation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8d97dbb2-290a-45bd-ae91-22127eab1f0e · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5f35d001-1f42-49c9-9a27-df899b525c8d · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Visplay: Self-evolving vision-language models from images
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7422b71b-1af5-43aa-8217-53cd58313dae · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Lora: Low-rank adaptation of large language models.Iclr, 1(2):3
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 00663b6f-c180-4836-998f-9ced3017c357 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b41f7614-9b3e-4d97-a544-1e7f5cd34ca3 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Large language models can self-improve
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 57d026c8-67e1-4277-af20-7b32c25b4dfe · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Likelihood- based reward designs for general llm reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 28de626f-df92-4a20-975f-61ed2752f138 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Mm-zero: Self-evolving multi-model vision language models from zero data
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 16caeec2-af6f-44dd-a223-a6cb3ee161ec · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0d67e572-25c0-4647-95ca-7e215ba60fe2 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Learning to solve and verify: A self-play framework for code and test generation.arXiv preprint arXiv:2502.14948
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 87a2e22c-a01c-4e04-affe-7eec69fa1f5e · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Spice: Self-play in corpus environments improves reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1f0d9409-00c5-41f6-9c9b-641cfa2b5484 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Mmc: Advancing multimodal chart understanding with large-scale instruction tuning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation fbe44300-a37b-428e-8437-8a4214bfca15 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Nover: Incentive training for language models via verifier-free reinforcement learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bf0ff9b5-e832-49a3-9024-be7fed9b2648 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Efficient paths and dense rewards: Probabilistic flow reasoning for large language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 83dfbca5-754e-4658-bc09-993207f89f4a · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Understanding R1-Zero-Like Training: A Critical Perspective
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8f05cf74-1770-4e58-ab9c-1663b97e08a5 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Direct preference optimization: Your language model is secretly a reward model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5514be62-3bec-4cb7-9ecf-8b7dc3c10321 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Can large reasoning models self-train?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 27103efc-20fa-49ba-9545-5f22c98f9388 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8ca4eef0-e1bd-4d8f-92e3-b750f6589665 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a998ee1d-6dcd-49e0-ada2-e292fee87927 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Ai models collapse when trained on recursively generated data.Nature, 631(8022): 755–759
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5a81e153-6fde-4b76-89e0-dd9a1e24c30f · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Large language models for data annotation and synthesis: A survey
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 73f7c354-7309-4110-bfb0-f21c8baf3e9d · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data A Survey on Self-Evolution of Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation cc76e99e-c79d-4848-a707-cde085ffddd0 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8494457b-3874-4abb-8e43-6112c20455ef · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 732e4e6d-91f2-4048-8869-e1e5ea745ace · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Associated with the WaltonFuture GeoQA-8K-direct-synthesizing dataset release
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b2592111-9a7d-4043-969d-5886727b9b42 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data A systematic survey of self-evolving agents: From model-centric to environment-driven co-evolution
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 140089a8-92bd-4ec4-829c-5d4be0611c28 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Reinforcement learning with conditional expectation reward
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ee43180b-99cc-437e-894f-55660022b433 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Qwen3 Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 287766b0-65f8-4c3b-9523-f4e7edd5de08 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data A survey on recent advances in LLM-based multi-turn dialogue systems.ACM Computing Surveys, 58(6):1–38
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 85edc4c9-1f3d-4cd6-af42-7efe2ab62f6f · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 23f2d637-7416-46f1-805d-7293ba346d2e · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Guided self-evolving llms with minimal human supervision
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation cbea89e8-be55-4407-a99a-d11bd8775cdb · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4b644172-508e-49e9-a4a0-88c20667348c · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Instruction-Following Evaluation for Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e01d6ce9-375f-44d7-b0c2-6248af681050 · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Reinforcing General Reasoning without Verifiers
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d6a93fdf-ad16-42ef-be51-5363fbcb7baa · outbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Self-Challenging Language Model Agents
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8e636395-6e99-48a7-bad2-df97c3671a52 · inbound
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops G-Zero: Self-Play for Open-Ended Generation from Zero Data
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 110f500a-8414-4793-aa78-b431ad1af83f · inbound
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning G-Zero: Self-Play for Open-Ended Generation from Zero Data
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.