Pith. sign in

Paper Citation Record · LEDGER

G-Zero: Self-Play for Open-Ended Generation from Zero Data

As of 2 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2605.09959.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09959 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:39:40.780801Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T18:33:28.493350Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:45:55.708592Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact28
  • verified fuzzy11
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0bca19f-eea2-4e01-97e0-3a4d26c32514 · outbound

This paper cites SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.388622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:793f36583db959fd1b3d004edef5bd6397172c843e1562a96501d8a545a926f5

Observation 9fbd0733-d719-4076-b473-b0e906b14996 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:00:22.024625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:9118330601c396b281271d67b5cde565624bc89ed709462dc184377cd9a1a421

Observation 31459089-a7f6-43eb-9146-0701db6d5c45 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:13:05.519270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:f454bf92c39f0e05883794439fb3979e5f6f362bc96aeb79de60ac740c1703eb

Observation c16415e0-110b-4953-8524-66b384c9247a · outbound

This paper cites Serl: Self-play reinforcement learning for large language models with limited data.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Serl: Self-play reinforcement learning for large language models with limited data

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.384121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:4df560102423a1d1c7c49e22d172549699b5010a8f2ef49a44cdb153b08b8237

Observation 831c6896-fa2f-4a9c-89d4-62112c5abc83 · outbound

This paper cites The Llama 3 Herd of Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data The Llama 3 Herd of Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.398856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:d8b47b00b8938467685797cff0452ba86e5be0fdfd5cc11847a6178e18e107df

Observation c0b57059-355f-42d3-afb8-5e498ca61b36 · outbound

This paper cites A survey on LLM-as-a-judge.The Innovation.

G-Zero: Self-Play for Open-Ended Generation from Zero Data A survey on LLM-as-a-judge.The Innovation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.687577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:7e5e5a9d05ab3e30cd287b2ae2da44ed2acbf593cf31b094cce626a4fdb9cd91

Observation 8d97dbb2-290a-45bd-ae91-22127eab1f0e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.407524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:c3ecbdf56e1d8572866b0bedbb2b805dc9e0cc1822b23f0cd0e218b39a27df5e

Observation 5f35d001-1f42-49c9-9a27-df899b525c8d · outbound

This paper cites Visplay: Self-evolving vision-language models from images.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Visplay: Self-evolving vision-language models from images

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.411974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:6d2364241702c57bd21a450e650b97dfb9adc8edb0ff287b9d37257e1fbe26be

Observation 7422b71b-1af5-43aa-8217-53cd58313dae · outbound

This paper cites Lora: Low-rank adaptation of large language models.Iclr, 1(2):3.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Lora: Low-rank adaptation of large language models.Iclr, 1(2):3

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.684032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:77deee2822586fbe4b6fb19e3335cce361578321ad62af558969a1eb471daf9b

Observation 00663b6f-c180-4836-998f-9ced3017c357 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

G-Zero: Self-Play for Open-Ended Generation from Zero Data R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:a33195859fb040eb3db93f28498e26066733f8209eb01ee1eda40abfb8e4adc0

Observation b41f7614-9b3e-4d97-a544-1e7f5cd34ca3 · outbound

This paper cites Large language models can self-improve.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Large language models can self-improve

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.675439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:99d45aedf5816608ed8cdd4c34c2743bea316a18488212ea38aec953582754c4

Observation 57d026c8-67e1-4277-af20-7b32c25b4dfe · outbound

This paper cites Likelihood- based reward designs for general llm reasoning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Likelihood- based reward designs for general llm reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.504310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:2db50ab0091b983ac705e4fb43ec2454d1a3ae40cc769a45d883b84159f95d14

Observation 28de626f-df92-4a20-975f-61ed2752f138 · outbound

This paper cites Mm-zero: Self-evolving multi-model vision language models from zero data.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Mm-zero: Self-evolving multi-model vision language models from zero data

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.486410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:1b2e661f8c01e8987b4450d977687a1c95633ce6b50a82fef431cf080a9ac2f3

Observation 16caeec2-af6f-44dd-a223-a6cb3ee161ec · outbound

This paper cites SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.431684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:8b6ef346dbf5a86a7066c4bc42dd3aca178c8585e9a8ce59ed7b15d6611135c1

Observation 0d67e572-25c0-4647-95ca-7e215ba60fe2 · outbound

This paper cites Learning to solve and verify: A self-play framework for code and test generation.arXiv preprint arXiv:2502.14948.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Learning to solve and verify: A self-play framework for code and test generation.arXiv preprint arXiv:2502.14948

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.473026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:edc03a522f82a710b76159f940d1466231d14569b6fd8940cafc8455ae3d54b5

Observation 87a2e22c-a01c-4e04-affe-7eec69fa1f5e · outbound

This paper cites Spice: Self-play in corpus environments improves reasoning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Spice: Self-play in corpus environments improves reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.453062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:03df093341293bc5e22382f2d554153c3ade0306c71bf012237feca43305f3ec

Observation 1f0d9409-00c5-41f6-9c9b-641cfa2b5484 · outbound

This paper cites Mmc: Advancing multimodal chart understanding with large-scale instruction tuning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Mmc: Advancing multimodal chart understanding with large-scale instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.666766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:57e71318f61a5f55379f6b87724211ba0ab82ad829d7285ea8e0bc8536e6090a

Observation fbe44300-a37b-428e-8437-8a4214bfca15 · outbound

This paper cites Nover: Incentive training for language models via verifier-free reinforcement learning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Nover: Incentive training for language models via verifier-free reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.670797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:a5bd96a8c9d01ae1c278d333b37a420496ba17e81f4012590979d7f50b36ebde

Observation bf0ff9b5-e832-49a3-9024-be7fed9b2648 · outbound

This paper cites Efficient paths and dense rewards: Probabilistic flow reasoning for large language models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Efficient paths and dense rewards: Probabilistic flow reasoning for large language models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.511550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:6f9d287c77573e7248778a14f8b662f13ca8fe8e6a0428047aa7924955cb67f3

Observation 83dfbca5-754e-4658-bc09-993207f89f4a · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.553716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:40c427952d51568ba4111450d93a43f9a8e72bd92d03a4ccbeda24fc3ed5733c

Observation 8f05cf74-1770-4e58-ab9c-1663b97e08a5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Direct preference optimization: Your language model is secretly a reward model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.660181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:e68dc915eb1e67a56a138484baa2c1d2b9f21799183ce258b7c59d4e6b08163f

Observation 5514be62-3bec-4cb7-9ecf-8b7dc3c10321 · outbound

This paper cites Can large reasoning models self-train?.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Can large reasoning models self-train?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.464065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:9074407c9108b06a6b87a6e1208eca6a393db85caf322dff6c79017b4c296f50

Observation 27103efc-20fa-49ba-9545-5f22c98f9388 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.437860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:8deb270a448292b1f43ae2ba34c5080ed56c5ccffcd4925a769881d9449d0c55

Observation 8ca4eef0-e1bd-4d8f-92e3-b750f6589665 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:55a22ac258a371c2913bb929b0d6b24f7ecc0f6ab9b2075c75330f470424591b

Observation a998ee1d-6dcd-49e0-ada2-e292fee87927 · outbound

This paper cites Ai models collapse when trained on recursively generated data.Nature, 631(8022): 755–759.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Ai models collapse when trained on recursively generated data.Nature, 631(8022): 755–759

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.680122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:0a513269b58fc506f9071d1460ac0ecef416d06e298676544a507fb366b97778

Observation 5a81e153-6fde-4b76-89e0-dd9a1e24c30f · outbound

This paper cites Large language models for data annotation and synthesis: A survey.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Large language models for data annotation and synthesis: A survey

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.655845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:ab2d1fa366da958670c10c90ec152fdb705bcc96f6b900f02a5d9fb2feccd28b

Observation 73f7c354-7309-4110-bfb0-f21c8baf3e9d · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data A Survey on Self-Evolution of Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.492516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:f5afd875265795a92f292744cfbd0585eacd0f39eda0cca2cf48c9eb8e4ce646

Observation cc76e99e-c79d-4848-a707-cde085ffddd0 · outbound

This paper cites Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.516462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:bc4ec900b2cf94f6d6e5acd29e7713d6853191447f78ed8c8be43372558a132b

Observation 8494457b-3874-4abb-8e43-6112c20455ef · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.643799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:c679da125d4e2b771b9c4b05744bdec9ae4b7b1e1dd0c43b8bf19baf0b8de5a4

Observation 732e4e6d-91f2-4048-8869-e1e5ea745ace · outbound

This paper cites Associated with the WaltonFuture GeoQA-8K-direct-synthesizing dataset release.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Associated with the WaltonFuture GeoQA-8K-direct-synthesizing dataset release

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:24.548729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:e2cb82b16a5725f138c83c9def4d77b6d809259682d63984e944388fbc1f26f5

Observation b2592111-9a7d-4043-969d-5886727b9b42 · outbound

This paper cites A systematic survey of self-evolving agents: From model-centric to environment-driven co-evolution.

G-Zero: Self-Play for Open-Ended Generation from Zero Data A systematic survey of self-evolving agents: From model-centric to environment-driven co-evolution

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.647847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:be2bbe288291a8e6f97a2084fc88ac2e2e11e916982e65acbae28cfc89ea79de

Observation 140089a8-92bd-4ec4-829c-5d4be0611c28 · outbound

This paper cites Reinforcement learning with conditional expectation reward.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Reinforcement learning with conditional expectation reward

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.534561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:533f4ec2f51407a047384854e01f3a4fef4ab60869b2a79dc3c2db3bd77850d2

Observation ee43180b-99cc-437e-894f-55660022b433 · outbound

This paper cites Qwen3 Technical Report.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Qwen3 Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.477601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:a91921d5148845c71549fbcec912f1d6b353b8093fcb06e06dcff69775ed20a7

Observation 287766b0-65f8-4c3b-9523-f4e7edd5de08 · outbound

This paper cites A survey on recent advances in LLM-based multi-turn dialogue systems.ACM Computing Surveys, 58(6):1–38.

G-Zero: Self-Play for Open-Ended Generation from Zero Data A survey on recent advances in LLM-based multi-turn dialogue systems.ACM Computing Surveys, 58(6):1–38

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.651818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:37063bff23c4d76b28356ae5884bd49dfda7b98013bec08b37e7a4bede57d91c

Observation 85edc4c9-1f3d-4cd6-af42-7efe2ab62f6f · outbound

This paper cites RLPR: Extrapolating RLVR to General Domains without Verifiers.

G-Zero: Self-Play for Open-Ended Generation from Zero Data RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.558508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:3e29c746fe5a5848c9675bd27aa0c09f1ccddcc6cf55e408319f9e02deb1ef5b

Observation 23f2d637-7416-46f1-805d-7293ba346d2e · outbound

This paper cites Guided self-evolving llms with minimal human supervision.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Guided self-evolving llms with minimal human supervision

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.457773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:9836349b103f4d40c7d67aa2a074d838aa2b017fef38d5b75df7e3ffcb15f9af

Observation cbea89e8-be55-4407-a99a-d11bd8775cdb · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:23:09.397455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:c2aeb40e420995dd26472d53391615cb48e99404509353bb6e7a139cb970f2c3

Observation 4b644172-508e-49e9-a4a0-88c20667348c · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Instruction-Following Evaluation for Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.444360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:2c8d710f8505462d2ba78ea54e545a4e4b717ce781504744c7943b76e9ca48ee

Observation e01d6ce9-375f-44d7-b0c2-6248af681050 · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Reinforcing General Reasoning without Verifiers

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.522825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:244af752f064b7febf29c047479384c1a61aa380ea0aa1e244c789dd5dacda0a

Observation d6a93fdf-ad16-42ef-be51-5363fbcb7baa · outbound

This paper cites Self-Challenging Language Model Agents.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Self-Challenging Language Model Agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.540151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:484cb6daa0d8aa8f8aa4df0e08365ebba9d709bc5420200d2223c3e5819753bf

Pith citing papers

Observation 8e636395-6e99-48a7-bad2-df97c3671a52 · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops G-Zero: Self-Play for Open-Ended Generation from Zero Data

Reference 128

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.709795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:8a871d98307c50883ef00e888d6f4006b0dac5abc35a6ae849c5e1eeefd47f64

Observation 110f500a-8414-4793-aa78-b431ad1af83f · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning G-Zero: Self-Play for Open-Ended Generation from Zero Data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-30T18:33:28.493350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:33:28.493350Z digest=sha256:e9ee05d4895a245cd8274f217f125a5266790c98892970e5bac26b42b5028fe1