Pith. sign in

Paper Citation Record · LEDGER

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis

As of 4 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2605.14392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.14392 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T20:58:07.921910Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:53:20.110535Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact22
  • verified fuzzy35
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6229edd4-ab9c-4f3b-b013-05b86abd056b · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.699251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:218124cc367c8752f7ed704f4ea92d7fef1a9b9df09a756aca16b221e1b22450

Observation 76173681-2e83-493c-b411-d7a9b04c8d2b · outbound

This paper cites Safe and scalable web agent learning via recreated websites,.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Safe and scalable web agent learning via recreated websites,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.323202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:2cd7ef739b9e74a885b0671ecb48007073565b37dbb5c88ddd75618e66a144f8

Observation 534311ea-f701-4f01-91e4-2f631e4e18e8 · outbound

This paper cites Safe and scalable web agent learning via recreated websites.arXiv preprint arXiv:2603.10505,.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Safe and scalable web agent learning via recreated websites.arXiv preprint arXiv:2603.10505,

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.678952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:c00402200792b47a0ed1e788ef43318f121bc71d7a207237da89c922ceb8f992

Observation 7271659c-099f-44de-8eb9-84f76207546c · outbound

This paper cites Spc: Evolving self-play critic via adversarial games for llm reasoning, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Spc: Evolving self-play critic via adversarial games for llm reasoning, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.301211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:d9dc6a2d0b846862a67be38c39d3bf245bc8ab6221ff3efab8ba6c3aa7e390d5

Observation ebc4e32f-3ffc-4138-8d81-064b0023cc21 · outbound

This paper cites Self-questioning language models, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Self-questioning language models, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.322176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:cf1f1d426ad56ee62077c30025bda0a79bbd477dcdbacb5ad34bdc207f182325

Observation 708291ee-4cbd-4a10-8e8f-035d81cdd489 · outbound

This paper cites Multi-agent evolve: Llm self-improve through co-evolution, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Multi-agent evolve: Llm self-improve through co-evolution, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.328834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:669dd37a346659b6f35434dfb9386b10abd64709f26399e276c974716a6d0859

Observation 85e338d2-36fa-418b-8ea9-595baf7361c0 · outbound

This paper cites Scaling agent learning via experience synthesis.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Scaling agent learning via experience synthesis

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.685698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:7fa156b8b2a3922cc25711d6bd40f0d9913f15d4ad9eee0f49e3c83712b7e15b

Observation 96221136-8f61-411f-8323-74bef630b744 · outbound

This paper cites Self-play fine-tuning converts weak language models to strong language models, 2024.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Self-play fine-tuning converts weak language models to strong language models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.306741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:a17c384e350fdece42381efe01f2118369111df03b08d46b8ad352a2759186c6

Observation ba801b0e-2ca4-4b2c-a504-9e5fcc862940 · outbound

This paper cites Webevolver: Enhancing web agent self-improvement with coevolving world model, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Webevolver: Enhancing web agent self-improvement with coevolving world model, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.309307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:a3cb9320d52944058fb2d49b8aca598e584bfc3cbd6d1b63ace36089397982f3

Observation 073a69ef-2661-491d-a163-4eada381642a · outbound

This paper cites Serl: Self-play reinforcement learning for large language models with limited data, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Serl: Self-play reinforcement learning for large language models with limited data, 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.293376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:c17f0237eef7c291173ac14429b5546e9f271eb032a52202c259a97b0b238b5a

Observation 41c86e56-5b54-4f8b-86b3-70099a394c40 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.289853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:ae11057f9350cd61a2ee5c9b0a7dd5d1e370407fe1d8e3d05afe6602c98fe5c0

Observation 61fbf321-5f2f-4c21-a9aa-9c01446f1947 · outbound

This paper cites How far can unsupervised rlvr scale llm training?.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis How far can unsupervised rlvr scale llm training?

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.291648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:be7539ae0fc82ea75c3fd6bf465814bce6237cafe57b1f26216af6999c823308

Observation 5ee682ff-afc2-40fe-9cd1-0271d1c66469 · outbound

This paper cites V-star: Training verifiers for self-taught reasoners.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis V-star: Training verifiers for self-taught reasoners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.334707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:65e486aefd54df1e57069fdfa9ffb3b9a952230daf02d9125005e2f0b30d60cf

Observation e35b5a32-7a9b-4e6d-a465-a6fe29d55074 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.695285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:8561d8ea6ac51abc27807b05edaa1cf2b5fce5ab79a1b53d1bf26b8cf4035879

Observation 9c15f3a0-2ffd-4d11-b266-141a6e55571a · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.670936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:c7523a32fc912a7324d80208e81d1501832e1b19243a8bcb4a62a936c0c07c5d

Observation 3304a84f-7a9d-4e82-9560-55cc4517cdb6 · outbound

This paper cites Language self-play for data-free training.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Language self-play for data-free training

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.673560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:1dd839c49f06ee4246e235f8098e5fa2c95ac122b82c611f25b69945dbc5590e

Observation 6565bd01-f291-4c52-b2d1-223a27fc1a18 · outbound

This paper cites Opensir: Open-ended self-improving reasoner, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Opensir: Open-ended self-improving reasoner, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.324046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:8be1cc87f65632d5c93323b0b2dd86ef51623c6095a38c62cf24982bb7ce2fe9

Observation 61336724-994b-4a19-807a-49254166e8e8 · outbound

This paper cites Embomatrix: A scalable training-ground for embodied decision-making,.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Embomatrix: A scalable training-ground for embodied decision-making,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.336670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:484dabb1d7fa9c5790fd7289819c6fb877f3b1a8b781d62091cb0f464c9f1a1a

Observation 055c6546-859c-4831-8d3b-39bdc6635db0 · outbound

This paper cites an unresolved cited work.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Unresolved cited work

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.708147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:a81e6e6252a70f7aa8716e043a28afc9525ef0e8e8274d0b69e32e5b37b31a68

Observation 731dccc9-fc72-47cc-b9a2-9d29d58767e9 · outbound

This paper cites Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.341517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:7e804c371f3838a04af44d83831adc9082e9fcf1bf59ac699bdfce9802612f4e

Observation 35590a8e-073e-4df4-afc3-be32921ca8ec · outbound

This paper cites Spice: Self-play in corpus environments improves reasoning, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Spice: Self-play in corpus environments improves reasoning, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.315957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:4fed24932285868a417b52b5bd70caba1e769be1987b7d7948ee598d6fa7a453

Observation 7d8425b6-52cf-4752-8c9c-a96a22796687 · outbound

This paper cites Chasing moving targets with online self-play reinforcement learning for safer language models, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Chasing moving targets with online self-play reinforcement learning for safer language models, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.317977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:5ca6f8654b73bec8e33316818c9397459169f783fa2153bf7524985122c26875

Observation ee23a353-2fbd-4293-8198-f0873e7045ad · outbound

This paper cites Prorl: Prolonged reinforcement learning expands reasoning boundaries in large language models, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Prorl: Prolonged reinforcement learning expands reasoning boundaries in large language models, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.326203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:c4342c2a66e8e838428c74ad253c9b08a82662f18d01c0a8d9ef675cf15cbd29

Observation 5bc10d6a-fe1a-4b9f-979c-7130db21a115 · outbound

This paper cites Search self-play: Pushing the frontier of agent capability without supervision, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Search self-play: Pushing the frontier of agent capability without supervision, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.311818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:dbf063dc38ea633569dfc43f15944a8ba693b8453ded04ac5bd4e62e1b3796ba

Observation a02f7a9d-22f3-46ff-8b41-75a19803e652 · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:05:04.665768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:9892b9e79289b8e56c39fdb0af0c0b6c205f3f18ea2f40ca2581db16dc200dee

Observation e8f135e0-7446-4dca-8629-f53d78934288 · outbound

This paper cites Self-consistency preference optimization, 2024.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Self-consistency preference optimization, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.288472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:75e24c5ab7c8414e480491b0838950ab74689b033669eb0573f9c8957fb37c0e

Observation 82603b9c-0cca-4cd2-ad53-2bc396cb2b6c · outbound

This paper cites Scaling synthetic task generation for agents via exploration.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Scaling synthetic task generation for agents via exploration

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.668535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:762b42ccc65bb608007e96cd086c63c8ba2b04ad97bb26bbcf26beb2c0f361d9

Observation 083f6318-4345-4c8a-9c20-5e034afd9b66 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.676192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:5f862ee3b134f46914fc1ef1e665ba12de8d174f13bbc7330ce26db409ec5e1b

Observation aa23d257-2368-4a5f-85c9-9cd5e799e508 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.681675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:c8f8bfcb47664f65653accef54a661600d28bee9107fd95c14a60aeb67713bf6

Observation 2657400e-367e-498d-826d-b06ac8058d7d · outbound

This paper cites Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.689661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:52ee8357e2983ef97d8933c51f9fd50b266d234e824ee639d9e4c455fef17bd0

Observation a66a1492-666e-485e-b4ba-70d2b61d3e55 · outbound

This paper cites Spurious rewards: Rethinking training signals in rlvr, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Spurious rewards: Rethinking training signals in rlvr, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.340322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:af1a507ee6aacb61219b29473e9e695d7c501b20768eb5cafd460efbdce145f2

Observation 971f2fc9-c1ad-422a-8557-274969bf2473 · outbound

This paper cites Beyond human data: Scaling self-training for problem-solving with language models, 2023.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Beyond human data: Scaling self-training for problem-solving with language models, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.314896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:f6a284079b344c45d9b3aad6eeef856cca938881bbdd2badb35e35d318c37a7d

Observation 91ef19e0-e861-4395-b37a-c4efee98d8ff · outbound

This paper cites Envscaler: Scaling tool-interactive environments for llm agent via programmatic synthesis, 2026.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Envscaler: Scaling tool-interactive environments for llm agent via programmatic synthesis, 2026

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.284851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:5cc0f0a72818f6ebf3ee292d546978f196308b9220a33fdf6312cb819567f8d7

Observation 07a92b4f-ac8b-4c18-ade0-7a3e0d053a27 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Kimi K2.5: Visual Agentic Intelligence

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.692501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:7c79a7d1a36c02ea06d5e23d8101ddae2fef75efd72d14ad1a5255bae0f23b4d

Observation 2cc281f4-124d-45af-8cae-f7292f9acbfe · outbound

This paper cites Nemotron-cascade: Scaling cascaded reinforcement learning for general-purpose reasoning models.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Nemotron-cascade: Scaling cascaded reinforcement learning for general-purpose reasoning models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.702250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:ae5c260e2125503762f2bc7cffbc3cd2ee758747c191dee9a2fdb7e09642e33e

Observation 7d7bbe2b-24c3-4856-96b2-e8b3e9f44049 · outbound

This paper cites Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.705245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:94bbf1cd97dd934323dd6a5e28049e07d0e0dc366fb382e0e3becef74c804d9c

Observation 465d29ec-7be2-4c9c-8e51-18a92bc40378 · outbound

This paper cites Socratic-zero: Bootstrapping reasoning via data-free agent co-evolution, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Socratic-zero: Bootstrapping reasoning via data-free agent co-evolution, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.285924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:9b4496233c2c5d41d36d6d2220a43f98af3c4d12b44486aee73c6186f8f5bdf1

Observation 5cdf91e2-4317-4ed6-a3a6-41411a5bf475 · outbound

This paper cites Llms as scalable, general-purpose simulators for evolving digital agent training,.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Llms as scalable, general-purpose simulators for evolving digital agent training,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.321013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:bf18e20c38a88ee8fe78ab354df5780890dce4b65d864279d55bb3e6deaf50ba

Observation 95874242-244c-45dc-85fd-daf6ccb5bf29 · outbound

This paper cites Llms as scalable, general-purpose simulators for evolving digital agent training, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Llms as scalable, general-purpose simulators for evolving digital agent training, 2025

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.660361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:6ba37002354a41c8f58df854efb5f2f376c8d4083146232ff516f11d3b202fb1

Observation 2c92b5f2-1f1d-416a-908e-5706f091332d · outbound

This paper cites Toward Training Superintelligent Software Agents through Self-Play SWE-RL.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Toward Training Superintelligent Software Agents through Self-Play SWE-RL

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.649387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:a13bef999aa998430fd47bd754ac8538f0ac0299c2d484698e0faab32102c433

Observation 6bf40004-0e52-4f1a-958a-38305042e275 · outbound

This paper cites Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.652492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:4ec7a88471b7d602f346f3bf7e6718ff44fe03d2c60a861ff9fe6b09f3933354

Observation cbe8ff7d-a908-4a56-b56c-516515d8bdfd · outbound

This paper cites Autowebworld: Synthesizing infinite verifiable web environments via finite state machines.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Autowebworld: Synthesizing infinite verifiable web environments via finite state machines

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.657651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:bdcfa883d7dfc8cf344a143dcb39bc95641d75e706bf0d1b435e2b0f5213f8c4

Observation bcad3a92-2457-4911-9568-0ef50ec95c26 · outbound

This paper cites Agent0: Unleashing self-evolving agents from zero data via tool-integrated reasoning, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Agent0: Unleashing self-evolving agents from zero data via tool-integrated reasoning, 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.307666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:7d5a5f84729568589f0e7d3ce3b7a23874a7069957de6798729793a6e5c063d8

Observation d40fc918-a190-4212-b794-e2f871807885 · outbound

This paper cites Qwen3 technical report.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Qwen3 technical report

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.287976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:86d30b61165689b45865bde9e3cd1a3bfbbcbb8da7e5d6318d497b8f28bdca92

Observation f92b34e1-1251-4615-a232-ce6434f01949 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Dapo: An open-source llm reinforcement learning system at scale

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.300324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:76e789583b43eefc84ec8b0ec35e7d1ce7e408bfe4e9866b4609cbfe628b429d

Observation 4889e128-47c4-41c4-9667-cdc5e2e3a076 · outbound

This paper cites Self-rewarding language models.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Self-rewarding language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.343517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:3b33da9ef40590e8716fb5ca6e983a5519d5b8f027c5f7ffeffbe975d0f3341e

Observation 7ed36f63-9743-4266-9c43-9595912d209d · outbound

This paper cites Star: Bootstrapping reasoning with reasoning, 2022.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Star: Bootstrapping reasoning with reasoning, 2022

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.279876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:ca80f3a020aee7df63ebc0bac9ec92aee56ad82f69683a2223244d7d402d3fb3

Observation c56b6159-8dca-4d3b-b3f3-75f35508a601 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis GLM-5: from Vibe Coding to Agentic Engineering

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.662785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:406cf27db35039633f041d90010f7840428e6f60f4246f6439880159ece1db17

Observation a31563ec-0660-463c-bd50-6a643bd37c69 · outbound

This paper cites RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.654889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:d90f50805811e345896a08e1c4c208a32720391b2113d460465ee746e128c5c0

Observation 408213ea-bd1e-49b4-a78d-60dd186b9c68 · outbound

This paper cites Darwin gödel machine: Open-ended evolution of self-improving agents, 2026.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Darwin gödel machine: Open-ended evolution of self-improving agents, 2026

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.298524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:6efe1e6b7329ea4457d1c130037f712150e023de785380f967a1bf1e9a156553

Observation 673c081c-6d39-4ca1-847c-542af2027d1d · outbound

This paper cites Right question is already half the answer: Fully unsupervised llm reasoning incentivization, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Right question is already half the answer: Fully unsupervised llm reasoning incentivization, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.316874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:8f2cd3a52175477b1efc65c916a6c7c6c8182ce3913b508c0d04ffed96f28f57

Observation 188e7641-ca90-427a-b0d0-40165e43ba86 · outbound

This paper cites Better llm reasoning via dual-play, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Better llm reasoning via dual-play, 2025

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.334355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:07d74b8736829464a0ae013e53e3c7416b01bf5cf205b264c46ba2b607868424

Observation 75f2629a-0279-4185-8ff5-6a08a5fa9bf6 · outbound

This paper cites InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-01T02:17:18.656844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:d5758b3b72f00ca2c73791b239f22d815ee21d31045770d18f840d17e6f37022

Observation 58d684c2-18b4-4f52-9de9-9412e5345e42 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.644104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:1c147fac1a1fb27ff21159b5c33bf8af89b8fa588cef033f331fc4bc2faaf767

Observation 94ad5bf9-665f-4960-894e-1ab6e992db1e · outbound

This paper cites Learning to reason without external rewards.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Learning to reason without external rewards

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.338109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:8cbed04803f40b604c4e55a85dd10ace4716896de65270d508bac8898e065e07

Observation 0494669c-c16a-4189-a7c3-1e4690d5f8b7 · outbound

This paper cites Self-challenging language model agents, 2026.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Self-challenging language model agents, 2026

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.336472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:ab5b91fc1aea25eb2d0b5ab2e93df5c9e080298eaf1453557994fe5ef9317c68

Observation 9895ff99-bc01-4371-beca-e90cd763047a · outbound

This paper cites Evolving language models without labels: Majority drives selection, novelty promotes variation, 2025.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Evolving language models without labels: Majority drives selection, novelty promotes variation, 2025

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.333061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:a0218d1900d5490e9737ddb2d103585c3510b089a00a211fa391340c4b25ef03

Observation b5b51823-a1ca-4c77-a74d-c59e6cf188c7 · outbound

This paper cites why is X happening.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis why is X happening

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:05:04.646962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:4aa21888e432cea8454e23db87710c99d8f7af8c48676fc77f7eefb57dcdd5d7

Observation 0aa5b604-a8ab-4b4d-bc9a-57451e340413 · outbound

This paper cites Given the multiset {S}, find a nonempty.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Given the multiset {S}, find a nonempty

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:24:01.339078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:44aed5928b972c7d8f106f611d9ffe9547d095b48d4e96fa76dfff6145ff8031

Pith citing papers

Observation 9c955222-a3f8-46fe-bfe1-b548ba89c74c · inbound

SkillMentor: LLM Agent Self-Evolution via Learning Blind-Spot Diagnosis cites this paper.

SkillMentor: LLM Agent Self-Evolution via Learning Blind-Spot Diagnosis Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T08:53:20.110535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:53:20.110535Z digest=sha256:39cb013b86e88486755e4004780306d409fe0152a50af5a7cb5d35fc9c2c4f8f