Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T20:58:07.921910Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2605.14392.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T20:58:07.921910Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T08:53:20.110535Z
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6229edd4-ab9c-4f3b-b013-05b86abd056b · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 76173681-2e83-493c-b411-d7a9b04c8d2b · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Safe and scalable web agent learning via recreated websites,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 534311ea-f701-4f01-91e4-2f631e4e18e8 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Safe and scalable web agent learning via recreated websites.arXiv preprint arXiv:2603.10505,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7271659c-099f-44de-8eb9-84f76207546c · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Spc: Evolving self-play critic via adversarial games for llm reasoning, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ebc4e32f-3ffc-4138-8d81-064b0023cc21 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Self-questioning language models, 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 708291ee-4cbd-4a10-8e8f-035d81cdd489 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Multi-agent evolve: Llm self-improve through co-evolution, 2025
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 85e338d2-36fa-418b-8ea9-595baf7361c0 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Scaling agent learning via experience synthesis
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 96221136-8f61-411f-8323-74bef630b744 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Self-play fine-tuning converts weak language models to strong language models, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ba801b0e-2ca4-4b2c-a504-9e5fcc862940 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Webevolver: Enhancing web agent self-improvement with coevolving world model, 2025
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 073a69ef-2661-491d-a163-4eada381642a · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Serl: Self-play reinforcement learning for large language models with limited data, 2025
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 41c86e56-5b54-4f8b-86b3-70099a394c40 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 61fbf321-5f2f-4c21-a9aa-9c01446f1947 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis How far can unsupervised rlvr scale llm training?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ee682ff-afc2-40fe-9cd1-0271d1c66469 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis V-star: Training verifiers for self-taught reasoners
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e35b5a32-7a9b-4e6d-a465-a6fe29d55074 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9c15f3a0-2ffd-4d11-b266-141a6e55571a · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3304a84f-7a9d-4e82-9560-55cc4517cdb6 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Language self-play for data-free training
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6565bd01-f291-4c52-b2d1-223a27fc1a18 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Opensir: Open-ended self-improving reasoner, 2025
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 61336724-994b-4a19-807a-49254166e8e8 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Embomatrix: A scalable training-ground for embodied decision-making,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 055c6546-859c-4831-8d3b-39bdc6635db0 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 731dccc9-fc72-47cc-b9a2-9d29d58767e9 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 35590a8e-073e-4df4-afc3-be32921ca8ec · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Spice: Self-play in corpus environments improves reasoning, 2025
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d8425b6-52cf-4752-8c9c-a96a22796687 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Chasing moving targets with online self-play reinforcement learning for safer language models, 2025
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee23a353-2fbd-4293-8198-f0873e7045ad · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Prorl: Prolonged reinforcement learning expands reasoning boundaries in large language models, 2025
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5bc10d6a-fe1a-4b9f-979c-7130db21a115 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Search self-play: Pushing the frontier of agent capability without supervision, 2025
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a02f7a9d-22f3-46ff-8b41-75a19803e652 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e8f135e0-7446-4dca-8629-f53d78934288 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Self-consistency preference optimization, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 82603b9c-0cca-4cd2-ad53-2bc396cb2b6c · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Scaling synthetic task generation for agents via exploration
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 083f6318-4345-4c8a-9c20-5e034afd9b66 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aa23d257-2368-4a5f-85c9-9cd5e799e508 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2657400e-367e-498d-826d-b06ac8058d7d · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a66a1492-666e-485e-b4ba-70d2b61d3e55 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Spurious rewards: Rethinking training signals in rlvr, 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 971f2fc9-c1ad-422a-8557-274969bf2473 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Beyond human data: Scaling self-training for problem-solving with language models, 2023
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 91ef19e0-e861-4395-b37a-c4efee98d8ff · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Envscaler: Scaling tool-interactive environments for llm agent via programmatic synthesis, 2026
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 07a92b4f-ac8b-4c18-ade0-7a3e0d053a27 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Kimi K2.5: Visual Agentic Intelligence
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2cc281f4-124d-45af-8cae-f7292f9acbfe · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Nemotron-cascade: Scaling cascaded reinforcement learning for general-purpose reasoning models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d7bbe2b-24c3-4856-96b2-e8b3e9f44049 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 465d29ec-7be2-4c9c-8e51-18a92bc40378 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Socratic-zero: Bootstrapping reasoning via data-free agent co-evolution, 2025
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5cdf91e2-4317-4ed6-a3a6-41411a5bf475 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Llms as scalable, general-purpose simulators for evolving digital agent training,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95874242-244c-45dc-85fd-daf6ccb5bf29 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Llms as scalable, general-purpose simulators for evolving digital agent training, 2025
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c92b5f2-1f1d-416a-908e-5706f091332d · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Toward Training Superintelligent Software Agents through Self-Play SWE-RL
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6bf40004-0e52-4f1a-958a-38305042e275 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cbe8ff7d-a908-4a56-b56c-516515d8bdfd · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Autowebworld: Synthesizing infinite verifiable web environments via finite state machines
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bcad3a92-2457-4911-9568-0ef50ec95c26 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Agent0: Unleashing self-evolving agents from zero data via tool-integrated reasoning, 2025
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d40fc918-a190-4212-b794-e2f871807885 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Qwen3 technical report
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f92b34e1-1251-4615-a232-ce6434f01949 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Dapo: An open-source llm reinforcement learning system at scale
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4889e128-47c4-41c4-9667-cdc5e2e3a076 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Self-rewarding language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7ed36f63-9743-4266-9c43-9595912d209d · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Star: Bootstrapping reasoning with reasoning, 2022
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c56b6159-8dca-4d3b-b3f3-75f35508a601 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis GLM-5: from Vibe Coding to Agentic Engineering
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a31563ec-0660-463c-bd50-6a643bd37c69 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 408213ea-bd1e-49b4-a78d-60dd186b9c68 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Darwin gödel machine: Open-ended evolution of self-improving agents, 2026
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 673c081c-6d39-4ca1-847c-542af2027d1d · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Right question is already half the answer: Fully unsupervised llm reasoning incentivization, 2025
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 188e7641-ca90-427a-b0d0-40165e43ba86 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Better llm reasoning via dual-play, 2025
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 75f2629a-0279-4185-8ff5-6a08a5fa9bf6 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 58d684c2-18b4-4f52-9de9-9412e5345e42 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94ad5bf9-665f-4960-894e-1ab6e992db1e · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Learning to reason without external rewards
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0494669c-c16a-4189-a7c3-1e4690d5f8b7 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Self-challenging language model agents, 2026
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9895ff99-bc01-4371-beca-e90cd763047a · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Evolving language models without labels: Majority drives selection, novelty promotes variation, 2025
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b5b51823-a1ca-4c77-a74d-c59e6cf188c7 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis why is X happening
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0aa5b604-a8ab-4b4d-bc9a-57451e340413 · outbound
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis Given the multiset {S}, find a nonempty
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9c955222-a3f8-46fe-bfe1-b548ba89c74c · inbound
SkillMentor: LLM Agent Self-Evolution via Learning Blind-Spot Diagnosis Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.