Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T06:04:14.799234Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 13 inbound Pith citation observations for arXiv:2504.19162.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T06:04:14.799234Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.285830Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T02:26:27.242958Z
74 of 74 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c4818701-5414-400a-b2b0-3172ff58f326 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Chain-of-thought prompting elicits reasoning in large language models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ce3b57-5804-4880-a878-40df3552ad78 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df0345db-cf64-4aab-a358-bd463743294f · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a7747f8-1245-4caa-ae39-9760be751873 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Language models are few-shot learners
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e5305829-e3d4-45f2-93a8-630535500090 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning GPT-4 Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8336a4cf-a945-4d4b-9779-6e19f95b6cc9 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning GPT-4o System Card
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98957b60-8a51-4a93-9994-04fe5aab8d3a · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Gemini: A family of highly capable multimodal models, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 126b6d6c-fd87-4997-b62c-e67baad1b90d · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Introducing the next generation of claude, 2024
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b22498d-faa8-4b2b-9cb2-cd8533744021 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71a068a0-3664-4b5a-aefd-7e082174e3b3 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3940a032-2ae3-4bf8-8e44-47a30610ba77 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Qwen2 Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b8881a-ede2-457a-8d78-14d3f6f15d67 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Qwen2.5: A party of foundation models, September 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 607bad48-0883-47b1-9c40-e25489c5fb88 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning DeepSeek-V3 Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3985ae24-e812-43ad-93f1-4933b7637891 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Training language models to follow instructions with human feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce7406d9-8776-4bc6-8d1c-8bd8d96dfce0 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Scaling Instruction-Finetuned Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de4e5c97-5d48-444f-a6d3-4b79199e8eb8 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Fine-Tuning Language Models from Human Preferences
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c82e5a0-b269-412f-9aee-d82439a078bf · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e9d7c7-abc3-4954-b1a5-30cf1d1eb324 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 659adce1-6566-4e18-9750-2302a2330d78 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Prover-Verifier Games improve legibility of LLM outputs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382acda4-e014-4b37-b9b4-75f5e90ccbda · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Openai o1 system card
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9ef91387-af08-46eb-acbf-d943021de027 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 123b766a-2c5b-449d-b8ad-84cc6dfb3316 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91671b48-b76b-42f4-ba6c-5f99bb9d76e2 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Let’s verify step by step
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14020d1c-0b0a-47ed-a7cc-a95d351da66d · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Skywork-o1 open series
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e1e8737e-ea83-413f-99b7-d0ae2ae338c9 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Solving math word problems with process- and outcome-based feedback, 2022
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6110111-5df8-4562-b836-aa4c29c75581 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e45cfc4-fcc0-49da-9398-c466e1de19c5 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning ProcessBench: Identifying Process Errors in Mathematical Reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64bbadf7-c265-4275-a40f-5cea1e0f3438 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Generative verifiers: Reward modeling as next-token prediction
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e96f158f-e359-4e4f-bac7-52b63dee4c60 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c250693-50eb-4c2d-9664-cc2b2f285829 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27a71eaf-1c26-4cbc-b64f-d6301da3b8f1 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee8533a-66dc-448b-8027-995dbdfc8a08 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac1f8437-b4dc-4094-9155-267366176814 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7a22f54-3bf6-4827-9f33-52d5703701f3 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning O1 replication journey – part 2: Surpassing o1-preview through simple distillation big progress or bitter lesson? Github, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dfb04b07-7747-45b5-955b-0e03d499233f · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Can large language models detect errors in long chain-of-thought reasoning?, 2025
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7fc92d19-3714-444a-bfb4-c76381a12a9e · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f58e72-2a5d-441e-9b80-8c9f6b497f49 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Aime 2024, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eae23f51-f64e-4364-94da-f84539ee27fc · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Introducing meta llama3: The most capable openly available llm to date, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a0e24e55-c82a-446e-8918-7537347e26c7 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24849fa7-1d53-4019-b291-6823dd354168 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e1b4268-29d2-486c-b4fd-3ede57d6e91e · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Tinyzero
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8716aa30-5ffc-42c8-966b-7b000b68704b · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning The lighthouse of language: Enhancing llm agents via critique-guided improvement
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c49bf63-b568-44fd-9b63-e71d0ac3af5b · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Some studies in machine learning using the game of checkers
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90787b21-6b19-4cd9-96d8-029faaa6ee36 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Temporal difference learning and td-gammon
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7ced5c1c-4dc2-4e17-8a53-fcf4c196b2b7 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation def38d47-ab18-456d-8563-61ab31747f66 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Mastering the game of go without human knowledge
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e49a1add-e3fc-4072-9728-f90a8444246c · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-playing adversarial language game enhances llm reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 587296c3-cfe3-49ca-8001-53afc87b5fdc · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7a9601-fa4a-477f-94ce-b72c0b3f0ad5 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 523096d3-8559-4589-8215-367a6c681402 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-Play Preference Optimization for Language Model Alignment
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68a38343-2479-4304-9cb9-667ff43ea9f8 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 363c2bc3-6819-426d-ad6e-5e376d9c5d98 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b017651f-4cc6-4ed3-bad1-c3d8be3080c5 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47dbf16d-0a09-4396-9816-f17e01c1f0f4 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bc964d7-207f-4dbf-babb-3dc0fd28d14e · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Qwen2.5-math-7b, 2024
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1c590def-d60e-478d-9be4-2ce55403d821 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Training Verifiers to Solve Math Word Problems
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe913151-b115-4f67-a1a4-4f9ece4fe7df · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8016ad59-934f-4767-9b6c-82862739b37b · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3b4ea95-52ea-4912-8ff1-901514054820 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Note: For each question, you will be given a reference incorrect last step
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6bd5e618-e267-4f57-aae1-f049e5c14b2b · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning At the end of the response, output \\boxed{{Correct}} or \\boxed{{Incorrect}} to represent the correctness of theLast Step
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8538f9b0-462b-4ed8-83a6-e40eed9d9f8b · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning If the Draft Critique includes this analysis, you can directly summarize from it
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 37bb3787-2a1a-40a7-88fc-da73680c846a · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning You should write a new version of brief critique here
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 53131bf5-e9cd-4903-a610-04691b79be61 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Please draw a conclusion about the correctness of the Last Step
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 31d32af9-ab56-49a3-813d-f9cff886c7f9 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning the critique
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bc56fce3-dcb1-4c2e-a818-b110ffd114db · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bd740ae5-ee52-47df-9226-cf956d234b14 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Instead, in the Critique, you start with an analysis of the Last Step
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 340cf9aa-4dc5-4041-bbd2-66e65078d28f · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Y ou only need to focus on the correctness of the Last Step
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation abdcb106-5899-4b5b-99e2-b66576ea4bda · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Clearly explain the solving process in the last step
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e1c3e9be-5bd2-4dd4-8aea-8b3ff0618aaa · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Predefined Error Types
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 63eaa358-2ed3-4195-8007-0df80ecc3b84 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c17ea9ab-5827-4b49-9985-6adc1f3e8fe3 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 950be891-8172-4be9-b83b-1a8be41985c5 · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c351eec6-d67b-4299-8ee4-b40abf1eabee · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning You should write a brief critique here
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e47bb238-85ec-44f9-9cce-35007ee680ad · outbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning At the end of the response, output <Answer>Correct</Answer> or <Answer>Incorrect</Answer> to represent the correctness of the Last Step
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8f069e31-eeec-4285-8eab-5a9ed739cc23 · inbound
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 177
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baa57ae7-86c2-4a8e-90df-38e7964957d1 · inbound
Lifelong Safety Alignment for Language Models SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60935ee-dde6-4400-8c29-27d6730082b3 · inbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 449b5695-ac43-4588-9625-8b28fde75fd8 · inbound
RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a359f50-1ac9-471b-9300-3a38ec70366e · inbound
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adff3f02-7490-446b-9e66-2fd40970239c · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c95307ca-c6ad-4781-96d8-a1e6fefae90c · inbound
SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b8b22f47-daa5-455e-80bb-bc43f4b79248 · inbound
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a0bca19f-eea2-4e01-97e0-3a4d26c32514 · inbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 63701c8b-1f32-45b6-84b9-d8b8cb3173f4 · inbound
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 72136545-3aff-48b4-82af-61d346e61e27 · inbound
Pseudo-Formalization for Automatic Proof Verification SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 33f4a043-fb6c-4b9f-bbc3-3c9edc09d1b8 · inbound
Pseudo-Formalization for Automatic Proof Verification SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 84fa483e-fc08-4289-b1c0-d2a7ebe9195a · inbound
Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.