Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:24.353008Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 13 inbound Pith citation observations for arXiv:2506.08989.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:24.353008Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:17:38.690339Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T18:17:33.758346Z
93 of 93 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e21e5946-6e2c-4958-aa70-05157d3d6c78 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21aae386-2ef6-4aa4-8d78-b48c9020da11 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f6efc6-d15b-4409-b085-79fcfb820213 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 795cfb9a-3440-4cc7-b00a-dfeb9e47aef8 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444bad16-8172-439e-af77-37911851bead · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Nearest neighbor pattern classification.IEEE transactions on information theory, 13(1):21–27, 1967
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8494ff6c-dfad-47cd-b313-b7eec9c5ea07 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Process Reinforcement through Implicit Rewards
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d618a45c-bbd8-44c1-b0e7-8ceaa4afb208 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65d781ef-83f7-4f1c-8f7c-0d316c9fd3c2 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e02db3-49d4-4303-b2bb-2dae77510a5a · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a6c548-6f5f-46b1-9fa2-e82483501392 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d58dc3c4-d2cf-4252-a6cd-fa0e7567d73e · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Olympiad- bench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d429c5a-2a21-4c17-9e9d-333aca6d0183 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Skyworkopenreasonerseries
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acdea8a8-df57-4221-90d7-adb65cde3e04 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Measuring mathematical problem solving with the math dataset.Sort, 2(4): 0–6, 2021
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fdc6bfa-d305-422f-9113-6ae52e74720b · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 408a24fd-cfcc-4564-b8ff-ff9c928c58d3 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81d44e76-1f7a-41b7-a26f-723b456eaa4f · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning OpenAI o1 System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aa4b2e8-32b4-4c68-bc0e-e7d79fd00ed8 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Knowledge- augmented reasoning distillation for small language models in knowledge-intensive tasks.Advances in Neural Information Processing Systems, 36:48573–48602, 2023
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dffebfb5-adc4-4029-9d9b-9559ba82485f · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Gonzalez, Hao Zhang, and Ion Stoica
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6e8813c-4150-4f42-b6e7-46e2507e6ec1 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35: 3843–3857, 2022
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8f00e26-8ce6-46d5-8511-c212abc361c4 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Common 7B Language Models Already Possess Strong Math Capabilities
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00847c0-fec3-4aee-bc29-da9ab492b095 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning From generation to judgment: Opportunities and challenges of llm-as-a-judge.arXiv preprint arXiv:2411.16594, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b2ffa4-f2eb-4445-bf66-67a400742f9d · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning LIMR: Less is More for RL Scaling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c93907a-1052-42cc-a055-453eef601138 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0527b2e1-f295-4287-a70d-9bbb64f064f3 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning From System 1 to System 2: A Survey of Reasoning Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e47cd1a-4d78-45bc-ada1-90e2f1782f55 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Task Oriented In-Domain Data Augmentation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0173a58-a462-435d-ab26-f430656a5616 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Let’s verify step by step
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c1205cd-def1-4f25-bf6d-668cc55db813 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Augmenting math word problems via iterative question composing
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b1eb7ed-159f-4a59-980c-567cb2b38c6b · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36f27c6a-6278-4eaf-b813-a876df3e527b · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426ff685-7541-4cb4-a2f6-2a7abf44b8c2 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9c5958-01a1-44c0-8ace-e662bbc76d89 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1ee18de-6c4a-4f6f-8ec9-ad19fbc7fa63 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c33ad2-872d-43ac-8203-ccf0d3296030 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning American mathematics competitions (AMC 10/12)
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3ea30b0-2db3-41fb-b297-c75c47c18ae7 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning American invitational mathematics examination (AIME)
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e2c2ddc-9a31-4f0a-bdc1-83ad9eb5649e · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning s1: Simple test-time scaling
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb1e8ecc-1bb2-448b-9ec6-5a8c2ea6156b · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e828667-bcf3-48a7-ad23-a9bd85c9dcfd · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730– 27744, 2022
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54129201-6111-4211-851a-87ceeab991c2 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 544df257-d2cf-496f-a09e-d4f549208a86 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Proximal Policy Optimization Algorithms
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8736f22e-dae6-4c9b-8faf-a1c93d6de001 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24662ffb-c264-4631-af2f-7f3314206f60 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f93e7aa4-8f1a-4dd0-8290-698a6f6bfce2 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f461f536-92b8-46b9-a689-3f5ae74fa4e8 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af7c036a-e277-46cf-bd1e-95ab1ad8d0ef · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Large language models for data annotation and synthesis: A survey
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3680da20-6462-4d17-83d2-5ab0e4357c1a · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Mathscale: Scaling instruction tuning for mathematical reasoning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fbe9324c-2fb3-4fd4-b76d-3acff595b7cf · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e925c12-d9b7-4362-8e23-4fa43bbb41d0 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e2b9996-07ec-4440-b989-c866a5647eec · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving.Advances in Neural Information Processing Systems, 37:7821–7846, 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 84ceaeff-0d18-45aa-8f79-707c113c317c · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Openmathinstruct-1: A 1.8 million math instruction tuning dataset.Advances in Neural Information Processing Systems, 37:34737–34774, 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ead05707-8f16-40f7-9349-cb3cef9a9d93 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b30c748d-c604-49f7-9c3c-aa054484011c · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Explore the Reasoning Capability of LLMs in the Chess Testbed
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2afe12b7-5507-4f09-9017-3eca11d63e40 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Examining false positives under inference scaling for mathematical reasoning.arXiv preprint arXiv:2502.06217, 2025
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf457c15-fa3a-4d33-a7ad-bcdc482e83be · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3792c869-4ccb-4a8d-b2fb-37e61dfa4142 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962e3e5d-9e77-4b1e-aa22-9c407beb8a47 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be76fbae-93bd-43b9-9954-a0c0b08d0d4f · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e5be5d-b6b1-4df6-9000-713d118aac47 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Qwen2.5 Technical Report
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cdadef9-18d5-456d-ab77-de1091d4c497 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49f89318-4f59-4b71-b1cb-89104f99129c · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning LIMO: Less is More for Reasoning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9f9b1f8-3f34-4f5c-b985-d526825b24e3 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Demystifying Long Chain-of-Thought Reasoning in LLMs
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1649c7b-2163-4771-b4ba-dd88ccf752bb · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation decbd585-0da8-4383-9b40-8653b184b4ee · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b9d5fd-007a-4d0f-8532-9d229f7347a2 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 357d701a-568f-4f66-9ec6-53b779eeae76 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9d31dd3-77a0-4581-918f-ffdecd649ba1 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89a8ed69-2f13-41d2-b7c5-05d6d8c1b80b · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3a71132-7d0a-4ed9-8b16-0e1a34f67b49 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c633bf1-ae61-4c18-a6d4-4047f6ea3b29 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Rest-mcts*: Llm self-training via process reward guided tree search.Advances in Neural Information Processing Systems, 37:64735–64772, 2024
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 228a7f3b-d51f-4453-892b-52afc8a2d0bb · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Balancing speciality and versatility: a coarse to fine framework for supervised fine-tuning large language model
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b8e0183b-dd8c-426c-b887-fb5a31b60961 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Process-based Self-Rewarding Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aad07c4-6676-4390-ac88-39428790589a · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3777fc28-c445-4af4-a4f2-55d684f645d5 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 593556dd-f49d-456d-9bcf-f6e075cae8d2 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Promptcot: Synthesizing olympiad-level problems for mathematical reasoning in large language models.arXiv preprint arXiv:2503.02324, 2025
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b7f6c82-2274-4773-bfb0-c0f83c69af82 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Fine-Tuning Language Models from Human Preferences
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608a7665-1ac5-4b48-9ee8-a877948827b9 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning TTRL: Test-Time Reinforcement Learning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ac589d-1f86-4088-b7ac-c1ef40899a16 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Let’s think step by step and output the final answer within “∖boxed{}
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4acc41aa-0735-45ab-9090-ced5baf37279 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3893a523-8b05-4b9e-b204-8ab6420d5ab1 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 906fa85b-88d0-4375-baeb-cf59e85ddbc5 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 765d8801-9e27-4e21-a3eb-90cc2a0225de · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ac286313-68a0-422e-8629-f9813c2d82fc · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 488ecaf7-548b-4632-aa20-a3822232e1b7 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning ### Output Format : - First , provide your brief outline and planning for the question design
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7d10985f-22f1-41e3-89f4-74447f4600e0 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc6fcc66-609d-4c2d-b844-5f0cc2574456 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9f239392-f56b-4ff8-8874-8024e6abbd73 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ab9efd30-3f79-4d89-b83c-a3af5ccf71fd · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9b151e98-c755-4475-86d1-5df25b4d329a · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1211b350-b2b1-4cc7-9083-ef9bcb61a25e · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a7b43e2e-2ffb-437d-9b50-cef7c5b7c573 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 72a69c35-1e4e-4365-87e9-1ef6359e27c4 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation db2a18fe-e500-49f5-a2c1-614161821f13 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 864def1b-ffd6-49e7-ab28-d8d31f5552df · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 944059f0-1116-4255-9878-3e7afa29d930 · outbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ce4099e-64e1-4d5f-bb86-20d23d32d5b1 · inbound
Libra: Large Chinese-based Safeguard for AI Content SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aada9b5e-40ca-4420-be05-92b822b7db85 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 193
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 33f605ae-eee2-4ce0-8cbf-94d14178471f · inbound
A Survey of Reinforcement Learning for Large Reasoning Models SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 299
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6776643c-97a0-4199-9eb2-0fd15a3547b6 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fbca871-fffc-476e-ac85-e9ceb31743d2 · inbound
SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70ec431f-26c0-44d1-a8bf-90c3a40bb5b1 · inbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 182
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 015fd05d-9846-488c-9baa-c6ea3f247377 · inbound
From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 202
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation baf5f158-3f8c-4031-9a2e-2ea31bf95a55 · inbound
Confidence-Aware Alignment Makes Reasoning LLMs More Reliable SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16caeec2-af6f-44dd-a223-a6cb3ee161ec · inbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1240ba79-4253-49fc-98b1-5c45025dfec3 · inbound
REVES: REvision and VErification--Augmented Training for Test-Time Scaling SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9c0575ca-1c65-4e26-a911-8601810c1a2a · inbound
From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9211432e-5b7f-4b22-82cc-94924bc16437 · inbound
Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320d73ca-b445-4c57-8b2c-a322c9e2aee3 · inbound
Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.