Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T19:51:04.779597Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 61 inbound Pith citation observations for arXiv:2504.20571.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T19:51:04.779597Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:55:28.597027Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
69 of 69 outbound references displayed
External citation measurements
7
pith, observed 2026-08-05T02:28:24.338817Z
Observation 8463121b-b1d5-43f8-bcd5-c0ca222f5325 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Learning to reason with llms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 843fc5fc-219d-41b6-8d90-e2ff72eedd0a · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5411608b-f66e-48e2-bd8e-c4f42fd216c6 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3d72fe2e-f683-4894-93c7-68d3f028cfd3 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example On Designing Effective RL Reward at Training Time for LLM Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87b81421-b53a-40ba-91d2-106b8c5f8f5c · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc6ddccd-e9fc-452d-b0ec-e119f35f12a8 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 36c6a299-11ef-4e56-89d3-bb545b01c390 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Proximal Policy Optimization Algorithms
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b51d1bfe-0d26-4b9a-b229-dc4efe278036 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c437fa1-e43b-4fe2-9ac4-1f9243aeb264 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fac2e47b-4999-44bc-8193-6a25d9126a7b · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ce80f86d-eeb6-4241-9a09-21e0f4e91a77 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cee44b9b-7798-4947-ac7b-86376f501c6d · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aabc6d2c-69c3-44b2-9775-d671920d3e6b · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Understanding R1-Zero-Like Training: A Critical Perspective
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b0bc476b-3f85-4b2a-a01e-ed8bf304e5c0 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Deepcoder: A fully open-source 14b coder at o3-mini level
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4fe3e517-7543-4635-a352-ec1c2a920bfa · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2979b2dd-a0bb-412b-8df2-a5ccdccbded3 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Srpo: A cross-domain implementation of large-scale reinforcement learning on llm
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13576480-c537-4a47-a2e0-917124a0ea70 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Numinamath
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ea1dd23a-3b73-4412-8edb-4727ac7f3c05 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1e0c9d2a-a9ab-408f-9fea-7357485475cd · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Limr: Less is more for rl scaling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23003ecf-b9d1-4578-aa82-057a56a42395 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6e0d4383-d98e-4329-bae7-b28ca559b054 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Rethinking Reflection in Pre-Training
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fcf169b3-66db-437c-be42-8238dcde9af3 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example HybridFlow: A Flexible and Efficient RLHF Framework
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84cc12e7-1966-4fc6-98ae-125f2872f1da · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example What makes a reward model a good teacher? an optimization perspective.arXiv preprint arXiv:2503.15477
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1abeb772-11dd-479a-a13d-9435f7103148 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Qwen2.5 Technical Report
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ddd9e76-925a-49f3-aa88-d8bee5e54ec8 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5d24e96d-549a-4e6e-85c2-fc981f7cfb4f · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example The Llama 3 Herd of Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3de5f4a-ed65-41d2-a7d9-77010e4b3036 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Measuring Mathematical Problem Solving With the MATH Dataset
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b67c5265-ec96-4f5c-aeb9-ffb0ac0c2141 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Gonzalez, Hao Zhang, and Ion Stoica
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13bec755-0736-4380-83aa-244082484b84 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Let's Verify Step by Step
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e9ef792-e0c8-4d1c-b4a1-086e03aa070c · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Aime problems and solutions
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a805ea32-f4e7-46af-aa54-bf931c2ec4fe · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Amc problems and solutions
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4f18ef18-be0f-473d-80ca-cb437e8b5a9e · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation efde09b4-e18c-4b22-986a-b3394130ce5e · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 11cfa2c8-a7ed-474b-9e2d-375d91a32853 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b1f6322-5128-4e25-a161-11cfdef840cd · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8d9ccd95-80b0-4530-a4a4-716df1d57e35 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 56670751-48f9-4bc0-90c8-2dd89773d875 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Deep Grokking: Would Deep Neural Networks Generalize Better?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6275d4fe-b964-45e3-97f4-58ec2de7eee8 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Progress measures for grokking via mechanistic interpretability
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 704c7c1b-a68f-4074-9204-915195302826 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Towards understanding grokking: An effective theory of representation learning.Advances in Neural Information Processing Systems, 35:34651–34663
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d49d725-10f8-4629-9d37-103edb2dd138 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example The Complexity Dynamics of Grokking
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6bad6672-554e-47bb-ae4a-f37db1727bef · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Grokking at the Edge of Numerical Stability
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3a7af56d-e27e-457f-9d2e-58d10736a594 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 89e88595-6b82-4a51-ba4e-fabbdd4e0f50 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4132bbca-c88a-47bb-8343-c8be44aaf6cc · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Fastcurl: Curriculum reinforcement learning with stage-wise context scaling for efficient training r1-like reasoning models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 24c1b142-0484-4d58-9960-dbe23dc2652d · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f872eae1-12ef-47fa-8932-3b10e4dffb34 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f86d6208-7dff-4623-ba6d-f875029cd0b8 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example TTRL: Test-Time Reinforcement Learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ad7eead-b72b-4edc-ab2d-e2114b9f9814 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Large-Scale Data Selection for Instruction Tuning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c577a14-1fa0-4612-9bb3-c3b69a252768 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Alpagasus: Training a better alpaca with fewer data
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23422147-dbc9-45c9-b1c3-d713982f9adb · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Smith, Hannaneh Hajishirzi, and Pradeep Dasigi
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2b50c2b9-3990-4065-b781-9c8d0fa75229 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example LESS: selecting influential data for targeted instruction tuning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 300baaab-6177-4ddf-ace3-da23afa1e48f · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Active preference learning for large language models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b88e2ac5-cdc8-4102-964c-61babca6a388 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c1a0ce5-1a1c-4302-ad9c-3deb7fd487ab · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Active Preference Optimization for Sample Efficient RLHF
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2db93978-383a-4694-b1b0-05266e629434 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33596df9-c1c4-42b4-95db-a08d683a0a96 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Concise reasoning via reinforcement learning.arXiv preprint arXiv:2504.05185
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c05cdff-c670-497c-ba7f-d2446bb8316f · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Schulman
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e341750-2138-46c8-9ec1-da57a5347267 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9c3e60c-fd4d-446a-b4af-14053b61484e · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b349fb69-6cc8-4a0c-8902-777cef9d4dda · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Skywork open reasoner series
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 80b4ccd6-9441-4721-a1b9-ca8877e95559 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example A Survey on LLM-as-a-Judge
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b92831f0-a2d8-4795-b403-82ec72eaa5b2 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5875cb2f-c345-4953-8204-9f78a78aee60 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example A Survey on In-context Learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e70c6b57-36fa-4b3a-8c0f-fcadd44d65b0 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Deep Learning is Robust to Massive Label Noise
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa047314-872e-4088-a1b4-22ba7582fd50 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Deep double descent: Where bigger models and more data hurt.Journal of Statistical Mechanics: Theory and Experiment, 2021(12):124003
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c635af5e-0242-4d74-826a-1916b9bc5f1a · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4576612c-ff1c-49f0-9e02-b233b0e11fba · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Smith, Benoit Dherin, David G
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 783b5164-98ac-4685-bebd-3fd3e8f8b850 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Acemath: Advancing frontier math reasoning with post-training and reward modeling
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61f95623-acc5-4a72-a3f6-ce1371155490 · outbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example L′ PG-GRPO(·, θ) +βL ′ KL(·, θ, θref) +αL ′ Entropy(·, θ) # ,(3) where β and α are hyper-parameters (in general β >0 , α <0 ), and “·
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6037df06-3d9a-4400-98df-705d2b3222b2 · inbound
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ddb875a3-dde7-4df8-9d81-f33f23f2a9a1 · inbound
The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df1f535c-3613-4506-b444-84dd3fb8bcab · inbound
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f0840b4a-20cd-425a-bafe-30c7d1e93f58 · inbound
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 90e8a6b3-0f95-4722-878e-f9142e2cb60c · inbound
Hierarchical Reasoning Model Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97238aa9-57a3-4003-bb15-0ac9bee3ee39 · inbound
Generalizing Verifiable Instruction Following Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc458fcf-f3f6-4432-8fef-b3bd6e6ba1de · inbound
The Serial Scaling Hypothesis Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 25803ded-892f-4eb0-8dfd-08c9acb5dbfa · inbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e3b7d35-cd44-463f-816c-8df179fb8adf · inbound
Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3956e337-5322-46db-8da2-37a0c21ed02f · inbound
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c47e22c4-2d7e-40b8-8d5e-d98405d5ef39 · inbound
XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8483788b-ad20-4095-8610-8555719921a9 · inbound
Base Models Know How to Reason, Thinking Models Learn When Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea65617-6196-4d4f-b4a5-6101a2cf85e6 · inbound
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d9e54dc-572f-4166-88d6-ec1dd1e6bd86 · inbound
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53825121-318c-45f8-bb7e-7751ca1e669c · inbound
Sharpness-Guided Group Relative Policy Optimization via Probability Shaping Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e67659e-2962-4cb4-b066-412c13cbcf6b · inbound
Differentiable Evolutionary Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8e1f2154-235d-4662-a4c5-651862baaaac · inbound
Learning to Discover at Test Time Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b947e40-a8bb-4b4d-94d6-e2c353ac01c7 · inbound
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a2ac95ec-0a92-453c-8de6-d1fffd89561d · inbound
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e02310b-c160-401d-bfba-c18ca0d98734 · inbound
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 70b8840c-b6a1-45be-bf21-4a5bd9a7919e · inbound
Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb29fe4b-54ac-46ad-98ed-3658784c21f7 · inbound
On the Emergence of Implicit Curriculum in RLVR Learning Dynamics Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e95cbf-e28e-4198-8946-86cd7942f5d0 · inbound
LLM Reasoning with Process Rewards for Outcome-Guided Steps Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 35d66e6c-ec7c-4bb5-8dac-5e3c9cdfb688 · inbound
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3a50fc4-141d-4be0-b4d1-641d27cead18 · inbound
EvoRAG: Making Knowledge Graph-based RAG Automatically Evolve through Feedback-driven Backpropagation Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c07c4949-cf9b-460e-8ba6-b2caae029302 · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f85758cb-421a-4647-89f9-8b07480f7489 · inbound
Infection-Reasoner: A Compact Vision-Language Model for Wound Infection Classification with Evidence-Grounded Clinical Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b790f4ce-5587-4045-9221-7d056d9ed2d5 · inbound
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6fdd3abf-f46e-40fa-bd5e-85ade0fe9de3 · inbound
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fadce319-d934-46f0-8fd7-16be1facb75e · inbound
Cost-Aware Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5bc1f85a-1970-4c0b-81f5-2f27ea2b4094 · inbound
Selector-Guided Autonomous Curriculum for One-Shot Reinforcement Learning from Verifiable Rewards Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff2b9387-b00b-4e4b-8b53-46ad76ec21e7 · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1358d144-2135-42ce-bcdf-bbe2578bd8eb · inbound
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3dce258f-0d01-4839-8fc7-223aaee4a1a6 · inbound
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7228c3f8-c4ae-447c-b9cb-859a9fdd0000 · inbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3e41791-5ee7-48b4-83ad-651960f75e7f · inbound
Gradient Extrapolation-Based Policy Optimization Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ecfe04e5-a551-4d23-b476-380bd5e7b7c8 · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5791f443-db00-45ab-a72f-23b3cb693270 · inbound
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 11f12a6c-80c8-4004-8747-c40e4f230b73 · inbound
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 153
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50a71dc9-cb22-4650-8e27-9eb12bd1e9a5 · inbound
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0429e0cc-5908-40f8-bcef-8d0ec3ef85da · inbound
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation faf44c2e-2dfa-47c6-b0b4-10407353a272 · inbound
Reasoning Can Be Restored by Correcting a Few Decision Tokens Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3eb97e40-a127-4330-bf81-4d1738208df2 · inbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b90a97b5-82a6-4100-81b2-cd3ab7c3b356 · inbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c28d3c4-e146-47c6-a316-0e2d7420c068 · inbound
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bdfe8046-8767-4940-bd26-4f8f068069a7 · inbound
Hide to Guide: Learning via Semantic Masking Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e13180f9-a5d3-4ef0-889f-94b394a5a387 · inbound
Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 56b0ad6b-cd6a-47cc-877f-7df9ddc22f11 · inbound
On the Generalization Gap in Self-Evolving Language Model Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4bb9c7f-1c18-4bd3-8f4c-0b2b03593d09 · inbound
Trust Region On-Policy Distillation Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc009b3b-49f5-433c-b105-9600622c3361 · inbound
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 37259542-0c6b-45a4-8074-17d174f934df · inbound
Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f495052-3063-4e91-a5b3-fbd4725fbd91 · inbound
Select and Improve: Understanding the Mechanics of Post-Training for Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f1bbcac1-8a10-4861-89e9-2b8793c6b517 · inbound
How Post-Training Shapes Biological Reasoning Models Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ca75886f-c9a8-4ae9-a44c-325a75358f3e · inbound
Continual Self-Improvement with Lightweight Experiential Latent Memories Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29db8f2f-1ba7-4855-8206-960c5dcf44e1 · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92c70472-23b5-4c9f-a708-67a086483e8f · inbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2dabf5c4-f7db-4e0c-9067-85ec7897adfa · inbound
Prompt engineering using order-of-addition experiments: An application to generating two-level fractional factorial designs Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b01ef1b-9cf8-4dbf-9158-8d55ba54e72a · inbound
RLVP: Penalize the Path, Reward the Outcome Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e762360f-4ca1-45cd-88a9-6af9be4ffd7b · inbound
TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9b913fd-6f47-4257-9f9b-58bbc8cb07c8 · inbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 1945
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5beffa7d-7656-4943-b9ae-fb73df30609f · inbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.