Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:21:08.269218Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 44 inbound Pith citation observations for arXiv:2508.10751.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:21:08.269218Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T22:03:28.124820Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T21:00:08.457909Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 20fe78a0-c3b7-4dba-9e54-00d8c30a84d5 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f65368-aa33-417e-a8cc-39db8d830f48 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Aime2024, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 311c3a56-09ae-41db-bcc8-c81911226bee · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Aime2025, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 62f72c5a-6235-4654-bc26-fe632c7dce51 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models A Survey of Exploration Methods in Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c1322f-379a-4b7f-a5f3-bdd4942d70a7 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3343a24b-c34c-4c73-b8dd-5dcf52e709bf · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fd4291e8-8156-4eed-a22b-59ee16b57e10 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42cc8885-3790-443d-b3e1-dab8a68371a9 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Improving large language models via fine-grained reinforcement learning with minimum editing constraint
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a097a0b-09ff-4ebf-8733-eafab4a5893d · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99805c56-6e20-4578-aacb-2b4bb2e1fcf4 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Reasoning with Exploration: An Entropy Perspective
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7a6e9fb-9cdc-4e8f-81d3-5e95ae19aa19 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Thinker: Learning to think fast and slow.CoRR, abs/2505.21097, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1139fa17-a670-4e6a-8954-5c8efaa7a883 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d842488-652c-4c9c-8f1b-b8d7477819f7 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c23da36-7ce8-4391-813e-3aa311c32bc2 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c2e054-71b3-4084-bd1a-8c9069851b83 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Stochastic first- and zeroth-order methods for nonconvex stochastic program- ming
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7084e6e2-2365-41dd-a044-67fafd312f2b · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Bootstrap resampling methods: something for nothing?The Annals of thoracic surgery, 77(4):1142–1144, 2004
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation baf71de0-f6fb-4fa1-b7ac-c21f72141e0a · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Seed1.5-VL Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 628887d4-4e15-42f6-bfb4-b9b8d72704ec · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Skywork Open Reasoner 1 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245b89ef-4157-44b5-b6b4-b08d30903107 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac01cefd-0396-434f-892a-51412916c5c1 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e500108-796a-4d7f-aa74-1553e62ba1c5 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ee29c13-b5f9-431d-a2c7-d3841c64ad5f · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Test-Time Learning for Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c53f383-d82f-45cc-81e8-90294d83ced7 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models OpenAI o1 System Card
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84b9ea5-a56b-42ca-a579-831457015b47 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Mistral 7B
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 280a78ad-13ba-4915-9c95-3bed5a25b21d · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ff1928-5d33-47cd-8da2-fce57a9ac1d9 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e909de1-669d-4dfc-90ad-4dee4fa37e3b · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models PP-PG: combining parameter perturbation with policy gradient methods for effective and efficient explorations in deep reinforcement learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation de686465-cf53-4b9e-89c3-2a3fdf6af3de · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31165b7e-a287-40f2-af57-a8ecee7d2dbb · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0e5f048-ce60-4b9d-8933-0fe3d1f6116c · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Understanding R1-Zero-Like Training: A Critical Perspective
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 135cdd04-f9ce-4df2-a4c2-2a1b862eefb8 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Inference-time scaling for generalist reward modeling.CoRR, abs/2504.02495, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63245719-83c1-421f-965f-4e7e4baeb7c0 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Learning from Peers in Reasoning Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eaa3cc7-49dd-4760-9e89-0b0c97512861 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f9c201-1872-4147-bbf2-d17c4396691d · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Nesterov and Vladimir G
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c3ddb77-da42-4ce6-a580-5cc87dd8b440 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 66f0cbd1-9156-4f7c-be0c-512a26e701f1 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Proximal Policy Optimization Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc96f34-873c-4b75-ae35-482082edbde9 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06d7b561-f235-48b0-80e6-e4599b64f9c2 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Spurious Rewards: Rethinking Training Signals in RLVR
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d6b4713-7f46-4386-95c7-6c6a7dfa789f · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02db5137-7dcd-4074-b5f8-5b25a7ef2719 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 871c4546-f675-4c08-b313-9a27eaa3a600 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e997e85-8055-440b-a0ec-1515799f464c · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 628a0a5e-3563-4f67-926e-7662a64adb4c · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d71cc55-8d76-47da-a7e8-3bd0ded1ff70 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Reft: Reasoning with reinforced fine-tuning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96da9560-e314-4886-91b2-7cde506d392d · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f7412a-5d75-498b-b50c-d766cf13e929 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8c9b5026-2071-446b-9b77-2a7369af0836 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bbbfd40-7616-42d9-8616-39e8c6566b31 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Williams
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e5fc915-e57f-440b-96a8-326d3280e5a2 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models ARM: adaptive reasoning model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b52b9ed-e82c-42e1-bdb1-a84c211552fe · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e394acda-e293-45be-98a7-f889ebaef45f · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Qwen2.5 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7ddb09-536d-440d-9862-378a9c939715 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c396b7-6c36-4336-b590-bd8dffcba871 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de3cb2e2-c7c8-4fea-b012-22ad3ae2bae9 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d6ef5ab-efdf-4d0c-b7e4-1c4240bdf064 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Boning, and Dina Katabi
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcac152f-aad7-4d74-a694-d4bd78ebca09 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 635baa46-0aa2-4c0c-83e7-d034e1aa98a7 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af82a7c5-9336-4f8c-a046-38fcc0f9ade4 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7303c30-23fe-4d3d-afb2-afc9e91ecb71 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models The surprising effectiveness of negative reinforcement in LLM reasoning.CoRR, abs/2506.01347, 2025
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daf99f1d-3ef1-47e9-b666-abb418f71612 · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fac393c8-196a-40fc-aff7-6f40df93604c · outbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models TTRL: Test-Time Reinforcement Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c40f1373-96d7-47d6-bf88-a85fa7fb7c7d · inbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 799af097-4862-4a75-989d-849b0b503152 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d36f566e-f01d-46dc-9179-90ccd70433a4 · inbound
Outcome-based Exploration for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b8b181-2879-4fdd-9eca-4e0f414cdbfa · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88eb6109-4516-431f-a947-b6785bebb238 · inbound
Emergent Slow Thinking in LLMs as Inverse Tree Freezing Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 191f6cdb-d00d-4350-822a-e95904af13b6 · inbound
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce897c27-c9b9-450c-adff-e09bdc6df006 · inbound
Beyond the Sampled Token: Preserving Candidate Support in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0672f5f-5a20-4c0a-a0c0-18135e3042d1 · inbound
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 365887bf-2e01-45cf-b80d-f1c14d8774aa · inbound
Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad4def12-7ced-403f-bfb8-dd4bfafe21ec · inbound
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 493c6899-d3c8-491a-8f31-7eff9bddad66 · inbound
MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ffb3f9a7-bae6-4387-a39d-f03a74925aa4 · inbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f9a438e-d767-480b-9e90-66e45d68f524 · inbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3a3e0fe9-f636-46dd-a3e8-249b9504625c · inbound
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 51162d63-f47d-4295-aad7-d13027ce7f2f · inbound
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b4bac14-8d4b-4e1b-99f9-62665a56480d · inbound
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation daa2a35e-d39a-4ae5-b48f-d4d49bffec98 · inbound
PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2396714-c593-4147-b06c-da2776362bf8 · inbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a6579577-05d4-4a44-a098-bf48c50f0bf6 · inbound
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed419b70-ffc5-4eae-bf39-c060302b0ee9 · inbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4efd0409-3d61-4fff-90f7-0358f16c0b19 · inbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a1a4a03-0132-4c2f-b4ba-3a3c836b1e46 · inbound
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f26dcbd3-d925-49b1-91fb-13e1c3be1936 · inbound
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dfe95361-ed25-42cd-8246-efab17eaac41 · inbound
Finite-Time Regret Analysis of Retry-Aware Bandits Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f35f1a59-919b-493e-b88b-69a2bfa9def1 · inbound
Finite-Time Regret Analysis of Retry-Aware Bandits Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42e13782-b429-419b-8dc1-0dce22bdd636 · inbound
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97047d31-dc3f-44fd-ad7c-ac45de92f9ca · inbound
Residual Skill Optimization for Text-to-SQL Ensembles Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f655d136-4f42-443e-b1fa-a80c77dff8a4 · inbound
Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc4b7227-bd3d-41aa-811f-d456642b1cf9 · inbound
Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19363c93-9a7d-400c-9e42-7115da93686e · inbound
Retry Policy Gradients in Continuous Action Spaces Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b02f5bbd-79a9-4cf7-8857-b3f6bccb225e · inbound
On Advantage Estimates for Max@K Policy Gradients Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 742e7457-58dc-4252-90bd-f8328b9c11b5 · inbound
OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eeba6203-dd94-4814-ae96-c7756f3bf42c · inbound
SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 31da5b42-3f07-437e-aab9-7e9d4d976486 · inbound
REVES: REvision and VErification--Augmented Training for Test-Time Scaling Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a43c264f-4835-47fe-934d-7d24b231d5d5 · inbound
Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation daa031ff-e73d-4e70-88eb-e64ef10d3383 · inbound
On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3aac177-1822-4da3-9390-aeb872f3e1f0 · inbound
Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65a40f4f-3a54-45ee-94f8-cb0002af75a6 · inbound
DecompRL: Solving Harder Problems by Learning Modular Code Generation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 735e13ff-2756-48e3-88b5-344b01ad6948 · inbound
Spectral Rewiring for Exploration, Purification, and Model Merging Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11a1f182-415f-45d8-a3ec-9b9c8e7a7bd0 · inbound
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e1b247-a286-4779-8f1e-f44820c54e7c · inbound
Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53d17f62-7644-4900-8a92-e26913548d27 · inbound
TaPR: Test-Aware Policy Refinement for Feedback-Conditioned Code Generation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f503e5d-6900-4a9f-9df7-401fa10d48df · inbound
Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53186b75-65a6-4686-bc35-4da457bbb35e · inbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.