Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:38:47.682695Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 7 inbound Pith citation observations for arXiv:2505.18086.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:38:47.682695Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:09:39.636062Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T22:46:20.251611Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c9a7d3d8-c24b-4338-a122-6e5ccd4fa911 · outbound
Stable Reinforcement Learning for Efficient Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe93923-50d0-41cd-9438-741854849ba4 · outbound
Stable Reinforcement Learning for Efficient Reasoning Scaling Laws for Neural Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ca8bd60-b550-43a9-b38e-1fa1042351e7 · outbound
Stable Reinforcement Learning for Efficient Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecb8034f-13b8-4298-b5b8-19ee820e5f47 · outbound
Stable Reinforcement Learning for Efficient Reasoning Qwen3 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad775c01-2cc7-44cc-8f75-7f57e29df208 · outbound
Stable Reinforcement Learning for Efficient Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513424c0-8561-4ecf-946a-cf6e63e53b55 · outbound
Stable Reinforcement Learning for Efficient Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f10c50-543c-4187-9aba-4dc66166503e · outbound
Stable Reinforcement Learning for Efficient Reasoning Training language models to follow instructions with human feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d224063-e1c8-4345-b429-4bca1e0e5344 · outbound
Stable Reinforcement Learning for Efficient Reasoning Proximal Policy Optimization Algorithms
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 945cf57e-8403-4d7f-b180-3496e2dbec1a · outbound
Stable Reinforcement Learning for Efficient Reasoning Group robust preference optimization in reward-free RLHF
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2de9ee18-9ab1-4ab1-a720-4f5ab0da98dd · outbound
Stable Reinforcement Learning for Efficient Reasoning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09bed22b-4186-4cc2-9e0c-4b544f4540bf · outbound
Stable Reinforcement Learning for Efficient Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a05546bc-dc99-4b26-8b01-9c0784e78b85 · outbound
Stable Reinforcement Learning for Efficient Reasoning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b91f5d9f-a288-40d2-8523-5518f8dd2466 · outbound
Stable Reinforcement Learning for Efficient Reasoning The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8b41cd-4bc4-4d55-a372-8f305df9f6cb · outbound
Stable Reinforcement Learning for Efficient Reasoning When More is Less: Understanding Chain-of-Thought Length in LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88487cad-08c9-4ba8-96d8-177b0f0632c0 · outbound
Stable Reinforcement Learning for Efficient Reasoning Dynamic early exit in reasoning models, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6ebe74d-386b-4f74-9ac0-e03020456fba · outbound
Stable Reinforcement Learning for Efficient Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaaf58df-ce8c-448f-a445-8adae971953e · outbound
Stable Reinforcement Learning for Efficient Reasoning S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab6bda3-e684-4120-8286-37f5ceb9ed7b · outbound
Stable Reinforcement Learning for Efficient Reasoning Training Verifiers to Solve Math Word Problems
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43b2069d-6689-4980-978e-fa6673d4508a · outbound
Stable Reinforcement Learning for Efficient Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f63de3ea-d04c-4cd4-ac5d-e3a980e3cb62 · outbound
Stable Reinforcement Learning for Efficient Reasoning Aime problems and solutions
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f898c5e0-3baa-4776-a942-9b7e0c2f03a1 · outbound
Stable Reinforcement Learning for Efficient Reasoning Amc 2023, 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af2ae48e-fb79-4463-9aec-7e25a8fbac07 · outbound
Stable Reinforcement Learning for Efficient Reasoning Measuring mathematical problem solving with the math dataset,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37b165fe-3c43-458e-999e-5afc1c52a00f · outbound
Stable Reinforcement Learning for Efficient Reasoning Learning to reason with llms
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71d0c005-ec14-442e-885a-eded273fb183 · outbound
Stable Reinforcement Learning for Efficient Reasoning Training language models to follow instructions with human feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0da1767a-63d3-4bbe-acac-d4186ae8de5c · outbound
Stable Reinforcement Learning for Efficient Reasoning On Designing Effective RL Reward at Training Time for LLM Reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143a71ca-8286-4c91-9cea-81e9229fdd6f · outbound
Stable Reinforcement Learning for Efficient Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff791d8f-0479-4963-a508-1360be495822 · outbound
Stable Reinforcement Learning for Efficient Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6d38db-9d42-4123-b1f9-235377011f6b · outbound
Stable Reinforcement Learning for Efficient Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13461ca3-1766-4edb-a8b9-c4016710b253 · outbound
Stable Reinforcement Learning for Efficient Reasoning Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca3c6a0b-18fe-4b0e-a2cf-36a820f24d15 · outbound
Stable Reinforcement Learning for Efficient Reasoning Training language models to reason efficiently
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 616be23c-0f78-40a6-8881-8a637368c55a · outbound
Stable Reinforcement Learning for Efficient Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 445767f5-5ddd-4cc7-9d3a-4bb0c17866a8 · outbound
Stable Reinforcement Learning for Efficient Reasoning Adam: A Method for Stochastic Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456d2ebb-02e8-4cfb-b296-56ce53d3e3d9 · outbound
Stable Reinforcement Learning for Efficient Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac5f1b6b-6f50-4cac-89de-0faf5272030e · outbound
Stable Reinforcement Learning for Efficient Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df1278cb-0481-40ff-81f4-9e0813881153 · outbound
Stable Reinforcement Learning for Efficient Reasoning Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d005cb89-ec4f-4219-824f-e9cb83aa415a · inbound
Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security Stable Reinforcement Learning for Efficient Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83c4f67c-abf3-4632-be6f-9ddc3cb4bbdb · inbound
Gradient Extrapolation-Based Policy Optimization Stable Reinforcement Learning for Efficient Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32eafdbe-ed52-49a4-b269-446d843604d6 · inbound
CLORE: Content-Level Optimization for Reasoning Efficiency Stable Reinforcement Learning for Efficient Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6967ca88-cf8c-4aac-a953-bb8e16f81ff5 · inbound
Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning Stable Reinforcement Learning for Efficient Reasoning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2163f506-be2c-4659-b578-8de3e92f0899 · inbound
Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af84df61-620c-4d7d-b0a9-b19de84984bb · inbound
Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef30e1dc-c058-4e94-95c1-46f0948119e0 · inbound
Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.