Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:04.508009Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 16 inbound Pith citation observations for arXiv:2505.15612.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:04.508009Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:08.292468Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:19:43.881659Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 28e84fc4-fc33-4282-ba75-0814213730e5 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cdb23a8-2f9d-4a88-b874-010ff7f73628 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Arora and A
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee27329-6959-4997-b431-f94e55efbdc8 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad411e10-4157-4f26-81fe-82d48ecf5652 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b32e92-0252-450b-8771-e9bd405a767c · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Gandhi, A
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3d3cf30e-68a5-4ad6-8741-267b5a3249cb · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Training Large Language Models to Reason in a Continuous Latent Space
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b40c05d-5a62-4179-9807-fb0ab96c2062 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af46f05c-9118-4372-a39c-0345ca05e32a · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Measuring Massive Multitask Language Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b241c72e-f219-4882-a4c3-24328accca6a · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Hendrycks, C
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b913ec90-d3c6-4440-b6c0-178707ad69a4 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd11b564-bd11-4335-97e6-3cf7d9d1af9d · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954cf2c1-8e1d-48fd-b239-df1e25e27ff8 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1d8b928e-392b-4957-aeb1-1c5fee32c0aa · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping s1: Simple test-time scaling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 288839eb-69ee-4030-a74f-913083a2f4e8 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Self-Training Elicits Concise Reasoning in Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d730ab99-18ba-4bbf-b24a-fe68527707d4 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping OpenAI o1 System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c36ddf6-0e6d-4bf7-b3dd-4867468d26a6 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d08d321-e896-4279-95d7-dd0755e96a78 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b27b57e0-8994-4bc7-874c-23c3118b2407 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb88eb2-64be-4b8b-a1f7-82935f1ad521 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Proximal Policy Optimization Algorithms
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35381598-08c1-420a-bc5c-2d6a5cc66f97 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc3775df-960a-4b0a-baaf-26fe0d688ef5 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping HybridFlow: A Flexible and Efficient RLHF Framework
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474c8158-6c00-4b95-a7e6-e9f2fa52867b · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a519d24-8e85-4981-8899-c72fa7b103f2 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bcae8ac7-cbd8-45d8-8249-037a2be47198 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 515b6d5c-4d55-4420-a9a5-dfc1e798df75 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 20148de8-f039-471e-bae1-d6fbfbd23660 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 41384598-14da-41dc-ae07-b941ab3cd98d · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dfbb9aa-1fd5-4817-b7d3-6c317d6bb337 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a01d5355-f674-45c6-923f-75fd472e93ab · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping <think>...</think>
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 577e7356-dde9-457a-becc-c5fcc438a479 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Wait, subtracting a negative is like adding the positive, so that would be 3 + 6, which is 9
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 01a5be31-685e-4785-a785-45382d5c0368 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Calculate \( f(-1) \):\[f(-1) = \frac{3(-1) - 2}{-1 - 2} = \frac{-3 - 2}{-3} = \frac{-5}{-3} = \frac{5}{3}\]3
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e37737d0-0717-4936-91d9-e1195dd1ca07 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80ecce7-12be-4b1d-ab64-ae428b6d2493 · outbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Demystifying Long Chain-of-Thought Reasoning in LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8215d2d0-315b-4add-9aa8-a54da7b0598d · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e2806cf4-e1a3-4291-8d88-34263ede61be · inbound
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb0dbfa-7cc7-4266-9804-628e6c489f8c · inbound
Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db48e9a-15f4-4c36-8b2f-f582bbe38f7d · inbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 57e59cd1-dfe9-492f-b79f-d81a8bceeacb · inbound
Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d3c71b-ce7b-40f6-9879-0ea2168f9d51 · inbound
T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation decffc5e-5dc7-49d2-9828-01f11c1e8ba2 · inbound
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 244
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 621aa321-98df-4cf1-8d43-93fc91d81e1d · inbound
Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1bd3322e-6d0f-47ed-aae6-2d7d2b22e475 · inbound
AIPO: Learning to Reason from Active Interaction Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b3bda08-724c-48d2-8257-5898f42dd122 · inbound
AIPO: Learning to Reason from Active Interaction Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0f8f4104-ab31-4fd6-a7e2-35d3b24f9e7b · inbound
LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9473f81e-b1c1-4bbf-b0f4-d52050b385e9 · inbound
CLORE: Content-Level Optimization for Reasoning Efficiency Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5f010621-c287-4485-9c86-3830856bcd32 · inbound
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c6993438-c84d-4363-b3cd-d682e4cf7a33 · inbound
SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4422fa3-a96f-48cf-aaa0-9b986e39183d · inbound
CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5c542a3d-e8a9-416e-a70b-6da1d6e9e7af · inbound
Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.