Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T00:19:22.140009Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 58 inbound Pith citation observations for arXiv:2503.04697.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T00:19:22.140009Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T00:18:54.255096Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T12:15:01.137692Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4fa81d11-7996-46ed-8743-588e7aa90531 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Let‘s sample step by step: Adaptive-consistency for efficient reasoning and coding with LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bb9e500a-512d-4bcd-90f5-e9a1159aa4cc · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning doi: 10.18653/v1/2023.emnlp-main.761
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0175b9e3-d0be-475d-a0ba-0bfcab57c2ae · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Training language models to reason efficiently
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8457cd01-76f5-4779-8d85-06dfbc2b3392 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Precise Length Control in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 693e5b07-7cca-48f0-b816-fb3b36404d09 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1dd3c76e-bd63-468f-9143-dfa452f4a92c · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 906baf4e-c8d2-41ed-9100-66a0d22d47f7 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ca3a1670-4346-4d94-9258-7c191e6aa7e9 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Measuring Massive Multitask Language Understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T01:08:06.256034+00:00.
Observation f679074e-85d3-4ca5-9804-e9bb25ce8be0 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Self-Refine: Iterative Refinement with Self-Feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 970e4475-3336-4207-bfad-bb842a42f4d7 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b03cabc5-fb79-45e5-b738-f67594039a59 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Cand`es, and Tatsunori Hashimoto
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6889fe3d-f329-40cc-ad02-9442b8617701 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning s1: Simple test-time scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 51bca921-17ac-442e-a70b-b5b802bdf172 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Qwen2.5 Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f6542abc-4698-49bf-908a-6df105a069b8 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ca35c44e-b644-488a-abf7-f99c6e96115a · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b57e7a6a-64d3-4f8a-b014-b0291c7ecc5c · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b9360124-0598-4a90-8084-6954ed8f2422 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 36eda330-1dec-4839-b241-50e9e4941b91 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c1ca1916-90d6-425e-9400-bf4524171c3a · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 293b0fbd-a5c3-4c8b-891c-47ef59a14796 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T01:08:13.648188+00:00.
Observation 758d6a05-6046-44ac-83b5-af68e33c09c1 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2be1dd12-3a24-4ba3-bc55-51dda81dff4c · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 12234fc8-e450-4c76-b010-59c7d10c6538 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T01:08:14.037057+00:00.
Observation 1d7e3f5f-15e0-41f1-bb05-ad519c6b1cbb · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Following Length Constraints in Instructions
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 16882dc6-a0b9-46f2-89f8-3a2bbe4acf2b · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9958e3e8-5516-4bad-a013-495a4c099ce4 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 22f01b3f-12f3-42f8-b33d-5fc46c21dac5 · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4faff704-0b5c-4030-bbd5-fe6f709c50ef · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning These results highlight the effectiveness and potential of LCPO to scale to even larger reasoning models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e9e9129a-40a4-46fb-b373-6b9e9ab6df1b · outbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Thus, the maximum real part is approximately 625.6
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2d8dbdd4-e35d-448a-9846-9108b89328fc · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2f529aae-b5d1-4972-941d-275476032a54 · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1a32c312-b75e-4471-b2cd-b264538e1092 · inbound
UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 817ac0f4-1ee8-47a8-befb-7b032f40701f · inbound
Reinforcement Learning from Human Feedback L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 192
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 18446a07-4e0b-493c-94a6-4ccc267d50cc · inbound
The Serial Scaling Hypothesis L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1fd82ba4-ebcb-4b60-997f-7bf0b3e2a295 · inbound
Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b2ec4e48-900c-46f7-9d15-d3368ee9c60b · inbound
Self-Aligned Reward: Towards Effective and Efficient Reasoners L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1107291a-a3b1-42ef-968f-dd80ab2f007a · inbound
A Survey of Reinforcement Learning for Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5aeae402-b595-4ac6-9205-079dbeecfc93 · inbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff693d59-b172-4e5b-a32a-fe37107681d4 · inbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cd61e5c2-b858-4ead-b612-3acf27a2c984 · inbound
TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c6678ef3-6f58-47d7-84ec-448e66f9b3fe · inbound
Think When Needed: Model-Aware Reasoning Routing for LLM-based Ranking L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6924bb75-9a4e-4f58-9547-652c19bc877e · inbound
ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1323ee15-722b-4546-b5d3-f3b96a4aae2a · inbound
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1644e321-7682-4755-bb86-b97c71eed2ad · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f59d64a-cd9f-4bd5-ae3a-7e17876ea507 · inbound
CRISP: Compressed Reasoning via Iterative Self-Policy Distillation L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0c860a27-a842-400f-b339-ab2fce5648d3 · inbound
Efficient Reasoning on the Edge L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3c2db13-f82f-4ffc-8961-46c6d0502710 · inbound
TiCo: Time-Controllable Spoken Dialogue Model L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 39be36f2-4c63-4c89-ba81-63a78e240b29 · inbound
Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e3c3b6-7117-43ec-ba3a-ec6518019675 · inbound
Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ae0c2924-0e93-4178-9fae-b430b5453e30 · inbound
AI Achieves a Perfect LSAT Score L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9e3c924e-b2de-40c9-b9df-3535548e64b9 · inbound
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fffc6d1a-6ad4-448b-88b6-d16ea9b7d2f2 · inbound
NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 803a5ec3-ef33-4874-93f2-2af4b5f69543 · inbound
Reasoning Compression with Mixed-Policy Distillation L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation df7b3897-72d3-407c-ae08-636e4bc3c5c5 · inbound
LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation aca0d2ea-0c44-49e1-b212-ddc0ad5e0407 · inbound
Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5a6d6e28-7551-4c74-8fea-92619d3c1a59 · inbound
STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 023a4cb0-f4b3-4b31-bb7f-d58fe5f2b03a · inbound
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d78ba50b-a751-4e15-abc0-ebbc5afdbf05 · inbound
Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 312bb44a-2a3e-4037-8b2b-3602306f0a2f · inbound
Mem-$\pi$: Adaptive Memory through Learning When and What to Generate L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f2b77133-3e36-4d22-a3b9-9d75b5f433a2 · inbound
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 696136e2-793a-455b-b747-29311ac43f46 · inbound
CLORE: Content-Level Optimization for Reasoning Efficiency L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0c77cd8e-4cbd-49a9-9019-d8dfe8a8da12 · inbound
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f264eb7c-f68a-442e-a821-c30051363312 · inbound
Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f89ebc58-fb66-410a-a04a-f1c084358a2c · inbound
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0fe43f3e-450b-47fe-9999-357dd4ae2610 · inbound
Selective Latent Thinking: Adaptive Compression of LLM Reasoning Chains L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bcb92eb5-13f0-4a01-bddd-76825724e788 · inbound
SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1a962ff9-7b82-4e6a-acc6-3009f4b6c7ae · inbound
Trust Region On-Policy Distillation L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 245
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7ffc2a5e-d264-46fb-b4b3-cd022fbf7bae · inbound
Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 27c7c1cc-dd59-47ac-8d31-8150c27874e9 · inbound
Adaptive Latent Agentic Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation da594ade-0083-4092-aa11-9d7cec5744ae · inbound
Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6a17d518-577d-485f-93ec-a358ec9531a0 · inbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a9b2f759-910e-4fb0-a6a8-1ff9d0254e77 · inbound
SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ede45ea7-fef4-4441-b872-83bbe816f0e5 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fbabde79-d719-4723-9e30-c79d3e012503 · inbound
Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e700ee3d-95c6-466e-b68e-4c00cc452bf3 · inbound
Finding the Time to Think: Learning Planning Budgets in Real-Time RL L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 88044151-cb10-4bd7-91d0-dd242c5ba572 · inbound
Finding the Time to Think: Learning Planning Budgets in Real-Time RL L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4ea7314e-e216-4389-88cf-15489acf6310 · inbound
LASER: Load-Aware Serving with Early-Exit for Reasoning LLMs at the Edge L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3b3ce192-d2c2-4d9e-a1a7-9850e6fcaf07 · inbound
CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fc8a6eaa-0aec-49cf-8452-22b4f1c2fea2 · inbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f61310-7bfc-4e5a-98d5-f0c3f2e26e29 · inbound
Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e2a4b3ad-323c-44bc-a3c7-dcde0069c807 · inbound
Length Penalties Make Chain-of-Thought Less Monitorable L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c5c03a6-9d1f-48da-8fd4-5589d1c98456 · inbound
Length Penalties Make Chain-of-Thought Less Monitorable L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca6bead-6994-420f-a141-59d3a0c82426 · inbound
Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41008c75-52f7-41f4-80a7-895bdf913391 · inbound
Is Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum Alignment L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a3a126-65b9-4054-bf34-fadf898fb3d6 · inbound
LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d07c7f28-408c-4824-8cbd-92ddd48df36d · inbound
Masked Distillation: Internalizing the Chain-of-Thought in Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f23f478a-906b-4dba-8968-13c4034c7211 · inbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.