Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:44.311377Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2507.15758.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:44.311377Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:04:38.759385Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T19:58:53.662557Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 63efb1d1-3055-422a-bacc-3f2a525f4774 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2ae177-ebe1-478c-a172-115c2560f6d9 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Thinkless: LLM Learns When to Think
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc79b765-1550-4f56-949e-23c0b32a0c8c · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615c17f0-de06-40d5-aa00-90439ad8af3b · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Token-Budget-Aware LLM Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b866691-815f-4263-bdb7-5c74eda232b8 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc127bb-2f9d-4e59-97fc-161b990fc451 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Measuring Massive Multitask Language Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c80df40-f92a-4d48-9390-0400681285f8 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Hapo: Training language models to reason concisely via history-aware policy optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5628c428-fb7c-4db6-843d-608cd7a3b9a0 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Overthink: Slowdown attacks on reasoning llms
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40826666-312f-482b-9b61-0a3ca07a52ce · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e7f71cc-cee9-4d46-84ad-e4eb1a61b23c · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 031ce5b9-3f8a-48e5-a7ff-64797412ca69 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe021649-68fc-4f31-969b-6cb2d99e2b6e · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization s1: Simple test-time scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca9559b-1e6c-4fd5-938d-e6263e64c308 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization RouteLLM: Learning to Route LLMs with Preference Data
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd09387e-8315-438e-b042-0a555fa54d84 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Concise: Confidence-guided compression in step-by-step efficient reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64582e5a-def5-485d-9a17-831dec3742b0 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16de9413-fb86-45b2-9678-08447fd5fd0d · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 970106e8-2575-4f0f-b7b5-89e0db836bf9 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d660cd-42a0-4987-9d75-5cd43b6ef237 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88529ffc-21a4-4fa3-a5bc-debeba036124 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f248839-594e-4aa9-a5cc-a8be95e04fe2 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e39cad-253f-40a9-9eae-35936d284e34 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Tokenskip: Controllable chain-of-thought compression in llms
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdba8db0-4a46-44b5-93ec-e6e4ed21c265 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Chain of Draft: Thinking Faster by Writing Less
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a6ea9b0-bdc5-43fc-a0ff-04b2443f1d8b · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization AdaptThink: Reasoning Models Can Learn When to Think
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 429ce97d-9577-489b-b6e0-91645c229703 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e228e5ed-4276-4971-b943-3801b8bf9c36 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4002c37a-0be0-465f-ae43-b3bb56358af2 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Demystifying Long Chain-of-Thought Reasoning in LLMs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ca42b8-4c01-42fa-b60b-a0580f4d084d · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Learning to Route LLMs with Confidence Tokens
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7167743f-36ee-4791-80e2-9212c2d9a929 · outbound
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe9cc27-c640-41c8-b2d0-218cc82279ba · inbound
BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79333e2d-250d-4ace-aee3-0a39cbf8206b · inbound
StaRPO: Stability-Augmented Reinforcement Policy Optimization LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 66042c5c-5030-4a85-830e-16688a93f121 · inbound
Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.