Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2502.11886.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:32.385679Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:27:36.759210Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e60bd50d-4186-4064-8860-005c7aed1ee1 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models LIMR: Less is More for RL Scaling
Reference 276
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8653614e-8e49-4678-92cc-08bf6d9a94ad · inbound
ToolRL: Reward is All Tool Learning Needs LIMR: Less is More for RL Scaling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9418ddc7-8284-4590-81cc-d8bc205438cb · inbound
The Hallucination Tax of Reinforcement Finetuning LIMR: Less is More for RL Scaling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0dcaede-b418-4ca1-974f-bb7ca676e1a3 · inbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMR: Less is More for RL Scaling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f9b73f4-d090-4ab2-93a1-d8eea78a0da2 · inbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning LIMR: Less is More for RL Scaling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b3dd73-f9b7-4c0c-abc1-d33bcdacbf4e · inbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning LIMR: Less is More for RL Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3903c7a-4f40-4b10-bb8d-1917083a9ea1 · inbound
SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning LIMR: Less is More for RL Scaling
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba80164b-d1a3-4ad6-b274-f52d49fe5572 · inbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMR: Less is More for RL Scaling
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 622b3240-6a49-4b5d-ba52-4a2c48cb0687 · inbound
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs LIMR: Less is More for RL Scaling
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07818c08-b6e3-4a43-a672-edcb1c042834 · inbound
SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis LIMR: Less is More for RL Scaling
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c790190-4f4c-4b4f-a4a2-43d7f28d5c46 · inbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMR: Less is More for RL Scaling
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b8cbe4-6afc-4a59-89fe-b898730394fb · inbound
How Far Are We from Optimal Reasoning Efficiency? LIMR: Less is More for RL Scaling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b2ffa4-f2eb-4445-bf66-67a400742f9d · inbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning LIMR: Less is More for RL Scaling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 960841b0-a8eb-43b4-ae90-d6e7ef6b0273 · inbound
Test-Time Scaling with Reflective Generative Model LIMR: Less is More for RL Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd2e522-1f8c-4017-9ab6-ba400e554300 · inbound
PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization LIMR: Less is More for RL Scaling
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b894d50c-c028-4d0c-9c0d-fb4703238dab · inbound
From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization LIMR: Less is More for RL Scaling
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd345c14-e1f0-4cdf-9d3b-80b3094c52e1 · inbound
Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization LIMR: Less is More for RL Scaling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b49024a-7409-4cfe-968d-bb5055889cf4 · inbound
FormaRL: Enhancing Autoformalization with no Labeled Data LIMR: Less is More for RL Scaling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7d6496c-3f15-4a00-b1e7-a273af4f0373 · inbound
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning LIMR: Less is More for RL Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dcc5459-bfb3-4a16-b94e-a0db283a9778 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models LIMR: Less is More for RL Scaling
Reference 289
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3560fc99-bac1-426f-a0dc-af3069e0e8f4 · inbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts LIMR: Less is More for RL Scaling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeaa8e7a-bb26-4076-88c4-0d3f498b73b9 · inbound
ISExplore:Informative Segment Selection for Efficient Personalized 3D Talking Face Generation LIMR: Less is More for RL Scaling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 79bc85bb-ab27-4304-95fe-6380e188238e · inbound
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models LIMR: Less is More for RL Scaling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b585bddf-8bd8-4bdc-bbcd-5039a2452c41 · inbound
Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing LIMR: Less is More for RL Scaling
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a9a9ae24-6dca-4b71-9d82-5ae649740511 · inbound
Can LLMs Learn to Reason Robustly under Noisy Supervision? LIMR: Less is More for RL Scaling
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f9d1155-8352-4010-98af-7708be565f98 · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment LIMR: Less is More for RL Scaling
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 11c49c28-d055-4fb3-83ec-ea8d36bce8ba · inbound
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes LIMR: Less is More for RL Scaling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1509c011-042e-4c26-97d2-6a9e59acfd0a · inbound
POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation LIMR: Less is More for RL Scaling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 27709112-a375-4f60-bd93-a7fc88a51927 · inbound
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e3478c03-e531-46c3-bb5a-fdfe21ce7e40 · inbound
TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning LIMR: Less is More for RL Scaling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3f4abc9e-1b48-467f-8ca3-87fd222fcaf4 · inbound
TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning LIMR: Less is More for RL Scaling
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61c068fe-f799-4ba1-a511-d6928464fd8e · inbound
Gradient Extrapolation-Based Policy Optimization LIMR: Less is More for RL Scaling
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91d7a693-0bc5-4807-b903-a3f5c693b9df · inbound
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards LIMR: Less is More for RL Scaling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9a08857-37e8-4d7c-bba7-23faace8eda1 · inbound
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation LIMR: Less is More for RL Scaling
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fed1f1c4-eb3c-429c-a80c-6f075a26b65a · inbound
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance LIMR: Less is More for RL Scaling
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e06bc65f-d8c7-46c4-9278-272df101a202 · inbound
Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection LIMR: Less is More for RL Scaling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e27e7df0-802f-4eea-a729-f5b9523d4823 · inbound
Trust Region On-Policy Distillation LIMR: Less is More for RL Scaling
Reference 285
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 83f0ada9-65b0-4860-b6ed-92bbdbe41d9c · inbound
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning LIMR: Less is More for RL Scaling
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 89a16f38-dd58-40f5-96c6-ac16c8f5eeeb · inbound
Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning LIMR: Less is More for RL Scaling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3843b153-84fa-4fee-90af-fe3a27f6222c · inbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR LIMR: Less is More for RL Scaling
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd9d861c-ef6a-4fd8-b5e1-be805572c6a1 · inbound
Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning LIMR: Less is More for RL Scaling
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac8c575-74f0-4695-83a4-4668ce82174b · inbound
A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) LIMR: Less is More for RL Scaling
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.