Pith. sign in

Paper Citation Record · LEDGER

LIMR: Less is More for RL Scaling

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2502.11886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.11886 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:32.385679Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:36.759210Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e60bd50d-4186-4064-8860-005c7aed1ee1 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models LIMR: Less is More for RL Scaling

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.483767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:46fbc4c0ded93e0f9c6354b3f2f22dc5bbf2dcff30574341a5c8731a375a012a

Observation 8653614e-8e49-4678-92cc-08bf6d9a94ad · inbound

ToolRL: Reward is All Tool Learning Needs cites this paper.

ToolRL: Reward is All Tool Learning Needs LIMR: Less is More for RL Scaling

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T00:26:48.445619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T00:26:48.291431Z digest=sha256:04c95ccd9d3f5dd0f14812ce050efe53d95aa1f4570fd9e76f59e465c7438b55

Observation 9418ddc7-8284-4590-81cc-d8bc205438cb · inbound

The Hallucination Tax of Reinforcement Finetuning cites this paper.

The Hallucination Tax of Reinforcement Finetuning LIMR: Less is More for RL Scaling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.385679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.385679Z digest=sha256:1211678bebb8bc286b7e5ddc734a93ad28dd4a072bd607e8b698cd6575207dfe

Observation d0dcaede-b418-4ca1-974f-bb7ca676e1a3 · inbound

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning cites this paper.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.361350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.361350Z digest=sha256:5e6f46373c100a4a16fc3e854dc1cbe22f5fd027287895d91b2621424234cbd3

Observation 6f9b73f4-d090-4ab2-93a1-d8eea78a0da2 · inbound

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning cites this paper.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:41.222191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:41.222191Z digest=sha256:15784539dac4578f75fab765a118731f675e4bbbd5de5493609f0756f067eb8b

Observation 12b3dd73-f9b7-4c0c-abc1-d33bcdacbf4e · inbound

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning cites this paper.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning LIMR: Less is More for RL Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.358276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.358276Z digest=sha256:e5c8a2fb6b0c1a5ded12395185f27982e530b46630a75160dbe6e36dc2f0666f

Observation d3903c7a-4f40-4b10-bb8d-1917083a9ea1 · inbound

SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning cites this paper.

SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning LIMR: Less is More for RL Scaling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:32.954737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:32.954737Z digest=sha256:00880179c82a75c7918084bbd4391a1a5dbb989f2185f386417b252fcac3c1d9

Observation ba80164b-d1a3-4ad6-b274-f52d49fe5572 · inbound

HardTests: Synthesizing High-Quality Test Cases for LLM Coding cites this paper.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMR: Less is More for RL Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.430228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.430228Z digest=sha256:070bbd85a96f6f7e320e3379d4547437ca1ef82c193d723feceedaf3eb784c65

Observation 622b3240-6a49-4b5d-ba52-4a2c48cb0687 · inbound

Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs cites this paper.

Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs LIMR: Less is More for RL Scaling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:06:25.160677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:06:25.160677Z digest=sha256:3a7482283521277b1bed96173efa5e74e171942f132737c1f211616bf2dc3aef

Observation 07818c08-b6e3-4a43-a672-edcb1c042834 · inbound

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis cites this paper.

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis LIMR: Less is More for RL Scaling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.327280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:35:45.327280Z digest=sha256:da03a7e4889e13d9630f689f169b00afc7cc403abf96e9dd0e55f9e0e0bcdc6f

Observation 5c790190-4f4c-4b4f-a4a2-43d7f28d5c46 · inbound

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts cites this paper.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMR: Less is More for RL Scaling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.876617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.876617Z digest=sha256:8aaf668abd5efe25a7511b21710a073a5c4e14c0326b2e17fe614e6a24bdbbcc

Observation 86b8cbe4-6afc-4a59-89fe-b898730394fb · inbound

How Far Are We from Optimal Reasoning Efficiency? cites this paper.

How Far Are We from Optimal Reasoning Efficiency? LIMR: Less is More for RL Scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.467346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.467346Z digest=sha256:6049a6bc5d07ae4aba7d28b86499aba11d8d856f49c243ebb1ee5efa9086b373

Observation 67b2ffa4-f2eb-4445-bf66-67a400742f9d · inbound

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning cites this paper.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning LIMR: Less is More for RL Scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.145061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.145061Z digest=sha256:bae26e4aeabf0eb1e56115030ae386496765a0f95e1f1f4a16b7262914c140d1

Observation 960841b0-a8eb-43b4-ae90-d6e7ef6b0273 · inbound

Test-Time Scaling with Reflective Generative Model cites this paper.

Test-Time Scaling with Reflective Generative Model LIMR: Less is More for RL Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.775814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.775814Z digest=sha256:8998c3661110c11b96ab4a04c36a8d015a1e05815061aa30a7987601d46009da

Observation 0bd2e522-1f8c-4017-9ab6-ba400e554300 · inbound

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization cites this paper.

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization LIMR: Less is More for RL Scaling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:12.582797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:12.582797Z digest=sha256:ca238b15990243ad3a55d8ff86c73ff802ba976e446f9c3796e0fd79153b5c8c

Observation b894d50c-c028-4d0c-9c0d-fb4703238dab · inbound

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization cites this paper.

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization LIMR: Less is More for RL Scaling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:36.105216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:07:36.105216Z digest=sha256:f1d3da0009474c72f55fdb8c8c5c478fcf22417a434578974945d5525c64a25c

Observation fd345c14-e1f0-4cdf-9d3b-80b3094c52e1 · inbound

Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization cites this paper.

Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization LIMR: Less is More for RL Scaling

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:31:53.188452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T22:26:52.748349Z digest=sha256:ef6a2d778a392e293c10363a08713616c1a7e1d869709c6b671533d0c8b2b340

Observation 9b49024a-7409-4cfe-968d-bb5055889cf4 · inbound

FormaRL: Enhancing Autoformalization with no Labeled Data cites this paper.

FormaRL: Enhancing Autoformalization with no Labeled Data LIMR: Less is More for RL Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:11:23.150929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:11:23.150929Z digest=sha256:2a53e492d200225ed06b0a7df298055c7c1f75abd8598b0288890f024c923720

Observation b7d6496c-3f15-4a00-b1e7-a273af4f0373 · inbound

Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning cites this paper.

Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning LIMR: Less is More for RL Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T22:56:02.527244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:56:02.527244Z digest=sha256:0a8ac5ac3aed11e433ccf110a580bd9479a5987e20ca6a252750501c445d5e8b

Observation 8dcc5459-bfb3-4a16-b94e-a0db283a9778 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models LIMR: Less is More for RL Scaling

Reference 289

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.776325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:c2b3ba7f8bc07eb0fdf5d6da29c88f467857eb4ff8736aae00ce4019502a0a95

Observation 3560fc99-bac1-426f-a0dc-af3069e0e8f4 · inbound

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts cites this paper.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts LIMR: Less is More for RL Scaling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:46.397811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:46.397811Z digest=sha256:3ac4877f66f4fb56c1707c73ef3ae87a68a89b44d5d21f36a03246c2aa6d72ac

Observation aeaa8e7a-bb26-4076-88c4-0d3f498b73b9 · inbound

ISExplore:Informative Segment Selection for Efficient Personalized 3D Talking Face Generation cites this paper.

ISExplore:Informative Segment Selection for Efficient Personalized 3D Talking Face Generation LIMR: Less is More for RL Scaling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:10:32.307963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T00:08:28.730095Z digest=sha256:533759ecdea8cf5981ef850c4e5f3f2226f86b81b264c1314686ba19534ca5cf

Observation 79bc85bb-ab27-4304-95fe-6380e188238e · inbound

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models cites this paper.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models LIMR: Less is More for RL Scaling

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.194777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:09:26.842696Z digest=sha256:c28433bcad658c4af1dee57f276933f017834eb68cc2adadb5206afe505a26bf

Observation b585bddf-8bd8-4bdc-bbcd-5039a2452c41 · inbound

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing cites this paper.

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing LIMR: Less is More for RL Scaling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:27:36.744086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:24:50.793194Z digest=sha256:c0d45c3c37ba5a057c214684df01ae73a7592f0fd2b9546b860105cd0440e86c

Observation a9a9ae24-6dca-4b71-9d82-5ae649740511 · inbound

Can LLMs Learn to Reason Robustly under Noisy Supervision? cites this paper.

Can LLMs Learn to Reason Robustly under Noisy Supervision? LIMR: Less is More for RL Scaling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:01.331952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T16:58:42.129870Z digest=sha256:931a1a6202057d05ca19148ff39870e18fb13c91f0df123c50981453d79a9867

Observation 9f9d1155-8352-4010-98af-7708be565f98 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment LIMR: Less is More for RL Scaling

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.517501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:616f782187cbb5af0c075d82c0b4e9188b15b6ef5da7b2c4878c87a273bad7b4

Observation 11c49c28-d055-4fb3-83ec-ea8d36bce8ba · inbound

Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes cites this paper.

Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes LIMR: Less is More for RL Scaling

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:10:09.661693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:55:18.468593Z digest=sha256:f2772abed7ad7bb7f4bd86b10767801f2287c3497988e94b47b5c65698e9f08e

Observation 1509c011-042e-4c26-97d2-6a9e59acfd0a · inbound

POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation cites this paper.

POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation LIMR: Less is More for RL Scaling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:15.529008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T04:37:41.629935Z digest=sha256:a2bac4115289c1284ece0dacdf94f735711127be0bb68df4df99b27d68810461

Observation 27709112-a375-4f60-bd93-a7fc88a51927 · inbound

Cost-Aware Learning cites this paper.

Cost-Aware Learning LIMR: Less is More for RL Scaling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.935572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T05:11:01.131590Z digest=sha256:a57a3a319d0216c14aaf492d069e775917b4ff0422d71ee96fd1c9def56b584a

Observation e3478c03-e531-46c3-bb5a-fdfe21ce7e40 · inbound

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning cites this paper.

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning LIMR: Less is More for RL Scaling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:16:54.879375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T07:15:28.648869Z digest=sha256:f6cbb719829e577cfd0c2b48b53cef9781f036ac570ca84faed0263dcd2ddf0b

Observation 3f4abc9e-1b48-467f-8ca3-87fd222fcaf4 · inbound

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning cites this paper.

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning LIMR: Less is More for RL Scaling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T02:26:14.089922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:26:14.089922Z digest=sha256:99bf139998bd7e99bed287c398ab32c3cbf953c2eff3ab0b9398110a5cf2f05b

Observation 61c068fe-f799-4ba1-a511-d6928464fd8e · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization LIMR: Less is More for RL Scaling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.798009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:8227dbbce43da6da4af504d359aff7ac9cc3d720cceccf9034dd2f420ef8827d

Observation 91d7a693-0bc5-4807-b903-a3f5c693b9df · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards LIMR: Less is More for RL Scaling

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:30.076748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:b8e16a5678f3350f4e789dab5aa885c0028be28ccf40e2c5fefe84baa70d98f7

Observation d9a08857-37e8-4d7c-bba7-23faace8eda1 · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation LIMR: Less is More for RL Scaling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.290990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:6504ac848e405b35c24ce44c0a26fb186dfff2d5e035a6c3e8f5714e89f7065a

Observation fed1f1c4-eb3c-429c-a80c-6f075a26b65a · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance LIMR: Less is More for RL Scaling

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.101924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:1b2fad2a860cd3161ed24f0c3588938871f53861bdc9d51977cdd8ea83e6d8ad

Observation e06bc65f-d8c7-46c4-9278-272df101a202 · inbound

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection cites this paper.

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection LIMR: Less is More for RL Scaling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:13:30.088668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T14:08:40.968105Z digest=sha256:179a0a9d1c4a7bad9cb77dece29dfa92c7b5703d59a2098e5bedad762a0d534a

Observation e27e7df0-802f-4eea-a729-f5b9523d4823 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation LIMR: Less is More for RL Scaling

Reference 285

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.626999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:fec79ba0cca298a213b1dedc9553409d9aab76c4762fc3ee1b6d0852d7915d98

Observation 83f0ada9-65b0-4860-b6ed-92bbdbe41d9c · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:27:36.760518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:e7ccc2a1cb4ac0a53527dae0c95db54642c2ea0968a46baecc220981bffd1c0c

Observation 89a16f38-dd58-40f5-96c6-ac16c8f5eeeb · inbound

Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning cites this paper.

Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:14:27.040572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:14:27.040572Z digest=sha256:70d82ba2f24b6ecc5828d6b4e3bbfa3264a5ecb16a0895c4bd14f89cc213983d

Observation 3843b153-84fa-4fee-90af-fe3a27f6222c · inbound

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR cites this paper.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR LIMR: Less is More for RL Scaling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.453562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.453562Z digest=sha256:162e7a69134f7e42e94aa4b6b615de17c54cd321ef40a63c8b217f0ef9bbff2c

Observation fd9d861c-ef6a-4fd8-b5e1-be805572c6a1 · inbound

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning cites this paper.

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning LIMR: Less is More for RL Scaling

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-31T01:31:11.009307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:31:11.009307Z digest=sha256:82dd808cc6e0e367eb075dc5f199fe38726f34ba3ae42476caded9f166b7b23e

Observation 7ac8c575-74f0-4695-83a4-4668ce82174b · inbound

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) cites this paper.

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) LIMR: Less is More for RL Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T01:07:16.546910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:07:16.546910Z digest=sha256:081d9ee310ffc39b25f8a179b794ce5d372cab42a0e0df0e3c26f18a547e196c