Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning for Reasoning in Large Language Models with One Training Example

As of 5 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 61 inbound Pith citation observations for arXiv:2504.20571.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.20571 v3

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T19:51:04.779597Z

measured 130 of 130 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:55:28.597027Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact40
  • verified fuzzy22
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch6

External citation measurements

7
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8463121b-b1d5-43f8-bcd5-c0ca222f5325 · outbound

This paper cites Learning to reason with llms.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Learning to reason with llms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.177303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:5501fe61ee587cc7bb23490318f94e287a06b573243503beb34d7467a3e690e4

Observation 843fc5fc-219d-41b6-8d90-e2ff72eedd0a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.844671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:16d5095869ae1a58da2cc37a1edbfbd80d884d3dde1a138cff3222dd1fcd4f4b

Observation 5411608b-f66e-48e2-bd8e-c4f42fd216c6 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.852051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:ef56f30f674d5885f302a1ac62daace5e6876732276dd727137aff2a51f623c3

Observation 3d72fe2e-f683-4894-93c7-68d3f028cfd3 · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:04.859644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:d2230feef497adb99c63bbb83bcfc61353cd2f9665a0a9018294eefa188f9e57

Observation 87b81421-b53a-40ba-91d2-106b8c5f8f5c · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T19:51:04.866291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:b583af29a3955d5e9c49e74d990317b94cdd26f17bee996b218629bd9a259031

Observation fc6ddccd-e9fc-452d-b0ec-e119f35f12a8 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:40:33.380427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:7eac1be9c39c059f45628bf886bf3cd2a475f8538780b254a6b912542c722490

Observation 36c6a299-11ef-4e56-89d3-bb545b01c390 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Proximal Policy Optimization Algorithms

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.879886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:b2390e4eef47490421425178889d3c2e13daaddcb1c7987bc0217a932b0f1347

Observation b51d1bfe-0d26-4b9a-b229-dc4efe278036 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.884980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:13a60f9cec00086c9c52aeb2d13362bb0e5ef3de0fec3e567a230dacb9aa77d3

Observation 5c437fa1-e43b-4fe2-9ac4-1f9243aeb264 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:04.893040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:4f0ccb958906823a552206b5ace1ebc4c8b98b8fd6b75fe6c44ff1651fcc3545

Observation fac2e47b-4999-44bc-8193-6a25d9126a7b · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:04.900537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:7bc273e587d2a866e58923176da8df5336ae291dcdde38793310fc5d255dd743

Observation ce80f86d-eeb6-4241-9a09-21e0f4e91a77 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.906397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:9344a57c9a7f1979478025906627c8a87f50b0d1513c81c5fc8d885f1c8c6e72

Observation cee44b9b-7798-4947-ac7b-86376f501c6d · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.931329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:fda6b93e0855cd3956a2394444cee272e50cda70c19e7fc11f317a6a4be0f37e

Observation aabc6d2c-69c3-44b2-9775-d671920d3e6b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Understanding R1-Zero-Like Training: A Critical Perspective

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.937895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:d14b966a60cccd618f811f09ed49bde86fc515519f74dccded9cef665025dbb2

Observation b0bc476b-3f85-4b2a-a01e-ed8bf304e5c0 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.140747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:fb52468c6ebc654a5c364cd46d2d3e5fea86f88938b2493e5bd662e31e21c084

Observation 4fe3e517-7543-4635-a352-ec1c2a920bfa · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.086065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:440da6ced8e18e2f1aacbf9b88aa2df73489a76619b333e6024a883b3f3f735b

Observation 2979b2dd-a0bb-412b-8df2-a5ccdccbded3 · outbound

This paper cites Srpo: A cross-domain implementation of large-scale reinforcement learning on llm.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Srpo: A cross-domain implementation of large-scale reinforcement learning on llm

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.150567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:feaee7350fa501bb9a7116a9a7f345a29807d9ecb6f597e5856ac56fa2b87d87

Observation 13576480-c537-4a47-a2e0-917124a0ea70 · outbound

This paper cites Numinamath.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Numinamath

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.154941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:9900bdf2a4c1a5b742934b29d73da52c7a0d5a1265f902f3509bc486498d7dbc

Observation ea1dd23a-3b73-4412-8edb-4727ac7f3c05 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.159497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:3c36e961bb6710bf7d7ddc8c1f2f20e7ebf0e43a60fddf861c4973e5149827d1

Observation 1e0c9d2a-a9ab-408f-9fea-7357485475cd · outbound

This paper cites Limr: Less is more for rl scaling.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Limr: Less is more for rl scaling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.164285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:409efebad7448894978e75498b2465fd5bf7a3ae4abdeef450dbd0b945b43dce

Observation 23003ecf-b9d1-4578-aa82-057a56a42395 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.091802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:92bc5ebe29ccaa9b6172b6d58be6828a1d6ceee4334aa351e7d7f3702066b290

Observation 6e0d4383-d98e-4329-bae7-b28ca559b054 · outbound

This paper cites Rethinking Reflection in Pre-Training.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Rethinking Reflection in Pre-Training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.097590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:128b71315e6bed17b34b46aa14c0b0a4b57cf9cfef6d502188bc76653a2273a4

Observation fcf169b3-66db-437c-be42-8238dcde9af3 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example HybridFlow: A Flexible and Efficient RLHF Framework

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.102918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:530022529986bd4a9757158cd15c6a90119c35f1bf09eff84b9345936f6dd533

Observation 84cc12e7-1966-4fc6-98ae-125f2872f1da · outbound

This paper cites What makes a reward model a good teacher? an optimization perspective.arXiv preprint arXiv:2503.15477.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example What makes a reward model a good teacher? an optimization perspective.arXiv preprint arXiv:2503.15477

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.108901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:f77ad8c2108323c3be671163b3f6a7c304c4aa90a0675388514a57dd8e19dda6

Observation 1abeb772-11dd-479a-a13d-9435f7103148 · outbound

This paper cites Qwen2.5 Technical Report.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Qwen2.5 Technical Report

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.114377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:4fe67fd0f13220c34b8a6a958a663e4ba7b234f19fbaa1dd6dd0f89154ae4560

Observation 4ddd9e76-925a-49f3-aa88-d8bee5e54ec8 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.120340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:37d55de7d4451f28245a028114bb5a2af25526ee49132d506edcd7e00b66dadb

Observation 5d24e96d-549a-4e6e-85c2-fc981f7cfb4f · outbound

This paper cites The Llama 3 Herd of Models.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example The Llama 3 Herd of Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.125944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:71e0ecd7e8d15402f2c4ad1114340befc4ffc0f6a6b46ff734b8aa4abeddc51d

Observation c3de5f4a-ed65-41d2-a7d9-77010e4b3036 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Measuring Mathematical Problem Solving With the MATH Dataset

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.131015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:5b05a5d1cdf35723488bd553910ac7843ed389702099a3fbf3888ebd0d0b5cd3

Observation b67c5265-ec96-4f5c-aeb9-ffb0ac0c2141 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Gonzalez, Hao Zhang, and Ion Stoica

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.203363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:32a1c4699923867543c8d373b7a9ded5f1961a6b9b0d6977f5946f371ff63bba

Observation 13bec755-0736-4380-83aa-244082484b84 · outbound

This paper cites Let's Verify Step by Step.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Let's Verify Step by Step

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.135648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:9fedd3a07682d0d724440a32e43eb611f00f3a10f2643d35896fd943fd866a58

Observation 3e9ef792-e0c8-4d1c-b4a1-086e03aa070c · outbound

This paper cites Aime problems and solutions.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Aime problems and solutions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.211331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:2e9bb9325c04b7a5f2b8985f7167c834f0cf2d9cd50e5481738e229b3ea6c442

Observation a805ea32-f4e7-46af-aa54-bf931c2ec4fe · outbound

This paper cites Amc problems and solutions.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Amc problems and solutions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.215330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:b9cf7edc9c3d0f1c4a54916ef9ef8b0b1c271740b5c4905001e0d8c63b4be70b

Observation 4f18ef18-be0f-473d-80ca-cb437e8b5a9e · outbound

This paper cites Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.220054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:9a5aca650c6450f9c8760c9ac9645fcac2331ce9e1a6b2e48d1e3ab9099c6e18

Observation efde09b4-e18c-4b22-986a-b3394130ce5e · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.944360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:12ae4ca79fe9f28af4535fca81cce9caac4ae59d80fe481b3c7e0be0b1652eb7

Observation 11cfa2c8-a7ed-474b-9e2d-375d91a32853 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.952016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:f726afa6de39d73b0fa96e803f76af88bc1de81bdf2e44fc55796f7f0302319b

Observation 1b1f6322-5128-4e25-a161-11cfdef840cd · outbound

This paper cites EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:04.958598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:101a969c3381a845a1e8e469670abede0423b45ed7f9593afb7bb1339848e4dc

Observation 8d9ccd95-80b0-4530-a4a4-716df1d57e35 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.965010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:5febad2a369b0c7b7a768228a7dfb38f2fb91d5d48a98f8d709d1972a0d798aa

Observation 56670751-48f9-4bc0-90c8-2dd89773d875 · outbound

This paper cites Deep Grokking: Would Deep Neural Networks Generalize Better?.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Deep Grokking: Would Deep Neural Networks Generalize Better?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:04.972881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:25e894309d2045eb7588f2cc3c35f1c4b841d855111c9cb640820736fd7030fe

Observation 6275d4fe-b964-45e3-97f4-58ec2de7eee8 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Progress measures for grokking via mechanistic interpretability

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.977899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:f26e625ff068323c1601e188dceb4448b77af980877afac7d519b69aad062381

Observation 704c7c1b-a68f-4074-9204-915195302826 · outbound

This paper cites Towards understanding grokking: An effective theory of representation learning.Advances in Neural Information Processing Systems, 35:34651–34663.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Towards understanding grokking: An effective theory of representation learning.Advances in Neural Information Processing Systems, 35:34651–34663

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.173391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:c704e39b584989f1b75b89cb313dee1ecab39eb6aae58028d92e46449ed22e54

Observation 4d49d725-10f8-4629-9d37-103edb2dd138 · outbound

This paper cites The Complexity Dynamics of Grokking.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example The Complexity Dynamics of Grokking

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:04.983727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:b1b70bdfadd592b7c2a09d4a9009fccff38e944929f01e6c8fc27892a86647be

Observation 6bad6672-554e-47bb-ae4a-f37db1727bef · outbound

This paper cites Grokking at the Edge of Numerical Stability.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Grokking at the Edge of Numerical Stability

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:04.990155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:efdecdb48a3eb950a663eac831315bbd6df8660734100c4b63a799fc811b833b

Observation 3a7af56d-e27e-457f-9d2e-58d10736a594 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.995171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:0760ba88adf84d878021900388052939e280fea10fafcef6033f0ce99ca62d37

Observation 89e88595-6b82-4a51-ba4e-fabbdd4e0f50 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.002577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:48e52bd26d6097f11dbcbba42151e45640d652b60e8919705780446f62208931

Observation 4132bbca-c88a-47bb-8343-c8be44aaf6cc · outbound

This paper cites Fastcurl: Curriculum reinforcement learning with stage-wise context scaling for efficient training r1-like reasoning models.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Fastcurl: Curriculum reinforcement learning with stage-wise context scaling for efficient training r1-like reasoning models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.008534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:24b21660556eeccac69312a6eb481d5d2924df5f57499c2e6cdd47f6411fa4ab

Observation 24c1b142-0484-4d58-9960-dbe23dc2652d · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.014418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:e5baf696b5c82dcd0dc0a075e13ef3ef85091ed76d0b288b58d10de85d7f4535

Observation f872eae1-12ef-47fa-8932-3b10e4dffb34 · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.021005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:effeffd7864265d029c54d8695027466e656be38b8e1629b369a11ed56a1d10d

Observation f86d6208-7dff-4623-ba6d-f875029cd0b8 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example TTRL: Test-Time Reinforcement Learning

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.027001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:9e9ff6e79f9eaf3eb5307ee9474f9096175c68f766dc0b64a21d82f28f694f94

Observation 7ad7eead-b72b-4edc-ab2d-e2114b9f9814 · outbound

This paper cites Large-Scale Data Selection for Instruction Tuning.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Large-Scale Data Selection for Instruction Tuning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.034374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:734a4c03e432c4d94e735b469f50f41ab9f63e9a7e4d442dbbe869d4b48468c5

Observation 0c577a14-1fa0-4612-9bb3-c3b69a252768 · outbound

This paper cites Alpagasus: Training a better alpaca with fewer data.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Alpagasus: Training a better alpaca with fewer data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.233005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:8800792c16794b9e82161bd6da007749ef43378e8f4598e3bad424e59ada90ea

Observation 23422147-dbc9-45c9-b1c3-d713982f9adb · outbound

This paper cites Smith, Hannaneh Hajishirzi, and Pradeep Dasigi.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Smith, Hannaneh Hajishirzi, and Pradeep Dasigi

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.236850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:b1420e25fd77b89fcfc6586f4ff0209b4a20e47942917154384389ea0d746b49

Observation 2b50c2b9-3990-4065-b781-9c8d0fa75229 · outbound

This paper cites LESS: selecting influential data for targeted instruction tuning.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example LESS: selecting influential data for targeted instruction tuning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.145356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:80c08339d74ac6227da15714fc7eb8d9cbc3a65faeb841e66bee8279df61d7e9

Observation 300baaab-6177-4ddf-ace3-da23afa1e48f · outbound

This paper cites Active preference learning for large language models.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Active preference learning for large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.168830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:c6145349e07c5284930ee4b291283fdd5c83b31db892b0fd79db5c36bf5da00d

Observation b88e2ac5-cdc8-4102-964c-61babca6a388 · outbound

This paper cites Enabling Weak LLMs to Judge Response Reliability via Meta Ranking.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Enabling Weak LLMs to Judge Response Reliability via Meta Ranking

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:04.837297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:5e399cefe90da037e22353a262778f09e3b779943a7a296674adb498c9d46cdb

Observation 3c1a0ce5-1a1c-4302-ad9c-3deb7fd487ab · outbound

This paper cites Active Preference Optimization for Sample Efficient RLHF.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Active Preference Optimization for Sample Efficient RLHF

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.040421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:74b33dc651e7a0c042d238699381ef67c8eaf3e141e52413ee97feb4c1bb9483

Observation 2db93978-383a-4694-b1b0-05266e629434 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.186078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:26e82fcc76954e1d8af466bc6e4ecf48f56e55d854bc078a59c5bf1f52daf0b7

Observation 33596df9-c1c4-42b4-95db-a08d683a0a96 · outbound

This paper cites Concise reasoning via reinforcement learning.arXiv preprint arXiv:2504.05185.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Concise reasoning via reinforcement learning.arXiv preprint arXiv:2504.05185

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.047054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:f7035750174e802e26aa2c569ed5ba95b46509a1b9fdd21ff2ec86c86564bc81

Observation 6c05cdff-c670-497c-ba7f-d2446bb8316f · outbound

This paper cites Schulman.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Schulman

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.194758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:af0c05ac563aa454320a3355a4e3a0b3a74e4312c693fb43b2cb81c0fe5f60be

Observation 2e341750-2138-46c8-9ec1-da57a5347267 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.052431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:1562ca3303a4cd6ea3a8759aaead3c53665ee308c0ee7ed9e9a6f4bd8f40e9f5

Observation c9c3e60c-fd4d-446a-b4af-14053b61484e · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:35:31.553327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:537375da4649b19affba3c7a154bce30dcb74175522da97a2071582f1dadaf6f

Observation b349fb69-6cc8-4a0c-8902-777cef9d4dda · outbound

This paper cites Skywork open reasoner series.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Skywork open reasoner series

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.224160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:1b667e240d2cdc506cf291a07e9f769bd1634b21b0130f6b131a19feae9da97a

Observation 80b4ccd6-9441-4721-a1b9-ca8877e95559 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example A Survey on LLM-as-a-Judge

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.063839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:c01ab3016162d0286d8a7592b632c3057748ec6264b57516f4d1419f3bf74403

Observation b92831f0-a2d8-4795-b403-82ec72eaa5b2 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.181718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:f7c9db6d4bfcd60572ed68f538de7b926a7d4589444e63cb79059672e7152499

Observation 5875cb2f-c345-4953-8204-9f78a78aee60 · outbound

This paper cites A Survey on In-context Learning.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example A Survey on In-context Learning

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.068857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:b44693737de8b2a515b8831f7dae43238c3be48e9704554e4d734d71ca42842b

Observation e70c6b57-36fa-4b3a-8c0f-fcadd44d65b0 · outbound

This paper cites Deep Learning is Robust to Massive Label Noise.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Deep Learning is Robust to Massive Label Noise

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.074776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:cc04908973e4cf84ebaae1cae77a82e935211281e123108414116e2fd4847a9e

Observation fa047314-872e-4088-a1b4-22ba7582fd50 · outbound

This paper cites Deep double descent: Where bigger models and more data hurt.Journal of Statistical Mechanics: Theory and Experiment, 2021(12):124003.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Deep double descent: Where bigger models and more data hurt.Journal of Statistical Mechanics: Theory and Experiment, 2021(12):124003

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.207458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:67d3fced55482c883b85910e606ee8ed8e89401dc435bc3456e97759945807bb

Observation c635af5e-0242-4d74-826a-1916b9bc5f1a · outbound

This paper cites On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:05.080527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:fcb26f88969d47f3d8a48520c7632f1f5e0632bc71bcf6bd78acdbd2ca972d2d

Observation 4576612c-ff1c-49f0-9e02-b233b0e11fba · outbound

This paper cites Smith, Benoit Dherin, David G.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Smith, Benoit Dherin, David G

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.190750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:32b624dafc31b533ce9f8037af787cb2d94cca18c080ca43a65f1c2cf3e80fa2

Observation 783b5164-98ac-4685-bebd-3fd3e8f8b850 · outbound

This paper cites Acemath: Advancing frontier math reasoning with post-training and reward modeling.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Acemath: Advancing frontier math reasoning with post-training and reward modeling

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:05.199248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:c4d7dea926f1ceea166cd5757a29ba5630b45b350abd24edeeaa0f888d60de30

Observation 61f95623-acc5-4a72-a3f6-ce1371155490 · outbound

This paper cites L′ PG-GRPO(·, θ) +βL ′ KL(·, θ, θref) +αL ′ Entropy(·, θ) # ,(3) where β and α are hyper-parameters (in general β >0 , α <0 ), and “·.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example L′ PG-GRPO(·, θ) +βL ′ KL(·, θ, θref) +αL ′ Entropy(·, θ) # ,(3) where β and α are hyper-parameters (in general β >0 , α <0 ), and “·

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-05-15T19:51:05.228629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:883f608dd56159c8543c347a770893b04cac110e20ac5a5da124d4d74f1a8384

Pith citing papers

Observation 6037df06-3d9a-4400-98df-705d2b3222b2 · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 165

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:42:11.062930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:a921718276b628284782cb837814df523d533c7363d4668691630aa88e56405f

Observation ddb875a3-dde7-4df8-9d81-f33f23f2a9a1 · inbound

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning cites this paper.

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:58:33.535702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T15:58:33.219451Z digest=sha256:8c46796567f5e984e0129358df7b6a82c1caa78f6053e27026c62b6d9915abfd

Observation df1f535c-3613-4506-b444-84dd3fb8bcab · inbound

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning cites this paper.

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T12:12:08.724844Z digest=sha256:8a7631f23f6b02d0a3f51bdda0254e089ff6828d5fa949fedebfd8ee82a0d724

Observation f0840b4a-20cd-425a-bafe-30c7d1e93f58 · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:05:47.611002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:0527c20ae5458706163c80b5296b781dcf6dddcae7aabe38f70008110a6611a0

Observation 90e8a6b3-0f95-4722-878e-f9142e2cb60c · inbound

Hierarchical Reasoning Model cites this paper.

Hierarchical Reasoning Model Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T04:58:03.283911Z digest=sha256:cc9e3d878bb183f65f3880f81ee7dfbe6b5ab652e7f7663dc12d75ecfe3c45d6

Observation 97238aa9-57a3-4003-bb15-0ac9bee3ee39 · inbound

Generalizing Verifiable Instruction Following cites this paper.

Generalizing Verifiable Instruction Following Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:42:05.989110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:39:56.467519Z digest=sha256:96fa52819cb52179297f63deeda80a1aad0c2ba45c23c816980d7cb79ef707d5

Observation cc458fcf-f3f6-4432-8fef-b3bd6e6ba1de · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 123

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:12:02.419405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:9c3b0cc2036ffa268288548a5976d22d4c50a9a3b257b162c776b0e5cfcb35cd

Observation 25803ded-892f-4eb0-8dfd-08c9acb5dbfa · inbound

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization cites this paper.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:16:57.104689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:d74837a11fb082d6227c7ba1923d72b18aab0d18bdfe87f1b18b8a3545480bc2

Observation 2e3b7d35-cd44-463f-816c-8df179fb8adf · inbound

Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned cites this paper.

Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:28.597027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:28.597027Z digest=sha256:8f5a24284d78e8b276d26a1d8bce6fbd34aabe3729d2f211e91407e4c1f72489

Observation 3956e337-5322-46db-8da2-37a0c21ed02f · inbound

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search cites this paper.

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:12:35.904803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T12:12:25.437344Z digest=sha256:7ef55fdd999d4b1073acc74f5bfc47de0c6efb74ac81667b5a49b1b0e5f187db

Observation c47e22c4-2d7e-40b8-8d5e-d98405d5ef39 · inbound

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation cites this paper.

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:11:12.717090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:11:12.717090Z digest=sha256:479620c413acbd8fc70d339abbc3c885187017073aa0add062ea7501579dc9c2

Observation 8483788b-ad20-4095-8610-8555719921a9 · inbound

Base Models Know How to Reason, Thinking Models Learn When cites this paper.

Base Models Know How to Reason, Thinking Models Learn When Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:02:19.585843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:02:19.585843Z digest=sha256:6134c55cbe21ccab81fc8a749c168f915c29bea2c1dda2852608b981faad1712

Observation 8ea65617-6196-4d4f-b4a5-6101a2cf85e6 · inbound

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning cites this paper.

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T07:21:04.419242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T07:20:01.505216Z digest=sha256:3a0a72bf21d6378f1bf67e548a1a652a8c0cac02381d04db380772301fcf98ea

Observation 0d9e54dc-572f-4166-88d6-ec1dd1e6bd86 · inbound

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting cites this paper.

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T08:49:34.341309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:49:34.341309Z digest=sha256:5f6ddb44c58c7595e4748e866cc1bdc10f671596125f7c8149ab7e1cdaf85dfc

Observation 53825121-318c-45f8-bb7e-7751ca1e669c · inbound

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping cites this paper.

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:12:22.151814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:10:57.839146Z digest=sha256:bcc65355dfa3ebf1535654d1e3008cc42bab5378b929e4c80894cf4cba7edfcd

Observation 2e67659e-2962-4cb4-b066-412c13cbcf6b · inbound

Differentiable Evolutionary Reinforcement Learning cites this paper.

Differentiable Evolutionary Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T22:38:37.697896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:9d20f29632606966dd81d61e79d70a36cc4efa969895eaa0f2dc63fe3f3f0d63

Observation 8e1f2154-235d-4662-a4c5-651862baaaac · inbound

Learning to Discover at Test Time cites this paper.

Learning to Discover at Test Time Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:16:04.253884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T05:16:04.001700Z digest=sha256:cd9ad5bea167a0be5a7d5558e5ec470b47ddbcb05dce8cc68042bd536872c3fc

Observation 0b947e40-a8bb-4b4d-94d6-e2c353ac01c7 · inbound

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning cites this paper.

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:27:44.451881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T10:26:45.961575Z digest=sha256:666b22754b58788759c7f34510e9294bbd78317adcf1aa4395d8ad654bf4097e

Observation a2ac95ec-0a92-453c-8de6-d1fffd89561d · inbound

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation cites this paper.

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:07:43.364006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T10:04:44.379128Z digest=sha256:8e8473646ac8d6934d07666bca388a1e4e563d85e3df217d72f4088eec2034be

Observation 9e02310b-c160-401d-bfba-c18ca0d98734 · inbound

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models cites this paper.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:12.992968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:09:26.842696Z digest=sha256:f07435ddd52fc7c6670214c56d8f969794fe87d1ce667ce043ab762c790d9da7

Observation 70b8840c-b6a1-45be-bf21-4a5bd9a7919e · inbound

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing cites this paper.

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:27:36.817013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T08:24:50.793194Z digest=sha256:b2a128b35fdc83e9b6cc5e9e2c7e0eff0c2c3426c918638c0e23e1facafa1bf7

Observation cb29fe4b-54ac-46ad-98ed-3658784c21f7 · inbound

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics cites this paper.

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T23:11:56.603670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:11:56.603670Z digest=sha256:88dda3341f7d1951289d8bbfd0da1d768b95e123d10370fc4a3f29cbf134775a

Observation 35e95cbf-e28e-4198-8946-86cd7942f5d0 · inbound

LLM Reasoning with Process Rewards for Outcome-Guided Steps cites this paper.

LLM Reasoning with Process Rewards for Outcome-Guided Steps Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:47:28.370284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:43:19.869111Z digest=sha256:a03bac2bd13843bd5e55576797b2e7425a07fa7184fa1686ae1d929fbc6d6a06

Observation 35d66e6c-ec7c-4bb5-8dac-5e3c9cdfb688 · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:a2ef219cc99d6b3e5bbc5f03676e60fbab33f3e3fde0b3fd2462ffc13cb2240d

Observation b3a50fc4-141d-4be0-b4d1-641d27cead18 · inbound

EvoRAG: Making Knowledge Graph-based RAG Automatically Evolve through Feedback-driven Backpropagation cites this paper.

EvoRAG: Making Knowledge Graph-based RAG Automatically Evolve through Feedback-driven Backpropagation Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:59:40.497067Z digest=sha256:ff4afaa3701b6deac279e8cb2bfa796b8d5d5e89e13e7f734405c903e3f9b3dc

Observation c07c4949-cf9b-460e-8ba6-b2caae029302 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:96fb8f73388eb1ef8d4f2f9d40ce7c55ba1bc4cf5159efe15478320113a4fda0

Observation f85758cb-421a-4647-89f9-8b07480f7489 · inbound

Infection-Reasoner: A Compact Vision-Language Model for Wound Infection Classification with Evidence-Grounded Clinical Reasoning cites this paper.

Infection-Reasoner: A Compact Vision-Language Model for Wound Infection Classification with Evidence-Grounded Clinical Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:14:24.453124Z digest=sha256:b6babef15cadc1c69c275215c36afeaf08a6f2b2952aefb68beb29231fbb65e7

Observation b790f4ce-5587-4045-9221-7d056d9ed2d5 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T07:00:32.206081Z digest=sha256:822e4428213971a5ce04d6099dedb77fb71d96d2a34bd41ca38850c3d6b9ad29

Observation 6fdd3abf-f46e-40fa-bd5e-85ade0fe9de3 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:49:14.924649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T23:47:53.282259Z digest=sha256:427c42a0ec1f8038ea2c116d2ec74b53bb6210b84b3df3c9480365e0641f8820

Observation fadce319-d934-46f0-8fd7-16be1facb75e · inbound

Cost-Aware Learning cites this paper.

Cost-Aware Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T05:11:01.131590Z digest=sha256:207eefd6f2034f19a367f2983d789f9328af22adf10649394a8a9f081bf66dbe

Observation 5bc1f85a-1970-4c0b-81f5-2f27ea2b4094 · inbound

Selector-Guided Autonomous Curriculum for One-Shot Reinforcement Learning from Verifiable Rewards cites this paper.

Selector-Guided Autonomous Curriculum for One-Shot Reinforcement Learning from Verifiable Rewards Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T17:29:18.280200Z digest=sha256:a7b19bf9eac4599893f78c3200c58c098b4d50a65b0b75a0bedf42f5388f73d9

Observation ff2b9387-b00b-4e4b-8b53-46ad76ec21e7 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:c2834e6095758c0f73ea141a4f20870ff4c6b0456dce082dfd0465f1bf57e6a3

Observation 1358d144-2135-42ce-bcdf-bbe2578bd8eb · inbound

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning cites this paper.

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T10:33:02.741857Z digest=sha256:fccbeaa5fa794b9955499dd1195b773745d9d45eb40620f198f2cfe86ee7c13c

Observation 3dce258f-0d01-4839-8fc7-223aaee4a1a6 · inbound

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning cites this paper.

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:10:54.313745Z digest=sha256:dc89b234468f882cde345c769ee1668b2c48dc8852e89b0c5d5c409bf4876fe2

Observation 7228c3f8-c4ae-447c-b9cb-859a9fdd0000 · inbound

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients cites this paper.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:767e981a483041c525964cdc1fba9b13d2aa7715987d343fc3f7090708a1ce1f

Observation f3e41791-5ee7-48b4-83ad-651960f75e7f · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:bde6cd87b2ffac12e3a2ba3a977b4d3fbea7580d2cd1f58ca66fb3fdc32fe7fe

Observation ecfe04e5-a551-4d23-b476-380bd5e7b7c8 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:069ec0c889f7e45bf6b460361bce6239569d4910a7faa7d88a2f78fa3a192fe3

Observation 5791f443-db00-45ab-a72f-23b3cb693270 · inbound

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control cites this paper.

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:50:42.836549Z digest=sha256:4a078c7a652815ab7c10992928c708e211169e750b5c479141ea93ccc8706436

Observation 11f12a6c-80c8-4004-8747-c40e4f230b73 · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:38c534302d0ccb708cf73f39b5e89787f6bd1055dec4a18ae95789441cc72e8b

Observation 50a71dc9-cb22-4650-8e27-9eb12bd1e9a5 · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:dece893ad1527a68abebd5e5023c1109c5b0eca8c80e32fe56da2a6b27b626c9

Observation 0429e0cc-5908-40f8-bcef-8d0ec3ef85da · inbound

Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting cites this paper.

Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:03:59.293800Z digest=sha256:7ae3c4c9f1105c6af4339c8f732e73d335fcdf881b2be246f4d07c0ec2ca4bd0

Observation faf44c2e-2dfa-47c6-b0b4-10407353a272 · inbound

Reasoning Can Be Restored by Correcting a Few Decision Tokens cites this paper.

Reasoning Can Be Restored by Correcting a Few Decision Tokens Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:57:46.832471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T20:56:48.771058Z digest=sha256:e69a3808040e39ee516268273da9542b3f834d861e8b4c7b56f16729aa97bebb

Observation 3eb97e40-a127-4330-bf81-4d1738208df2 · inbound

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning cites this paper.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:14:03.324992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T08:10:14.041279Z digest=sha256:3fe84455d96b5d40ed9aaf6da20fd75fc37f676799f55213d8438dff6a9c8606

Observation b90a97b5-82a6-4100-81b2-cd3ab7c3b356 · inbound

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning cites this paper.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.284952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:5a4ffd7b8d806dbe545c40585aede484b23d812677dea66982a5ab4383d486e5

Observation 5c28d3c4-e146-47c6-a316-0e2d7420c068 · inbound

Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning cites this paper.

Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:44:38.431639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T05:44:12.807291Z digest=sha256:d402d3fe1f4f47ed4e5e067e4528d6bbd6dba8540f92b3faebe946aa2ab78937

Observation bdfe8046-8767-4940-bd26-4f8f068069a7 · inbound

Hide to Guide: Learning via Semantic Masking cites this paper.

Hide to Guide: Learning via Semantic Masking Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T12:34:39.101921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T12:16:12.108715Z digest=sha256:ddd4c279eb4d8acfbcc56186e7bfc5d2a8158d0b97f30a7b47c2bc3461fc42eb

Observation e13180f9-a5d3-4ef0-889f-94b394a5a387 · inbound

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs cites this paper.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:13:26.565970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:e173c7179f69246aa991c056b0d9714fc7b73a0a8703f706b7a17611fc1da855

Observation 56b0ad6b-cd6a-47cc-877f-7df9ddc22f11 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T17:22:24.111531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:88be55425a455c50682d51dd0ccddf54bb18ec6f744e42ebf14369995849fb7d

Observation d4bb9c7f-1c18-4bd3-8f4c-0b2b03593d09 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.476262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:943d01f980053f252304a71c8283013d69d17faff97e02f211437861f21d066e

Observation dc009b3b-49f5-433c-b105-9600622c3361 · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:27:36.741801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:fb41f55b89baac0667ca3b7f5b634961d42f4a41c6f2df58c001b8966a88db7d

Observation 37259542-0c6b-45a4-8074-17d174f934df · inbound

Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization cites this paper.

Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:27:56.834469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T09:59:45.526226Z digest=sha256:592a030d849595465e2938b315741375112b8f4f7f78572ad8da0710ea32f479

Observation 8f495052-3063-4e91-a5b3-fbd4725fbd91 · inbound

Select and Improve: Understanding the Mechanics of Post-Training for Reasoning cites this paper.

Select and Improve: Understanding the Mechanics of Post-Training for Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:08:21.748795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T07:14:47.302277Z digest=sha256:b98b2e1e20885eae1c3e9d58a0f73184e6aeed098bc018414282f111f9e592a2

Observation f1bbcac1-8a10-4861-89e9-2b8793c6b517 · inbound

How Post-Training Shapes Biological Reasoning Models cites this paper.

How Post-Training Shapes Biological Reasoning Models Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:55:31.008268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T07:48:31.110861Z digest=sha256:a1873a25285919cd97fa15232966af02aad5b3dce529a06860f157329fa175da

Observation ca75886f-c9a8-4ae9-a44c-325a75358f3e · inbound

Continual Self-Improvement with Lightweight Experiential Latent Memories cites this paper.

Continual Self-Improvement with Lightweight Experiential Latent Memories Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-03T19:18:54.805783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:56:44.269388Z digest=sha256:8604a586a20a47d397b97132ec35432f81015e9bc82673d816ee788c31bdbbbd

Observation 29db8f2f-1ba7-4855-8206-960c5dcf44e1 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:38:55.987264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:8ef5e46a85e3a027b6f76e32bdaba01fce6e3e30dc296f0008a04c22a345ae95

Observation 92c70472-23b5-4c9f-a708-67a086483e8f · inbound

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently cites this paper.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:09:44.762957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:48a5ce234cffaf270d9690a970b7dd39478dd85fe049cd9f0265728da8762355

Observation 2dabf5c4-f7db-4e0c-9067-85ec7897adfa · inbound

Prompt engineering using order-of-addition experiments: An application to generating two-level fractional factorial designs cites this paper.

Prompt engineering using order-of-addition experiments: An application to generating two-level fractional factorial designs Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T06:09:35.110633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T06:09:35.110633Z digest=sha256:f9ec61ac453d537daab7128518063a2e0d53bc61d1ea5885e96ddfe6b1d27d89

Observation 5b01ef1b-9cf8-4dbf-9158-8d55ba54e72a · inbound

RLVP: Penalize the Path, Reward the Outcome cites this paper.

RLVP: Penalize the Path, Reward the Outcome Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.387942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:9ce8c3183630e5e6aab11fbc0ee677b6b94485ff9e54b8175960720b4daeeb8d

Observation e762360f-4ca1-45cd-88a9-6af9be4ffd7b · inbound

TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning cites this paper.

TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:54:29.121056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:54:29.121056Z digest=sha256:a373191490cea0c0e2e166ff1a7717d9eda3870a5d1395c935eee5729962ddbc

Observation d9b913fd-6f47-4257-9f9b-58bbc8cb07c8 · inbound

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR cites this paper.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 1945

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.474113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.474113Z digest=sha256:23f26ef1379aac28b576f63ccd8cd7161f48ea221925aea3d37f6f782b63fe7b

Observation 5beffa7d-7656-4943-b9ae-fb73df30609f · inbound

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute cites this paper.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.337255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.337255Z digest=sha256:860a5312dc465ecd1da560cc3f4fa3707b6f7803e3a554a8d87d7970d0b0b7c8