Pith. sign in

Paper Citation Record · LEDGER

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

As of 19 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 69 inbound Pith citation observations for arXiv:2505.24864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24864 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T20:52:33.620041Z

measured 134 of 134 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 69 of 69 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:46:57.916405Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact10
  • verified fuzzy47
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation db55c592-fee7-489c-917f-8a715bd1b81f · outbound

This paper cites OpenAI o1 System Card.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models OpenAI o1 System Card

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:52:33.737148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:f5fe0ee5094b882030e5b832e5aa75087682d3ec98ee199aa9e1de9f6eb3589f

Observation 0064d6c0-ef0b-4571-b920-68cd5c3407bf · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:52:33.724965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:28cade4cfc1062b7d435f265c7ed9625e8fe2987833036e8930655305d13d2bf

Observation fe294fc9-23aa-471b-be50-bca517e8ecf6 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.942017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:a2939c96b16a732caa2345ee9601a03be2def0ce27c50e40b452373605a65e79

Observation 60023cbc-ca21-4ba3-8268-b2039fe12f5b · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Dapo: An open-source llm reinforcement learning system at scale

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.945910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:959b741283f09e1f8c2f85d0ef2ae085da28abae6b67f085a7a035c4c4535ac0

Observation 78dc3227-3d7b-4b4b-8625-fa8c2ce84328 · outbound

This paper cites Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.949604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:e903b73bd63761b7d474bdf6052d4142665d07ef97e996af9e81c3b8c075536a

Observation 26154e3d-2f9e-49bf-8b85-972a14d5fbd6 · outbound

This paper cites Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.953023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:2110897dbc08239b2563c1d861306ce3088e03a53de218280753f4dab6c201f0

Observation 32789a36-fa2d-4743-bda5-459f1daaa73e · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.956593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:4d5976833207b688fda58eb59bb96fb9cd393ace2454052a7e5a07f6c68b18bd

Observation 8da7fc17-dcab-46f2-9cdc-dd5cbbd752b0 · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Code-r1: Reproducing r1 for code with reliable rewards

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.960415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:8318a8557c1bf2450c1220b3fa34a630d4f0a216a23ec5e82df49cc29684dda3

Observation 26f2d617-b911-4953-83c6-a852e7f74399 · outbound

This paper cites Concrete Problems in AI Safety.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Concrete Problems in AI Safety

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:52:33.694039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:7570b521e1ff9ad936be1935443acc1875c5c45922501c0278ffefa3912c6a50

Observation 778cbe5b-0bdc-498f-8034-e4b8237e28b7 · outbound

This paper cites Reward hacking in reinforcement learning.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Reward hacking in reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.963310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:207e1a8a0816df6de5f2453489c879197b17b18adf0a511aeef6e31109c6c237

Observation e03fd634-d6af-4fc2-bc8c-b5d0a3ede3ab · outbound

This paper cites Language Models Learn to Mislead Humans via RLHF.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Language Models Learn to Mislead Humans via RLHF

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.686569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:a1704d29683cfebe26bf45c34558faa784b754a9480fa28a4276442836cd5a73

Observation 6d433f32-0dac-46e1-92cc-be5f136bd599 · outbound

This paper cites Ai as humanity’s salieri: Quantifying linguistic creativity of language models via systematic attribution of machine text against web text.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Ai as humanity’s salieri: Quantifying linguistic creativity of language models via systematic attribution of machine text against web text

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.966236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:25f8f90156447ad6532d86af64d3b176c77e6ba3ee546330239d92ce264b708a

Observation 5838bba3-2103-4dd2-84e2-86795e131750 · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.969536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:d0d0ed65dbe8d77db72676a63a5e7d3308eb77b2d38afef6f637a1890dee6677

Observation 048a6162-14fa-4d30-b7fb-99493383e734 · outbound

This paper cites Zico Kolter, and Aditi Raghunathan.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Zico Kolter, and Aditi Raghunathan

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.742234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:d34dee5bc09cb716d6a7b59c5c11088d6113927aa618367e6efb9f067a2441ef

Observation 0a4f8708-77ed-4e95-a57e-adc1e70dad1f · outbound

This paper cites Echo chamber: Rl post-training amplifies behaviors learned in pretraining.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Echo chamber: Rl post-training amplifies behaviors learned in pretraining

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.746777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:162a7f2277beaebdec410f604b381eea53f327f73bd4a6ed9fc42db7d05d441c

Observation c014312b-3f37-474a-a862-3c91589e9c6a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:52:33.705777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:e944ee234c4ce3c799b0e3c4345c5c4aa9fa389b94a4a709b7810b294979ed8f

Observation 32b866aa-1325-4416-9ff9-42a67780c45d · outbound

This paper cites Proximal policy optimization algorithms.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Proximal policy optimization algorithms

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.751463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:72d485226901d26d7f7a7db00607a1bc852449fcb96c5d787b3958bc6da08a6f

Observation c29eb7d5-5b95-4c6b-b1a7-ebf6863e36db · outbound

This paper cites Skywork open reasoner series.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Skywork open reasoner series

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.755535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:777119dd058648edbb5a8abc80bb8a5369e7d36a7cc0553689f708aa078e9826

Observation a6e6dd55-2273-4630-9067-7ca35653cae1 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Hybridflow: A flexible and efficient rlhf framework

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.760509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:13f76be122179214646ee099d2cbd26888ba7b91c5feb4d3e754551e75a6850f

Observation c1356bb4-87b9-48a0-8078-329a91fad184 · outbound

This paper cites Decoupled weight decay regularization.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Decoupled weight decay regularization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.768334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:2cd007de4459ac3615e6535a3bcd4d28306b607b61e9f83537ae1d0d598362df

Observation a664540d-7487-4b8d-8502-cf019f896747 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.772425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:61873f4b87b0384d842ae02c90a084a0d3dfa1bb5242ecbb488c7d5d00c24e01

Observation 032fb386-a6fa-4320-b3e0-463613682f23 · outbound

This paper cites American invitational mathematics examination - aime.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models American invitational mathematics examination - aime

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.776150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:b705ae71cf0f8fabda3f4e11a8ce6df9f4841a8f017a0081363e80868639949c

Observation 5769f3b9-e674-44aa-b035-0d3a24e4ed02 · outbound

This paper cites American invitational mathematics examination - aime.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models American invitational mathematics examination - aime

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.779815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:ce36a3a97e9fa45882326b887012b9496604e16f4725133745338edc1a0eae8e

Observation 4c78e818-9456-4aad-89d9-c237e427d7da · outbound

This paper cites American mathematics competition - amc.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models American mathematics competition - amc

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.783256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:0cb8ee3c699c05b1d273f335b072ea299bee08d514826562d9bfc15e3f66f2a3

Observation ad0d9e4a-d37f-4c2d-8129-926b53fcc9fd · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Measuring mathematical problem solving with the math dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.789102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:75cb78aeba5dd5dce46e81dcdd59718b5e9a067a5047bc63f835983c0f32343f

Observation 346425fb-271c-451b-b981-072f5b73c05a · outbound

This paper cites Solving quantitative reasoning problems with language models.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Solving quantitative reasoning problems with language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.794425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:061a328b6c8efb955d41c0e7dc6b78a33f54be30c6d1cf84de059021d404b571

Observation 6f05626e-695f-4967-a5d4-1377e44d62e3 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.800330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:649b7fe94159bd37b1c82866b3f65e4eb01779ba7da884f4b921d1e6021ab6b2

Observation 5c417ec9-007a-476d-bba7-1922037368d0 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Process Reinforcement through Implicit Rewards

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:52:33.699361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:d56782bbb566b7fd1eaa6477e759a4df6dd81413fe3ae5e63e9ee4a9d09df308

Observation 91521688-f5ce-4ac0-ae3b-4e9de7211fc7 · outbound

This paper cites Measuring coding challenge competence with apps.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Measuring coding challenge competence with apps

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.804471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:117b79fc95d62f86b8c84a7d6760e5d0d936d721891b4fa05242643e740338bc

Observation 26835d4f-2536-4123-aad7-88ce73a16de3 · outbound

This paper cites Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.809575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:d89e5afeb06c42b0b081e6e3cbd867a10348f9fe8790c51bd3840604aebcc39b

Observation a9b6e0cf-c930-4050-ab40-b9e3d51a231f · outbound

This paper cites Taco: Topics in algorithmic code generation dataset.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Taco: Topics in algorithmic code generation dataset

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.816463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:34cda5517a44461cd06b96bbaef7a192e98da0dae40da20b63ce77c4b9bb33fd

Observation 2ffcc323-b873-41c4-9b99-42734d9af7e8 · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.820615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:72a31c429c041664f4dbd0c69bdd9c5d43aa249f78e8d90f87ef8bf523d20804

Observation 164d5c1f-a245-4616-ac35-257d945aac9b · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.824701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:e31b4ce3458507523739277a3e5d7c421b89c14018b25fa1cda0a97dcb921cce

Observation 530d317b-ed23-4f40-bb8d-e586549599df · outbound

This paper cites an unresolved cited work.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:52:33.829282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:2accb2bc1cd3e8022b1f2845b2820b303fd7d33843ad17f464ca73c30e71d6f4

Observation 10cc91c2-8100-410c-9168-3e433426bccf · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Online difficulty filtering for reasoning oriented reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.833898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:4139871cac04b5759f502d1bfc703c009b6c447944dd35655e22cbe160cc1aaa

Observation 2c2a71e3-91f5-4af9-a98a-ac712c3c4df7 · outbound

This paper cites Instruction-following evaluation for large language models.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Instruction-following evaluation for large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.840161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:975bb3c6ad4503d8689c1492b5415b74b1451aaec520b1a61baf0dc613bf5655

Observation c77aede0-9205-417f-a444-6bc3afe13e97 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.844846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:4075b2ea4eda93f78c574830605730d1c66000208fcf42ef2a191a1e4ed58a67

Observation 8d17720d-5fb7-42db-9826-95e81a245a76 · outbound

This paper cites The curious case of neural text degeneration.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models The curious case of neural text degeneration

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.848529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:49aa5318d9058ba084ffd2ea15ef32cbb604d3a16253a375115ef336dd35a456

Observation 0f67f247-b569-499c-a439-156277766e15 · outbound

This paper cites Stop overthinking: A survey on efficient reasoning for large language models.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Stop overthinking: A survey on efficient reasoning for large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.852649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:65a81a9af9755dafea80d3861f1c49f5482fe22100155ec02057b358a83865dd

Observation 203275a8-477a-4007-8762-adad24320538 · outbound

This paper cites AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.719088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:25fc6aa636d79d6611deecfcf7e3ac60bd7622087dbf3db6a20eeb22306c5463

Observation 76225005-8ccb-4468-b2d9-d0cf9f0ad41e · outbound

This paper cites Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind 12 Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind 12 Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.858216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:854da7981e4d915aa3b9270cf7d8a2a0aecc767874ff0bea0c0176daf26f3e48

Observation aee5fc0f-31e1-4b00-b5e9-49dd4ef86a5d · outbound

This paper cites Learning to reason with llms, September 2024.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Learning to reason with llms, September 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.864157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:378fa3ecc20917e3c2b8a357fa08f645bec1f625d335d06cba4402a511eb2a7e

Observation 45d5ab71-34b9-48c7-bf52-b7cd86a5e60d · outbound

This paper cites an unresolved cited work.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:52:33.868519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:a366868dbcfadcf9bc7844b4d5d6f87dc0f1eab1a90474d7325ab69e4fdd68f8

Observation f63224ff-cf95-435d-b7b6-48b88c80fd25 · outbound

This paper cites an unresolved cited work.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:52:33.872232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:50b9ac7f85b5c55e276d844ad6f8d6c6c89f22a7b0605e7585a8e6a3f0cd20ec

Observation 79d7094c-b39f-4277-a9de-180580feadcf · outbound

This paper cites Mirror descent policy optimization.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Mirror descent policy optimization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.876291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:e689572b93531558c99d2d070dc19b56799db7b3fd93818951d7769bcebb6718

Observation d466648d-2945-453f-b53e-4d6781f8629a · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.880257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:8f2cd27aef350d4450147e1ebadb7572d27ca178e7e329bacb6fb2a595ef58dc

Observation 759eee3b-d102-43a2-b0bb-61c78a1f5e64 · outbound

This paper cites Scaling llm test-time compute optimally can be more effective than scaling model parameters.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Scaling llm test-time compute optimally can be more effective than scaling model parameters

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.884171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:a4a84800290562d56cc79eaaacb232db719bb341e69273b484395a0065eccbb0

Observation 2b8cff30-f1a7-4cab-a1c9-4b1857646c86 · outbound

This paper cites Reinforcement learning enhanced llms: A survey.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Reinforcement learning enhanced llms: A survey

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.887556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:036f04a6db61739f0586427bed5e161f4b020b8a049a9f99a83ddf5f0ae3391d

Observation 19d77b62-eb97-41df-913c-99317297aae1 · outbound

This paper cites Playing atari with deep reinforcement learning.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Playing atari with deep reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.890735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:621847e86d28b2c038c14fd9ae99279be311eae05cfd7cc1efbf5fef507e15c9

Observation 1477d07b-b4a8-4f43-8183-ee41aac0f9c4 · outbound

This paper cites Human-level control through deep reinforcement learning.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Human-level control through deep reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.894533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:516ff8a3a5f1bbe8eeabdb7acf4200e8831131818c15f9ea4c7dfc12031d41a5

Observation 8bdf552e-43e1-41d6-b700-0c12848922cb · outbound

This paper cites Mastering the game of go without human knowledge.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Mastering the game of go without human knowledge

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.898429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:eef171f554ddd0723c2fb6aad4963dfddeab3b6946a4c838293d96c358174898

Observation 9b6e9b70-3ad0-4aa2-8d87-7e3fd4ddb4c5 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:52:33.711723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:840f126d9fbc98e63dd792e592e1cd072279707562a5aa4aaa2d95fc9ca5bda3

Observation 50ec815a-887f-41b1-b2a4-faf5a47ed89e · outbound

This paper cites Llms are greedy agents: Effects of rl fine-tuning on decision-making abilities.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Llms are greedy agents: Effects of rl fine-tuning on decision-making abilities

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.901749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:893ab1a4af24113d0b0968c1afae283541235af0fb8e54275822ea04338e4afa

Observation e4515ad7-1ba7-4bd8-acd5-42d3be0fa1b1 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Star: Bootstrapping reasoning with reasoning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.905512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:e7c1ab3784dac97b7d0eb6481fd17c8a3a43483cefccfa600a32fa106d98129d

Observation 261f2417-ad85-4c6d-be4f-b9651063f3b6 · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:52:33.730999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:9a17c0518eee72ff0b5f0d852b5bb8b6c406b636e9e16aa73428cd67431802fd

Observation ddfc39ac-8803-4be4-9a5a-f6fb9d0158f9 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:52:33.673649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:b26a1e0bfd47048f2a250441ee11fe2d07c981866088a123bc59bbe399ae377c

Observation 0e56cc4d-93a4-4104-9845-e4ea560b15bc · outbound

This paper cites Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.909671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:df181af0e5b2c930cfe40bcf96d465b9b271bcdfac43f12a86bbfc803bae6d97

Observation 3d9b0d6c-2395-4ac2-99f5-7f1d45d95b39 · outbound

This paper cites 0": 1, "1.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models 0": 1, "1

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.913331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:06526fdc130b37169377e9456e360838edfe3f996527b342e0e56674e05dd0ea

Observation 34f53166-71e9-43bf-85f6-fdcae4bbe61e · outbound

This paper cites However, the final answer shoule be a list of action plans for multiple steps.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models However, the final answer shoule be a list of action plans for multiple steps

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.917445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:7108d1b8e5ed16547a8ffe8ac7820fd2d42e3703e054c5cf39f87d283ff579fb

Observation 30701deb-379d-4ed2-a5dd-683c7f9bb4bb · outbound

This paper cites Agent[x, y].

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Agent[x, y]

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:52:33.920950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:078be97cc93b06bff7cd32e96c3130d194d164afa3c0f0558de78dd74416d506

Observation 56225a52-a014-4aff-ad30-f681612009aa · outbound

This paper cites an unresolved cited work.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:52:33.924262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:8635d240b5a2df367c676866f868b3b670ed8b89339d2fe80193434305831d37

Observation 453d596d-c12c-4d87-9695-564725174cea · outbound

This paper cites an unresolved cited work.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:52:33.928050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:589e0b7acde57da899df67bdff57a98d8d38160d18a3e36545b1dcd3c9812968

Observation e5bd4e70-742b-4608-a5e1-f5b034d18afa · outbound

This paper cites an unresolved cited work.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:52:33.931597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:5ebe6052154949979f280c2530902e8675f7804ae3b22a617a2c450f99ce1eed

Observation 739902a1-7692-4b82-a9a8-7948808a1042 · outbound

This paper cites an unresolved cited work.

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:52:33.934957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:026406dec2124d62b3b0bb07f35084de99786b771862397d676784d53c35bd17

Observation c4baeffc-828d-43e0-b236-0b5f2df58d98 · outbound

This paper cites Agent[0.5, 0.5].

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Agent[0.5, 0.5]

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T20:52:33.938690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:52:33.620041Z digest=sha256:f4a276256143fb106c361babab770d9b94181245d3fe0a999c9abd120d7e2d97

Pith citing papers

Observation 2992c4e0-dd2d-42da-8445-a7a0f67d9f0e · inbound

Flow-GRPO: Training Flow Matching Models via Online RL cites this paper.

Flow-GRPO: Training Flow Matching Models via Online RL ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T18:45:16.641012Z digest=sha256:603cca46975b616aaa0fd23005f59a5659cdf666a3399a9a18f009e1ec15fcf2

Observation 6535f3c3-f3f3-410c-addb-de6dcf3795fe · inbound

Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions cites this paper.

Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:19.371511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:01:19.371511Z digest=sha256:6c2ef67eb00803a8a3f41d8e17de51ac517832a079ab1b9eeb3bb88da4882ea9

Observation cfe38482-8f06-465a-97f4-0d92e1ae9e09 · inbound

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs cites this paper.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.538629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.538629Z digest=sha256:8333a2d4ecd4ac9055a8b48af86521ed0d569711058ee9643d3c7e17833265c1

Observation 9945ac94-99ee-4cd5-bc09-518bca2d4311 · inbound

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy cites this paper.

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:35.394606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:35.394606Z digest=sha256:cca2f01806fcbb2d959b8f179cd7af1de480df36d43bc4a51407f92cb7009700

Observation fb35d26b-a10e-49c6-a693-7e217cb88bdb · inbound

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization cites this paper.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.916405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.916405Z digest=sha256:c76e3053da1fde2da66395fdebd0539df594b7e4aac03aab1fff88ee096d5008

Observation f20fba18-42b1-4061-b4f8-a82e513065ec · inbound

DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation cites this paper.

DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:33.390202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:33.390202Z digest=sha256:8c4d4eeb1ee8e847c5e6dff4f42ad3bbfebbc0a86db8a047500c7c5a646756b8

Observation eae0d52c-25d5-417a-9a7f-e699cecff859 · inbound

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling cites this paper.

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T23:40:46.416012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T23:39:39.018498Z digest=sha256:775bcfab093121abada8a55b6c5a9aa518c88863b5ab0f8dc6e72c673dfc70c4

Observation fd7fd3a8-3250-479c-9eee-d7e955d58426 · inbound

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization cites this paper.

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:17.317668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:17.317668Z digest=sha256:0073a159a6e54d212fdac658cf02309d61cfe8bd59889c75398c62eaf901f8b6

Observation da5b0f7a-d27a-42ab-b494-5c6f70157773 · inbound

Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback cites this paper.

Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:36:05.421844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:36:05.421844Z digest=sha256:c3dd4592001c2a2323107907d346e8fb1484ae2505d97a0d9386d41ba909bebe

Observation 0a4edfb0-8eae-4bc5-b7ee-22f942df262d · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-21T23:24:26.257703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:53702d50fbc4c83c560cf8d2bab39db9c120637a693f5f80b1f66ca714b864b5

Observation 1cd65a90-dbba-4744-bbdb-cdc1931fba4c · inbound

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny cites this paper.

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:14.585900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:20:14.585900Z digest=sha256:6f1364ee68dda0a496405c0091f247fb983c46f3a13ad2ad4fdaa55fe5acb38e

Observation 3dd56042-bd00-4b39-9500-092af60b8646 · inbound

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning cites this paper.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.512694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.512694Z digest=sha256:bbbefbdb94fabc6f8e647f15ea968f82f50cf43d3ece3ea4faff0857b3f575f5

Observation 9a4b9eac-311d-45a1-96ca-8af558b20872 · inbound

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation cites this paper.

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:04.044724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:43:04.044724Z digest=sha256:80faa742967c002c85a4090953daf622e4af11a1567080da485412e0b2cd2a90

Observation 0b894881-4190-409d-9d68-3950e4797a83 · inbound

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR cites this paper.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.116441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.116441Z digest=sha256:5614910bd147138d4977809a962a726fc6787241f82ab99ede3ca03b5058e474

Observation de686465-cf53-4b9e-89c3-2a3fdf6af3de · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.634233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.634233Z digest=sha256:bc396fdc2fe65872349d7f880f94f9ccbef98a9961a02ca1cd384939d07fd7ca

Observation 71057d06-6966-4285-b979-283f21bb2a11 · inbound

CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention cites this paper.

CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:36:21.460084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:36:21.460084Z digest=sha256:075ce6e35189d264acb6680fe557729c31a7407f6c793a091fbe63a51ae74502

Observation 55cc0bf5-2d45-4161-b7db-c60155982b52 · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:a3dcfa202b245a30060d9f3f964744cc94cb7eed2402f5410264b51cdc995c6a

Observation aa9b1cef-8270-4200-9f95-350fa9bf2956 · inbound

rStar2-Agent: Agentic Reasoning Technical Report cites this paper.

rStar2-Agent: Agentic Reasoning Technical Report ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.432202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.432202Z digest=sha256:d71f658d751cdb3bbe1398fa99cb13faa014cffe2c9315b8183ce9e7f053a768

Observation 185b89bf-3df2-45a3-ada4-b3de776752b5 · inbound

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning cites this paper.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:23.530235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:23.530235Z digest=sha256:37ad020901266df8e2a37a772ced3a107d46641fc3265fc700f4b1319b13c350

Observation 7031dc5b-c461-444e-a4c3-f9889dbcefcb · inbound

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR cites this paper.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.833521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.833521Z digest=sha256:5b2a4f2bde258313ba39e493616cfc71cd8da1b09c2671bf59f3e71ab52ad8c3

Observation c5a2fb91-77c4-41df-a9a1-7264959a1673 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:1d4f3aa4819d9b8bb42b095d714f7bd2735c05bca436091a05a2b615e6d20179

Observation b6ae3d25-f8a1-4a16-9eda-4c4e34c279a1 · inbound

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers cites this paper.

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:52.919834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:52.919834Z digest=sha256:3f3787a0059ed227cd8942d7cd66a20a337592dd40f658200345717108517fab

Observation e5c41cfa-13ad-4851-be0c-ddde473625ce · inbound

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning cites this paper.

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T08:02:11.189795Z digest=sha256:f62c88c15585a589cc3a748bb3b7b73b6274da6eb2298c8903fc1231973747e3

Observation ed1f0e17-23bb-40a1-9489-9665a333a24a · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:34.887140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:34.887140Z digest=sha256:5643cdf145f28a483e31b1e7fe4efeac6fae461fae541a15cd22f34666e7ce3f

Observation 753d33ce-83a0-434f-8b0c-52661879e75f · inbound

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search cites this paper.

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T12:12:25.437344Z digest=sha256:b51e604d694bb678e23ef119711940e09582d4878e3d8cbb46bb175c8d5a850d

Observation 9619e9b8-149d-40fc-b8af-29a8728bebe8 · inbound

When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL cites this paper.

When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T20:30:35.530451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T20:29:25.874620Z digest=sha256:9a74f703fb18ac846938dcb2c812ee1987448044dd30b6771dfb7549a52067e7

Observation 20c5a131-08e4-4437-a6d9-a166faa0df27 · inbound

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective cites this paper.

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T07:37:12.501489Z digest=sha256:728eb6cea709d7cbd7caffdb54c02306a0b5c6115a169d76d6461b9a9bf190be

Observation 8768ed37-c50e-4366-b06c-99eaf0815fb1 · inbound

Representation-Based Exploration for Language Models: From Test-Time to Post-Training cites this paper.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.979065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.979065Z digest=sha256:e6691ee8e7557d21eacbc5af11361d5d4f61fa44997bea990107b7cfa4703f4e

Observation 150fcd04-9d9e-427d-9dba-c88f8f72ac38 · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:0eb60f78a314de1aa7c3c45da2449ebe72c1cf70e3b4d6f2af0d844ca4ec8a2a

Observation bac8f873-8339-4a8f-9c0d-b90c6164c2da · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:53.349181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:53.349181Z digest=sha256:df46dd41a8aa8461ef4562ffd9682ccfaaa468ada8f2e2790ecbe79805aa7010

Observation f8b28bf1-c0ac-4f36-805a-8d86fc864981 · inbound

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments cites this paper.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.481047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.481047Z digest=sha256:adac61c1bf9d93ae73b11bdfc3d57b67ed414c0d5d1e8d8fc2767a60f9bfdf81

Observation fd087267-0ab3-4580-ab85-1469ec6bfc2b · inbound

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning cites this paper.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:1ea9d61f2ad6d70d90f6e53b5ed74e9e323b9e97a077178aec3cba5ebb5ee96d

Observation 19cb6ca4-e3c7-42a9-886a-29c14fdec193 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.861548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.861548Z digest=sha256:8538d1adb0d7a273478fa1435efddd28a8e55b922bdf9be65f5520edde0f5e3f

Observation f81db7c7-8ba2-4c8a-910f-01f72ebd2662 · inbound

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models cites this paper.

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T09:03:18.768478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:03:18.768478Z digest=sha256:aab4ff131c9060b04feabe5138ab6b3d97fc0c30a53cf266f1ab9081b472770b

Observation 839e7f3d-c551-46b8-96f9-79f52ad59d89 · inbound

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning cites this paper.

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T10:26:45.961575Z digest=sha256:b57e6c7278b03b581ef5c52adb90803214702b0d71a8b0447e0625d6aebf5e57

Observation 23075704-eeea-47cc-b239-ed2782eeeeb0 · inbound

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models cites this paper.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.046565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T14:09:26.842696Z digest=sha256:7f366ff032b82d2fa97dd6aeb06b129be60fe07801716841f366ad89b043a9f0

Observation 91d121ed-b446-499c-8dbd-dc90fbd3dc0c · inbound

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics cites this paper.

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T23:11:55.011109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:11:55.011109Z digest=sha256:a4d547024ce1ec0a4a2b30477c682b14fd8c73cb7bddcac807b5c8b4cae4f1e7

Observation 4bfa553f-d1b8-4998-8295-0797b2c3c779 · inbound

TiCo: Time-Controllable Spoken Dialogue Model cites this paper.

TiCo: Time-Controllable Spoken Dialogue Model ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T00:38:52.182973Z digest=sha256:4fd60341a8388420c12e59087c4126e46052bd34249ca27bef9b653443dba20c

Observation b34124ea-dc9e-4db3-9eb5-0a692f4ec903 · inbound

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy cites this paper.

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T14:28:35.916911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:28:35.916911Z digest=sha256:f8377a7f1c9534691492766c2a4be05a7dec12629cdcbdbdc204dd679e3fb752

Observation 3bd4f56d-4dc1-463d-9567-a72a04affa89 · inbound

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models cites this paper.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:c6d1bde6734beba6abcb37fce67d1d8f876861e617096a0f442875400191627f

Observation 03a8463b-9a94-4055-8605-f1a3cbbc8500 · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:da018adc1559c4ce927ec5075e1112ea20c838d5b6f2682c46ee3f7da54f0ffb

Observation d5663aad-3b44-4559-b8d2-59f0324274cf · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:890c80f2c355189d7e1c52c86cbaa58dcc4ad1ef063e697cff43148d41ec77c2

Observation 67539e43-f4b6-4bd8-a62c-aea2828bf0e6 · inbound

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent cites this paper.

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-05T15:11:10.854096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-05T15:03:50.420072Z digest=sha256:ae4038aeaf5f388af4d09c84d9067706ff748fca3b66e64a27c13eaabaa2f733

Observation f0b0da62-cbde-4ef9-a153-fc8902bd412d · inbound

Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs cites this paper.

Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T01:10:03.459912Z digest=sha256:91cabeece627c8aa6a68eb8a0825e6223090337655f72e2beaf778c33f61e235

Observation 849fb43e-80fd-4b73-9eb0-8242fbb32d33 · inbound

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning cites this paper.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T01:57:30.877915Z digest=sha256:870aa8266c261ce3ae1faf1f8bed614e9741044c55516aa1f49cbbf0323e9f74

Observation 844b31fa-83ab-44a1-8b2e-f7d5ebd5d7fd · inbound

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning cites this paper.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.664024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:1541f6dcf5746823829f20098baeeb6cd4eb99325925e58ed01c1b1d679d76ce

Observation 79c5031c-f0ec-4fe3-9a85-60c213abefb0 · inbound

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization cites this paper.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:2db204bbe3e26cc9f9e9c77a0cb62cfb4bb13d2919b55b0b534b0bba64c28ea4

Observation 8046a357-0c77-4e27-bfa6-a6774ffc35ed · inbound

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning cites this paper.

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T01:19:49.761472Z digest=sha256:9afe2fe925d0be37e3042adb0421cac16062609a307a2640c5f6475979568119

Observation 407e4e2c-5068-4de7-ae40-e55547c2d641 · inbound

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy cites this paper.

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T05:50:13.653022Z digest=sha256:e9c11cfbca89d988f769aac87323516824ba47af69a0d396c66271a78b01a719

Observation 16cd8fdc-8158-4c29-9b74-d5e80f31b0cc · inbound

Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR cites this paper.

Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:18:54.379011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:18:52.845177Z digest=sha256:9bca950aacb86b2a1b7645af3a30d5965416601ab926675b67403484fc73d0ae

Observation c1d5640c-0a04-4079-85cc-f7cb910bdb2b · inbound

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning cites this paper.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:14:03.312265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T08:10:14.041279Z digest=sha256:9413faccffa3507f6aec7681e7fb97692420ce6b51f87e71430b32778fb8febe

Observation 05bf0cfb-0094-4d53-977e-58f0dab65e35 · inbound

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning cites this paper.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.267554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:dddfcbd54bed8c7cc0e006c6606fa894f913a25c9bde1e0e4ec12f5d807d0ceb

Observation 649240f5-9b95-4ceb-a52d-9f92302ee866 · inbound

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination cites this paper.

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:12:47.100067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T23:08:05.597810Z digest=sha256:de28eba08774ffa905eb8bab53d6f67b5a2ccca0df7ee2aa8498875c2a40f67f

Observation 818e8dfd-a375-4be8-b94c-906bad8d0842 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.820455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:a18623244341eb5708a8b9cc07794e50023f420c2b1769160566bdcdef27660a

Observation 74d0f968-bad8-41f0-ae77-f77721cb68cd · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 225

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.555191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:67354318ef7c69f6397e88dab835cab521abfe31f299c1c1c974064bc6fe4627

Observation 92d3cf5e-fff5-4c56-872e-5a3bdd08238e · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:13.376818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:47139dfd1a4997f4fb7a2003dad06d8cd1755ef4fa1c68246a5d90954053f152

Observation ee6a4c90-954b-4c34-bb64-172387919d13 · inbound

Robots Need More than VLA and World Models cites this paper.

Robots Need More than VLA and World Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 163

Resolution
verified exact
local_arxiv, observed 2026-06-28T01:11:28.844134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T01:01:33.530167Z digest=sha256:1fa974787b83fb735d848de636a7353dfd61f54bf72933bbc8aab4fd0960d22d

Observation 84a3e069-37aa-4a20-9957-5255bebcc66f · inbound

Teacher-Free Self-Training Amplifies but Does Not Compound: A Pass@$K$ Crossover on a Free-Verifier Domain cites this paper.

Teacher-Free Self-Training Amplifies but Does Not Compound: A Pass@$K$ Crossover on a Free-Verifier Domain ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:47:09.127737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T22:26:50.223393Z digest=sha256:e52a59df469591ce30066448935b6f43a6a419a61550741d34f2d7b42f9f3937

Observation 2b9d8924-e10e-4c4d-885e-763294d3745a · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:27:36.689469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:9cfc678ef04ea6c55319435aab556d5b58f8d2f6fefbfcef4331b50e47cf782b

Observation 940785f6-1537-4491-92c1-fa81b0913bec · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:20:20.555477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:2374b20472adb533431b0b1daaab015a9a6f2437052203736128263d92b0bc1f

Observation a8904ec3-a755-437d-a82f-547e68bfb96d · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:56.207246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:5a5671816a001c30badf2aa65d618cbd291ad81c534a634236306aeeded57092

Observation 5793ee26-189c-44b3-8447-eb8716704650 · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 127

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T08:57:48.118568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:d7290d413b554a6b90c50d0fd888bdb8b4f16fc33a8cc036029f7696b611469a

Observation f5c25b2d-c968-4333-b469-e58d2bcee014 · inbound

RL Post-Training Builds Compositional Reasoning Strategies cites this paper.

RL Post-Training Builds Compositional Reasoning Strategies ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.490553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:c7fa869254df986a06ac65b34c76bcb5d060833b920a2374d336c3c87674c394

Observation 99c812be-c174-4e18-a448-0094ca043f57 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:38.160357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:38.160357Z digest=sha256:f8c50760313cd236aa273b773c135421a6df2aabfce117e649f2f352c593c5e3

Observation 486c41ab-4829-414a-898a-1eb672dc75eb · inbound

ISO: An RLVR-Native Optimization Stack cites this paper.

ISO: An RLVR-Native Optimization Stack ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T12:52:00.138146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:52:00.138146Z digest=sha256:2f8f55dc1777fc49e99242deb03449a26adf0f32d0dc7fcaa6076e862f3ce249

Observation 03a342d7-a398-48f5-8ca1-385ab2de3150 · inbound

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR cites this paper.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.796151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.796151Z digest=sha256:ce0fcecd1e623702b7461b4f6ac7cc9eeb580eb8373ab6709c0886a63bc5db1c

Observation 7df0f2db-156f-4f71-9816-5ef0710f6e4a · inbound

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning cites this paper.

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-31T01:31:11.492967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:31:11.492967Z digest=sha256:eda56699b5af02730649bf7877901f8f9dd0ffcb647168327aa16c66e06e9390

Observation 4dcbe759-1785-4b01-93f2-7865b7afdc9f · inbound

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics cites this paper.

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T00:08:25.365379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:08:25.365379Z digest=sha256:a92029c6aa85ecbe373299a56e53e571fc6cc48193321f744b8e00e3a65941af

Observation 469e5597-9969-4364-bdd8-aefef327d7dd · inbound

Parameter Exploration for RLVR via Variational Learning cites this paper.

Parameter Exploration for RLVR via Variational Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.916702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.916702Z digest=sha256:2b144ed94054b000ef1532b8159436bede8871f058ce88fc9bcc167b011d546c