Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T20:52:33.620041Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 69 inbound Pith citation observations for arXiv:2505.24864.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T20:52:33.620041Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:46:57.916405Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
65 of 65 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation db55c592-fee7-489c-917f-8a715bd1b81f · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models OpenAI o1 System Card
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0064d6c0-ef0b-4571-b920-68cd5c3407bf · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fe294fc9-23aa-471b-be50-bca517e8ecf6 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 60023cbc-ca21-4ba3-8268-b2039fe12f5b · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Dapo: An open-source llm reinforcement learning system at scale
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 78dc3227-3d7b-4b4b-8625-fa8c2ce84328 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 26154e3d-2f9e-49bf-8b85-972a14d5fbd6 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 32789a36-fa2d-4743-bda5-459f1daaa73e · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Deepcoder: A fully open-source 14b coder at o3-mini level
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8da7fc17-dcab-46f2-9cdc-dd5cbbd752b0 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Code-r1: Reproducing r1 for code with reliable rewards
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 26f2d617-b911-4953-83c6-a852e7f74399 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Concrete Problems in AI Safety
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 778cbe5b-0bdc-498f-8034-e4b8237e28b7 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Reward hacking in reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e03fd634-d6af-4fc2-bc8c-b5d0a3ede3ab · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Language Models Learn to Mislead Humans via RLHF
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6d433f32-0dac-46e1-92cc-be5f136bd599 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Ai as humanity’s salieri: Quantifying linguistic creativity of language models via systematic attribution of machine text against web text
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5838bba3-2103-4dd2-84e2-86795e131750 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 048a6162-14fa-4d30-b7fb-99493383e734 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Zico Kolter, and Aditi Raghunathan
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0a4f8708-77ed-4e95-a57e-adc1e70dad1f · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Echo chamber: Rl post-training amplifies behaviors learned in pretraining
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c014312b-3f37-474a-a862-3c91589e9c6a · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 32b866aa-1325-4416-9ff9-42a67780c45d · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Proximal policy optimization algorithms
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c29eb7d5-5b95-4c6b-b1a7-ebf6863e36db · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Skywork open reasoner series
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a6e6dd55-2273-4630-9067-7ca35653cae1 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Hybridflow: A flexible and efficient rlhf framework
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c1356bb4-87b9-48a0-8078-329a91fad184 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Decoupled weight decay regularization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a664540d-7487-4b8d-8502-cf019f896747 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 032fb386-a6fa-4320-b3e0-463613682f23 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models American invitational mathematics examination - aime
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5769f3b9-e674-44aa-b035-0d3a24e4ed02 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models American invitational mathematics examination - aime
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4c78e818-9456-4aad-89d9-c237e427d7da · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models American mathematics competition - amc
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ad0d9e4a-d37f-4c2d-8129-926b53fcc9fd · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Measuring mathematical problem solving with the math dataset
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 346425fb-271c-451b-b981-072f5b73c05a · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Solving quantitative reasoning problems with language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6f05626e-695f-4967-a5d4-1377e44d62e3 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5c417ec9-007a-476d-bba7-1922037368d0 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Process Reinforcement through Implicit Rewards
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 91521688-f5ce-4ac0-ae3b-4e9de7211fc7 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Measuring coding challenge competence with apps
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 26835d4f-2536-4123-aad7-88ce73a16de3 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a9b6e0cf-c930-4050-ab40-b9e3d51a231f · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Taco: Topics in algorithmic code generation dataset
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2ffcc323-b873-41c4-9b99-42734d9af7e8 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 164d5c1f-a245-4616-ac35-257d945aac9b · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 530d317b-ed23-4f40-bb8d-e586549599df · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 10cc91c2-8100-410c-9168-3e433426bccf · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Online difficulty filtering for reasoning oriented reinforcement learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2c2a71e3-91f5-4af9-a98a-ac712c3c4df7 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Instruction-following evaluation for large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c77aede0-9205-417f-a444-6bc3afe13e97 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8d17720d-5fb7-42db-9826-95e81a245a76 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models The curious case of neural text degeneration
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0f67f247-b569-499c-a439-156277766e15 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Stop overthinking: A survey on efficient reasoning for large language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 203275a8-477a-4007-8762-adad24320538 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 76225005-8ccb-4468-b2d9-d0cf9f0ad41e · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind 12 Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation aee5fc0f-31e1-4b00-b5e9-49dd4ef86a5d · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Learning to reason with llms, September 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 45d5ab71-34b9-48c7-bf52-b7cd86a5e60d · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f63224ff-cf95-435d-b7b6-48b88c80fd25 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 79d7094c-b39f-4277-a9de-180580feadcf · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Mirror descent policy optimization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d466648d-2945-453f-b53e-4d6781f8629a · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 759eee3b-d102-43a2-b0bb-61c78a1f5e64 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Scaling llm test-time compute optimally can be more effective than scaling model parameters
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2b8cff30-f1a7-4cab-a1c9-4b1857646c86 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Reinforcement learning enhanced llms: A survey
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 19d77b62-eb97-41df-913c-99317297aae1 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Playing atari with deep reinforcement learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1477d07b-b4a8-4f43-8183-ee41aac0f9c4 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Human-level control through deep reinforcement learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8bdf552e-43e1-41d6-b700-0c12848922cb · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Mastering the game of go without human knowledge
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9b6e9b70-3ad0-4aa2-8d87-7e3fd4ddb4c5 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 50ec815a-887f-41b1-b2a4-faf5a47ed89e · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Llms are greedy agents: Effects of rl fine-tuning on decision-making abilities
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e4515ad7-1ba7-4bd8-acd5-42d3be0fa1b1 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Star: Bootstrapping reasoning with reasoning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 261f2417-ad85-4c6d-be4f-b9651063f3b6 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ddfc39ac-8803-4be4-9a5a-f6fb9d0158f9 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0e56cc4d-93a4-4104-9845-e4ea560b15bc · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3d9b0d6c-2395-4ac2-99f5-7f1d45d95b39 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models 0": 1, "1
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 34f53166-71e9-43bf-85f6-fdcae4bbe61e · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models However, the final answer shoule be a list of action plans for multiple steps
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 30701deb-379d-4ed2-a5dd-683c7f9bb4bb · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Agent[x, y]
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 56225a52-a014-4aff-ad30-f681612009aa · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 453d596d-c12c-4d87-9695-564725174cea · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e5bd4e70-742b-4608-a5e1-f5b034d18afa · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 739902a1-7692-4b82-a9a8-7948808a1042 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c4baeffc-828d-43e0-b236-0b5f2df58d98 · outbound
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Agent[0.5, 0.5]
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2992c4e0-dd2d-42da-8445-a7a0f67d9f0e · inbound
Flow-GRPO: Training Flow Matching Models via Online RL ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6535f3c3-f3f3-410c-addb-de6dcf3795fe · inbound
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfe38482-8f06-465a-97f4-0d92e1ae9e09 · inbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9945ac94-99ee-4cd5-bc09-518bca2d4311 · inbound
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb35d26b-a10e-49c6-a693-7e217cb88bdb · inbound
OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f20fba18-42b1-4061-b4f8-a82e513065ec · inbound
DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eae0d52c-25d5-417a-9a7f-e699cecff859 · inbound
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fd7fd3a8-3250-479c-9eee-d7e955d58426 · inbound
CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da5b0f7a-d27a-42ab-b494-5c6f70157773 · inbound
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a4edfb0-8eae-4bc5-b7ee-22f942df262d · inbound
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1cd65a90-dbba-4744-bbdb-cdc1931fba4c · inbound
Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dd56042-bd00-4b39-9500-092af60b8646 · inbound
Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4b9eac-311d-45a1-96ca-8af558b20872 · inbound
Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b894881-4190-409d-9d68-3950e4797a83 · inbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de686465-cf53-4b9e-89c3-2a3fdf6af3de · inbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71057d06-6966-4285-b979-283f21bb2a11 · inbound
CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55cc0bf5-2d45-4161-b7db-c60155982b52 · inbound
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation aa9b1cef-8270-4200-9f95-350fa9bf2956 · inbound
rStar2-Agent: Agentic Reasoning Technical Report ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 185b89bf-3df2-45a3-ada4-b3de776752b5 · inbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7031dc5b-c461-444e-a4c3-f9889dbcefcb · inbound
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a2fb91-77c4-41df-a9a1-7264959a1673 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b6ae3d25-f8a1-4a16-9eda-4c4e34c279a1 · inbound
Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c41cfa-13ad-4851-be0c-ddde473625ce · inbound
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ed1f0e17-23bb-40a1-9489-9665a333a24a · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 753d33ce-83a0-434f-8b0c-52661879e75f · inbound
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9619e9b8-149d-40fc-b8af-29a8728bebe8 · inbound
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 20c5a131-08e4-4437-a6d9-a166faa0df27 · inbound
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8768ed37-c50e-4366-b06c-99eaf0815fb1 · inbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 150fcd04-9d9e-427d-9dba-c88f8f72ac38 · inbound
The Art of Scaling Reinforcement Learning Compute for LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bac8f873-8339-4a8f-9c0d-b90c6164c2da · inbound
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8b28bf1-c0ac-4f36-805a-8d86fc864981 · inbound
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd087267-0ab3-4580-ab85-1469ec6bfc2b · inbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 19cb6ca4-e3c7-42a9-886a-29c14fdec193 · inbound
Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f81db7c7-8ba2-4c8a-910f-01f72ebd2662 · inbound
The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 839e7f3d-c551-46b8-96f9-79f52ad59d89 · inbound
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 23075704-eeea-47cc-b239-ed2782eeeeb0 · inbound
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 91d121ed-b446-499c-8dbd-dc90fbd3dc0c · inbound
On the Emergence of Implicit Curriculum in RLVR Learning Dynamics ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bfa553f-d1b8-4998-8295-0797b2c3c779 · inbound
TiCo: Time-Controllable Spoken Dialogue Model ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b34124ea-dc9e-4db3-9eb5-0a692f4ec903 · inbound
Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bd4f56d-4dc1-463d-9567-a72a04affa89 · inbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 03a8463b-9a94-4055-8605-f1a3cbbc8500 · inbound
Characterizing Model-Native Skills ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d5663aad-3b44-4559-b8d2-59f0324274cf · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 67539e43-f4b6-4bd8-a62c-aea2828bf0e6 · inbound
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f0b0da62-cbde-4ef9-a153-fc8902bd412d · inbound
Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 849fb43e-80fd-4b73-9eb0-8242fbb32d33 · inbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 844b31fa-83ab-44a1-8b2e-f7d5ebd5d7fd · inbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 79c5031c-f0ec-4fe3-9a85-60c213abefb0 · inbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8046a357-0c77-4e27-bfa6-a6774ffc35ed · inbound
Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 407e4e2c-5068-4de7-ae40-e55547c2d641 · inbound
When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 16cd8fdc-8158-4c29-9b74-d5e80f31b0cc · inbound
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c1d5640c-0a04-4079-85cc-f7cb910bdb2b · inbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 05bf0cfb-0094-4d53-977e-58f0dab65e35 · inbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 649240f5-9b95-4ceb-a52d-9f92302ee866 · inbound
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 818e8dfd-a375-4be8-b94c-906bad8d0842 · inbound
On the Generalization Gap in Self-Evolving Language Model Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 74d0f968-bad8-41f0-ae77-f77721cb68cd · inbound
Trust Region On-Policy Distillation ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 225
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 92d3cf5e-fff5-4c56-872e-5a3bdd08238e · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ee6a4c90-954b-4c34-bb64-172387919d13 · inbound
Robots Need More than VLA and World Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 84a3e069-37aa-4a20-9957-5255bebcc66f · inbound
Teacher-Free Self-Training Amplifies but Does Not Compound: A Pass@$K$ Crossover on a Free-Verifier Domain ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2b9d8924-e10e-4c4d-885e-763294d3745a · inbound
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 940785f6-1537-4491-92c1-fa81b0913bec · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a8904ec3-a755-437d-a82f-547e68bfb96d · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5793ee26-189c-44b3-8447-eb8716704650 · inbound
Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f5c25b2d-c968-4333-b469-e58d2bcee014 · inbound
RL Post-Training Builds Compositional Reasoning Strategies ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 99c812be-c174-4e18-a448-0094ca043f57 · inbound
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486c41ab-4829-414a-898a-1eb672dc75eb · inbound
ISO: An RLVR-Native Optimization Stack ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03a342d7-a398-48f5-8ca1-385ab2de3150 · inbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df0f2db-156f-4f71-9816-5ef0710f6e4a · inbound
Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dcbe759-1785-4b01-93f2-7865b7afdc9f · inbound
Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 469e5597-9969-4364-bdd8-aefef327d7dd · inbound
Parameter Exploration for RLVR via Variational Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.