Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:51:43.099838Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 9 inbound Pith citation observations for arXiv:2507.12507.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:51:43.099838Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:04.024925Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T20:56:13.554863Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 01702ac7-f90e-41df-ad01-e3366eed0b87 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training OpenAI o1 System Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b91c1ff4-f3e1-4e84-836e-c7ef764722fb · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cd6446e-a56c-4d76-832e-a6d630ed29b0 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Tang, Manan Roongta, Colin Cai, Jef- frey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae443173-a672-4bb6-87ec-464a477d3bef · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Dapo: An open-source llm reinforcement learning system at scale, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72de7c8-7f04-4e43-b906-ee8f2bc20d39 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Process Reinforcement through Implicit Rewards
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3eded81-4d23-4ab7-abbf-2568a743311f · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7202fe5-07b3-4c21-8477-6aec33c4b1fc · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2d45004-e28f-4f90-b895-4a4bf75b0781 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c2e0656-9b38-4af2-8d61-d9e6fb88c713 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Instruction-following evaluation for large language models, 2023
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73910816-2529-43d4-ba57-7b29a62e8611 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a9ed1a-7369-4098-9a0b-75cbc62aa494 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Proximal policy optimization algorithms, 2017
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a62020-97b4-44bf-b019-693d73a66d65 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Approximating KL Divergence
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9f94756-acd8-4fa3-9f76-d3a664121d0b · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Deepcoder: A fully open-source 14b coder at o3-mini level
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5036a461-add3-44f6-971b-82c784a0fff8 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks, 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06296c41-5472-4ca1-a519-acb330813ae0 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Skywork open reasoner series.https://capricious-hydrogen-41c .notion.site/Skywork-Open-Reaonser- Series-1d0bc9ae823a80459b46c149e4f51680, 2025
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9fb672b9-d13d-4e17-97ff-e2bb93fea91c · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Hybridflow: A flexible and efficient rlhf framework
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f6e9583-fd27-417e-ae19-f98d530e3e09 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Decoupled weight decay regularization, 2019
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4371198a-3b07-42c5-a489-2af2baca32d0 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.https://hkust-nlp.notion.site/ simplerl-reason, 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ec5ac18-d477-4187-bdce-9003b7eb8312 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training American invitational mathematics examination - aime
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62f2ad05-1b60-44e6-b401-05a5ffa833e7 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training American invitational mathematics examination - aime
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea63ccc2-ada3-4e43-b86d-47132b004ac6 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training American mathematics competition - amc
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7028bb1-2669-48a3-b27d-c54cbecb0cb6 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Measuring mathematical problem solving with the math dataset, 2021
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde776d0-36a9-4c99-a458-dfee812f1963 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Solving quantitative reasoning problems with language models, 2022
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58383f93-4bd2-4994-9394-11a5d15a3604 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43299953-156f-4e44-91f1-fd7833a838ea · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Measuring coding challenge competence with apps, 2021
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31cdb266-0bbe-4b21-ba0e-157ff5892789 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12403a07-828b-4577-9316-ef7bc464d2d8 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Taco: Topics in algorithmic code generation dataset, 2023
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a14543-e3ae-4cb7-bccf-8b9d918c770c · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae44f13a-84c5-4e58-82bc-276933a6febe · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202e4379-5d88-4739-af38-9804f644e500 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0df78d3-4be0-4360-8fef-94fdf4fc1630 · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Online difficulty filtering for reasoning oriented reinforcement learning, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bba4a0ec-2391-4fa3-a991-b13230e25b7a · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Gonzalez, Hao Zhang, and Ion Stoica
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3573b1db-f3f2-4d53-8cdc-f1b92d7d675d · outbound
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training The curious case of neural text degeneration, 2020
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74b92fd2-80fe-457e-8dfc-f7d363cc0b3b · inbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da2f79da-b1fe-4b5c-bc8a-065e3102426e · inbound
Learning to Reason Efficiently with Discounted Reinforcement Learning Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3310297e-1a88-4046-94cf-424c1b424a51 · inbound
Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ddab7b49-4cd9-47a9-8686-07871448f7bd · inbound
Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19e525d1-700f-4d79-b634-ea8bad51050a · inbound
Generalization in LLM Problem Solving: The Case of the Shortest Path Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5e3e00a-54c4-4420-88fc-74d3b10c6268 · inbound
DISA: Offline Importance Sampling for Distribution-Matching LLM-RL Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 087dd9df-be03-46a5-a491-625313e27927 · inbound
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85b190e6-9899-4bb1-95db-2d694c32d13a · inbound
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 373a9187-b9f6-4da7-8a24-f1581a8da92d · inbound
Trust Region On-Policy Distillation Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Reference 226
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.