Pith. sign in

Paper Citation Record · LEDGER

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

As of 7 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 9 inbound Pith citation observations for arXiv:2507.12507.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12507 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:51:43.099838Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:04.024925Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.554863Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01702ac7-f90e-41df-ad01-e3366eed0b87 · outbound

This paper cites OpenAI o1 System Card.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training OpenAI o1 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.308878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.308878Z digest=sha256:38a3cce2ccec02cd98bb6a35928906495d553eca41adeabf0be22a196160d418

Observation b91c1ff4-f3e1-4e84-836e-c7ef764722fb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.379398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.379398Z digest=sha256:2ed0e4c5ac7d184f3c133f7f3c7fc64f59338025a034ac49b620a509ab7bdb3f

Observation 9cd6446e-a56c-4d76-832e-a6d630ed29b0 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jef- frey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Tang, Manan Roongta, Colin Cai, Jef- frey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.536753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:41.478402Z digest=sha256:e6bd7d615df97707243e31bed33028b74a44c54e61d83fa720c7a4ca2b4e7127

Observation ae443173-a672-4bb6-87ec-464a477d3bef · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.605352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.605352Z digest=sha256:f6a86e8c3d8b383d22bf351475f104efb42dc7908e6e3feeb4697fe5fdd1b127

Observation d72de7c8-7f04-4e43-b906-ee8f2bc20d39 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Process Reinforcement through Implicit Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.675432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.675432Z digest=sha256:8285105c9a66bb8e56645c03cb2ca0fe10e67f6f233bd5bc104eba6dc58161d8

Observation e3eded81-4d23-4ab7-abbf-2568a743311f · outbound

This paper cites Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.754532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.754532Z digest=sha256:8dd919cf7792d26a20d5243f779d21541120daf25c9978f7db4cc5204b82ffa0

Observation b7202fe5-07b3-4c21-8477-6aec33c4b1fc · outbound

This paper cites Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.440037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:41.839172Z digest=sha256:dd03b2c33f3430e1f3bf2291831ef5a7e0b00b63c5436d688f9d06f8524660f6

Observation e2d45004-e28f-4f90-b895-4a4bf75b0781 · outbound

This paper cites an unresolved cited work.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:51:44.361140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:41.927847Z digest=sha256:dda40905acb135241cee0b43984f0cc62f4eaeb39be12013d978e72bf552a5e4

Observation 6c2e0656-9b38-4af2-8d61-d9e6fb88c713 · outbound

This paper cites Instruction-following evaluation for large language models, 2023.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Instruction-following evaluation for large language models, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.986503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.986503Z digest=sha256:5003e33af237fd4d1550c78aafde87ea2a38ac27ad8d61c1ae39a4b69c31afdd

Observation 73910816-2529-43d4-ba57-7b29a62e8611 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.072510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.072510Z digest=sha256:f701d9204d4b40f92343fca46af1c3bf591e3fac2e4e3b33ab343418cd8d2681

Observation 38a9ed1a-7369-4098-9a0b-75cbc62aa494 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Proximal policy optimization algorithms, 2017

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.132236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.132236Z digest=sha256:366eea59586a40e67cb15297a1e54cf705847b9a1f6edf330c01709948450772

Observation 44a62020-97b4-44bf-b019-693d73a66d65 · outbound

This paper cites Approximating KL Divergence.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Approximating KL Divergence

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.270150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:42.192659Z digest=sha256:fbc005b31df7339b97f5e68b47a6c484754ee0931bd3a95844a57da3b67bc7fb

Observation e9f94756-acd8-4fa3-9f76-d3a664121d0b · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.190811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:42.236452Z digest=sha256:2b0e0f26db9b4298de1f9f297a4bab94d9b8b95c89ebba46bf7e6a3c4ed62a97

Observation 5036a461-add3-44f6-971b-82c784a0fff8 · outbound

This paper cites Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.120691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:42.281094Z digest=sha256:e8eb9ec4c5a8253557f508f672e37f8fed447cee924eda1562077b977e0113e7

Observation 06296c41-5472-4ca1-a519-acb330813ae0 · outbound

This paper cites Skywork open reasoner series.https://capricious-hydrogen-41c .notion.site/Skywork-Open-Reaonser- Series-1d0bc9ae823a80459b46c149e4f51680, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Skywork open reasoner series.https://capricious-hydrogen-41c .notion.site/Skywork-Open-Reaonser- Series-1d0bc9ae823a80459b46c149e4f51680, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.066802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:42.312913Z digest=sha256:fa91c87912120b8b94455a7955e2d107a631993edf77ec170fe47045a3cddce0

Observation 9fb672b9-d13d-4e17-97ff-e2bb93fea91c · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Hybridflow: A flexible and efficient rlhf framework

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.928142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:42.366724Z digest=sha256:493853940cecfb54d2ae229a04e52881a4f442632638b022df27fb4976eedb27

Observation 1f6e9583-fd27-417e-ae19-f98d530e3e09 · outbound

This paper cites Decoupled weight decay regularization, 2019.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Decoupled weight decay regularization, 2019

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.428615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.428615Z digest=sha256:3bddf78bac44d5ae5758e9d896414fd72a45213852b7b1019586c96919f9514c

Observation 4371198a-3b07-42c5-a489-2af2baca32d0 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.https://hkust-nlp.notion.site/ simplerl-reason, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.https://hkust-nlp.notion.site/ simplerl-reason, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.806690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:42.481817Z digest=sha256:efd019522c89d8f7402d5241100d53493ce21c611b1bd52d61677cd73273a69b

Observation 9ec5ac18-d477-4187-bdce-9003b7eb8312 · outbound

This paper cites American invitational mathematics examination - aime.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training American invitational mathematics examination - aime

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.711163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:42.511720Z digest=sha256:a3f0d68964428f2be7a257839c7cd6683c7ffc8536e91a679edda80b0892f8b2

Observation 62f2ad05-1b60-44e6-b401-05a5ffa833e7 · outbound

This paper cites American invitational mathematics examination - aime.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training American invitational mathematics examination - aime

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.608952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:42.546714Z digest=sha256:47cb9f5c3161bd7be493027c5bab00f87ab326867d624abb009b3241204fb23a

Observation ea63ccc2-ada3-4e43-b86d-47132b004ac6 · outbound

This paper cites American mathematics competition - amc.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training American mathematics competition - amc

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.543312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:42.576951Z digest=sha256:e001a339fecb583338b7cfc5aae0b0588307b8569fa85c3642ca40abef7f69aa

Observation c7028bb1-2669-48a3-b27d-c54cbecb0cb6 · outbound

This paper cites Measuring mathematical problem solving with the math dataset, 2021.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Measuring mathematical problem solving with the math dataset, 2021

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.599645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.599645Z digest=sha256:537f0162f65cac94d384c931363e29b6a1097282127b0c4f7a4d1e4c704c11dd

Observation bde776d0-36a9-4c99-a458-dfee812f1963 · outbound

This paper cites Solving quantitative reasoning problems with language models, 2022.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Solving quantitative reasoning problems with language models, 2022

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.633994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.633994Z digest=sha256:45fb380943532c4eed4089a517a13ebccf80103ff6ed38fde21ef682da9d0c36

Observation 58383f93-4bd2-4994-9394-11a5d15a3604 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.657244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.657244Z digest=sha256:89523cf05a8679cf04dc7bde2807d4a4b37af11ad28738ffe7436eba05900a22

Observation 43299953-156f-4e44-91f1-fd7833a838ea · outbound

This paper cites Measuring coding challenge competence with apps, 2021.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Measuring coding challenge competence with apps, 2021

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.719030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.719030Z digest=sha256:a4fbc6f8926ec4e04a712579e4d089f70cd115c145de7d06c6c66dcb23d2e2cb

Observation 31cdb266-0bbe-4b21-ba0e-157ff5892789 · outbound

This paper cites Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.778292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.778292Z digest=sha256:6cd6f97556b465d1f9c6aece78941839325fdf7630f51884ea316ffe115e7d7f

Observation 12403a07-828b-4577-9316-ef7bc464d2d8 · outbound

This paper cites Taco: Topics in algorithmic code generation dataset, 2023.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Taco: Topics in algorithmic code generation dataset, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.831655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.831655Z digest=sha256:c02909ac55a046fd7e11da0dce1838a584c700c2559f9be28d345bce70b1de8e

Observation 98a14543-e3ae-4cb7-bccf-8b9d918c770c · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.418470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:42.871217Z digest=sha256:bf8cfd79f85f4fbaf863760f4e0a6136bf38fb435f25c7fcee8a95a8c55edf88

Observation ae44f13a-84c5-4e58-82bc-276933a6febe · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.887053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.887053Z digest=sha256:6b8ea65d83abca530f7cb0c728e6b88dac61a91902f25fc95ecde61d7d1e236d

Observation 202e4379-5d88-4739-af38-9804f644e500 · outbound

This paper cites an unresolved cited work.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.908581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.908581Z digest=sha256:8a92cf3f82c70ac03d2577e413e3f99fe6477f434efccebd06bfc954fc10541e

Observation e0df78d3-4be0-4360-8fef-94fdf4fc1630 · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Online difficulty filtering for reasoning oriented reinforcement learning, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.984504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.984504Z digest=sha256:18abd141211ac43f44049b9f86b98d3a9465e6e445489e3c7e1ff5fb613f6bde

Observation bba4a0ec-2391-4fa3-a991-b13230e25b7a · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Gonzalez, Hao Zhang, and Ion Stoica

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.341473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:43.036697Z digest=sha256:1832820f1268d2db3c654b783dc5a92044f137af6b4ca2fac7074a17221d6cc0

Observation 3573b1db-f3f2-4d53-8cdc-f1b92d7d675d · outbound

This paper cites The curious case of neural text degeneration, 2020.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training The curious case of neural text degeneration, 2020

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.283087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:51:43.099838Z digest=sha256:ba389d2d2d1b1f36935b338219add023ce0a8547400c36f974c72d9923034356

Pith citing papers

Observation 74b92fd2-80fe-457e-8dfc-f7d363cc0b3b · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:04.024925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:04.024925Z digest=sha256:dd887c03fd80b3aedd37017cde7e88d1d1e694870031f871b3309cdd68f493ac

Observation da2f79da-b1fe-4b5c-bc8a-065e3102426e · inbound

Learning to Reason Efficiently with Discounted Reinforcement Learning cites this paper.

Learning to Reason Efficiently with Discounted Reinforcement Learning Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T07:59:53.046706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:59:53.046706Z digest=sha256:d1ffb00273318e1515c2c0c1646c4b83dfe18f74d8f4936c080a393a6c23d9c1

Observation 3310297e-1a88-4046-94cf-424c1b424a51 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:25:59.026410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:03:15.222571Z digest=sha256:61ec5657307034c0f2b8090bcbc4ee176eb1e7627784d6feaacb06147038e68a

Observation ddab7b49-4cd9-47a9-8686-07871448f7bd · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T22:48:45.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:48:45.647588Z digest=sha256:19dae64e3f3db116422e5c632ca8c86f6e4ccba2425c04ad44b4dc1bbb589fb4

Observation 19e525d1-700f-4d79-b634-ea8bad51050a · inbound

Generalization in LLM Problem Solving: The Case of the Shortest Path cites this paper.

Generalization in LLM Problem Solving: The Case of the Shortest Path Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:39:38.040670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:37:45.355872Z digest=sha256:7b23ecd960b0de76e4196528929a6031c74c0f12181f4e19527facb265c0a6b0

Observation e5e3e00a-54c4-4420-88fc-74d3b10c6268 · inbound

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL cites this paper.

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:24.864105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T15:11:43.235574Z digest=sha256:ba91a40a884f10b982af7ff173ff8198b46e5e41885ab64d40dcc0fc0a0c937c

Observation 087dd9df-be03-46a5-a491-625313e27927 · inbound

ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning cites this paper.

ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:35:21.594882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:31:37.641598Z digest=sha256:2c783298001e30b1ea0d098596e5f4cb91dc92679ab32d715b6ef89aeeff3666

Observation 85b190e6-9899-4bb1-95db-2d694c32d13a · inbound

ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning cites this paper.

ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:24:55.362639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T16:21:40.582248Z digest=sha256:b376860dbfb908fab4a9872108cd207ccaa62b60860a5238e1137de79641b1e4

Observation 373a9187-b9f6-4da7-8a24-f1581a8da92d · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 226

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.556324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:68ae35d79bbcd7cc66e4898fcf8b7fcedde96dc77c319d9f382cc16197daf886