Pith. sign in

Paper Citation Record · LEDGER

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 32 inbound Pith citation observations for arXiv:2506.20512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20512 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:52:05.518110Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:13.384509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:59:51.967723Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 929091d6-a92c-44a4-9033-f3f2ca02a490 · outbound

This paper cites Phi-4 Technical Report.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.294425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.294425Z digest=sha256:442658ce05c90bbe28f290c23de67535fb7a51c704211dd85bb3e0541ca7e065

Observation caadb374-c96d-43b2-b30f-7b6d072ec2ad · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.299864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.299864Z digest=sha256:fddde49ae5592d7e023de3c29647da30ad94cd201ee2b2e935da3ae5fa668fda

Observation 58ed5096-4ca2-4ee0-a4b6-47b71c506583 · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.304965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.304965Z digest=sha256:e8afef74539aaeb34e63f7f2865635d2e02ce76bd540a6b33378d78e3db81832

Observation 2888274e-5224-47db-bc55-5cd049667693 · outbound

This paper cites Mathqa: Towards interpretable math word problem solving with operation-based formalisms.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Mathqa: Towards interpretable math word problem solving with operation-based formalisms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:52:06.360449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T22:52:05.309417Z digest=sha256:2d8dbe45c9939273ebe73f2a9a8d5c100f8bf5892697c980008b1d1213c14002

Observation a51e9972-3454-49ef-bab2-f8a1ceb5ff57 · outbound

This paper cites Jiang, Jia Deng, Stella Biderman, and Sean Welleck.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Jiang, Jia Deng, Stella Biderman, and Sean Welleck

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.314312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.314312Z digest=sha256:70386327217c15ff33717989a337970194bcc48563f851366d934c0eac7c749f

Observation c9d7d25c-fadd-425f-8568-f3ad62498718 · outbound

This paper cites Puzzle: Distillation-Based NAS for Inference-Optimized LLMs.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Puzzle: Distillation-Based NAS for Inference-Optimized LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.318809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.318809Z digest=sha256:9f93e62ac6c51f8994551b301179ba1a9154ed6acdf514da02ad1ca5e109a7eb

Observation d657da24-bb83-4cbf-a411-07bce32b6ceb · outbound

This paper cites Llama-nemotron: Efficient reasoning models.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Llama-nemotron: Efficient reasoning models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.323803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.323803Z digest=sha256:1aa0aeda4f9c503c764173d70676f26e3dcffaef716bf28a1fab2fc1c89d3709

Observation c67a416f-8d7f-4307-8a13-004bc6d008b9 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.327980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.327980Z digest=sha256:4b8d3425d7bbccf1de3ed428b2b93502b5ecb6ec04e6b19044c343bced24f899

Observation ff991753-d2ba-4dc4-8cd8-4f15f47cc95d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.332462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.332462Z digest=sha256:fb1ce7bb2094496678ce96d745157ee8a618c85b9851ec9293aeba27094f308f

Observation e1a23e6c-7bb1-43ee-92c8-2c30c0b7df0d · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Enhancing chat language models by scaling high-quality instructional conversations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.336845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.336845Z digest=sha256:6b10d7fa6438fbefa6cfd90bf8c04e890a9ebfa205afd8fcd48da4c773875021

Observation 61ecd802-5387-4e14-b5a6-6073b91ed178 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.341465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.341465Z digest=sha256:ce77fdb90ac0575775964eaab9f58255bab6499384fd8dccf8937aa687869026

Observation 3421b259-5927-4b0a-a9a3-2ddc0dec2d6b · outbound

This paper cites Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.345605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.345605Z digest=sha256:eb237b3f2941cf47d5983768f581c3213391d7589de9b2b8b7e230bfe8f8d6f0

Observation 911403fa-58dd-47a8-b391-83cee0c03df9 · outbound

This paper cites The Llama 3 Herd of Models.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.349613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.349613Z digest=sha256:bcc148e48302e6b517b234e67d5563a4d86ca4d0f6e3c40a495e0d2ae93ab5eb

Observation 412ca939-8e09-423b-8593-cb07098072d2 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.354143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.354143Z digest=sha256:d3b00a9edfff3ee048b0d6896650e23b3b26490dcbf0f114538c650e99f30f21

Observation 9b624733-24ad-41fa-a483-0ad778281b9f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.358503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.358503Z digest=sha256:90e96bdd286ec69079fa465f3c11b4739f275e215c6baa9a95e562e176dbb8b5

Observation 40e64994-d9c5-4039-a752-06cfd274aa5a · outbound

This paper cites Infi MM -webmath-40b: Advancing multimodal pre-training for enhanced mathematical reasoning.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Infi MM -webmath-40b: Advancing multimodal pre-training for enhanced mathematical reasoning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:52:06.337771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T22:52:05.362497Z digest=sha256:ebfef73e51ecad6993c48a2f830ba8b51a89c0749cc9876546021255add46f41

Observation fa3d8727-6f3e-4250-ade8-c5f6b732c88f · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.366697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.366697Z digest=sha256:d1b8153bfa26f7336f42a8a1f50c4909223181d102e28938916b4767237ff567

Observation 71197d56-ceca-403a-ae74-519e4a708844 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Measuring mathematical problem solving with the math dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.370603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.370603Z digest=sha256:5906caa21ecf6d26c5a7b7215415293957731d9d0dca15216c2a74e3ca7a5c46

Observation 517401c8-bb18-401d-a19c-32522e9b8007 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.374846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.374846Z digest=sha256:a3a80f597d114f74daebe7ec1b485147353bef69423542262a7450ce55175023

Observation f383e86c-955c-4243-92b1-b1a511ae4f9c · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.378916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.378916Z digest=sha256:6bbcce98767ad0bacdc1284fa9cd07b91128da87edb7db89e69e0656abe593a7

Observation 0773918f-4cf9-4f32-b3c3-20a092f0ad9e · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.383515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.383515Z digest=sha256:d11c3038232aa5289222f457bda4420afde2fe953e0b928f175af89993be5e9f

Observation 1157c3ec-3dfb-4399-a34a-854b1e7e3c5f · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.387635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.387635Z digest=sha256:63a1d4bb953d7f5713d3789ae9498ce804a034980ce9354e9bbdbb02545679ee

Observation 292af8cf-d9ac-4c0d-81f2-200803b4fa1d · outbound

This paper cites Mawps: A math word problem repository.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Mawps: A math word problem repository

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:52:06.296015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T22:52:05.391813Z digest=sha256:1013dcba675b290717b1b2ecf84b0d2ef42c080c57c9d3dfb38afa5b2fec3c9f

Observation 9563caa2-1a9e-426e-8c31-6e698813945f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.395822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.395822Z digest=sha256:d13f3b260d3e3d68878e0ead8c17b2ce7f89bc96e1a83afa515a24a5f780f1ec

Observation 9b8feced-0230-4c83-9391-6ec317cb6918 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.400522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.400522Z digest=sha256:bb3db78dbb7fc9aa55a94f1e56b43b99b07a1c68feacd5f3c331d4d76232ee00

Observation daa28f16-e706-4fe5-a113-4efd81675bce · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman - Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur - Ari, and Vedant Misra.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman - Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur - Ari, and Vedant Misra

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:52:06.282462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T22:52:05.404810Z digest=sha256:30bc961fccb879068a16af0e80a2e0ed5e202bf65806c0e1c9583b39b09dffd0

Observation 89bb13e7-368a-4004-9562-02310017278c · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Datacomp-lm: In search of the next generation of training sets for language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.408671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.408671Z digest=sha256:865912e8cba2168daef475d4ce026473d20570f1d1f461d0fe0078fe21d25648

Observation 65cb26dd-5b41-4fec-8575-8a00428e08bf · outbound

This paper cites Let's Verify Step by Step.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Let's Verify Step by Step

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.412514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.412514Z digest=sha256:c442337292a3ae083a8a864067945f2fa6d3ccb836b600b054cb9460a3fb234d

Observation 1251579e-ea1f-48d4-92bd-5b7c03c374fc · outbound

This paper cites DeepSeek-V3 Technical Report.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling DeepSeek-V3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.416710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.416710Z digest=sha256:9c455afebd1c9ec3d2473d3863400d313f675bddd5b8e460d6b92b243ed44154

Observation da330cee-41dc-4db6-b959-64dfd496c82b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Understanding R1-Zero-Like Training: A Critical Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.420871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.420871Z digest=sha256:1b609b599765b7ad6296a20cba1053c3faf1a36bf12a963aadefd23252922c40

Observation 2fbc964f-05cd-437c-91da-87bcd0c8392a · outbound

This paper cites Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:52:06.259901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T22:52:05.425103Z digest=sha256:0bed64d27caff328065eef7528a95dd57a5812b97bad22bd08c719a332bf8df0

Observation dfc3b450-3ab7-4840-85aa-39c4a4ad266d · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:52:06.246156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T22:52:05.429064Z digest=sha256:68564ddef465ad21b093a1c4738253f77662ab99d7a75791e1eb21bffe40aee4

Observation 8928c801-905e-420a-91e4-ecaed0a0b48f · outbound

This paper cites The Llama 3 Herd of Models.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Llama 3 Herd of Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.433019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.433019Z digest=sha256:ceb15415e1d073ef71e6746c05888f6202a09e68e53b00ac33d8c0eff59aa236

Observation c5f940e9-4831-4e9f-a37e-64b29f917ca5 · outbound

This paper cites A diverse corpus for evaluating and developing english math word problem solvers.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling A diverse corpus for evaluating and developing english math word problem solvers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:52:06.232659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T22:52:05.437072Z digest=sha256:00f519edfc6e2477a5fe49cd02591e501250610d410061c1e17c60a1c66460da

Observation 71955f78-bbb2-49d3-973e-4d8085020009 · outbound

This paper cites 2 OLMo 2 Furious.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling 2 OLMo 2 Furious

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.441159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.441159Z digest=sha256:42b11898e1c051298c3670cd0a3a99294f26222931d23fd957db6bf53e5a44be

Observation 0aaa0c9a-4954-4ec5-9107-6ce4b8f7816f · outbound

This paper cites Introducing openai o3 and o4-mini | openai, April 2025.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Introducing openai o3 and o4-mini | openai, April 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:52:06.219472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T22:52:05.445297Z digest=sha256:53c3a948a3310810eb7647a216066a561083c80f4b8439acc68ef189763b7500

Observation cc9157ba-2e02-422b-b451-c9724d4b5243 · outbound

This paper cites OpenAI o1 System Card.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling OpenAI o1 System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.449157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.449157Z digest=sha256:cc129e1362dcb8aaf9449bcdf4800129889c132cb00f41dae311468d48d742ba

Observation ac5fe4e2-e053-4b46-88d0-fea47c453f0a · outbound

This paper cites Openwebmath: An open dataset of high-quality mathematical web text.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Openwebmath: An open dataset of high-quality mathematical web text

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.453673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.453673Z digest=sha256:e3c67eea91272ed0f437dd101718248a4f8ae2cdcc5cbe9c8846d0edbd253296

Observation e15ece67-2dbf-4414-b18a-b61522ce84a9 · outbound

This paper cites an unresolved cited work.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:52:06.197571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T22:52:05.457662Z digest=sha256:84121adcdf7499af6cbf871d5b4e672877898ebbc0a2a13e18c9438f894521d3

Observation 4c0b8fcf-9b99-4446-8d7e-3af410306ce9 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Spurious Rewards: Rethinking Training Signals in RLVR

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.462042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.462042Z digest=sha256:d10f42911272a89649a4f3948d060b0245a8ca969ee63dc6759467a2ff88fb9e

Observation a920ac3e-ef93-482e-a0cc-1a2b5cf27fc6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.466700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.466700Z digest=sha256:f7b40ba89f8fc3396df9588e6b22076e262ed51ec634dd555922d39ec5899756

Observation 9e4dd77a-3d4d-4073-bdba-4025517f6ad1 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling HybridFlow: A Flexible and Efficient RLHF Framework

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.471516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.471516Z digest=sha256:3c08e740fc0c4fcbdd15e942a8c3c5a851f97b921b8eeea54dbf8f26eed67222

Observation 2a5d9039-bf41-4e53-b2ad-6382f8516dc2 · outbound

This paper cites Yi-Lightning Technical Report.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Yi-Lightning Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.475987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.475987Z digest=sha256:5745cda53bc52ad1b2173fbf02ede580a47dea11d4248af881573d9fe2fc6f90

Observation db169a52-7fcd-4069-9202-58c8caba5ea3 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.480386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.480386Z digest=sha256:84ce84143e9b035074e28218b3775f8f39a9cc73be4374bdabe2ddbd7b7a8c53

Observation 51ea2f7c-d7d4-45f9-b240-7a45841c4876 · outbound

This paper cites Mathpile: A billion-token-scale pretraining corpus for math.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Mathpile: A billion-token-scale pretraining corpus for math

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:52:06.184013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T22:52:05.484580Z digest=sha256:a4c4b3e14ec72e122fb6583345caa98aad0f39c84a6bf71ccf1e92132a460681

Observation 13e66756-f81a-4faf-992b-6f9c79c1b74c · outbound

This paper cites Chi, Quoc V.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Chi, Quoc V

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.488692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.488692Z digest=sha256:fcb8cc677622ef3ff95573e12f73ce9ff046b7f5bc1f773832ffa0f4c344670d

Observation f4dd5087-a61e-42fc-8bda-303ebd4f342a · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.492777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.492777Z digest=sha256:56ef3cc6cae02b67a0f424d455c94d6998d0a857664172894977b5787212ad02

Observation 988b1b55-d734-4e77-b00a-5a9d17fdda05 · outbound

This paper cites Qwen3 Technical Report.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Qwen3 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.497044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.497044Z digest=sha256:a67ce219e97acca5ec8cc21fd5eb8ceb8f1649958deb375330c34de95cea05c3

Observation 53ecc5cd-73f7-43d3-a95f-f1d81640c141 · outbound

This paper cites Qwen2.5 Technical Report.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.501310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.501310Z digest=sha256:eb637ada50ce6c47c646d18a63a3636feeac57dfedff5e662a019cc07821623a

Observation c65920a3-dddc-4552-83ad-b0475a4caafa · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.505403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.505403Z digest=sha256:d63e7ce73911423d3636cd7b055219e9d709ca286c1920f0e1ef511d42077d3b

Observation 62f44857-132a-43e2-a473-329d03f21ba2 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.509858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.509858Z digest=sha256:dca4dd9c98167fce60f513978fc599a39b83ecef5ca04775a025c16b6a153e9c

Observation b2d41fb7-5123-41f5-a502-a652bd9d0f90 · outbound

This paper cites Wildchat: 1m chat GPT interaction logs in the wild.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Wildchat: 1m chat GPT interaction logs in the wild

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.514023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.514023Z digest=sha256:0c950ffb836e26ce40778e28ee21603e44ebb6455a0161c136c16116863a0623

Observation 5362a314-eefa-40b4-be62-a8ef38279548 · outbound

This paper cites MegaMath: Pushing the Limits of Open Math Corpora.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling MegaMath: Pushing the Limits of Open Math Corpora

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.518110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.518110Z digest=sha256:4382bfdf878b9c8816234b4cf43161320a2d1fe11516eb3087c844822e32cf26

Pith citing papers

Observation 18ce7cce-fbc7-45bb-96c2-2bb845d70217 · inbound

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling cites this paper.

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:40:46.412977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T23:39:39.018498Z digest=sha256:2a0db7905f5e615db49cbf592f6c5e9fd7b4c4ec60a48efc19e250705aaffeea

Observation d8941b72-95c4-4586-a154-9027845a5560 · inbound

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models cites this paper.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.423521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.423521Z digest=sha256:90e20d0c1994d1d5aea841687d7cd72b29d3e234d93f9c02ff64c667e50f30a7

Observation adbfa44c-b68c-4764-a063-580da13d60af · inbound

LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points cites this paper.

LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T05:48:28.554889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:48:28.554889Z digest=sha256:8d81e5949eb5920f33c1ff6263ba56023e8af7be7945ecccdfb080798596d1de

Observation b8317bc3-54de-4590-a75c-37df5c60a5d1 · inbound

Large-Scale Diverse Synthesis for Mid-Training cites this paper.

Large-Scale Diverse Synthesis for Mid-Training OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T05:43:18.003368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:43:18.003368Z digest=sha256:b9a2f63497ae70be679b2359256ed1688f82883ca59d611ab4b7b22202a6c37f

Observation ef6e6e4c-790c-4b24-8c58-4716de72ec9d · inbound

SSRL: Self-Search Reinforcement Learning cites this paper.

SSRL: Self-Search Reinforcement Learning OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.589663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.589663Z digest=sha256:10d154c5e5bd0c45dbd4e8fe16d5d402130ed434abc83ff54a034dd87206e751

Observation c659ab70-9778-4f58-b8fa-30abd22cbe4e · inbound

RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction cites this paper.

RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:59.788638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:32:59.788638Z digest=sha256:5b77dd98be3f4fa8057dea71fa49ed111e273537357632342ba39f8e50a9acff

Observation 13915e3a-a539-4303-9c73-15d01e11b28b · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:43.404800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:43.404800Z digest=sha256:6807182feee619539f32f23c5effdc348b9152a7ca9655a6e61b0d5e71ebbe52

Observation 67b20b37-35a5-48b2-8b64-50cd07d52306 · inbound

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation cites this paper.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.773353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.773353Z digest=sha256:6f7f96bc3721dcfd77a2c6ce3646ba4e1530d1cb511391c24c582efbf3c1f9c0

Observation 60c30425-b1ca-4821-bba3-baf13a75143b · inbound

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs cites this paper.

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:37.675267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:27:37.675267Z digest=sha256:063348dd73a41165a4655990556a14c4668790b0a7e48cd0ab27c38361ad4e1c

Observation 939ad9d8-6458-47d3-a519-4c336360443d · inbound

SAM 3D: 3Dfy Anything in Images cites this paper.

SAM 3D: 3Dfy Anything in Images OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:47:15.036120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T11:47:11.554134Z digest=sha256:945fb65ec48572d30b5ee19809992ceef152ebe2a71e7f8010f5aadcf99d9b39

Observation dc0cabda-655d-41ec-87ab-9ac6b21e9e92 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:23.124972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:23.124972Z digest=sha256:0096b15671fcafd238e6c461309390c74340e2c0ab370d1bba3cdb2bd0b01007

Observation 5d203758-1e76-44d3-9140-7655cad53b77 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.891598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.891598Z digest=sha256:fba775ff0ba38449fa016d8bfbb227bcfd1bf80053a1740ce8f1d6141b4d7379

Observation 5877dcad-ab33-452d-90da-745304b5d0a3 · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:21.405084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:21.405084Z digest=sha256:2b1801b13365068d72cf2c457ca51e2c2ac35d93348fa994d2aad12929c1ad61

Observation 2346ce8e-b95e-460a-a771-8ede395b587e · inbound

The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment cites this paper.

The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:15:55.899355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:36:44.401045Z digest=sha256:2a41a3b276e77a11f9be93aa546b03f4423bf458044eaba4718a1ea5683ae794

Observation d4d9d98e-bd5b-4311-80bc-c0906307defc · inbound

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space cites this paper.

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:41:04.004819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T12:50:57.603403Z digest=sha256:4a0bb1ee3d701b2af605e704e31db8d0278802fbcb228b9c2193e902e48e69f7

Observation 9d4b2c58-29f4-4393-b7cd-4a463c65bba8 · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.539437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:789ce5a7ca6ccd721283d5a95d67c3375e55c069fefd19efa0780cb10eadedb5

Observation bdbd9c5c-a6a3-438e-91ca-75d45231740d · inbound

Query-Conditioned Test-Time Self-Training for Large Language Models cites this paper.

Query-Conditioned Test-Time Self-Training for Large Language Models OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:07:54.089983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:04:51.248797Z digest=sha256:b718967364c2e1ac4861c2c2342124b0b1f9a50a70b31dcc1ff643b5727df062

Observation d703ccae-ce48-4c5e-8210-0a8781f004f1 · inbound

Query-Conditioned Test-Time Self-Training for Large Language Models cites this paper.

Query-Conditioned Test-Time Self-Training for Large Language Models OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:55:04.846371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T05:53:42.235498Z digest=sha256:b44d4ee1d56933fe3c39d6b0c39c6bc77047f09ccb3b9b6ecc189909a736a7e0

Observation 0fc009b1-53fb-4da9-ad04-80a7ab58f589 · inbound

Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains cites this paper.

Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:23:12.087277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T10:21:00.882064Z digest=sha256:fe3738014dc2afbaf8a8698408c1a567826ca3cc28c53e29f33c3b08b38314dc

Observation 6b5a8e6b-a893-4e3c-8a92-8a48f313ac67 · inbound

Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP cites this paper.

Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:33:03.987076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T05:31:11.191388Z digest=sha256:f84162c6c251864bee56024ba3d5f61bacfeb1457eb51c77b48e69358218f477

Observation 3b478f30-6f23-4581-a2b3-d0ad9651221d · inbound

The Future of Facts: Tracing the Factual Generation-Verification Gap cites this paper.

The Future of Facts: Tracing the Factual Generation-Verification Gap OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.419029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T18:31:04.169632Z digest=sha256:2ef6e7be5b103c47dab563106b876d7525c7fc006ff4c40d5f7903c008f68a7f

Observation e758cdd0-5a2c-4c81-b43c-3917ec7ce817 · inbound

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification cites this paper.

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:27.236550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T10:53:00.223228Z digest=sha256:fe8b87591034220261fefe4ac5f2d5e9274a31fa2ce0e5b0975f75bbd0d13db1

Observation a3e1adbb-b97d-4e67-86f4-3e734304c20b · inbound

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards cites this paper.

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:36:47.603371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T05:59:00.005336Z digest=sha256:29819d0a36fc1cbce05d6d7ae615617afa2de8d0267ad62cb5703452ac4e8be5

Observation 669a75df-30e2-4d7a-831d-0b8a5667faa5 · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.418404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:6a04ee82c26c90e1cdf07308a09ff4259a0fd9d31056af28fd72c880b0d164fc

Observation 32374512-3c5e-4da4-bf6d-26d1d08dc426 · inbound

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation cites this paper.

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.982648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T02:17:30.974692Z digest=sha256:cb11cd0836fd2e9bf9df1ff1eb905027abf01619287c99d27d31f76af390ce15

Observation 576f1bcd-e701-4314-a36a-0e07764004bf · inbound

Sumi: Open Uniform Diffusion Language Model from Scratch cites this paper.

Sumi: Open Uniform Diffusion Language Model from Scratch OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:09:19.753058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T20:34:43.697141Z digest=sha256:f5f8e6d5975edb6d929df6241b355d061db5b9c88346c92747002b66f9c158d7

Observation 9dbf31ad-e306-4508-9fc4-435d05ab4ea3 · inbound

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs cites this paper.

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:59:51.969197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T04:47:47.691913Z digest=sha256:9766e6a08e4c4a1a385fa23575a548d77653e1b9bb2cb601a12f2c7ccd07175f

Observation 55efe14e-984f-4d5d-a268-308f1c1c4626 · inbound

BashCoder-R1: Towards Robust and Explainable Bash Code Generation with Robustness-Aware Group Relative Policy Optimization cites this paper.

BashCoder-R1: Towards Robust and Explainable Bash Code Generation with Robustness-Aware Group Relative Policy Optimization OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:05:50.853682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T04:16:04.477464Z digest=sha256:587c20cfa631e579b0a980c84ef5876b328fc02b7239d2e47ca31c840e890dc5

Observation ba993dc2-d7c1-4805-bf0a-76acd9fd48c9 · inbound

Addressing Over-Refusal in LLMs with Competing Rewards cites this paper.

Addressing Over-Refusal in LLMs with Competing Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:05:29.233867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T06:59:12.695984Z digest=sha256:26c293ab3efcc9764a3f10341f2c2ecd5a3867ffc455207029d24928db711ed5

Observation d438acb0-8fb1-4806-bda1-c06e768a7a4e · inbound

When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors cites this paper.

When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:35:41.342530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T05:25:56.914399Z digest=sha256:eb077f40ac8e3e1e40f6bb84df0cd9bb116d2e1f1f915d0bf1b00c4b13cc29df

Observation a4ace8ce-f77b-4adc-84fb-f20c7b127fc1 · inbound

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data cites this paper.

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T07:01:43.604779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T07:01:43.604779Z digest=sha256:343fae079fb9359d34af0e7493c1bae83efd05ea7f199baf5c64ac1dbf4419aa

Observation da7b42ed-f429-4851-b35a-c8196afec172 · inbound

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure cites this paper.

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:13.384509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:13.384509Z digest=sha256:b821b9296adbd08a4fe48b670cfdd0afe073a513a7de2d34c32c3aaa4dd2c060