Pith. sign in

Paper Citation Record · LEDGER

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 7 inbound Pith citation observations for arXiv:2505.22653.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22653 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:08:26.418982Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:21:41.900632Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:06:40.819064Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 336360d1-d394-4743-b0ae-7be6a657ddb8 · outbound

This paper cites Rethinking reflection in pre-training, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Rethinking reflection in pre-training, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:20.774171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:20.774171Z digest=sha256:2400ba758fb8b46f9c89ca38a05c412d504bd94920274d4a9c26d59547bf2ce4

Observation cfee5b61-bfa2-41b1-8ec9-ebeaef0ecf13 · outbound

This paper cites Math- arena: Evaluating llms on uncontaminated math competitions, February 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Math- arena: Evaluating llms on uncontaminated math competitions, February 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:30.141674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:20.882511Z digest=sha256:d3a934393a594885a4208b1043ff666c2fcd6f63415b54620e8e7b7e8261640b

Observation 998bfda0-0c08-46be-aa8c-b9ec8d061431 · outbound

This paper cites Do not think that much for 2+3=? on the overthinking of o1-like llms, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Do not think that much for 2+3=? on the overthinking of o1-like llms, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:20.971903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:20.971903Z digest=sha256:47c421593ae645ea5ba46647283c7d29ceda5071243c99c833748ff0206ec2f9

Observation dd4d2aa7-3ebd-4178-84e4-761d8bc8e0d7 · outbound

This paper cites The accuracy paradox in RLHF: When better reward models don‘t yield better language models.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason The accuracy paradox in RLHF: When better reward models don‘t yield better language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:29.793026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:21.046797Z digest=sha256:9ee90e4a08a754e57d5baf91e87600dad21cd9f3d9d954ef7bb074ad80a50b6d

Observation 5cc1211e-9c9b-4cbe-a5b8-4b980717d800 · outbound

This paper cites Fortify the shortest stave in attention: Enhancing context awareness of large language models for effective tool use.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Fortify the shortest stave in attention: Enhancing context awareness of large language models for effective tool use

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:29.579980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:21.149444Z digest=sha256:a195c223b87ab53947bdcfccd2c7b15e6c9df8d188733522992ca658a49b6939

Observation fa6a6ec4-aa68-44e1-ae75-ea2c1ad58338 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:21.245093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:21.245093Z digest=sha256:5935688e875de6f95942e599db29d0ef53f1a0f3d6faca89ae28444b354ed947

Observation 2fbaa476-27b7-4e7f-a4ab-9f973e655974 · outbound

This paper cites Q*: Improving multi-step reasoning for LLMs with deliberative planning, 2024.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Q*: Improving multi-step reasoning for LLMs with deliberative planning, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:29.438633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:21.723778Z digest=sha256:ce8c7f95f04c74fc67dd273b027b6dbfefb64f3622420f62257aba8563bd2f81

Observation 3881cd0d-3e5b-43ba-989f-bbf78fe41ef5 · outbound

This paper cites Fleiss et al.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Fleiss et al

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:29.324041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:22.439678Z digest=sha256:60df93fceb86bf683ab264f44aed91947e18ec64684035067f337f02f74d11cf

Observation 2bca5e9a-1439-47f8-aaf6-7f87adf5a74b · outbound

This paper cites Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph E.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph E

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:29.144749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:22.712573Z digest=sha256:363d745242e2323eed8ac6768db1eae96b14d1ec9a38f993e0a9c7fc20194c60

Observation 64d7804c-fb4f-41bd-872c-ca23e7ddd1b6 · outbound

This paper cites an unresolved cited work.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:22.798068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:22.798068Z digest=sha256:09189de1e78e9ffce589e911438e0dccec9ffb26a338c311a74c3e6dc48fdc3a

Observation bd96b21b-2804-4f67-adb6-7cb18f974884 · outbound

This paper cites Training large language models to reason in a continuous latent space, 2024.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Training large language models to reason in a continuous latent space, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:28.939096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:22.889241Z digest=sha256:2d22edfaaab12ef3f92b5928715859eccc52f51e69d951c5f1552599ed2a3ee0

Observation 635020b9-ad11-4fdd-af1a-d3fb2e407619 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Measuring mathematical problem solving with the MATH dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:22.944756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:22.944756Z digest=sha256:005b3ee79c6efb8db182247c4cae42dbccdaba2da92832aa37e697de65808874

Observation ee38baa9-bbe4-4dfd-88bb-1f454b229ffb · outbound

This paper cites Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:23.015073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:23.015073Z digest=sha256:b9621aa40e55ec021e218c3cf229b36a6463f491ccbd5b555daf377fea473eef

Observation 378ba01a-8ced-4322-ad0c-92e08e401e32 · outbound

This paper cites Human-centric dialog training via offline reinforcement learning.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Human-centric dialog training via offline reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:28.733015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:23.127945Z digest=sha256:f29029b6842a063f9b4e6fc3981fd24c87cdd7af0d8e607c2c42caf94f74fcec

Observation cca45cb9-7636-447b-b605-72d834896154 · outbound

This paper cites Smith, and Hannaneh Hajishirzi.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Smith, and Hannaneh Hajishirzi

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:28.526514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:23.187072Z digest=sha256:414030e38a157fcda8813cd8d600845c482728868bf78fbc5c046ec435decba1

Observation a2855bad-a891-4e71-8de2-9b2da73099ff · outbound

This paper cites Skywork-reward: Bag of tricks for reward modeling in llms, 2024.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Skywork-reward: Bag of tricks for reward modeling in llms, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:28.347936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:23.332664Z digest=sha256:abef0db262f4e3293ec6938e6885ff9a00d506e8e82c80871b263b528c2effe5

Observation 9c252d4a-6eba-44fc-8507-e4a48c7916b7 · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:23.476989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:23.476989Z digest=sha256:769cc72ed6b7ab4fbf88a147614016b859d2bded13bad3097ecb4316bc01c0e7

Observation 14cce726-7ded-4d3d-8c40-62737ca8a218 · outbound

This paper cites RM-bench: Benchmarking reward models of language models with subtlety and style.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason RM-bench: Benchmarking reward models of language models with subtlety and style

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:28.032968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:23.564975Z digest=sha256:3d02d63ed62e250b69da9f0a8ebd2c427a7219c8119750f89c5b75f4ace51b47

Observation 8f5711fa-ead4-4460-b00e-d306e59b5d64 · outbound

This paper cites The llama 3 herd of models, 2024.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason The llama 3 herd of models, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:23.664904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:23.664904Z digest=sha256:2174c29b978e8e92c2970bd184e379ab062fd5b994318bb07de4fc8562e101b1

Observation 08f18d02-a882-4f7a-be83-c2d6097de2ab · outbound

This paper cites s1: Simple test-time scaling, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason s1: Simple test-time scaling, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:23.784375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:23.784375Z digest=sha256:2af07bff68ed55307542705c1163065502791b4c17bcf6e593e03d509e1e4095

Observation 69dee594-ee21-4766-be72-8080c5d2b93d · outbound

This paper cites Webgpt: Browser-assisted question-answering with human feedback, 2022.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Webgpt: Browser-assisted question-answering with human feedback, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:24.224696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:24.224696Z digest=sha256:4aa4c50475c2b20d074e5ecd5ffe00e97e02546858f90f6e9a089bf97cc90be9

Observation 10d11159-eb51-492b-9e73-18ce08a42abe · outbound

This paper cites an unresolved cited work.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:24.363161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:24.363161Z digest=sha256:ef1f757a916f81bd18f8c328ce042dc140ee6249180faff15244921a263e0d94

Observation 59acf32a-c8fb-4bb1-ae9e-927664cdeaf9 · outbound

This paper cites Tinyzero.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Tinyzero

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:24.515080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:24.515080Z digest=sha256:81ddfdb3155f02a6f94e170a83fee056cffea17357b4d5e97c83dc53f5fe3832

Observation 7d818213-3f27-4f53-bf0e-9dcb03f1d974 · outbound

This paper cites Lee, and Sanjeev Arora.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Lee, and Sanjeev Arora

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:27.645811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:24.675041Z digest=sha256:e51fe4709c173dfc547a5fbe0ba3419f23f052348836bbd4659d91cd9ba3706c

Observation f037cab5-2d76-46dc-a77e-88e8f44d0696 · outbound

This paper cites an unresolved cited work.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:24.812651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:24.812651Z digest=sha256:9328d4fa7b85dc690797ff259b57eb69a786dfca1e9c3fcb450dc480d4f381c5

Observation 9181ac7a-da72-464e-939f-190848bf5bc9 · outbound

This paper cites High- dimensional continuous control using generalized advantage estimation, 2018.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason High- dimensional continuous control using generalized advantage estimation, 2018

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:24.984971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:24.984971Z digest=sha256:2a0144ff2aed4cb2bad735778d68efc9ded1cc68f5a6c63545ba7dda12ce30e6

Observation 11c5364e-0c72-46a3-aeb0-994d299d28bb · outbound

This paper cites Proximal policy optimization algorithms, 2017.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Proximal policy optimization algorithms, 2017

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:25.215091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:25.215091Z digest=sha256:d07c70ae146869f33bf4fbf34ba09293e486f5379119fed13286dd2f8a1731a4

Observation f7c7bc53-22d7-49cc-9177-88a313d57cbe · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason HybridFlow: A Flexible and Efficient RLHF Framework

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:25.364234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:25.364234Z digest=sha256:33cf7dc661a2b590f25a69b0ebfadebef40964a3ab6216c3562946449de4a49b

Observation 4272dc37-bbc5-4ce1-aef3-aed663720f48 · outbound

This paper cites Kimi k1.5: Scaling reinforcement learning with llms, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Kimi k1.5: Scaling reinforcement learning with llms, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:25.504871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:25.504871Z digest=sha256:22aa1b95c917985c43b45599adc02e893ec8ff4467950523ba3a94356d5454bb

Observation 26abeed2-a18d-420c-9afd-e7d1a79adca9 · outbound

This paper cites Dedicated feedback and edit models empower inference- time scaling for open-ended general-domain tasks, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Dedicated feedback and edit models empower inference- time scaling for open-ended general-domain tasks, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:27.328579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:25.654842Z digest=sha256:cd6362374e5c64f86a69630ce414a735bfbd50158fcec472bbc78782065e1615

Observation 7b6fdd6e-43c2-4eb5-ae7c-4740d996a3fd · outbound

This paper cites Rethinking reward model evaluation: Are we barking up the wrong tree? InThe Thirteenth International Conference on Learning Representations, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Rethinking reward model evaluation: Are we barking up the wrong tree? InThe Thirteenth International Conference on Learning Representations, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:27.187371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:25.754909Z digest=sha256:ad27f2ba40b4620883a810aaede6c62dd6f7060a23e11b805341f3fac99578e1

Observation bea22113-c115-4ecd-812f-1e937335b71a · outbound

This paper cites Qwen2.5 Technical Report.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:25.864911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:25.864911Z digest=sha256:e9fa1de6726b935504973836dd9f4b47da5a34021ee79cb45158260f720ce11d

Observation ea6018df-d49a-4ef0-bbf8-b858be14404c · outbound

This paper cites Demystifying long chain-of-thought reasoning in llms, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Demystifying long chain-of-thought reasoning in llms, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:25.962237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:25.962237Z digest=sha256:7fcfce5f99902e6bec5c003a48041bf698a7c7290187be2dcf2c1063856096c0

Observation a4b3b8f4-9604-4a7b-b586-b5d512a31f90 · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:26.027884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:26.027884Z digest=sha256:1c92ea63aa8cb5804ae6cf21a78ae27c8c5f0d9b29e457db4e9bd614532bf43f

Observation 9eb3e823-7347-465f-93b5-749aa9c69481 · outbound

This paper cites ReST- MCTS*: LLM self-training via process reward guided tree search.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason ReST- MCTS*: LLM self-training via process reward guided tree search

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:26.063576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:26.063576Z digest=sha256:6cbe56ce93e76d9dc4da9084e5e9e026a5465a3ba27972b01c1f75ff89c76a2e

Observation 7f4098c4-46f3-4d53-a169-68d533074504 · outbound

This paper cites Found in the middle: How language models use long contexts better via plug-and-play positional encoding, 2024.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Found in the middle: How language models use long contexts better via plug-and-play positional encoding, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:27.018577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:26.156253Z digest=sha256:abceb6fa86baf72354d7e14cc572af8812edbcb01b192286df43116ff90c80bb

Observation 1396fc1a-f01e-4061-8aa8-07a35ce0d174 · outbound

This paper cites Rmb: Comprehensively benchmarking reward models in llm alignment, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Rmb: Comprehensively benchmarking reward models in llm alignment, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:26.853533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:26.285176Z digest=sha256:fcdf567aa0c1c1f49a57e3c3c57a5564225434a3761bf218dcfe338b29b94d78

Observation a4ece631-7ac0-429d-88eb-33ba21a03c33 · outbound

This paper cites Assistant:␣<think>.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Assistant:␣<think>

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:26.655809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:08:26.418982Z digest=sha256:e661c2a56e199014fa441f829519dcb984127b73bbc1bf73b785b44258fea144

Pith citing papers

Observation 2c78ad0a-163e-4860-87d2-2b82f402346b · inbound

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context cites this paper.

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:41.900632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:21:41.900632Z digest=sha256:112971edc01b1197b82d659943581b7bf1952fd90726e5e6a3ab1d82e2f0dcda

Observation c1e0f7d7-0436-40f6-9d38-34d50fec9880 · inbound

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason cites this paper.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.857127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.857127Z digest=sha256:a0c8d1e80ecd77273ce5ead7b5caab170af4e55ee5eaad245ac6c32fbbab9392

Observation 44a3cdc7-cb6d-4f98-8329-29478b16b20e · inbound

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions cites this paper.

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:35:50.973056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:35:50.973056Z digest=sha256:1b5de5e91d8d7df46a4865d4d5b79893679f97dac3146f9f7c66739669c04cd4

Observation 108954cf-c305-4ec2-8327-6849d809a05d · inbound

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR cites this paper.

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:53.070918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T18:52:52.969408Z digest=sha256:5e180838692bdb1a1f801a9e8b6a555949140c792e44fdff9aa74176f3dd2315

Observation 8ebcca75-35d3-4ee1-87a0-82b75a9e79f3 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.164977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:805b7fc585973df39f1b4c98af02751b2231c138688bff53a8fc6e67093e09d0

Observation b8673b48-961e-4c8e-a55e-11795d86cadd · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.695124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:9ce3b6051b9532b1d825978c964bf6d236987580c7e9a76c30f2000e5a6fce2b

Observation 7a0e6b5b-8f8d-4cde-9b40-2fdff5e0c9b4 · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:06:40.820599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:c342ba33488040c30c7c32b9f40083943171b05150433ae1fdb39bdec465d183