Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:08:26.418982Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 7 inbound Pith citation observations for arXiv:2505.22653.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:08:26.418982Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:21:41.900632Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T06:06:40.819064Z
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 336360d1-d394-4743-b0ae-7be6a657ddb8 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Rethinking reflection in pre-training, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfee5b61-bfa2-41b1-8ec9-ebeaef0ecf13 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Math- arena: Evaluating llms on uncontaminated math competitions, February 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 998bfda0-0c08-46be-aa8c-b9ec8d061431 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Do not think that much for 2+3=? on the overthinking of o1-like llms, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd4d2aa7-3ebd-4178-84e4-761d8bc8e0d7 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason The accuracy paradox in RLHF: When better reward models don‘t yield better language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5cc1211e-9c9b-4cbe-a5b8-4b980717d800 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Fortify the shortest stave in attention: Enhancing context awareness of large language models for effective tool use
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fa6a6ec4-aa68-44e1-ae75-ea2c1ad58338 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fbaa476-27b7-4e7f-a4ab-9f973e655974 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Q*: Improving multi-step reasoning for LLMs with deliberative planning, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3881cd0d-3e5b-43ba-989f-bbf78fe41ef5 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Fleiss et al
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2bca5e9a-1439-47f8-aaf6-7f87adf5a74b · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph E
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 64d7804c-fb4f-41bd-872c-ca23e7ddd1b6 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd96b21b-2804-4f67-adb6-7cb18f974884 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Training large language models to reason in a continuous latent space, 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 635020b9-ad11-4fdd-af1a-d3fb2e407619 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Measuring mathematical problem solving with the MATH dataset
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee38baa9-bbe4-4dfd-88bb-1f454b229ffb · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378ba01a-8ced-4322-ad0c-92e08e401e32 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Human-centric dialog training via offline reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cca45cb9-7636-447b-b605-72d834896154 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Smith, and Hannaneh Hajishirzi
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a2855bad-a891-4e71-8de2-9b2da73099ff · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Skywork-reward: Bag of tricks for reward modeling in llms, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c252d4a-6eba-44fc-8507-e4a48c7916b7 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14cce726-7ded-4d3d-8c40-62737ca8a218 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason RM-bench: Benchmarking reward models of language models with subtlety and style
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f5711fa-ead4-4460-b00e-d306e59b5d64 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason The llama 3 herd of models, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08f18d02-a882-4f7a-be83-c2d6097de2ab · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason s1: Simple test-time scaling, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69dee594-ee21-4766-be72-8080c5d2b93d · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Webgpt: Browser-assisted question-answering with human feedback, 2022
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10d11159-eb51-492b-9e73-18ce08a42abe · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59acf32a-c8fb-4bb1-ae9e-927664cdeaf9 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Tinyzero
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d818213-3f27-4f53-bf0e-9dcb03f1d974 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Lee, and Sanjeev Arora
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f037cab5-2d76-46dc-a77e-88e8f44d0696 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9181ac7a-da72-464e-939f-190848bf5bc9 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason High- dimensional continuous control using generalized advantage estimation, 2018
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c5364e-0c72-46a3-aeb0-994d299d28bb · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Proximal policy optimization algorithms, 2017
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c7bc53-22d7-49cc-9177-88a313d57cbe · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason HybridFlow: A Flexible and Efficient RLHF Framework
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4272dc37-bbc5-4ce1-aef3-aed663720f48 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Kimi k1.5: Scaling reinforcement learning with llms, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26abeed2-a18d-420c-9afd-e7d1a79adca9 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Dedicated feedback and edit models empower inference- time scaling for open-ended general-domain tasks, 2025
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7b6fdd6e-43c2-4eb5-ae7c-4740d996a3fd · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Rethinking reward model evaluation: Are we barking up the wrong tree? InThe Thirteenth International Conference on Learning Representations, 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bea22113-c115-4ecd-812f-1e937335b71a · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Qwen2.5 Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea6018df-d49a-4ef0-bbf8-b858be14404c · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Demystifying long chain-of-thought reasoning in llms, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b3b8f4-9604-4a7b-b586-b5d512a31f90 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eb3e823-7347-465f-93b5-749aa9c69481 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason ReST- MCTS*: LLM self-training via process reward guided tree search
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f4098c4-46f3-4d53-a169-68d533074504 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Found in the middle: How language models use long contexts better via plug-and-play positional encoding, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1396fc1a-f01e-4061-8aa8-07a35ce0d174 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Rmb: Comprehensively benchmarking reward models in llm alignment, 2025
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4ece631-7ac0-429d-88eb-33ba21a03c33 · outbound
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Assistant:␣<think>
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2c78ad0a-163e-4860-87d2-2b82f402346b · inbound
ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e0f7d7-0436-40f6-9d38-34d50fec9880 · inbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a3cdc7-cb6d-4f98-8329-29478b16b20e · inbound
Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 108954cf-c305-4ec2-8327-6849d809a05d · inbound
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ebcca75-35d3-4ee1-87a0-82b75a9e79f3 · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b8673b48-961e-4c8e-a55e-11795d86cadd · inbound
Trust Region On-Policy Distillation The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a0e6b5b-8f8d-4cde-9b40-2fdff5e0c9b4 · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.