Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T20:01:06.623978Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.04332.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T20:01:06.623978Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a220e5de-7d5c-4501-871a-ccaa756fe300 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models A Survey of Reinforcement Learning for Large Reasoning Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14b14a91-0652-45ef-89c9-6e2f347d73b7 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cdab405-7cf5-4d94-be43-5d3540f04d9a · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models It Takes Two: Your GRPO Is Secretly DPO
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8dfb08-5b5e-404c-9640-a71ee07a683f · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Outcome-based Rein- forcement Learning to Predict the Future,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72a7312b-b801-4abe-9926-cf3ff656fb01 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Firm or Fickle? Evaluating Large Language Models Consistency in Sequential Interactions,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fcdf683-f389-4078-b34a-6fd8beecd64e · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2105efa6-413f-40c9-8f38-ccba8458c4e2 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1ac6df-f903-4a0c-ab74-697491d8a9a5 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 552e4d26-c4c1-4602-b77c-a61592b04498 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models GPT-4 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 829a1b1c-21f1-481b-9f62-9c15450639b4 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models SFT memorizes, RL generalizes: A comparative study of foundation model post-training,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb3bfabd-0f07-425b-aea1-b5ac96dc0ca8 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models RL’s razor: Why online reinforcement learning forgets less,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20833cee-8b4b-4ffc-a583-9bc80c1e743d · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f41aa734-1a8c-4352-957a-802fff9735fd · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models SaySelf: Teaching LLMs to express confidence with self-reflective rationales,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f41b44a-aed2-453f-bdb2-d7fe5ee1d8ea · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Confidence Before Answering: A Paradigm Shift for Efficient LLM Uncertainty Estimation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9945b7-bd89-4aaa-b59f-77455df36b64 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c1dd897-8b16-4eb1-ac37-adabd9cc5d5d · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Admissible probability measurement procedures,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9727fec-f34c-4af3-890d-d4e977eaad9f · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc17b0a-6311-4668-89b6-f506f9bbeb86 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a48ec1e8-5be8-40b7-834a-7ea622dc476d · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Why Language Models Hallucinate
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ce190f-a24f-46e5-8dfb-87ec167a8392 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models HotpotQA: A dataset for diverse, explainable multi-hop question answering,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88bc4ac-058b-458a-800b-2418edf48972 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f640058f-4066-40c0-b8b7-8541b8dd9730 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b495b955-400d-485b-bd24-cc79af99dc45 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Qwen2 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 195e7627-36c7-40e6-a00c-1042e9f08ff2 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59fbb22e-1df8-4f57-ac4e-e9641c1fd19d · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Understanding R1-Zero-Like Training: A Critical Perspective
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed11d6c-ee92-440b-8fb4-94d7cf5a503c · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Qwen3 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe127a5d-f4d4-4a23-95c2-5cf57a3c2661 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaf0b481-ec22-4976-a8db-0cf8c7b45d8d · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models A General Method for Comparing Probability Assessors,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3efbe34a-f4f6-45fd-b435-8693e67a15da · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2bfca92-7de1-4bdb-898f-ca227b722aa5 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Strictly proper scoring rules, prediction, and estimation,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c213e2d3-1013-49eb-86aa-e42925393cce · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models LoRA: Low-Rank Adaptation of Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d486a7fa-41ef-40f6-972b-059fe2b112e7 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a8a7361-e3af-4335-aa39-bdd53b5ba9fa · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c259ad32-a8d4-4767-9fa4-d4faf44a274a · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models On Calibration of Modern Neural Networks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e8cbd5-89f1-4c0e-9127-07b3c1a6ac20 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models A statistical theory of target detection by pulsed radar,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ef4e4af7-e240-4f84-9688-f525108920a3 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models The theory of signal detectability,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f68b387-ed16-4195-b404-b2759d6b5463 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Verification of forecasts expressed in terms of probability,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20230dff-c10a-4d32-9cfe-d403097cc923 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models ROUGE: A package for automatic evaluation of summaries,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb4943ac-b91a-495b-b4df-7686b2808feb · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Available:https://aclanthology.org/W04-1013/
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e083bd8d-01fd-45b0-b92d-6a2988c40621 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Efficient Memory Management for Large Language Model Serving with PagedAttention,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec3a7e3-3fd6-44f2-af9d-435d2c2b9de0 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Decoupled Weight Decay Regularization,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fc90be2-76fd-43d7-8142-ddd6367fbbc9 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089d2199-ed5b-4cad-a769-2b66af852dd3 · outbound
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 884c7249-7ac2-4069-8493-394f70e9f434 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models The Llama 3 Herd of Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef90444-179c-411b-8300-920208accd13 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models For our experiments, the minimum possible representable confidence value is 1 202, hence confidence reward hacking cannot take place in Log- 1 ln 202
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e875adf7-587c-48a2-b0fc-13bdb7690e48 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models In Appendix D.1, we generalize this finding by proving a sufficient condition for a non-hackable reward confidence scheme to exhibit overconfidence or underconfidence bias
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a52d87f1-9274-4cd4-a338-e26bd7cf9c34 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Final Answer:
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc10e0de-61c2-4e28-bf46-4f84e7eb5743 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a37ac73c-4a2b-4f5e-a98c-5697128d98f1 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models For mathematical answers, answer in LaTeX format
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb868bef-2c00-46b5-8a2a-400ce35bbee4 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c11076fc-4da9-4b45-8a8a-d508b8ebb589 · outbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models The Grand Tour
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.