Pith. sign in

Paper Citation Record · LEDGER

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.04332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04332 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T20:01:06.623978Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a220e5de-7d5c-4501-871a-ccaa756fe300 · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models A Survey of Reinforcement Learning for Large Reasoning Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:5ece0da32fd44935b5d1956f2a9e2d01945e5dfa6b71943a1659d23bae364a12

Observation 14b14a91-0652-45ef-89c9-6e2f347d73b7 · outbound

This paper cites SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:9abaed4a174ba317b4838f8a8062cac4fbd85395964da3a465d4f43afa74b14e

Observation 8cdab405-7cf5-4d94-be43-5d3540f04d9a · outbound

This paper cites It Takes Two: Your GRPO Is Secretly DPO.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models It Takes Two: Your GRPO Is Secretly DPO

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:3bf5d1a268967b5df218297eb9a52e52b95d8e59acd9b624ac91176a421ac012

Observation 8d8dfb08-5b5e-404c-9640-a71ee07a683f · outbound

This paper cites Outcome-based Rein- forcement Learning to Predict the Future,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Outcome-based Rein- forcement Learning to Predict the Future,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:85bfec1031bd5b80429ee9304832ef4d67eb8b02360d20a596b69bf64deeac24

Observation 72a7312b-b801-4abe-9926-cf3ff656fb01 · outbound

This paper cites Firm or Fickle? Evaluating Large Language Models Consistency in Sequential Interactions,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Firm or Fickle? Evaluating Large Language Models Consistency in Sequential Interactions,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:738786a52d4fafae3189834299b460ad6fe2e499fecd2246f6f390e8f89d9dd0

Observation 9fcdf683-f389-4078-b34a-6fd8beecd64e · outbound

This paper cites Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:5d667afbf2b1d52e256b935a389aca176004342abb976505d032001975d1dc0e

Observation 2105efa6-413f-40c9-8f38-ccba8458c4e2 · outbound

This paper cites an unresolved cited work.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:17409d5168c244b09c40daabb96049add1be688005c97175030f0ccdcf85d0b1

Observation 1e1ac6df-f903-4a0c-ab74-697491d8a9a5 · outbound

This paper cites an unresolved cited work.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:3728925296ca08dbe222ad2b3519671717ffbacdebb66637366ee99e936813d6

Observation 552e4d26-c4c1-4602-b77c-a61592b04498 · outbound

This paper cites GPT-4 Technical Report.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models GPT-4 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:7f36b14478a6ba40b2b559dc451b4193ef6250e314f2341b77fcc2554b55d674

Observation 829a1b1c-21f1-481b-9f62-9c15450639b4 · outbound

This paper cites SFT memorizes, RL generalizes: A comparative study of foundation model post-training,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models SFT memorizes, RL generalizes: A comparative study of foundation model post-training,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:1f63fea7b34fbd96674312a60d051e1694bc58f4bf180db53a37aab825600f06

Observation eb3bfabd-0f07-425b-aea1-b5ac96dc0ca8 · outbound

This paper cites RL’s razor: Why online reinforcement learning forgets less,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models RL’s razor: Why online reinforcement learning forgets less,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:8e13d0c9e5307d9ec8d6f0d9097262582bf55c440633d8830656e3a0b7296c73

Observation 20833cee-8b4b-4ffc-a583-9bc80c1e743d · outbound

This paper cites Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:b47e824e5b351df2bb1475961e965b2f4281d1687229ee7671e6c0afe20d8d20

Observation f41aa734-1a8c-4352-957a-802fff9735fd · outbound

This paper cites SaySelf: Teaching LLMs to express confidence with self-reflective rationales,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models SaySelf: Teaching LLMs to express confidence with self-reflective rationales,

Reference 13

Resolution
malformed identifier
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:df88ad6d6fdddaca371158ec70fc67e48fa01da2b12ce3b36d25b2b31e9b9081

Observation 5f41b44a-aed2-453f-bdb2-d7fe5ee1d8ea · outbound

This paper cites Confidence Before Answering: A Paradigm Shift for Efficient LLM Uncertainty Estimation.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Confidence Before Answering: A Paradigm Shift for Efficient LLM Uncertainty Estimation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:3e90369cc3046f609cf0ce0c2afc1bf576f7e345f1723e852ae7159c98c35a11

Observation ea9945b7-bd89-4aaa-b59f-77455df36b64 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:86b37c8c9d1f9694fcc2814cdb6ed4742f95f85b2427c61a8a10758a2782f204

Observation 6c1dd897-8b16-4eb1-ac37-adabd9cc5d5d · outbound

This paper cites Admissible probability measurement procedures,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Admissible probability measurement procedures,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:4769f236cb9752470b83d37a653274d29aee9d2d62e0f81ab387710f88909ae0

Observation f9727fec-f34c-4af3-890d-d4e977eaad9f · outbound

This paper cites an unresolved cited work.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:b84e865c3fd2d2bc2a4d850f63597a0424637d7f82a462e5c1aad89bcc4298f9

Observation fbc17b0a-6311-4668-89b6-f506f9bbeb86 · outbound

This paper cites BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:d97d7012bd100ca91a9ecd52842531a6ebdab62358544ae5e145ed2e67a47d79

Observation a48ec1e8-5be8-40b7-834a-7ea622dc476d · outbound

This paper cites Why Language Models Hallucinate.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Why Language Models Hallucinate

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:a38f827402461e96c8701e542e34de85fbbcadbca245d0c063f3ba8dd29a338b

Observation 58ce190f-a24f-46e5-8dfb-87ec167a8392 · outbound

This paper cites HotpotQA: A dataset for diverse, explainable multi-hop question answering,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models HotpotQA: A dataset for diverse, explainable multi-hop question answering,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:2371ca593d27d567f64e57cb7a99b8d9c1dde9e504aa2149ab0b0d06acbc0b23

Observation b88bc4ac-058b-458a-800b-2418edf48972 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:6dc6e415e531cc15716d75f3dadca56b28cbd2bf274241684cdd76e7dbb1e0bd

Observation f640058f-4066-40c0-b8b7-8541b8dd9730 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:df0dccd81bf4cce23001d267849453b9c0b863a1e4d64982c22ae8cd886dfa97

Observation b495b955-400d-485b-bd24-cc79af99dc45 · outbound

This paper cites Qwen2 Technical Report.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Qwen2 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:df04325afa151b4a78ae56965f0fdc1d6fb54e50b868aed8678b38569cb6c567

Observation 195e7627-36c7-40e6-a00c-1042e9f08ff2 · outbound

This paper cites an unresolved cited work.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:012a2f6f2e9900c88fe7470c1c4911bc61cde616c557a266eb6e0384b3b0a58e

Observation 59fbb22e-1df8-4f57-ac4e-e9641c1fd19d · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:4958e4fa6de4d1e43f1191e059ddbdcea094e084132670afd177ff118446a9c6

Observation 5ed11d6c-ee92-440b-8fb4-94d7cf5a503c · outbound

This paper cites Qwen3 Technical Report.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:f151b0fbb854b4b0d9d66be6fae2bd21b0be9fd0d0e62779653b3ff6538cc416

Observation fe127a5d-f4d4-4a23-95c2-5cf57a3c2661 · outbound

This paper cites an unresolved cited work.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:40813af5b9b47e4d25613595e3c499d52affd33f63b3597fd2e10258451ccb9c

Observation aaf0b481-ec22-4976-a8db-0cf8c7b45d8d · outbound

This paper cites A General Method for Comparing Probability Assessors,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models A General Method for Comparing Probability Assessors,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:346fe933ecdd6c4d0aa4d67f2ff7ae36eb666894ca86674e111bbb3e644780d9

Observation 3efbe34a-f4f6-45fd-b435-8693e67a15da · outbound

This paper cites an unresolved cited work.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:46fc71e702ab78a9c6854234936f1a5477294484d84ae71331eee6a32f911830

Observation c2bfca92-7de1-4bdb-898f-ca227b722aa5 · outbound

This paper cites Strictly proper scoring rules, prediction, and estimation,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Strictly proper scoring rules, prediction, and estimation,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:8b7233acc65b114a50c584b269aca4dee31ecd2b28d3999b810936dce2c478ff

Observation c213e2d3-1013-49eb-86aa-e42925393cce · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models LoRA: Low-Rank Adaptation of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:ba0e02f1e8261a3a74d7cc920d6ec90e9befc4e7a4ed6893a3da6580167f03b1

Observation d486a7fa-41ef-40f6-972b-059fe2b112e7 · outbound

This paper cites Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:615edccf36c3155b624235445c65ebc2ac43f848d89c8c7035e630778aa0ad65

Observation 6a8a7361-e3af-4335-aa39-bdd53b5ba9fa · outbound

This paper cites an unresolved cited work.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:22d86fe174b32dcbe4d2f5b69487fb5c0ca903fe0ff8b17d848b11fef6c4ec5b

Observation c259ad32-a8d4-4767-9fa4-d4faf44a274a · outbound

This paper cites On Calibration of Modern Neural Networks.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models On Calibration of Modern Neural Networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:b5f66e29f4735ce11eb984432a19f6bbe7b5faa0706d85666c18e94d7c84bc20

Observation 33e8cbd5-89f1-4c0e-9127-07b3c1a6ac20 · outbound

This paper cites A statistical theory of target detection by pulsed radar,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models A statistical theory of target detection by pulsed radar,

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-11T20:08:14.251687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:99b26ab93d340c02ba64ea3bbd184ac4128ff21e2fcda305bca1dade213e2bff

Observation ef4e4af7-e240-4f84-9688-f525108920a3 · outbound

This paper cites The theory of signal detectability,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models The theory of signal detectability,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:6395c89dc6b3d767086d5f55b8a6595de27187b9bbc28318d8f13678bdf23ade

Observation 6f68b387-ed16-4195-b404-b2759d6b5463 · outbound

This paper cites Verification of forecasts expressed in terms of probability,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Verification of forecasts expressed in terms of probability,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:19d39ed2c9c6f648442bc2fe1d62e4e253a935b7c80fae6a5971b5a81c82b31c

Observation 20230dff-c10a-4d32-9cfe-d403097cc923 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models ROUGE: A package for automatic evaluation of summaries,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:98a84ce590d52f4fdb46ab45175b7c6bb7becf1431bd3f1bcd8aca063fc4f003

Observation eb4943ac-b91a-495b-b4df-7686b2808feb · outbound

This paper cites Available:https://aclanthology.org/W04-1013/.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Available:https://aclanthology.org/W04-1013/

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:dda2328b118e99715d0cc990b1eeff62850ef13ea48a1e01515c93402cd7007e

Observation e083bd8d-01fd-45b0-b92d-6a2988c40621 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Efficient Memory Management for Large Language Model Serving with PagedAttention,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:ea41e65aa5da2f7e7e9b3d705ef46ef6be25978424c48b6c9ca9694f61ff91f7

Observation aec3a7e3-3fd6-44f2-af9d-435d2c2b9de0 · outbound

This paper cites Decoupled Weight Decay Regularization,.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Decoupled Weight Decay Regularization,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:f4f2b6aba5ac95491f83a09d863ff3e10ed5b57a97e3ee02cb9e742e385e2c8e

Observation 3fc90be2-76fd-43d7-8142-ddd6367fbbc9 · outbound

This paper cites an unresolved cited work.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:b2bb4806aed617e68dc8b843c51673ebd10b6f5ff28231836e42fdf1dc575a07

Observation 089d2199-ed5b-4cad-a769-2b66af852dd3 · outbound

This paper cites Devic, C.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Devic, C

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:73b6ae207204952f51c27d3e6664d130029018d82ec83c9bab94870a78b5b48f

Observation 884c7249-7ac2-4069-8493-394f70e9f434 · outbound

This paper cites The Llama 3 Herd of Models.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models The Llama 3 Herd of Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:e10a2ec4f0d532490828fe8594cee294e33441051064e20da412fcce1453e4ef

Observation 8ef90444-179c-411b-8300-920208accd13 · outbound

This paper cites For our experiments, the minimum possible representable confidence value is 1 202, hence confidence reward hacking cannot take place in Log- 1 ln 202.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models For our experiments, the minimum possible representable confidence value is 1 202, hence confidence reward hacking cannot take place in Log- 1 ln 202

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:a8fde3e5297c2a3efac0eff490fe75b5efaad0d34743dd23204f57e1463883cc

Observation e875adf7-587c-48a2-b0fc-13bdb7690e48 · outbound

This paper cites In Appendix D.1, we generalize this finding by proving a sufficient condition for a non-hackable reward confidence scheme to exhibit overconfidence or underconfidence bias.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models In Appendix D.1, we generalize this finding by proving a sufficient condition for a non-hackable reward confidence scheme to exhibit overconfidence or underconfidence bias

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:4c4184638243d6fa3f34fe84db4d5625e1bc67853cbb710180195c2a92b967f5

Observation a52d87f1-9274-4cd4-a338-e26bd7cf9c34 · outbound

This paper cites Final Answer:.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Final Answer:

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:20cd63f909fb9821efe87d841f6d8594a90d8e8095a4f413f283e79bf22a34ad

Observation dc10e0de-61c2-4e28-bf46-4f84e7eb5743 · outbound

This paper cites an unresolved cited work.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:18f25085b465dbbcf69988b90d9a8938b0673ae343db19ceaf73cc9cb8441297

Observation a37ac73c-4a2b-4f5e-a98c-5697128d98f1 · outbound

This paper cites For mathematical answers, answer in LaTeX format.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models For mathematical answers, answer in LaTeX format

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:aad880f4633111dbb6c56dec193aaa575f3aefd6e3b9d604e6fd0a5393db94be

Observation cb868bef-2c00-46b5-8a2a-400ce35bbee4 · outbound

This paper cites an unresolved cited work.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:245b501417a4782773cb419cd616008982b8c62737e7fc0efae822e84f34f3a1

Observation c11076fc-4da9-4b45-8a8a-d508b8ebb589 · outbound

This paper cites The Grand Tour.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models The Grand Tour

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:9dd092435c5f7b75027b3dce9db66f13b61557b4a946e451eb12278f6b7fee7f

Pith citing papers

No inbound Pith citation observations are available.