Pith. sign in

Paper Citation Record · LEDGER

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration

As of 22 August 2026, this Paper Citation Record lists 100 of 108 outbound references and 0 inbound Pith citation observations for arXiv:2607.13753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13753 v2

Coverage vector

measured 100 of 108 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:56:28.073779Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 108 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 443ca3e7-4616-46c7-96cb-be21c0a9fc7c · outbound

This paper cites 2023 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2023 , eprint=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:16.720764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:16.720764Z digest=sha256:047fcb523e732263a4c0b50844d3b22fc8ac5e4af74f4526be771961ffe774fc

Observation 8137768e-02ec-4d3a-afb5-67e47d2f671f · outbound

This paper cites 2026 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2026 , eprint=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:16.883058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:16.883058Z digest=sha256:890e0ab1b91b5f1bcc0d711bb48e5d53714c587421b75e2563b6f57360b97cca

Observation b818d33e-8ba0-4a99-ba8b-dd3b1a83e111 · outbound

This paper cites 2026 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2026 , eprint=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:17.046894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:17.046894Z digest=sha256:a53b04bae582214c1498aec5aec49c89ce3ef220717307f8b98ea49476767e7f

Observation 995417c0-b4b1-426d-a748-ff81dcc1857a · outbound

This paper cites 2026 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2026 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:17.305342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:17.305342Z digest=sha256:efaa6c49e3dab458982611192a1683de39f96f168dac46e754a6913b418be6d0

Observation 4ee00f7c-e497-4795-8696-7bf719da8121 · outbound

This paper cites Advances in neural information processing systems , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Advances in neural information processing systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:17.525161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:17.525161Z digest=sha256:ab56f30f8eaabfeaece6623dcb8c93b690b88eea1cc49d5d5e3b629808046c20

Observation 303f3ccf-1834-4265-9893-814d146fa234 · outbound

This paper cites 2021 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2021 , eprint=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:17.635578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:17.635578Z digest=sha256:90a9d3de4dd54f52601e01f8ee13693b210cddcc0d51092a215ca809d88e63f4

Observation 9bbd247a-6fae-4182-af5a-2a9e713390b5 · outbound

This paper cites 2022 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2022 , eprint=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:17.746809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:17.746809Z digest=sha256:03a9affacf2dce20ab7280969e0533ef779b4d3d32945078f69a5e1c1e7224dc

Observation 481f8c39-189d-4581-8f4e-39db76134d38 · outbound

This paper cites Advances in neural information processing systems , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Advances in neural information processing systems , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:17.858651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:17.858651Z digest=sha256:cf08d9a542f1911b9d620ac3e174be7d73fd307c5c8c66d575d6d6f53fec83ea

Observation a5aedc63-fff2-49b5-82b3-d7e73f61193d · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:17.970359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:17.970359Z digest=sha256:b9b6299953b3f3940e2d25fa516fefb4af607e05cda624f6deb06e819e5efa45

Observation 8db46a12-1db6-41bd-9b33-2135fa964d2b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Advances in Neural Information Processing Systems , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:18.081629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:18.081629Z digest=sha256:601b83c6ac43e3fe7209a72325ac8c5338f16e98c474e4de9b7419dd6ce1f290

Observation 14a5ed35-9b7e-417b-9ebc-e76c0a4a07e8 · outbound

This paper cites Advances in neural information processing systems , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Advances in neural information processing systems , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:18.155745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:18.155745Z digest=sha256:6c618191109ca4c18e44083b21655b53dc0814158e512e3986801eaee4bfc830

Observation d5eaecc8-914e-4465-a95f-2b7cfeb643e1 · outbound

This paper cites International conference on machine learning , pages=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration International conference on machine learning , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:18.269368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:18.269368Z digest=sha256:8b88f79a0354e3a742ffb0a71a74019f7d98daeacf59d16cd844f201562253f3

Observation 70570ddf-5a85-4d37-877f-27f3bf62b6ed · outbound

This paper cites Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:18.344443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:18.344443Z digest=sha256:eecba8e7dad42b77a04fe4976a7da02e14d7a50ff7f6c21de15c5be651f6b388

Observation 6db761c1-bac1-4316-9bb1-20e4ce62ceb3 · outbound

This paper cites Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:18.454309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:18.454309Z digest=sha256:28c228fa469d9e686b19f1141383e0b9993104ac20060d79d757df06486fb1ff

Observation ae9d9765-a8e3-4738-9e46-66826995793e · outbound

This paper cites 2022 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2022 , eprint=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:18.671429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:18.671429Z digest=sha256:d8d02c5668957c23c29387848ea612a85731c047be13645d32e9f479f2fd1b38

Observation 1553f37e-77f9-475d-a7cd-62e039ee26cc · outbound

This paper cites International Conference on Learning Representations , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration International Conference on Learning Representations , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:18.942951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:18.942951Z digest=sha256:4b152f9feaeda0558294431745da63364717131baba1164229bb533160b88932

Observation 35799c2f-73dd-4fcd-957d-f4d650b189a9 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.030784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.030784Z digest=sha256:98fa9953d94a6fb4a4b4a162b61c6588a7518557bccee85f2cedc874a18e5121

Observation 0d752492-ea31-4e86-a8d5-2bb895aa4e13 · outbound

This paper cites and Hauskrecht, Milos , title =.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration and Hauskrecht, Milos , title =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.161410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.161410Z digest=sha256:13a3a55b9653d69e737cca04674667a5fa0cd2a16ef42afdbac2b9e951be7d0f

Observation bfa3907a-f566-461b-bf16-2ecafcc4d116 · outbound

This paper cites Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.236437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.236437Z digest=sha256:4b88ffc71235e77bf7e2391ce05cb7c17e35267c751bebdc47a0381dc7c0ac93

Observation 2d482648-1523-4db9-949e-092c661cb44d · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Transactions of the Association for Computational Linguistics , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.311265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.311265Z digest=sha256:b8af3dc7d1989201fe9a8a184b603da10124aee26c1a999bf9722199e88b58dc

Observation de624713-9517-46d4-8228-da746102f819 · outbound

This paper cites Advances in neural information processing systems , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Advances in neural information processing systems , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.357682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.357682Z digest=sha256:fbe1ce547ca63f4a9b56a9d7895a236a06a1030b37493246d52cba1f97e8e2eb

Observation 5872479e-bfa7-4646-9170-0f548b7c7fc8 · outbound

This paper cites Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods , volume =.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods , volume =

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.405028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.405028Z digest=sha256:d32cc40a5017cc99c73fc9579ec71607ed19467cc83cfc95a0e28d1eb2ac5b88

Observation 2931a62b-0358-4aef-a985-a22bbc4e8700 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.467260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.467260Z digest=sha256:cad16e241fbf6da0402d6eb6612f42c8f4275484874213450a2e65bbb1a09dd2

Observation 1d21ee5c-4e73-461c-9dd0-3a159e64e037 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.539479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.539479Z digest=sha256:dabfccc2297d1c81bd636e52c6280d6a3261b8f72ed1f3ecde26029306227a9d

Observation 14429ecd-e4fc-40fb-bdd2-dcf829f58a32 · outbound

This paper cites Advances in neural information processing systems , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Advances in neural information processing systems , volume=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.688306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.688306Z digest=sha256:4fa9ec3720966d2bc50d5f89758bdfcd92250b7273cf74690a6933bae5d3dc0b

Observation b8358bdf-efb0-4455-9865-cf423c91d429 · outbound

This paper cites Advances in neural information processing systems , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Advances in neural information processing systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.785245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.785245Z digest=sha256:51d2a44e05a473ed6ed87f37a4b11870dc8021cd06d50ac14e401df395c279ac

Observation a8f3c383-7b93-48bf-8d1b-36fbef7d563e · outbound

This paper cites Advances in neural information processing systems , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Advances in neural information processing systems , volume=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.841972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.841972Z digest=sha256:7b6ddad81337226b891876d9376717a00fd6540315c03ed48003f96b2f2a11a0

Observation 3cfe850b-d513-4f74-ba7d-7fe46af8b78d · outbound

This paper cites Advances in neural information processing systems , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Advances in neural information processing systems , volume=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.896437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.896437Z digest=sha256:7ba01e5d9c239e5fb2c7923dd7c00636ae395cf15964f357e5bc571478fdea11

Observation 2ec03989-89e6-4858-bd76-777660b7ef7a · outbound

This paper cites 2024 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2024 , eprint=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:19.967258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:19.967258Z digest=sha256:d7621b036771c4e2fe5278cdf0df64a8cf2469e17beff69de43ba8d7797b4bee

Observation 23f17299-ffe3-49a9-b8d2-093aca93aa4a · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:20.024839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:20.024839Z digest=sha256:d534f80511a0f31e887a6a392e617ecab0eed8c9e09ac124e6a8b39c634bebc7

Observation d0caa55b-9345-4fce-b06c-cf68ef045cb4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:20.086170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:20.086170Z digest=sha256:50ba1bc9dc44ccaae29b734dfff4ed0c000bf2b635512404c4c2af3b0cb2a4e0

Observation bbb41a7b-cb88-49c8-b141-d2a90c8abb8e · outbound

This paper cites 2025 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2025 , eprint=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:20.204278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:20.204278Z digest=sha256:62dee5e2a35d7912d2331cecdec0f8bd4c27761453769b7f65064d82b7ca6e51

Observation ba77f5a3-39ae-42c2-9381-757c04c03ad0 · outbound

This paper cites 2025 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2025 , eprint=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:20.271780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:20.271780Z digest=sha256:7993d34dff2a04af5aa31b6c55a53afea84fb6545a3c6c2d3a12978f62a5c36b

Observation f449d21f-b2be-439f-b653-db349c7fd0fe · outbound

This paper cites Qwen2.5-Coder Technical Report.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Qwen2.5-Coder Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:20.714439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:20.714439Z digest=sha256:85bf49a3655e939b357773daa9ff0ce6a42e5f83f7800c152ad7105cc5cc8686

Observation 896c379b-cc53-476c-90ac-d0aef515b035 · outbound

This paper cites Advances in neural information processing systems , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Advances in neural information processing systems , volume=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:20.799675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:20.799675Z digest=sha256:f7e8bb44e96928de570da1745451faed2982d786548183e5cec41fa12aad15b1

Observation 3de3af29-3ee5-4466-a8be-a08728f475c2 · outbound

This paper cites 2024 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2024 , eprint=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:20.880149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:20.880149Z digest=sha256:b4a6734cc972e01be39f84dda7341783905f43255b328d6f1fbde160dfee0332

Observation ce50c94b-a3a4-471f-be24-24441214db39 · outbound

This paper cites 2025 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2025 , eprint=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:21.053384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:21.053384Z digest=sha256:581c0fe0765350bc64291e17784f395767a8f8e989df0c425ab4394aa83361f1

Observation 7c694cdb-d3f0-43dd-8de0-b2ad25436cbf · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:21.137234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:21.137234Z digest=sha256:92a69f3162472582dade7c39d4fdd5fe381910cc0c1bb72b7b2e23db8412e63e

Observation a7dfab01-8fb4-4bbf-80ed-fc97908c1d1b · outbound

This paper cites 2024 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2024 , eprint=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:21.226495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:21.226495Z digest=sha256:85a5a1f66375fd72cd5b4dc7ac32cf2040b8de265b2c64e748021f625d7ec627

Observation 5dcdc66f-3e2e-4473-8687-5f67dabb6097 · outbound

This paper cites International Conference on Learning Representations , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration International Conference on Learning Representations , volume=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:21.778799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:21.778799Z digest=sha256:affda90998e2e160e47021dd4d96898d1174a75c66c833d7cd1e092f9e069ef0

Observation 0df156b7-8300-426f-a02a-3068a2c613f2 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:21.914841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:21.914841Z digest=sha256:783f12dcd8773b06b1132f79dda6b380f2f5384983f867ad757e55200419504d

Observation 4a335679-0cce-4c6c-b368-e2629f9fe916 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:22.005298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:22.005298Z digest=sha256:3012b4a3128d54b8c4a53400b7c0c9c993df4dd70964886d3d2823ee216ef9d2

Observation b80a5562-e38f-49c0-bb48-a47b19172bb9 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:22.088004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:22.088004Z digest=sha256:1027354a5f243e34758f0a94c78941d63fa9a596706dd1b49b665a3c1e88c285

Observation 81448863-aa35-4d29-8385-7335affaafb8 · outbound

This paper cites 2021 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2021 , eprint=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:22.483686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:22.483686Z digest=sha256:eb4c5067fa60051221bf52e7841bdf25f22700eb514b1e1ae76ec0ee01b897ce

Observation ca59b8ff-fdf7-4965-84ec-c4351306bedb · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Evaluating Large Language Models Trained on Code

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:22.639590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:22.639590Z digest=sha256:e44aa7f650f4bb8be4953770ce8cd80a6330bbb9803854657fbe4a769df88712

Observation 61fb4327-a01c-4cd3-a077-1fa48da83bf6 · outbound

This paper cites International Conference on Learning Representations , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration International Conference on Learning Representations , volume=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:22.760061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:22.760061Z digest=sha256:4c67c622c87f6b72037d1e0ed5aea5701dff6e9b8b209167ff9860dbf36d82af

Observation d3b39df6-4f93-4661-9ee9-3888b3ca71ef · outbound

This paper cites 2024 , eprint=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2024 , eprint=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:22.894503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:22.894503Z digest=sha256:fb04709e0a227065e707ea1e75dee661e879b885b2110136128f622a9cc96cef

Observation 348de46d-1511-4261-bede-1369d2b10dc5 · outbound

This paper cites 2025 , note=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2025 , note=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.030533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.030533Z digest=sha256:addf1995b89163b4e58875420ded829a487453f7ec710802a7b00f33d5912a04

Observation f888b210-f776-43c4-b9c2-38565231cac3 · outbound

This paper cites Notion Blog , volume=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Notion Blog , volume=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.120640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.120640Z digest=sha256:fb2710d91bfcee2c8e1363174327b719d1958c3d741d5c254f3aa4e06068015e

Observation 83eb78a5-a572-4d8c-8c7a-c698a76b7835 · outbound

This paper cites Proceedings of the Twentieth European Conference on Computer Systems , pages=.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Proceedings of the Twentieth European Conference on Computer Systems , pages=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.165656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.165656Z digest=sha256:2d0c78d70f9bc6e150031438e5b1c063a1fdb81b000aa286a1437f54b7fdbaac

Observation 67ad94e3-5616-4e9f-917a-d7a4dcb50dd3 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.249870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.249870Z digest=sha256:4a112f9b15b808ad19c06fecdc3fe1caf82157cdc2e45c5dcabfe9130d9312c7

Observation d5a324d4-d99d-4d29-b747-38185b5d4cd5 · outbound

This paper cites Suchanek, and Gaël Varoquaux.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Suchanek, and Gaël Varoquaux

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.361339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.361339Z digest=sha256:ca2c54a2556ce0cff0366695525f87ab198b832fe511df7ce486af56d4d49df5

Observation 3c453e4c-d9c0-4800-880b-d144d49c2f3d · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.440593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.440593Z digest=sha256:bd24605bd7a19580c7d7daf8aa30719a5c9340c05192cbcd643ec441d160e89e

Observation 4ef649a2-8b45-4b4d-8810-387b6f215161 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Training Verifiers to Solve Math Word Problems

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.497995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.497995Z digest=sha256:31b535aaceba249ee9c0a577888ec48a534a5e2e5dd1be1140c9723c8bfc6114

Observation ac7bb018-069e-414f-a716-39959d97cc90 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.548141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.548141Z digest=sha256:9dbbe2b2d8678c1a28770da16015612636c95c638e8b09379cf752a402a860b1

Observation feb18894-f6f1-442e-921f-1624600d7218 · outbound

This paper cites Deep Think with Confidence.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Deep Think with Confidence

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.603454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.603454Z digest=sha256:970e6c1d4987caa9d56dd532c2d4f5cbe516bf21ec2f3dba729ed7904341e346

Observation 3ee690ed-577e-4134-96b9-af4d7d7f4621 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.658276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.658276Z digest=sha256:dbc943554234be8516687d49af5be1677d448a26537794d5a3f74e5804e9c9e4

Observation b82e8989-5a63-4900-8338-5ac5bd326d37 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.714297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.714297Z digest=sha256:8f58349861d4cca72d045e3db09f737aafcd113c229befa8e66a9d9c4ae1d745

Observation 961b300a-ac3f-40e0-908f-beb82197fc5c · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.768342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.768342Z digest=sha256:aef625a90d5c3f7fd4573cf1237b00441d88e7f0f631797f631c3a5de4dda3c7

Observation fa844820-789c-4b45-98c1-a39b92167c1a · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.823431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.823431Z digest=sha256:c3a6e85b0522b9b69334f8070dbf04a7152fafb7873368f968c86a1dcac69aec

Observation 21d58873-741b-4fd8-a84f-6f553481f917 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Measuring Mathematical Problem Solving With the MATH Dataset

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.888612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.888612Z digest=sha256:6f59c5a41be19245b7c3900ad830b7668118550613442232a7744e7479f8f45e

Observation 799011ba-2eaa-4488-894b-d1c81aebb9ee · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Distilling the Knowledge in a Neural Network

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:23.929756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:23.929756Z digest=sha256:6cae018bfd387305ff8e8b1d3564a0e2027f3d4c98e0563b6270317f7b2286dc

Observation 0ff3e0b5-87f7-449e-b92d-40ba825e54a2 · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:24.014273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:24.014273Z digest=sha256:21a84c6cf2c40a740e5c5b1c551d83c5eb6d84166b44ceb9e54fac1c21c6eae1

Observation 75be82ed-5f57-4301-9cb3-29cd980988fa · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:24.068987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:24.068987Z digest=sha256:ce38ca1b0a19ea474b56b896b3347957fe4800cb3ff03e715c974767d4c15513

Observation 2f972ddc-0bc2-405e-989b-8f7482a4d691 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Entropy-Aware On-Policy Distillation of Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:24.076911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:24.076911Z digest=sha256:1e7877be6af4fc839351cc977c7c346d8f5b6a0080357c62d6d450cbd02c6d18

Observation 03eea6e4-c882-4cbd-872f-e4a7e795d3a8 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:24.096298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:24.096298Z digest=sha256:b4fb197958da22be692bfa4702907e17279ac5dd515d7e5854fa95829f67ba60

Observation 554be546-d273-4163-b323-fd0cd6788333 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Language Models (Mostly) Know What They Know

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:24.247165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:24.247165Z digest=sha256:da57d203a242c40f322c65210cd6348408a91a4cdb01a02f56869f4a83364e03

Observation afd62219-3532-4b9e-887e-6aa5939bf3d1 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:24.403912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:24.403912Z digest=sha256:7a97b0cd566e51249f1cd0476288c858eb6cd95e2fdabfd1fdf03acf3f1cf144

Observation 680bbf43-3e73-46f7-a35d-adc395118d28 · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:24.521676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:24.521676Z digest=sha256:c66a105d577e3aaf0c40f6fd9270f987f8999d958bb8d30ce3843d6c86ee7d4c

Observation c8716189-2874-4206-8d3d-81b5277bf5fa · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Solving Quantitative Reasoning Problems with Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:24.673658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:24.673658Z digest=sha256:b275dd729256813fee243dbb3748698b0990bc1cdf3aaf087a6664125a243690

Observation 3920efc6-6911-4955-a8a4-d4ed1b692a53 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:24.777752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:24.777752Z digest=sha256:bb99d48e4ecd847eb524ffff680ded5f0117b3a2cdcc01db439c22cbb5f1136a

Observation 8dc3cc09-663e-4bd5-aaf4-9188b42317c6 · outbound

This paper cites Teaching Models to Express Their Uncertainty in Words.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Teaching Models to Express Their Uncertainty in Words

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:24.978542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:24.978542Z digest=sha256:4ce6204250e6ff16062fb42ea1a342358f160e0557660583f1437ea0db9943dc

Observation d63560fb-7b10-4f43-baa7-64685e72286a · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:25.128541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:25.128541Z digest=sha256:db16f0748bff958540c2b5e8a83aee4c5c5f0cec33537b100893e369a70b1f2e

Observation 4b0f8d4a-8e97-48d2-bd37-9c2575b34c86 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:25.283578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:25.283578Z digest=sha256:829e5d7a69be1deecddec2e06b52bd2e0eb9e62d013df546b3ba527b94b70d01

Observation 823f7f7e-bf7b-4604-8f60-852d7fc3c112 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:25.392473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:25.392473Z digest=sha256:e1a3dab045233795d6a93050cb230d37bfd4caeb9ee562626a7ad78fb4cde959

Observation 0cb65d26-ae1a-4f91-a47f-594dcfa87961 · outbound

This paper cites Cooper, and Milos Hauskrecht.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Cooper, and Milos Hauskrecht

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:25.502679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:25.502679Z digest=sha256:c7440131e43dcd90b6309fcf8f001b2b5b26181e0000e0ce8ef9b556fb43d6cb

Observation 0b77afcd-db0d-4e7f-8d39-71960060d9f8 · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:25.601371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:25.601371Z digest=sha256:aa53af1491ba6195c22a8bbe57456a0c0470778298d966c5cbff11454701d1d0

Observation ee3d3407-7d5b-4cf3-9a81-0fbaf9356a30 · outbound

This paper cites GPT-4 Technical Report.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration GPT-4 Technical Report

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:25.765151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:25.765151Z digest=sha256:02e1127f08c4fdecbf9afaced7cdb5695ad38bf9492712c4eae98596486d535f

Observation 8dcf0309-6377-4327-b1e6-74cf0cf49094 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:25.929517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:25.929517Z digest=sha256:041d6a1b43e1ffa4cd627fae4bdb52a6f171c78f6d73644630ba46c1476561ed

Observation 23048bec-69ee-42ce-ae5d-6f87b4149b89 · outbound

This paper cites Qwen2.5 Technical Report.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Qwen2.5 Technical Report

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:26.085330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:26.085330Z digest=sha256:b955cb83cfc391d4eb5174e61a003ff788c2bc8588783b422b1c9ec0f61011d0

Observation cff3b64b-e71c-4d53-a4e5-2c8f3aed78e4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Proximal Policy Optimization Algorithms

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:26.201200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:26.201200Z digest=sha256:29102e4842938552215c72ba39d16eaf65e13a5d4e06323c07a543fbb4e51a47

Observation baca7525-b88d-4f3b-868d-4c077fe4f304 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:26.300681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:26.300681Z digest=sha256:300349ed0107a0a2a9372d04480d2c243b9d30ab612e71f142eb259ace4cf5d6

Observation 77d4108a-1dbd-40c2-9bc3-66605ce49084 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:26.472307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:26.472307Z digest=sha256:b2b6d8c40e2aa37e19fcee7bcc7374a0ae8683dba52311b1ca8bd32e2e82418b

Observation e0ff153d-3cab-4e56-aead-9b7e7bd48a79 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:26.701252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:26.701252Z digest=sha256:da83a4df0cb308d88392f331196074d0efec526ac29c39c6f9925bea2c3cd4ce

Observation 9c2de194-814c-47ec-aafd-045bfe1db7d1 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:26.864154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:26.864154Z digest=sha256:5a4871afe141f5a508a015db47e71f5e20818bf227e9a2a279607c1fbf0278f2

Observation dd6fe679-a26e-42ef-b5d6-8390dda94ef1 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:26.959184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:26.959184Z digest=sha256:e5785335d8d4cf0c96b99c9695b3b0eade0b14bed77051f59f372d58d0f474c6

Observation bfe0b2c0-1b95-40de-b0f3-e1591a58c3a4 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.043472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.043472Z digest=sha256:d92a336169419a6861565952e69cfae78545a1bbf3edd13f5243cc1bfb13b88f

Observation b3d05b55-f6b1-4e03-b256-80f2afb2143a · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Solving math word problems with process- and outcome-based feedback

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.103545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.103545Z digest=sha256:7f03192148e84afdd6c5dae8f97936d7f2f80a5a253c613717504fb37d9bd9b3

Observation 42bbac1a-559f-4d9c-b285-f70dfc65cc5b · outbound

This paper cites Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.159465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.159465Z digest=sha256:f5e00582dca6df0a35e3b12aeef3f84aaa0d603780755c181cfbd79f2b306505

Observation 18fd48cc-3fa1-4ebe-ba9b-d3f7a56333c9 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.235699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.235699Z digest=sha256:24a729747a010e2e6f1ed8be1557ddff57ed5030c925f39e908a2a28aff987f6

Observation 1e9aa211-3d85-4e7e-b7ee-9e22139d5307 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.321716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.321716Z digest=sha256:34d7fa5a1d6e79b8eb4e14c63a048b142af84dcfe36b4ab9b0aa3db7d7a34c66

Observation aee81d70-fd58-468c-b624-0423198e0d1b · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.406144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.406144Z digest=sha256:5c15ddf839a6aab80abc32243c0b53a221d34de8e48c7ac0d93b15a3bc153c4e

Observation 895e38b0-66f5-49e2-89de-35a5e31d6882 · outbound

This paper cites TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.461043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.461043Z digest=sha256:8588b2599561ddd291d4d0cb85da04a67617fcb6997cbe74d384af5994cd6ca3

Observation f6b27116-e6a6-4ab6-8f5c-b9a11e92cbd0 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.545672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.545672Z digest=sha256:3b15a3267a7a77c908f4a4e827f539245e399b39fcbd8a118c57e71a8bafb22a

Observation 0cba5a80-bd82-4469-ab58-fa7c906b6ecb · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.628364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.628364Z digest=sha256:687ce6d44b5873a78c59c815923909bc9595726f7c5d84a800cbdd4f957c7f16

Observation 67dd4cdc-1227-43f2-989e-788c9c1f206f · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.647136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.647136Z digest=sha256:0b169be2dae10e7a7e223c35ddc23c576ceae52eca7b10a2a536bcffdcd7ac21

Observation 7e058d93-dbaa-4d53-9c0c-4d9c58af3a61 · outbound

This paper cites Qwen3 Technical Report.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Qwen3 Technical Report

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.724507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.724507Z digest=sha256:93173687a55c023dad092187ef10fd68421e4c4432cf52e13196b33fa229c861

Observation 2a416eac-6665-4247-b66c-dc64f2311023 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.820539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.820539Z digest=sha256:b0512a1c12d1d04bc0b3152e4d28154987bd60f83df55cdd099ab2cb4de38c37

Observation db62ca64-039c-4e18-bb53-303a7ae71964 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.965607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.965607Z digest=sha256:5c37768355f9ca5b349642c07fd28b3d5ebe4f3fa19f37957bc7a1118a97b609

Observation 1df684f4-74c3-4f48-b436-f3abc4622833 · outbound

This paper cites an unresolved cited work.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Unresolved cited work

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:28.073779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:28.073779Z digest=sha256:d0bb8f8c6c95a8ad01b453140e5f8f6cdedf527a952d5d43a76e465fa9e95be5

Pith citing papers

No inbound Pith citation observations are available.