Pith. sign in

Paper Citation Record · LEDGER

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models

As of 21 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2606.11211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11211 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-04T19:41:53.668529Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T17:54:49.531461Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact7
  • verified fuzzy3
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f87eb00c-b55e-4074-aff3-ed7288f41cb4 · outbound

This paper cites S., and Robbins, H.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models S., and Robbins, H

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:20:09.135397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:80fc55f1b709239b21189841e47cd6efa59078e430e237b9c896a8799d5e1926

Observation d6739967-4412-4e2a-9d37-28c8c6d44b2f · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.131904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:9f802738367a80af91cb693120d995218c883c9ca45ea804da38c496ca65d950

Observation 2cde472d-7047-441a-a028-5554742abd39 · outbound

This paper cites and Durrett, G.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models and Durrett, G

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:20:09.148597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:f4d014656283a5913eed6a543ad59c517654fb03932640fd6e922b8d369a87c2

Observation 9117c447-94d2-4404-8c11-30177d24e448 · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Adaptive Computation Time for Recurrent Neural Networks

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.032546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:b2ef351bbb5e94e6609a1f468653b51f4bf225300b31c4acab3853a107471851

Observation 58863f1b-a17d-4ea0-aad8-f30999955fcb · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.130285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:ccafd6ef203ec875edb57f966160ff4e2111ab59b8f8ab569106dbdce661a708

Observation f5f11320-fba6-4092-8c37-1ae056d00f20 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.146722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:3757d0c4578f7515c98a97f6dee38a53f135e066ed6d4adf7bd326bcb28c8be7

Observation 37b82850-abff-4f42-9480-14f87e70ef49 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Language Models (Mostly) Know What They Know

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.026754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:61fce5cb0430fe00bab163a37c57836fc723f6f83357504f2e506a2482efc01f

Observation ed7b77ce-7994-4b10-bdae-54e15a80edb4 · outbound

This paper cites S., Reid, M., Matsuo, Y., and Iwasawa, Y.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models S., Reid, M., Matsuo, Y., and Iwasawa, Y

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:20:09.153386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:5ed9a4422e53826148276c7c87db341d23c9eaf834f2791e59cfd424350a8136

Observation 8313b1c8-b01b-4d4f-84be-d8627593d63a · outbound

This paper cites Let's Verify Step by Step.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Let's Verify Step by Step

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.016854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:1710c6400be2624817c8960b8b36f58c32dbea929e9417bbc1608f8e2dc2244c

Observation 38b7dd66-97df-498f-ae7c-2816b5a5c5d9 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.140748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:e9297d4fafc299c227ef8b3f773fe0a18c5618a6f364b2f6218b9695e2f84ec4

Observation 0fe2b956-6397-4cd3-b797-8236d84f9571 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.142893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:de852733c00f67ccb354cd551ff0dab5328174d5f099b0919f5232a72f15cf92

Observation 2ed1e3c5-4b8f-4763-bf78-62e6dae5475e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.019780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:d4a57e2d34982a9ae0745f45dfb046ec8e520cffe5593548cf5309a109d72fa2

Observation 3e18dc46-94a8-40de-818b-c45bd05f6cca · outbound

This paper cites Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.029851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:6af6fb62e502a3a4d95a38bcd816a2c30dc711c225760c4148cc0e4a1f3b7fe7

Observation ed5194e4-7b5e-4a83-a55d-206447994df0 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.144853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:6e9674fa9458db61b51f64349d726aa7ccff6043a7e51ff601910b0280b3c2b9

Observation b5230aa5-0d4a-4a49-98b8-714690069b05 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.013847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:4337fa0bf1886d0ecc124953412c1800ca68e99d049013ce0d93183ca4feea4f

Observation 2e6fd3ac-b290-4930-9c4e-029cc767df3a · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.136380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:4b2ba495ea9d169398504a8427e3b56692f5090d4e4883f61404974ac583e096

Observation 95b704c1-5d80-455c-b195-1aea19123b79 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.138840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:6bcf4b6959d5c1ca1628cc32b5f666facd8ed07ede6383fbd9ccf041fb11c419

Observation a6d37504-b47c-46b1-9600-88db77864a59 · outbound

This paper cites Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:50:10.023261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:85004bdbafa54aa6495b3a80dc372589ed995230c3ce47d2741e1d8b43a728ce

Pith citing papers

Observation d8879794-a8c6-45f1-8b95-c0b2815085c5 · inbound

Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents cites this paper.

Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T17:54:49.531461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:54:49.531461Z digest=sha256:59132b1c9b799e7bb6a0586414cd5c098e2372e3570c122408eb79e0d7b1af16