Pith. sign in

Paper Citation Record · LEDGER

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 5 inbound Pith citation observations for arXiv:2504.13367.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13367 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:14:48.950737Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:43.292037Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 43c2f155-1b99-4149-bfaa-030f7a739ab3 · outbound

This paper cites write newline.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.696075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.696075Z digest=sha256:a50dc7077d9325b6bd14e81ee8e13e3305c417be6a3dc86dce71a01113bbbbc4

Observation d4dc3494-2e7f-460b-89c0-2fc42f74b2ae · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.702938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.702938Z digest=sha256:78fdd14a93e76d6d65d9929e9c367f0f67c36ce8d9b303b082802a1ba5eb259c

Observation 131b31e2-6268-414d-a5fa-76286ca826cb · outbound

This paper cites Over-reasoning and redundant calculation of large language models.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Over-reasoning and redundant calculation of large language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:14:49.666718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T12:14:48.708938Z digest=sha256:1d0025382fca8f022b9fbaf3ee86ab6d5c17b53b987108392428d72dc21c142e

Observation 8e826a7d-29de-4795-90b3-c6956d153280 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.728995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.728995Z digest=sha256:f1f04383a1decc675b1a113a1f7985aedd914f6cd31c58225adfead47b6f5442

Observation 40b657e6-77ab-432e-bf09-ce7f7a9b9f7f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.765172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.765172Z digest=sha256:dbcf599fb0874f8f51126ce7fadab3a0d0b276e2709633e860698bb9c4a6dba6

Observation 89397f45-798e-4cc4-b659-dea66024f808 · outbound

This paper cites From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.806467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.806467Z digest=sha256:d7c6309ce950a15196b76ebe8443c544a483a91666e901a60daf9dc752bcec16

Observation 8138f147-c3ab-43d0-b091-4447a71f61e3 · outbound

This paper cites Mercury: A code efficiency benchmark for code large language models.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Mercury: A code efficiency benchmark for code large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:14:49.649927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T12:14:48.819812Z digest=sha256:fb738e79854787cd213567d7862d8d9c512d986f0b61979e45b0d90ca3e888f8

Observation ebc2cf72-08f3-4931-b578-8a19db43dc35 · outbound

This paper cites The Llama 3 Herd of Models.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.825269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.825269Z digest=sha256:4e6fd000b54386aaf60be455c6803c0f86262924670f30fcd2468c8eaf599585

Observation 890f2dd9-ab3c-4901-85ab-af273b6b27af · outbound

This paper cites Efficiently Scaling LLM Reasoning with Certaindex.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Efficiently Scaling LLM Reasoning with Certaindex

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.831035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.831035Z digest=sha256:c76e108be0a2c4a62d2778759a6c388bdb639f09dba68152a3be57e74f585e00

Observation 14891e74-26e4-459a-90ac-3eb649d0eb48 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Measuring mathematical problem solving with the math dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.836274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.836274Z digest=sha256:c149a9d64416a53082a16affc48c3d7b988c510a0396c30fc555a15287ea81ef

Observation a9687924-55a1-4f17-88e8-0144c8d3e4a6 · outbound

This paper cites OpenAI o1 System Card.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models OpenAI o1 System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.841825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.841825Z digest=sha256:f0153818e0ac01ee5c4e37dfa6a1bd8ed701ac4015bb8fbd0d2a6459b40fcc0f

Observation 9ea386ff-7360-4b18-b448-7968903ce8cb · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.847707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.847707Z digest=sha256:b0a96ccd2c045ab02aebb00610059937c8b85277d97a521bfb70173c2731d992

Observation 08a3d596-66ee-4b56-83f2-0580513257f2 · outbound

This paper cites Let's Verify Step by Step.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Let's Verify Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.853345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.853345Z digest=sha256:b876217dc89814d823e12ba7d90acfbe557a3f3f5680731381678514a06182e4

Observation 18416b25-db0e-4c2a-846a-818b13c5e27d · outbound

This paper cites Zebralogic: Benchmarking the logical reasoning ability of language models, 2024.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Zebralogic: Benchmarking the logical reasoning ability of language models, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:14:49.622475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T12:14:48.858819Z digest=sha256:cf4ba0f953620426bee117df0d0ad28a11a4b2bd1a54fdfa3c687e68b110acfd

Observation b0895ca6-2b2a-4fea-99eb-56b6826f249b · outbound

This paper cites Can Language Models Learn to Skip Steps?.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Can Language Models Learn to Skip Steps?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.864121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.864121Z digest=sha256:c2650add19d406c4584343582e2e53b35b1e55b9a67ba126dea473872ccc8275

Observation 18ef791c-9136-41eb-bb59-eeeaf2c62948 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Gemma: Open Models Based on Gemini Research and Technology

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.869627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.869627Z digest=sha256:2c5f13a7aa03b2c16841489c1ec188fb0068c616fb550c8ee42b81100eb6c97a

Observation 05040b97-1084-40e0-9fa5-d14a24de9d5d · outbound

This paper cites s1: Simple test-time scaling.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models s1: Simple test-time scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.875230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.875230Z digest=sha256:17756a1dfea42b53a5f63b6c8fd07d9533451185d8026fe3aa7c84a57ea50f8b

Observation 8da46fcd-b93a-47b4-bc08-1c3ed408f3ac · outbound

This paper cites Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:14:49.606259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T12:14:48.880594Z digest=sha256:ab735c7bd01323db640ab18f2e6bd4ade94472d9cfe5b7ed253e7a458c1c4dec

Observation 09fe61dd-0ab5-4525-816d-8cbe1ee31d2f · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.886922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.886922Z digest=sha256:175442d1d892601e8b8c19e3a560fcdbadeaec47e846668074a60f64479d9040

Observation f16e2643-33ed-428c-9448-276d20d50424 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.891900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.891900Z digest=sha256:28e4c08224605734bd82ac7a6a97f898bbbf6ecc7d334e9fa0f99427524ef1d3

Observation e839c21f-89ab-424d-abcb-1bda4628b216 · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.897470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.897470Z digest=sha256:462feeaa0019968c1de6baac41e13658346b2af6e675fb0267a79bacc359ee41

Observation 558dea1c-e9b8-4584-8e69-fd9aac865849 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.903190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.903190Z digest=sha256:8bc5ceff86ee0bf47163511c3c6479d0cce9cf45fb1537fbca11246293057642

Observation ac3de784-0977-4b60-8761-e7545f859678 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.908634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.908634Z digest=sha256:9d86ddbe5eea857fa3c6de3fb95ac8cba6e7ff495a0a3763c33660c882501156

Observation ba55d97d-57b5-48b0-b4ac-e7e288568540 · outbound

This paper cites Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:14:49.577326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T12:14:48.913689Z digest=sha256:fcc4bd1a92df73901ce536e240be0cfd6530dade80869f7aa812de77c3f2fbfa

Observation bb2eac34-96ff-4878-8cbb-57175fa68ac2 · outbound

This paper cites Qwen2.5 Technical Report.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Qwen2.5 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.919605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.919605Z digest=sha256:3c8d9667c88e62c27ab159d28f7a28d8896dfeb1d436784c1be8653fbd0756c9

Observation 0f3c8050-3965-40e9-931c-3b6c50b10716 · outbound

This paper cites Towards thinking-optimal scaling of test-time compute for llm reasoning.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Towards thinking-optimal scaling of test-time compute for llm reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.925257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.925257Z digest=sha256:2d7adf3f7aa9a9b885400ef9b2fe4e739181a70e116bbe1ee84c32c99cec7392

Observation 9f6f807f-5a06-4f50-a79f-b2d089b54981 · outbound

This paper cites Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.930518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.930518Z digest=sha256:7d867ecd02961f9d4559d4092d6bb690eb113dcb237f8d24806e47615be9933b

Observation 1cdeee7e-923d-4e0c-a08e-82d9643bf942 · outbound

This paper cites DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.935323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.935323Z digest=sha256:18f2ea0057970246e51ccdf68f97c4a71f7234aaf2abfa339ea0033eb514ddcf

Observation 52afbd10-403f-48a1-ae6b-0685b53e8fc9 · outbound

This paper cites @esa (Ref.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models @esa (Ref

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.940185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.940185Z digest=sha256:96c2792c2a2f305f94ced003bb01d83164494e0b1a73bbb9deffacc227e4ca35

Observation 40f9b3f5-fa2a-4150-b323-baef6c1e8ba7 · outbound

This paper cites an unresolved cited work.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.945251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.945251Z digest=sha256:7c7333b900749009745dd12f4f8c5545521eb4b5f1aef40c9d69fa9f47bfb103

Observation ddfa9544-8400-4079-bfc6-3dddd38848c9 · outbound

This paper cites JH cBP ]D aW Z *+b tW77wϙ 9w;Ri @ O.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models JH cBP ]D aW Z *+b tW77wϙ 9w;Ri @ O

Reference 31

Resolution
malformed identifier
no resolver link, observed 2026-08-16T12:14:48.950737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.950737Z digest=sha256:bb265f898e557af39b9d714b1bbdd3549ee2d053541ff32ca7ca4fc64b52222f

Pith citing papers

Observation 61691de3-89af-489a-b581-36ef3b0a6fab · inbound

AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time cites this paper.

AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:43.292037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:43.292037Z digest=sha256:339bb3ea7d6cd723d38713863718d12d830ea569a587f83ebcb0f18259d8c727

Observation 30c26e1a-4e14-4d49-b0b5-ba9f08e06c7d · inbound

How Far Are We from Optimal Reasoning Efficiency? cites this paper.

How Far Are We from Optimal Reasoning Efficiency? THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.203391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.203391Z digest=sha256:fdeeb892e793678fa5f72e1822dfe57a65dd381d888906afeb5510c588e036bd

Observation 965ee1d5-4697-42e8-951f-79d218f4d34c · inbound

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models cites this paper.

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:27:07.339264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-19T06:25:12.799097Z digest=sha256:f59e1f5655319b5644ea4e8105f2bcb495ffc2593cbede6c35d5e28c81168284

Observation 7efa6940-8689-4eaa-abf6-800d893ff98e · inbound

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens cites this paper.

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T17:04:38.732463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:04:38.732463Z digest=sha256:ecb309de13fe7eda4dda89ca5c5a2de5159b83cd88a222810274eb7aa54eb463

Observation 14627de6-5ed9-4683-92de-60bac275179f · inbound

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models cites this paper.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.464479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:05442935d6b5b3256eb425b2d1411e4c493eb6bec873489fe5acc9617b85f3f8