Pith. sign in

Paper Citation Record · LEDGER

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 10 inbound Pith citation observations for arXiv:2505.22113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22113 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:20:17.785500Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:24.782504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:27:24.497080Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved21
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c384100b-c8b7-4d1c-b309-d55a9f9c7178 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.917705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.730171Z digest=sha256:beb9695b05e468aa041ed084a1f933360b752ff81639a0efdecd7f921f6a33b4

Observation ac1eb0a1-e109-406b-988f-f90903cd7299 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.651898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.874746Z digest=sha256:caa94306bf5e31c8db1bc3fac95f15dd1fc324ebe35c181c2f4709cc1c022bd9

Observation 93ab8204-0496-486f-bde8-8f826a067208 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.425334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.990072Z digest=sha256:0f31edacd9f05baf6b8d8d57176c4849574c5a5518cb87bd6d328b4f29cfe086

Observation b24a198a-bec0-47a7-b87b-9f14c1907439 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.182153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.095136Z digest=sha256:0eb913d1266b9c5b4b2b5b69eaf0020c0d45b3d20b7cb0429c345daf20d5ed46

Observation e93cbc2c-dc6f-4737-836a-fe532de576a4 · outbound

This paper cites solution1.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models solution1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.896976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.203468Z digest=sha256:94eb088224d8efb27134935b918810f816eef18ed18ac47c24864a46d725c621

Observation b0690b8e-de93-40a9-ab65-9a913a55ab7b · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.332107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.332107Z digest=sha256:8ff8167999675b25875ca1cdde5cacffa6101eb6f1e4a002d221cdac815ddfb9

Observation b94e670f-245b-4ca8-b275-77d94a51c5de · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 10

Resolution
parse uncertain
no resolver link, observed 2026-08-07T13:20:15.410585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.410585Z digest=sha256:06caebdcaff7f299cda2cdb3fed7d0e23f3427877c62c26ad12fae215e6219e6

Observation a0f55e05-4ba8-4f10-b81a-5675b489ce45 · outbound

This paper cites step_index.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models step_index

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.723877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.541578Z digest=sha256:26057e2ff89e0b2aadb56174efb6fff345f5d8f1fcdc098e85f4f7a217ae3c3e

Observation 8d26cb6c-619a-45ca-9630-2699465649da · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:20.470923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.647111Z digest=sha256:603765cf3aa77cb5a5feb7441a87768298bbeafa49f01952ddc9ef133fe5f954

Observation 03e0ef78-1002-4995-902a-e9e024f821a1 · outbound

This paper cites correct_answer.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models correct_answer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.257129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.714787Z digest=sha256:97fc8c03d918294707b80ae99ee6533e904a0a253d609d80262e3d531af2ac15

Observation aef9f10d-7f23-4bb8-8aaf-f09ab6276f2d · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.844503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.844503Z digest=sha256:1ba1de6fe06e65dc8bd5146ea13cb0dce0a684a74b3fffb32587edd8fed0c961

Observation b14553b6-303f-45e2-8cb1-4b2373b2aad6 · outbound

This paper cites Match": Aligns with ground truth -.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Match": Aligns with ground truth -

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.984932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.969033Z digest=sha256:7ce1309bacb3394d0bf65dd93589d657117cb448e2edfdd6c02b803da8ba82e4

Observation 8cbf3a86-79be-4057-ae6e-b8ff97544a75 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.007807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.007807Z digest=sha256:4da837e8f2cf5c853070645984b11f1c69f89fb1c29438d1f2bade4684767472

Observation fb0856b0-a0ac-4267-b70a-af6bcb28aae7 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.097596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.097596Z digest=sha256:bcb6131fdf7eefd3c2fc7c689ff1e68ae0d10b6a602cf0e3b8c42c9fbd408807

Observation 40bbb535-5b02-4e5c-af68-02dfae855d04 · outbound

This paper cites Always include the final step that contains the answer.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Always include the final step that contains the answer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.720710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:16.225473Z digest=sha256:b84436048e20975e72e149b736cc623681a7581a7986b9ca651460734e068c2d

Observation f4f114d5-2c3a-40ea-8af4-a05d4cd84ba3 · outbound

This paper cites step_type.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models step_type

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.432023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:16.407694Z digest=sha256:975e53619129364a2c57350807c97ae796c2e581d33a02be9c9c90b0ad792130

Observation 74c47502-d618-4397-8d1f-e3632be3252b · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.505667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.505667Z digest=sha256:0eb999a363cc225156d04a160df539dbf33da73917d7651e6213e69a46775c76

Observation 839a257e-a783-4f7a-a736-932603356af4 · outbound

This paper cites 19 Invalid reflections include:.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models 19 Invalid reflections include:

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.156441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:16.674204Z digest=sha256:02fa49cb8195839b91948de234a92ea0e3d6b219614e4d2c85ad10c5d0eff6ea

Observation ad604f45-77b5-4b0c-819c-8c47d608303e · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.761555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.761555Z digest=sha256:02777a212446ebe9405b17105400ab1e359fba00c5bee35c07ed57d06d8d915a

Observation 838ee48d-2e6b-4dcd-8579-6b0d44f19a8a · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.841406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.841406Z digest=sha256:193e30e9b4dc5beedd7271afb54c56515f854d768e4e3651798d7ca2dc812b19

Observation 72863295-368d-4bbc-90bd-91da3ac42c22 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.934380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.934380Z digest=sha256:3541c608d4b6a98deaaef0461dc655712c8a7440df326d7b6b26ba21f46cd604

Observation 26672661-3b2d-440b-a9cd-7bf058668dfc · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:18.783558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:17.164581Z digest=sha256:67a1b629dcf36e9682b14ff406ef897d212709739d5d8dfa64da39653fc49954

Observation 89b1ac81-2722-4701-964c-bc09b596c553 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.242820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.242820Z digest=sha256:524f19d040a64832f0d4c8e9c04e2257bc69987aecbfb2da191fc137268fd439

Observation 92efe1ec-d154-4a5b-bb06-18470675cb3f · outbound

This paper cites conclusion.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models conclusion

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:18.406192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:17.367067Z digest=sha256:adf92032704d8265b19ed125643f963fb7eef416b4fb68fe0419617fb6dc6e41

Observation db905805-387f-4bb5-8f29-99dfd5e8faa9 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 28

Resolution
parse uncertain
no resolver link, observed 2026-08-07T13:20:17.467252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.467252Z digest=sha256:064fb5bdbabb45577665a8a8e87040e5532f91a07efb2eebbcb0bd17f9c70d0e

Observation 12026156-5ed3-45a3-b32f-5f98b0d84f8e · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.535059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.535059Z digest=sha256:0cb2bfbffb0ee44ae531e3a64d684b200151f0a12c7f358182d250282a6131fe

Observation cc4b721e-f123-4f09-ad8f-ed4e19d84568 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.616522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.616522Z digest=sha256:2dc2820eac61eff25a1f49c863fedd1895bddbb0a9336c0e718e204e1157bd95

Observation a50eea92-b63e-4e4f-b570-1d26d36d623f · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:18.118989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:17.785500Z digest=sha256:055863fd2909bfef3495bf3ed184693c0d1ad2a11155f0dc192910b32848335b

Observation 713c670a-8be1-4f03-a533-2e3109cfedac · outbound

This paper cites ARB: Advanced Reasoning Benchmark for Large Language Models.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models ARB: Advanced Reasoning Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:14.608540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:14.608540Z digest=sha256:d9e4bab8f5e3a3fc43f68d48efd599b34afa125f8170a5da9940459a6f853eba

Observation b00ac0cc-38c1-4948-b687-b6e90501a4a1 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:14.394659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:14.394659Z digest=sha256:ab41c33d0c8969a46b2341c137e865f01bfb235cc89b165b8bce970b07d3fa46

Observation 9edd3401-5a7a-4679-a16e-ab99d28e7a5a · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:14.455229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:14.455229Z digest=sha256:b1efbd3dbd6b733819ccbddb61a52beaebb7893efc8563a2c16452b88895b0f8

Pith citing papers

Observation a4fcd218-24f4-4bd3-89f4-9ed46bc3b58e · inbound

Strategic Reflectivism In Intelligent Systems cites this paper.

Strategic Reflectivism In Intelligent Systems THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:24.782504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:24.782504Z digest=sha256:42ffac42f0ff422f211dd95116f4a4db1d31ce5b3049c13f1459ee710cd365f7

Observation fcf09946-aa25-42e3-84b2-db2d72a9a230 · inbound

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models cites this paper.

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:27:07.313600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T06:25:12.799097Z digest=sha256:17a186c63820efe067e0ebb4850cd4ea6a34b02ca875663f902f014c2821a0f5

Observation 9f5055de-f5b5-42b8-8bc3-4c8f0321b50f · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.162463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.162463Z digest=sha256:984e93c43005ff08d95a2c3a2c4503072c9385bcf1eb03589e43f56311ec9be5

Observation 3753c99a-84be-4992-9af0-af1dace88f70 · inbound

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture cites this paper.

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:26.865491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T17:59:42.921867Z digest=sha256:9b497f4a57f0cd96cd4340c55d5125c8e8d142a3fe726742db9aa4a018da9cd0

Observation 823da83a-998a-40b9-9eb6-47c105ff82a4 · inbound

Structural Anchors and Reasoning Fragility:Understanding CoT Robustness in LLM4Code cites this paper.

Structural Anchors and Reasoning Fragility:Understanding CoT Robustness in LLM4Code THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:06.021781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:18:42.321975Z digest=sha256:9562d9fb12620e275bc790d6a0b28f28fcc77e3b9f1789a4e91dd9b1e2f28d56

Observation f24ec194-6adf-4669-a666-cc30c21e49b5 · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:09.086035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:ac3b29bac28d58985688fc26796d500bf14b07d2a3f0c24b163062ee55b2b2ad

Observation c7bc5be2-94be-4ac2-a0a3-5254c7bb8527 · inbound

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing cites this paper.

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.557913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T04:45:39.641535Z digest=sha256:32d82f84cefafe2e48fac64a51b8a940dc03f043c273bb25c6acf4497f4b4f16

Observation c28731c9-63ec-462c-a515-ebf21fbfbbb2 · inbound

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing cites this paper.

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:15:04.194995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:12:37.353935Z digest=sha256:d0e7db43136779d1d98bead7e6f30955762b5698b711f2692f69b483191cc622

Observation 1243446f-35a8-45ae-9684-bb4a006b1d33 · inbound

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks cites this paper.

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.499140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T19:45:30.671490Z digest=sha256:3ea41af48a9dc78e9333dd9a372f2f695de7516cfc1dde29737d9b77239bf934

Observation 864d426a-3310-4c57-a10c-2c3373f1cc2a · inbound

ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs cites this paper.

ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:24:32.223129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T09:23:26.386399Z digest=sha256:f12083e4a1a32fdf1c2a6591277eeada0ec356cf49f41a182cf9507840d77173