Pith. sign in

Paper Citation Record · LEDGER

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

As of 5 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.26102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26102 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:41:53.849176Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 16446460-0a45-4338-b6de-6041d2596f4d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:50.700365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:50.700365Z digest=sha256:163aa524552955c5c841744553041b8f8b55b73cbb287a0e68dafca23dd27105

Observation 2d31b836-828b-4711-9a36-b6d256bb1cc1 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Self-consistency improves chain of thought reasoning in language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:50.817858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:50.817858Z digest=sha256:f9ea6c03ba600809a49e06287ff440c7c60681368a0bfeb64f765fef0fdfb567

Observation 2ef6c58b-d963-4bf6-b26a-90781f73acfb · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of- thought prompting,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Language models don’t always say what they think: Unfaithful explanations in chain-of- thought prompting,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:50.893111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:50.893111Z digest=sha256:396890d5b1ec611c42d0b283bb176c77480a0a8c552b84e3220e388ea7e77302

Observation 06bc27f3-5fd7-42c4-ba11-503531c8f602 · outbound

This paper cites Federated generative intelligence for explainable and autonomous cyber defence in critical infrastructures,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Federated generative intelligence for explainable and autonomous cyber defence in critical infrastructures,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:50.975169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:50.975169Z digest=sha256:3c7e9ce70db54e39b7853d95e772ac14fc108b6d0c1e9a6f1a351b1f3ba94e35

Observation a1ea8b0c-5f36-40ed-b395-37e7d3f453b4 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.171583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.171583Z digest=sha256:f28c36100b2c0d5c4bb70416bda551da3c68bd11da21e434d3f7c24222f43b49

Observation 81e50c51-c0d7-49a5-83af-0160a179ab67 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.342440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.342440Z digest=sha256:b8a2e40cada321643103b1bfb03aca40fe947b6e6eb677d47b5b344879a5f059

Observation cb9e2ce9-471c-490d-ac5d-da24b6daa20c · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Measuring mathematical problem solving with the MATH dataset,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.449290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.449290Z digest=sha256:5a684be454e86e22888e3c26c930c84684394f5d1851dc6f5c296e90a2318be5

Observation 4a7da1d2-2edf-481c-8d62-2e0872c13e61 · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.475326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.475326Z digest=sha256:0a2356c1c02ce190c2f7cfda1976383dd377a4aa2ba0efb273ee0a29078da165

Observation f09d010b-6dda-448c-8409-984d68cec4cd · outbound

This paper cites Agentic AI framework for autonomous and self-managing cloud services,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Agentic AI framework for autonomous and self-managing cloud services,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.614360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.614360Z digest=sha256:f5a4ab76bc732998a0a6c90cb06933cae0c74c05dc8b204baaa1db1b456d5dae

Observation 5189926f-6de6-4c5d-9a3d-c72bf80fb01e · outbound

This paper cites On measuring faithfulness or self- consistency of natural language explanations,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models On measuring faithfulness or self- consistency of natural language explanations,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.750526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.750526Z digest=sha256:71ed23da184a04928dd4697a1b5992c78ac7815de4575055bf2f21a390784d60

Observation 76295650-45be-4ec4-a378-a6e94c3212b9 · outbound

This paper cites ROSCOE: A suite of metrics for scoring step-by- step reasoning,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models ROSCOE: A suite of metrics for scoring step-by- step reasoning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.903167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.903167Z digest=sha256:30dfdd01a0b5f730cf350cfae426848fcda3e142bb1128bf40defcf98de58a74

Observation fe00edda-5930-4fe9-b34c-524fc2528c07 · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with disentangled attention,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models DeBERTa: Decoding-enhanced BERT with disentangled attention,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.035866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.035866Z digest=sha256:24af2e882e45a401a42dc9ad31c4fad4d77314737d7bfb351b7c5a9c4ea34c65

Observation 2ad813b1-e3bd-45cc-8dcb-7fa791cc1138 · outbound

This paper cites ReCEval: Evaluating reasoning chains via correctness and informativeness,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models ReCEval: Evaluating reasoning chains via correctness and informativeness,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.122552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.122552Z digest=sha256:c5a458bacabe6ea6e4514402fd9fe0f52496bf59015d053bea4c759e40ac5c85

Observation 9c07f871-3937-45d9-8c67-17757a48e9ba · outbound

This paper cites On authentication schemes using polynomials over non commutative rings,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models On authentication schemes using polynomials over non commutative rings,

Reference 14

Resolution
malformed identifier
doi_truncated, observed 2026-08-01T02:46:11.242397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-01T02:41:52.304612Z digest=sha256:4d38540c3122b8ad5a1b197327418f35baf4e5ab14f6c9c97e5feff2f20de413

Observation b7716dde-f6c6-4e4e-966e-8158fbc0d3c5 · outbound

This paper cites A benchmark for verifiers of reasoning chains,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models A benchmark for verifiers of reasoning chains,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.458185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.458185Z digest=sha256:4ca461860b890bc8407f974c6ac04e6ff19e82a4a788f6c5924b3c124d820aa5

Observation 813cb8fc-300f-4cf8-9b68-dcd3c6aa2e23 · outbound

This paper cites Direct evaluation of chain-of-thought in multi-hop reasoning with knowledge graphs,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Direct evaluation of chain-of-thought in multi-hop reasoning with knowledge graphs,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.612889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.612889Z digest=sha256:49556dd3a65bb0bd94ecb41401df7f609139ef20ad3e68766c05370eee9f33dc

Observation 3d78d6be-d91a-4c28-8e1b-988c2ea84ad2 · outbound

This paper cites Self-adjointness of semi-relativistic Pauli-Fierz Hamiltonian.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Self-adjointness of semi-relativistic Pauli-Fierz Hamiltonian

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.825454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.825454Z digest=sha256:80426023f471291f26e966b2d805d9a852ea149f1efcbf442489b03a1988c357

Observation cadcf317-82bb-467d-a790-e4c1c12978f8 · outbound

This paper cites Making reasoning matter: Measuring and improving faithfulness of chain-of-thought rea- soning,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Making reasoning matter: Measuring and improving faithfulness of chain-of-thought rea- soning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.941251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.941251Z digest=sha256:be12c76236b378c70f0aa527930d7be603fd2a48372b51c840602297c42b303d

Observation 3d9b508b-ab56-488a-95ac-76cd8a60a83d · outbound

This paper cites Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.089640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.089640Z digest=sha256:497bd4bef1740fbe1c11fd5c785af4b2ce20c2a1110db5b4d5d18e9d2a52421e

Observation 44a42c15-a6d0-4a29-acbd-39bd6d3e271a · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.202479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.202479Z digest=sha256:0fa5cb2f29d35d89398a1136696281ca2d7fd1826ea2e34a32c76a92fa41e696

Observation 6339b64f-3029-4e7d-951b-f1526013122c · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Reasoning Models Don't Always Say What They Think

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.351709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.351709Z digest=sha256:79c8eff43875427e8552e7e3068926dc07aced5e4ce52143f882fff2c6d5fb04

Observation e7d0c5bc-7b1a-4d33-8a3a-a2e66e218315 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.515878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.515878Z digest=sha256:6a5b6f72e868050f3bfc854000b6032c09b179d378a5fde4ada217909bc59446

Observation e680009d-7091-4440-93d3-8ef717e4fa51 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.681811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.681811Z digest=sha256:6109379eb80739fbcc6367c6224c985e796df5db0c49adf2bd4750aa95c4ab14

Observation 9e6c368a-13f9-498d-8a55-51c82260292f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.849176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.849176Z digest=sha256:9539d5940c089e4a4a31d32f793bc75a62d07035e3d8b0c1fe9f67b475ce9b7a

Pith citing papers

No inbound Pith citation observations are available.