Pith. sign in

Paper Citation Record · LEDGER

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

As of 23 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.26102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26102 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:41:53.849176Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 16446460-0a45-4338-b6de-6041d2596f4d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:50.700365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:50.700365Z digest=sha256:a0464a18cebfe4e73852cbdba8e9be55a6d843000b404c232537b0467b5a0d2b

Observation 2d31b836-828b-4711-9a36-b6d256bb1cc1 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Self-consistency improves chain of thought reasoning in language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:50.817858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:50.817858Z digest=sha256:f76a1a2d8386e6477e485364e8b85412f18af4381c8870cd621417e17c906e3a

Observation 2ef6c58b-d963-4bf6-b26a-90781f73acfb · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of- thought prompting,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Language models don’t always say what they think: Unfaithful explanations in chain-of- thought prompting,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:50.893111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:50.893111Z digest=sha256:6a27e4873aa568b21e58e7cb86c28ccedc8a2a31f6afd4f361fc207681c0871d

Observation 06bc27f3-5fd7-42c4-ba11-503531c8f602 · outbound

This paper cites Federated generative intelligence for explainable and autonomous cyber defence in critical infrastructures,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Federated generative intelligence for explainable and autonomous cyber defence in critical infrastructures,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:50.975169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:50.975169Z digest=sha256:89bbd93b36af198b2eff75fd3eab64d275c4bebf28b952b18765ee872d1c7d5e

Observation a1ea8b0c-5f36-40ed-b395-37e7d3f453b4 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.171583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.171583Z digest=sha256:6dbb418a77504e9f8ebd73ae31314e8f581766932b81584004c51f5eeb86e92e

Observation 81e50c51-c0d7-49a5-83af-0160a179ab67 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.342440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.342440Z digest=sha256:415059a43a979c0b9d5617166d4b794c7d73571b53b8e73210c396f78e3dcd8f

Observation cb9e2ce9-471c-490d-ac5d-da24b6daa20c · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Measuring mathematical problem solving with the MATH dataset,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.449290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.449290Z digest=sha256:0e78b508eeedd24a4305b68d67346442dcc0aeb1054993b4ffec81d14d2f622b

Observation 4a7da1d2-2edf-481c-8d62-2e0872c13e61 · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.475326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.475326Z digest=sha256:95f13547d9da72c431f66a728595f61b1e4095367dc0f3cf2d8b6f52a8d36c36

Observation f09d010b-6dda-448c-8409-984d68cec4cd · outbound

This paper cites Agentic AI framework for autonomous and self-managing cloud services,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Agentic AI framework for autonomous and self-managing cloud services,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.614360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.614360Z digest=sha256:0fc8f2a59f9a1301a7f3ef597cf2fbb23827ff90b95494d1b09a2334b497fe0e

Observation 5189926f-6de6-4c5d-9a3d-c72bf80fb01e · outbound

This paper cites On measuring faithfulness or self- consistency of natural language explanations,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models On measuring faithfulness or self- consistency of natural language explanations,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.750526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.750526Z digest=sha256:1e6dd55459fffc8a09cd77ab69d33a19dff5dd06a5313bf584338a6f9280fee7

Observation 76295650-45be-4ec4-a378-a6e94c3212b9 · outbound

This paper cites ROSCOE: A suite of metrics for scoring step-by- step reasoning,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models ROSCOE: A suite of metrics for scoring step-by- step reasoning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.903167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.903167Z digest=sha256:1a11bc990703cfb87367d65eb991f4fc354ad918e7dfddc1d86973f64a1f3299

Observation fe00edda-5930-4fe9-b34c-524fc2528c07 · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with disentangled attention,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models DeBERTa: Decoding-enhanced BERT with disentangled attention,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.035866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.035866Z digest=sha256:6ce1afdc9b054c2eec44737c53a8425be5a854c5257dd84b369c7b04844cbb92

Observation 2ad813b1-e3bd-45cc-8dcb-7fa791cc1138 · outbound

This paper cites ReCEval: Evaluating reasoning chains via correctness and informativeness,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models ReCEval: Evaluating reasoning chains via correctness and informativeness,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.122552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.122552Z digest=sha256:9bb7cca9b7d960398b5c2adbe79fb2e82bbb2471093b6c4b5c13c998a0c483d7

Observation 9c07f871-3937-45d9-8c67-17757a48e9ba · outbound

This paper cites On authentication schemes using polynomials over non commutative rings,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models On authentication schemes using polynomials over non commutative rings,

Reference 14

Resolution
malformed identifier
doi_truncated, observed 2026-08-01T02:46:11.242397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-01T02:41:52.304612Z digest=sha256:dd64408f769029d42404bad8076c96e8d575150668dda932b1d245eeab42b413

Observation b7716dde-f6c6-4e4e-966e-8158fbc0d3c5 · outbound

This paper cites A benchmark for verifiers of reasoning chains,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models A benchmark for verifiers of reasoning chains,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.458185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.458185Z digest=sha256:e4998442241008a0d7111eeac04eba33957a06529d32a8b60b87c1eecabb41f8

Observation 813cb8fc-300f-4cf8-9b68-dcd3c6aa2e23 · outbound

This paper cites Direct evaluation of chain-of-thought in multi-hop reasoning with knowledge graphs,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Direct evaluation of chain-of-thought in multi-hop reasoning with knowledge graphs,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.612889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.612889Z digest=sha256:6e7661f949a49abe4012959c6a1cb0416538db2502ace1095bd1d536ecfa517e

Observation 3d78d6be-d91a-4c28-8e1b-988c2ea84ad2 · outbound

This paper cites Self-adjointness of semi-relativistic Pauli-Fierz Hamiltonian.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Self-adjointness of semi-relativistic Pauli-Fierz Hamiltonian

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.825454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.825454Z digest=sha256:8a33fb315b7edd611b4b1156abe118a51cd50df17770e5347b459cb38ebb839d

Observation cadcf317-82bb-467d-a790-e4c1c12978f8 · outbound

This paper cites Making reasoning matter: Measuring and improving faithfulness of chain-of-thought rea- soning,.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Making reasoning matter: Measuring and improving faithfulness of chain-of-thought rea- soning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:52.941251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:52.941251Z digest=sha256:635b8a65aa135f02952c6b3d333c13eecf01fb30e36055a6ce737d14ce98080a

Observation 3d9b508b-ab56-488a-95ac-76cd8a60a83d · outbound

This paper cites Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.089640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.089640Z digest=sha256:72cae3eebf446fc88db2a7058e79fce9918aacc2327588e80483d9614027abb1

Observation 44a42c15-a6d0-4a29-acbd-39bd6d3e271a · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.202479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.202479Z digest=sha256:3efe28e7f8f7005d1c839a782c6cff5d2640564f8a15bea50f814cfc8406d07c

Observation 6339b64f-3029-4e7d-951b-f1526013122c · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Reasoning Models Don't Always Say What They Think

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.351709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.351709Z digest=sha256:b70298352484bf10530655bd32d190fc23a8cec6268ba2287867de88c7f01f52

Observation e7d0c5bc-7b1a-4d33-8a3a-a2e66e218315 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.515878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.515878Z digest=sha256:febe7f34eba5be027d7df328c3476a260e0d37de80b81fad96472d197500dd31

Observation e680009d-7091-4440-93d3-8ef717e4fa51 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.681811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.681811Z digest=sha256:245a347d8d5fd1ff0d6e01127608f9bcac797574b4647aa8c160d10b09fcfcb4

Observation 9e6c368a-13f9-498d-8a55-51c82260292f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.849176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.849176Z digest=sha256:d5d218d7e3dade1580d0245b02e735576ded37a10da2f6165a1a772a94253a54

Pith citing papers

No inbound Pith citation observations are available.