Pith. sign in

Paper Citation Record · LEDGER

Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2502.12964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.12964 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:43:21.890291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:06:29.134779Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7fa80d3f-1441-4d8c-a625-80d56c7839ee · inbound

The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems cites this paper.

The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:21.890291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:21.890291Z digest=sha256:2d526619cf60f1a60273d26adedd23cafc0cf83b2342a421daef164105582cda

Observation 29f924ff-d768-4776-ab16-9827381fb2ac · inbound

A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations cites this paper.

A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:01:15.033088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:01:15.033088Z digest=sha256:dd2d2a546f71bded9d217f001c2c82d636c2a0c1c77eb9990880503e2a9a6c64

Observation 947c7c23-9bee-4ee1-a2fe-9b64140a0c9b · inbound

Tractable Asymmetric Verification for Large Language Models via Deterministic Replicability cites this paper.

Tractable Asymmetric Verification for Large Language Models via Deterministic Replicability Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:42.808795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:42.808795Z digest=sha256:e8299d0621ecefba3e719a25e0be8962d82043797af55c565cb94dcfda83950e

Observation fc089923-f062-4713-9b86-731138440fd7 · inbound

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting cites this paper.

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:59.437743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:59.437743Z digest=sha256:332fbb38efca5c914f2aa01d36ebe8477a11d58fc9c90c8ac45b6a67bcba3170

Observation fc80d791-d6d7-4c84-b149-2638c565961d · inbound

Ensemble-Based Uncertainty Estimation for Code Correctness Estimation cites this paper.

Ensemble-Based Uncertainty Estimation for Code Correctness Estimation Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:53:14.123949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T22:52:58.524934Z digest=sha256:7fefcff153cc7c4c8a8debab056e239004b232bdabe80ab764b0030b30e27ce5

Observation b3f1125c-65a0-49e9-8658-76ef33eaf6e3 · inbound

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence cites this paper.

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:53:11.900607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T19:48:13.133733Z digest=sha256:c53ecce584726f00c048b9dddd0f0c054630c1597d7f72efc6ccb44f912ea4b6

Observation 88621839-a4cc-4316-bb1f-e87c28b67e2d · inbound

Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification cites this paper.

Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.595642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:29:24.974157Z digest=sha256:e1b4900c64d86f1177908fcd2bb82710cf803c6c4f943513b988cdabcac67ea1

Observation 95aa4e56-a116-4659-82b2-cb422143cab2 · inbound

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? cites this paper.

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:24.409363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:18:36.854164Z digest=sha256:6f8da214635a03fb99c12ddf510f898216abf4ab392323fac3632b0eb877c5dd

Observation 7b4e0c60-d32f-44c6-bb31-7e93d808b41f · inbound

Quantifying Faithful Confidence Expression in Large Reasoning Models cites this paper.

Quantifying Faithful Confidence Expression in Large Reasoning Models Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:29.137832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:24:38.335417Z digest=sha256:460cec0c090b42d90d2a9bce583ebcd50a9dd3449d3a7f9f87bbcd39adc57561

Observation da2a5e8d-e1b8-4e57-aad9-11080228da7d · inbound

Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States cites this paper.

Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T05:45:08.896651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:45:08.896651Z digest=sha256:3e5a0946cfcf8baac4137be23601fb6f3625b43c1283a0e86e79926c2dc4329a

Observation 93b5537b-a67a-4993-a461-80ab3964f3ee · inbound

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models cites this paper.

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T09:19:46.058147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:19:46.058147Z digest=sha256:12eea461c69df46148d8f9fc6c93ed41c07e608142279734d5bd9e237f53a65d

Observation aaa44321-b620-45fd-8ab2-737f3cce57bc · inbound

$\Sigma$-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems cites this paper.

$\Sigma$-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T22:14:36.622744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:14:36.622744Z digest=sha256:471b659f87047e8e174b2e9b3f66530f0b4ed95be8e600a767e960c8dcdbb029