Pith. sign in

Paper Citation Record · LEDGER

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy

As of 21 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2501.11721.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.11721 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:00:32.747694Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:36:55.345947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:36:55.445243Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9af8672c-e08a-41b2-91b7-49dbd3c8b608 · outbound

This paper cites Introducing gemini: Google's multimodal ai model.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Introducing gemini: Google's multimodal ai model

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.129677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.642092Z digest=sha256:6f3db2e0acac935cc4662761b373a6d99c9c55df3359ce54a09c37e9347ec4b0

Observation 604d3217-b67a-4864-83c9-53b324ee341a · outbound

This paper cites Introducing claude: Anthropic's ai assistant.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Introducing claude: Anthropic's ai assistant

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.114815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.647171Z digest=sha256:4d94c7106271b3e424cd8e2d7ae47c70daedd4993dc9c0bf6df9563df586bd64

Observation 6c624207-06d6-40f1-b487-e8e488aa30e5 · outbound

This paper cites Explainability in ai: A survey.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Explainability in ai: A survey

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.100711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.651632Z digest=sha256:e25264afe5a45c5bce2a0e0d97d9a0e39dff72c2e99daed71aa3d707411a2d3c

Observation 4745ce75-e64c-4df0-bb58-1a58cf4b7fb3 · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 610--623, 2021.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 610--623, 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.085857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.656481Z digest=sha256:2ede142d3983840a45de8cc0ebe470e100b264dc62f56bd05b8f162ac3dcb9b0

Observation 38fa1fa5-51b2-4ff2-9a2c-49782fe47314 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy On the Opportunities and Risks of Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.660995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.660995Z digest=sha256:ee31cd15f3f541ab55ae70426ef74d36bfee0f324cb2b8dcef60699c2c766175

Observation 3f885034-6a26-45aa-a436-1f61cc575181 · outbound

This paper cites Language models are few-shot learners.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Language models are few-shot learners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.071485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.665972Z digest=sha256:afeca551c4a83ab54de6a29e9b6c37517ff3ddc6a0437d5370a0dda9535a85b6

Observation e7c0ecb4-fe86-414c-949e-8c707a9f56c5 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.057548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.671018Z digest=sha256:343db2da5fddfb98b2250338199b46b036f27a287e34b77d6bfd7813c4f2e69d

Observation 05283e8a-614a-4685-a41b-4bad9c1d6467 · outbound

This paper cites What if $\phi^4$ theory in 4 dimensions is non-trivial in the continuum?.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy What if $\phi^4$ theory in 4 dimensions is non-trivial in the continuum?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.675479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.675479Z digest=sha256:2f505197a47e6abbce2640c6f9bd80c445fda2b0e42bf94089728227ca23a8cf

Observation c537fa1b-b37e-464c-9e4f-0d61bb4d2692 · outbound

This paper cites Prediction of solar wind speed by applying convolutional neural network to potential field source surface (PFSS) magnetograms.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Prediction of solar wind speed by applying convolutional neural network to potential field source surface (PFSS) magnetograms

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T18:00:32.933485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.680162Z digest=sha256:9b6302624de76c419b3b0b73f144e38d15300b20117b0846cc7044d9c46f1804

Observation a725d6df-df7f-4458-add2-f214bc0435c1 · outbound

This paper cites Structure Factors for Hot Neutron Matter from Ab Initio Lattice Simulations with High-Fidelity Chiral Interactions.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Structure Factors for Hot Neutron Matter from Ab Initio Lattice Simulations with High-Fidelity Chiral Interactions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.684957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.684957Z digest=sha256:73e64f0328d0dcf1698b877f34e6867f708b5d8b64356edd549d8b18fe09ccc1

Observation d4f9bfda-7183-431a-8d62-7ff38e768035 · outbound

This paper cites XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:00:32.894951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.689053Z digest=sha256:b6ce23f9d0c2ec7f61fba2bdd01624f963115dbe2948c4155eaee9d1bd3b112f

Observation e7c513f4-769a-460a-86d0-0771b8c722e9 · outbound

This paper cites RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.693612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.693612Z digest=sha256:e98d18da792850d4d9ed47225c122975849c72041667e3d8b5674b67f8350e8a

Observation b06cca36-3cec-43ab-ba6d-052c9c216d86 · outbound

This paper cites Gpt-4 technical report.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Gpt-4 technical report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.697624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.697624Z digest=sha256:353bbe972ea36c1c824c1d88043b12a16219ef3be7b478af9b4a4a38a0289b68

Observation 29a10000-aa0b-4db0-ba6b-4d80c3c78123 · outbound

This paper cites Squad: 100,000+ questions for machine comprehension of text.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Squad: 100,000+ questions for machine comprehension of text

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.034560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.701626Z digest=sha256:7f9de89b6312c291ef1a952d5d6f7eb5ce2368b97939e418ff88533a0b2cf5f2

Observation 94671660-d958-470d-8465-82b5acbbfaf9 · outbound

This paper cites A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:00:32.852516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.705662Z digest=sha256:e98962ed2405b70093f7b0dd2d85a0852b96443216778942704bccd631d3df3f

Observation 152ce31f-476b-490f-b0de-7a40c0496714 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy LLaMA: Open and Efficient Foundation Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.709790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.709790Z digest=sha256:6cfd352059725678644ba5d2314c5040d6294fbc64796baf68b3833423aba16b

Observation 8cfcdcf8-cdd4-4a73-9b45-91dcf06e0348 · outbound

This paper cites On the fluid slip along a solid surface.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy On the fluid slip along a solid surface

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.714286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.714286Z digest=sha256:cab891767dec01795e4db7a070137d12f132399314df0a001f97951ff4e7276f

Observation 3250ac9a-1429-4ea8-a333-a1313e7004d2 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.718756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.718756Z digest=sha256:56b4854cdb64dc0a75576234d9a0de186cc084672867e60a055b5f01cd0d9416

Observation 14b953e3-a8be-47dc-998b-8f641d11d0b5 · outbound

This paper cites Teach me to explain: A review of machine learning interpretability through explanations.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Teach me to explain: A review of machine learning interpretability through explanations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.020554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.723277Z digest=sha256:a580d44f3675092d3a250f58bcb7a9377d291ac5bfe615e5882c348385f46a77

Observation 9cf92435-8d6a-446c-b62c-9041449ef5c9 · outbound

This paper cites Language Models can Evaluate Themselves via Probability Discrepancy.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Language Models can Evaluate Themselves via Probability Discrepancy

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.727925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.727925Z digest=sha256:3c6e7d0f8365eb2b374af52d4c03a9c8d3d5a71e93710cd3b3a7a8ade1302f37

Observation aac56f11-00f1-4206-b11f-996977b5673c · outbound

This paper cites Evaluating paraphrase sensitivity in large language models.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Evaluating paraphrase sensitivity in large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.006280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.732816Z digest=sha256:63928752716ef78021aab006f8248e30db098544561e1e1faca4f6330dccbf7a

Observation beac91c4-4df4-43ca-92f2-6504e68182b8 · outbound

This paper cites @esa (Ref.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy @esa (Ref

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.737661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.737661Z digest=sha256:297d18747308a89dbc9caa9ea11c1f1285ad50b3058a5053db8854ef632d4322

Observation ad4ab37a-ed28-4d22-bcd9-ffb4d37d84f9 · outbound

This paper cites an unresolved cited work.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.743000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.743000Z digest=sha256:7981babd0615b3c1c6ba9c2174683dad696eeea0e8d229f3a1d6c3cd25bd8746

Observation bded3b9a-c992-42de-ad0a-58f71e7b9702 · outbound

This paper cites an unresolved cited work.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.747694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.747694Z digest=sha256:2e9105ffd112d973f035128f6f151608cc702b0ddf92f7c29c0fdc657221dab2

Pith citing papers

Observation 4a41e5e2-253d-4009-bdaa-2fe060155d38 · inbound

Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI cites this paper.

Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:36:55.451186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:36:55.345947Z digest=sha256:68f005abcf9ed6143a1b26459e27389b61237c85db16f451945475eb55b125f4