Pith. sign in

Paper Citation Record · LEDGER

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy

As of 17 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2501.11721.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.11721 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:00:32.747694Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:36:55.345947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:36:55.445243Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9af8672c-e08a-41b2-91b7-49dbd3c8b608 · outbound

This paper cites Introducing gemini: Google's multimodal ai model.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Introducing gemini: Google's multimodal ai model

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.129677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.642092Z digest=sha256:7b7a01ef1dc9709f578c1aa6798de4c49951989f85c26a11cc90bd56dc7cea98

Observation 604d3217-b67a-4864-83c9-53b324ee341a · outbound

This paper cites Introducing claude: Anthropic's ai assistant.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Introducing claude: Anthropic's ai assistant

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.114815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.647171Z digest=sha256:8775e01510480109781b2b8ccc59ab09fc55f7dd48735f7f911a22d1aabd2ae1

Observation 6c624207-06d6-40f1-b487-e8e488aa30e5 · outbound

This paper cites Explainability in ai: A survey.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Explainability in ai: A survey

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.100711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.651632Z digest=sha256:f562e97e56c76868ff62faeaf2de1d8c182712ef7e487c4309009adba87a7a3f

Observation 4745ce75-e64c-4df0-bb58-1a58cf4b7fb3 · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 610--623, 2021.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 610--623, 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.085857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.656481Z digest=sha256:0d21a16a923d8b1a94d3279bcb115acb03c3cb557c5f3069c27ebe1fef821697

Observation 38fa1fa5-51b2-4ff2-9a2c-49782fe47314 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy On the Opportunities and Risks of Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.660995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.660995Z digest=sha256:afd4dd48e4870d5f7d82537233229f4a439e86b55fa2e4b190a596cff7f9c680

Observation 3f885034-6a26-45aa-a436-1f61cc575181 · outbound

This paper cites Language models are few-shot learners.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Language models are few-shot learners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.071485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.665972Z digest=sha256:3d0065decac24dff0e9f4a16cc9f97435a4111943dd2a2e1f6df7d01c7f4d10e

Observation e7c0ecb4-fe86-414c-949e-8c707a9f56c5 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.057548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.671018Z digest=sha256:a16d0340b74cb08830e78a762a1b12f7430565281ba6d7ae73d4c5ce600c0762

Observation 05283e8a-614a-4685-a41b-4bad9c1d6467 · outbound

This paper cites What if $\phi^4$ theory in 4 dimensions is non-trivial in the continuum?.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy What if $\phi^4$ theory in 4 dimensions is non-trivial in the continuum?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.675479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.675479Z digest=sha256:644fe0da536a989242fecbe47cae6769ae9fff1afb5e64cf99236f789ba2406e

Observation c537fa1b-b37e-464c-9e4f-0d61bb4d2692 · outbound

This paper cites Prediction of solar wind speed by applying convolutional neural network to potential field source surface (PFSS) magnetograms.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Prediction of solar wind speed by applying convolutional neural network to potential field source surface (PFSS) magnetograms

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T18:00:32.933485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.680162Z digest=sha256:116b493e52339a0ba39eda45619135534c22884f844c6d67f154b5a6a01514e4

Observation a725d6df-df7f-4458-add2-f214bc0435c1 · outbound

This paper cites Structure Factors for Hot Neutron Matter from Ab Initio Lattice Simulations with High-Fidelity Chiral Interactions.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Structure Factors for Hot Neutron Matter from Ab Initio Lattice Simulations with High-Fidelity Chiral Interactions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.684957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.684957Z digest=sha256:b279a7044baa543ba5ec53687e9d23bbf4faa525afbfd3ccac40457d818decfd

Observation d4f9bfda-7183-431a-8d62-7ff38e768035 · outbound

This paper cites XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:00:32.894951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.689053Z digest=sha256:9dd5692789db7776063839942f8616198e11831777bc0fcaebaed8a2eb4a1f23

Observation e7c513f4-769a-460a-86d0-0771b8c722e9 · outbound

This paper cites RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.693612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.693612Z digest=sha256:006101ac74b38d37a682e18beb0757509bf00283bef575bb6a7429c9c2de480e

Observation b06cca36-3cec-43ab-ba6d-052c9c216d86 · outbound

This paper cites Gpt-4 technical report.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Gpt-4 technical report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.697624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.697624Z digest=sha256:a198cb965118fac12560be0eecffa7aae2296ae2779b90612bafb58dbe6a1d19

Observation 29a10000-aa0b-4db0-ba6b-4d80c3c78123 · outbound

This paper cites Squad: 100,000+ questions for machine comprehension of text.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Squad: 100,000+ questions for machine comprehension of text

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.034560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.701626Z digest=sha256:7f6bea147d8d9849bd98b4113c9fadad8cb4ff908eb5d2769503be8e6d8f67c9

Observation 94671660-d958-470d-8465-82b5acbbfaf9 · outbound

This paper cites A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:00:32.852516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.705662Z digest=sha256:f59a9bf9f336b24d699035b56bfe1bdf599b2de3419457af12f7f7e48fba35d4

Observation 152ce31f-476b-490f-b0de-7a40c0496714 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy LLaMA: Open and Efficient Foundation Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.709790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.709790Z digest=sha256:e27b553549b5191d8ce1413d9890221fa739a66497479db956dfaed9f5554a5b

Observation 8cfcdcf8-cdd4-4a73-9b45-91dcf06e0348 · outbound

This paper cites On the fluid slip along a solid surface.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy On the fluid slip along a solid surface

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.714286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.714286Z digest=sha256:4b9426576320f7c907586b42231c80ec85af41a750f31b570264488da1935e4d

Observation 3250ac9a-1429-4ea8-a333-a1313e7004d2 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.718756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.718756Z digest=sha256:5182913a62fd7cdf2171277f8631ae80b58a03b5e389ce80c102f9c4ac70b447

Observation 14b953e3-a8be-47dc-998b-8f641d11d0b5 · outbound

This paper cites Teach me to explain: A review of machine learning interpretability through explanations.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Teach me to explain: A review of machine learning interpretability through explanations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.020554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.723277Z digest=sha256:813f0e3e8747056c300d004909e872daae8d5328b3765f108ea36aa47906fd90

Observation 9cf92435-8d6a-446c-b62c-9041449ef5c9 · outbound

This paper cites Language Models can Evaluate Themselves via Probability Discrepancy.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Language Models can Evaluate Themselves via Probability Discrepancy

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.727925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.727925Z digest=sha256:1db0d8db51785fd9786a9c3020ffa3779a5c58d1e480000d9e89b7095597705f

Observation aac56f11-00f1-4206-b11f-996977b5673c · outbound

This paper cites Evaluating paraphrase sensitivity in large language models.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Evaluating paraphrase sensitivity in large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.006280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.732816Z digest=sha256:50c7060cd9533aabd99bf5fa84bbc5655076f995e625000d65b554e98c140d5c

Observation beac91c4-4df4-43ca-92f2-6504e68182b8 · outbound

This paper cites @esa (Ref.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy @esa (Ref

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.737661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.737661Z digest=sha256:d64c0448685f60745813f13cb76acc1d744662c210c037c0d1983d84758eea16

Observation ad4ab37a-ed28-4d22-bcd9-ffb4d37d84f9 · outbound

This paper cites an unresolved cited work.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.743000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.743000Z digest=sha256:ed8a57cd3a84ac3aec1e0b74607f106b48e15327b93ee0e0daca6d50c11b7a1b

Observation bded3b9a-c992-42de-ad0a-58f71e7b9702 · outbound

This paper cites an unresolved cited work.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.747694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.747694Z digest=sha256:b53f3abe5b64e8a188743bb82cbc2797b34b081b3476fb689d5f2bcc1a77b7de

Pith citing papers

Observation 4a41e5e2-253d-4009-bdaa-2fe060155d38 · inbound

Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI cites this paper.

Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:36:55.451186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:36:55.345947Z digest=sha256:10a15b02ee03099796ff538cbd6dfe3c59fb2f7949eb61751746c3aeb4e29051