Pith. sign in

Paper Citation Record · LEDGER

Potemkin Understanding in Large Language Models

As of 10 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 5 inbound Pith citation observations for arXiv:2506.21521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21521 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:30:32.913920Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:20:14.933677Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T04:12:02.502599Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d63f77aa-6ef6-49ee-a283-79eb99afab8a · outbound

This paper cites write newline.

Potemkin Understanding in Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.435106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.435106Z digest=sha256:4dbf96b72b863a0cda6260785bb544caa8546d0cd05dcfb03dd698c81760d25a

Observation f1e6da37-5ba4-4f69-8fa9-0c59c9d581ff · outbound

This paper cites GPT-4 Can't Reason.

Potemkin Understanding in Large Language Models GPT-4 Can't Reason

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.514523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.514523Z digest=sha256:55f02c7f71614f62d605c530794abf256c7e4a4cc36d8a1cbd98520f385d46eb

Observation f2924307-bf4f-46d0-8620-46c553022461 · outbound

This paper cites Synthetic and Natural Noise Both Break Neural Machine Translation.

Potemkin Understanding in Large Language Models Synthetic and Natural Noise Both Break Neural Machine Translation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.626627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.626627Z digest=sha256:ca83701417f97148cb090bdfcd77a542ca9e7127d932a9cd2a826cf555931bd3

Observation da117aa1-2c46-4bef-ab57-ea86758a63f7 · outbound

This paper cites The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A".

Potemkin Understanding in Large Language Models The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.715765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.715765Z digest=sha256:d3fe1fba2dd0e8d05ab6fd0bc006aa659e0165fdf1ebf098df9f6e08562c4cab

Observation 2aadb985-3932-4443-9b5c-20ebf50d25b2 · outbound

This paper cites What Will it Take to Fix Benchmarking in Natural Language Understanding?.

Potemkin Understanding in Large Language Models What Will it Take to Fix Benchmarking in Natural Language Understanding?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.809315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.809315Z digest=sha256:da9545c33a5d2bd8035de0e51d969f9faac703d239576bc8040bf3a644c081f5

Observation acad158d-780c-4e01-9d25-2f7b9e40fe4f · outbound

This paper cites A large annotated corpus for learning natural language inference.

Potemkin Understanding in Large Language Models A large annotated corpus for learning natural language inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.890401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.890401Z digest=sha256:deff104dddff131157f5b80b728ce9c12967aef463df6fff46c846d3e51f02eb

Observation 2b23d7a8-3d0b-4c30-9521-1a1bfc1a288f · outbound

This paper cites T., Li, Y., Lundberg, S., et al.

Potemkin Understanding in Large Language Models T., Li, Y., Lundberg, S., et al

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:37.175863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:28.992268Z digest=sha256:192e23764767ef5ba8fd43055fc63ed702a25f52e3329e0f1a344514e896a3fb

Observation 3c13f625-6a91-4d56-bd70-58eeb0d90de1 · outbound

This paper cites With Little Power Comes Great Responsibility.

Potemkin Understanding in Large Language Models With Little Power Comes Great Responsibility

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:30:33.752650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.088514Z digest=sha256:d14831cb206e1ffc41388ae846e08fb87a214cbaa5f620afa35ee99ed7e8a52f

Observation 570370e0-e76e-4430-9284-c1068694b47e · outbound

This paper cites ChatBench: From Static Benchmarks to Human-AI Evaluation.

Potemkin Understanding in Large Language Models ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:29.181666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:29.181666Z digest=sha256:530428d6c35acdfc9aafbc1240a7a34ded7f9aa557648420841ebea16d2a1ac4

Observation ebe472ad-7448-4859-924d-4375114536cb · outbound

This paper cites N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J.

Potemkin Understanding in Large Language Models N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:37.015281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.278520Z digest=sha256:a13a6f918a5cd1c20aa0f2bc146c74a349aa66fb35a84e12293962652ff0a1e9

Observation f0562646-d1d9-4fd3-b47a-e047e37a74de · outbound

This paper cites an unresolved cited work.

Potemkin Understanding in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:30:36.872633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.378079Z digest=sha256:347644d67c818acbd4a87efe79ede43356da3f2f54cdd7aff0b52b902afe828b

Observation c9b3191d-80f6-4b06-9503-beaa37bb8d3c · outbound

This paper cites and Etzioni, O.

Potemkin Understanding in Large Language Models and Etzioni, O

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.720223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.481957Z digest=sha256:4732c8007c7573af18502244c8dfa59a206664749f8e662c12ad7041c0073fdc

Observation 1ff9d6c6-be20-479b-8761-9cabe6a83991 · outbound

This paper cites Recognizing textual entailment: Rational, evaluation and approaches--erratum.

Potemkin Understanding in Large Language Models Recognizing textual entailment: Rational, evaluation and approaches--erratum

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.600044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.559562Z digest=sha256:75e2f4f384111a4199434a54e0467fbba4c82cacf6fef87c6d941d1e4caaf6c2

Observation 4cb16b18-fd89-4569-94ca-464bf44350d6 · outbound

This paper cites Testing ai on language comprehension tasks reveals insensitivity to underlying meaning.

Potemkin Understanding in Large Language Models Testing ai on language comprehension tasks reveals insensitivity to underlying meaning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.453392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.621688Z digest=sha256:5abbd3926cfaf201f8822a1d5d21be5f5ea63ce57c6b1dba6d1143bf66391457

Observation a0ea739a-535a-4172-8f46-2181181d8d19 · outbound

This paper cites and Meurers, D.

Potemkin Understanding in Large Language Models and Meurers, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.269650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.666182Z digest=sha256:0eb26f5b7ab7c67c461ae5cee3be3d4a9eb94800cdcbbb8a00cf375319c176fb

Observation 6077b6e3-a40e-4ede-a49f-cc31384401c8 · outbound

This paper cites Measuring and improving consistency in pretrained language models.

Potemkin Understanding in Large Language Models Measuring and improving consistency in pretrained language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.131723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.750083Z digest=sha256:ed606ba3d1948fa58b6d9fade9d1063bd44f593f2ba7402aa82bdf9d1b7bff5f

Observation 795b1987-0c86-46cd-9f60-e2fbbe755337 · outbound

This paper cites Evaluating superhuman models with consistency checks.

Potemkin Understanding in Large Language Models Evaluating superhuman models with consistency checks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.945909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.825513Z digest=sha256:e89b3944a5d33b4dae4630cbe734326af618c97930cb7bae8e5d91a92dc479fb

Observation df50e5bb-3e3c-4b1c-88ec-df9aaa6f1fe7 · outbound

This paper cites W., Wallach, H., Iii, H.

Potemkin Understanding in Large Language Models W., Wallach, H., Iii, H

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.807415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.888225Z digest=sha256:4531b2fe411d2877c958caf9e70aac028e8f90370d05579ca6d0470475f78aa2

Observation 4ef75983-0583-4e5c-bdb5-f556c86df559 · outbound

This paper cites an unresolved cited work.

Potemkin Understanding in Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:30:35.689875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.949572Z digest=sha256:365f5e4414e8529b82b747b6fd7098ea9d9ad23b513b523d280adfba5155daa8

Observation 9d7a95cd-1d8f-4e3a-ad3e-c7466b06c8a0 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Potemkin Understanding in Large Language Models Transformer Feed-Forward Layers Are Key-Value Memories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.028454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.028454Z digest=sha256:78751ae464ee9d747a0d44915fa16f96f9d2c55aa620a287337827642894fa80

Observation ad7f7eb0-fd01-4e0f-82f0-18976d86aa83 · outbound

This paper cites Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space.

Potemkin Understanding in Large Language Models Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.090564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.090564Z digest=sha256:55cc0b3d793e13e5ae1c9fac710884a00b4d1e94858f8f3b42dd9130ab24eedf

Observation 882665b5-e390-4393-a700-f7ee81ed8000 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Potemkin Understanding in Large Language Models Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.137112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.137112Z digest=sha256:71a9162553f65fccf432f9ab94ffb659ae20796c8a1e3857a3a12156c52cae0b

Observation 57ecf5a2-459b-4cc0-8b2b-0afe2ce9ee84 · outbound

This paper cites What can large language models do in chemistry? a comprehensive benchmark on eight tasks.

Potemkin Understanding in Large Language Models What can large language models do in chemistry? a comprehensive benchmark on eight tasks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.569299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.184234Z digest=sha256:dd63a8b49adebdfca451ad5bf4f95a1a1a31dd36fca0f2301e0d1b8b66849e28

Observation bb395e0f-87b7-47c5-b22f-19332a00494b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Potemkin Understanding in Large Language Models Measuring Massive Multitask Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.241460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.241460Z digest=sha256:eecd9c962b3a89dc43e310105167f9eee1325bc116ebc8fadfe971c3903cfc58

Observation 9ed4d6bd-64af-43c2-8e0e-b91ee427f5ac · outbound

This paper cites Understanding by Understanding Not: Modeling Negation in Language Models.

Potemkin Understanding in Large Language Models Understanding by Understanding Not: Modeling Negation in Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.327528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.327528Z digest=sha256:3c8fba24f2233ed09d13665891c279c583f4a875115415968393e501a8306771

Observation f80f7806-56f6-4f3a-adde-ab4af24c94be · outbound

This paper cites Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models.

Potemkin Understanding in Large Language Models Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.380612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.380612Z digest=sha256:629eaa17ffafb0774f6e0cab915188b00c5b34a4ae04a29c4e5c1b07b4db2a6e

Observation 0a12bff8-1508-4e41-9956-e1d2c96a5a6b · outbound

This paper cites Adversarial Example Generation with Syntactically Controlled Paraphrase Networks.

Potemkin Understanding in Large Language Models Adversarial Example Generation with Syntactically Controlled Paraphrase Networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.427899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.427899Z digest=sha256:9ce833d984e7fb57a0a644d03284398014f78b06408b141fcf5fa5be3b58f69c

Observation bd4c4a24-c70e-48cf-a109-fd0779d6db79 · outbound

This paper cites Accurate, yet inconsistent? Consistency Analysis on Language Understanding Models.

Potemkin Understanding in Large Language Models Accurate, yet inconsistent? Consistency Analysis on Language Understanding Models

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:30:33.511891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.469419Z digest=sha256:f200217e443d5972ef8844242109b53f52d82a0b343a980ef66d37794cfca675

Observation 4ecb5384-98e1-4909-99a9-00cfd77fb317 · outbound

This paper cites S., and Lukasiewicz, T.

Potemkin Understanding in Large Language Models S., and Lukasiewicz, T

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.468299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.557719Z digest=sha256:3226902bf5e7e7421834d48ed774320467ba33ca092e10894322d4695d302fc0

Observation 68983c8b-c899-4210-9d69-74e87fd3f5a1 · outbound

This paper cites Consistency Analysis of ChatGPT.

Potemkin Understanding in Large Language Models Consistency Analysis of ChatGPT

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.615032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.615032Z digest=sha256:fc830f7442bc1ea82b616b7cbdee8b80318ed61b0d638c5b3f0246e97f40d6bf

Observation 354c96c6-6b8a-4b97-9c6b-027e9347c2bc · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams.

Potemkin Understanding in Large Language Models What disease does this patient have? a large-scale open domain question answering dataset from medical exams

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.375868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.655859Z digest=sha256:b841a09fd38f700bf67b3e7316425e6aaed611d2944d0cd7ca68c0a8b9fa4f5a

Observation 6c072be9-1d61-4f5e-97e5-5ac8b700aab8 · outbound

This paper cites Dynabench: Rethinking Benchmarking in NLP.

Potemkin Understanding in Large Language Models Dynabench: Rethinking Benchmarking in NLP

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.708684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.708684Z digest=sha256:0ee79d21258e9f9ba16f1913a819c2fe36a822ec55b92b96220024d013649682

Observation 0186b315-2370-42f2-bcec-d4c750206225 · outbound

This paper cites a ldchen, S., Binder, A., Montavon, G., Samek, W., and M \.

Potemkin Understanding in Large Language Models a ldchen, S., Binder, A., Montavon, G., Samek, W., and M \

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.273866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.785634Z digest=sha256:71d916d3c7e8ac6a774c663b730197f351efc45266b78ffed82d82dc0af7b463

Observation 61a4b5d1-70c0-452f-accd-f2ac542cbe02 · outbound

This paper cites Benchmarking and Improving Generator-Validator Consistency of Language Models.

Potemkin Understanding in Large Language Models Benchmarking and Improving Generator-Validator Consistency of Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.853855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.853855Z digest=sha256:efb071e4e4cfd6a4de6354a4310e3443de29209f4bb11962a9d6305cf377884f

Observation 34dd9cdc-8b85-4a39-9792-2ecfad2c4309 · outbound

This paper cites Holistic Evaluation of Language Models.

Potemkin Understanding in Large Language Models Holistic Evaluation of Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.910951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.910951Z digest=sha256:313f9aa2fe36c3e4d3268a7f903b31808699ad955772303cc68425d255a4b321

Observation 21d94355-3964-424e-89e1-753c74677884 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Potemkin Understanding in Large Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.965179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.965179Z digest=sha256:dafc567ae8edf920f9477859f2662378a5de16152331a5e40bef2f982148da07

Observation 2549b930-ffc8-4b3d-934c-a762e4745d65 · outbound

This paper cites MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark.

Potemkin Understanding in Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.067924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.067924Z digest=sha256:8b338b0195a6d4704d96ca1a5d337945ffb17e594106848681b6f9846de8ba2e

Observation be411334-3402-45b3-8c6d-ce53abf7f192 · outbound

This paper cites Sycophancy in Large Language Models: Causes and Mitigations.

Potemkin Understanding in Large Language Models Sycophancy in Large Language Models: Causes and Mitigations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.121504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.121504Z digest=sha256:72658b706ec97bdf7ef6f9a9c7525286aa74e197206a067948cbec45b37a78d9

Observation 091c2206-b1ef-4a53-bb76-a814cce90897 · outbound

This paper cites Locating and editing factual associations in gpt.

Potemkin Understanding in Large Language Models Locating and editing factual associations in gpt

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.195017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.176258Z digest=sha256:2305a9857f5670674f2ae877c10e38412dec898518b1df50ec82ff9eb1481293

Observation 6035467f-3a81-400a-b0d5-cc888ce0ad89 · outbound

This paper cites Fast Model Editing at Scale.

Potemkin Understanding in Large Language Models Fast Model Editing at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.224912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.224912Z digest=sha256:c539e5445151629680099974831a84305b63537e2e47d6a7421f3ca3545c43a1

Observation f5d35c09-92f8-4221-b27e-3e9898934714 · outbound

This paper cites Why AI is Harder Than We Think.

Potemkin Understanding in Large Language Models Why AI is Harder Than We Think

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.291743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.291743Z digest=sha256:3af9fd3ce805380bd123d15d03ba4c33674ce6ca13d621f27a28fda2e9caf518

Observation 3f0bdea2-db8d-4924-b014-b67fb3794e51 · outbound

This paper cites D., Bender, E.

Potemkin Understanding in Large Language Models D., Bender, E

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.113972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.379058Z digest=sha256:f93f8c88dc9ba482902f2565f72bbac83c18d4c34b3a26f17c434e175f940b3d

Observation 4d8bd2d3-03dd-41a4-b40a-9090106f3e7f · outbound

This paper cites Language Models as Knowledge Bases?.

Potemkin Understanding in Large Language Models Language Models as Knowledge Bases?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.430923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.430923Z digest=sha256:2f6cc306454c0db6d79991ff300c2e6cc45f5973783b8fac414d41ce452c9a3b

Observation c7fccb6c-5ef1-4c84-ad1c-c551ef3259a2 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

Potemkin Understanding in Large Language Models Measuring and Narrowing the Compositionality Gap in Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.483348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.483348Z digest=sha256:26b24b280f3e31380aabbe1ca420d6d07e0556da18e6eb1b23476cb98a73be1c

Observation 6c20d737-f59a-4e24-a439-fe41835ade21 · outbound

This paper cites AI and the Everything in the Whole Wide World Benchmark.

Potemkin Understanding in Large Language Models AI and the Everything in the Whole Wide World Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.598191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.598191Z digest=sha256:bc8c1f3aab6899f0ba5bf3ca203ded7884d103665be3c50a1546e6b3f9e0b261

Observation 8654d9be-99c3-404f-9f7e-8492fb9d6de5 · outbound

This paper cites Do imagenet classifiers generalize to imagenet? In International conference on machine learning.

Potemkin Understanding in Large Language Models Do imagenet classifiers generalize to imagenet? In International conference on machine learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.017358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.642476Z digest=sha256:679c01251d88a800bb3ebc5929a91c708059b6f2a9cffa4854ec188e5bfd794f

Observation efe6e90d-51f8-4d07-8a1b-e0d52e412b8e · outbound

This paper cites BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices.

Potemkin Understanding in Large Language Models BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.722951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.722951Z digest=sha256:f55d2f554f615cef3b4916809d7d051695017c12006bccef84fb7a18220171a4

Observation bf318aeb-57fe-4762-ba40-f4bb1662706d · outbound

This paper cites why should i trust you?.

Potemkin Understanding in Large Language Models why should i trust you?

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.941864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.791629Z digest=sha256:9a97fdfd34f2f74860441fbe4ba804e056a2a670b5ac7a77f45e9e885c9c2cf3

Observation 59283c74-da0b-44ab-9b13-74f0e61b7f41 · outbound

This paper cites T., Singh, S., and Guestrin, C.

Potemkin Understanding in Large Language Models T., Singh, S., and Guestrin, C

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.829170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.837348Z digest=sha256:7feab948046c02892e473795725099cc0858a7612010649ff0d0d163f9a525e3

Observation d13db075-ea71-430e-a719-e1007ab4d794 · outbound

This paper cites T., Guestrin, C., and Singh, S.

Potemkin Understanding in Large Language Models T., Guestrin, C., and Singh, S

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.720084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.909318Z digest=sha256:a880fdc3b9ab336ae7f5e7677e9eb2185610c854f09aa64d8e437d44abe36026

Observation 1bbb4541-0b04-4f46-83a6-dcf80abbd733 · outbound

This paper cites Beyond Accuracy: Behavioral Testing of NLP models with CheckList.

Potemkin Understanding in Large Language Models Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.993054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.993054Z digest=sha256:e868ad56c30c722ca299c2694d2429fe7228d92c180521b7078b83f0fc8e3b56

Observation 3e4ec196-89d8-4a30-a34d-ea2f8cffbdc8 · outbound

This paper cites Models in the wild: On corruption robustness of neural nlp systems.

Potemkin Understanding in Large Language Models Models in the wild: On corruption robustness of neural nlp systems

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.619733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.060512Z digest=sha256:171aecefdd465a076181f892109912dcd5caeff776f1a34fed8f9380cca8a207

Observation 756778bd-8b0b-448d-86bc-517d7c6e4844 · outbound

This paper cites LLMs' Understanding of Natural Language Revealed.

Potemkin Understanding in Large Language Models LLMs' Understanding of Natural Language Revealed

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:30:33.173737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.137478Z digest=sha256:5813f1b6646136c63cf0a9326326a9db11fd586b459f9ac25fecbd95a6e1b395

Observation 2cc69fd1-3dd3-41b0-864e-6fd6c07ae10a · outbound

This paper cites everyone wants to do the model work, not the data work.

Potemkin Understanding in Large Language Models everyone wants to do the model work, not the data work

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.533005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.192447Z digest=sha256:6b8b9a914d5c067e9679bd192e149de243b2f01c5d665dcf24725b2541d03a3d

Observation 216fad6c-c506-449f-8b23-6b8a3e88fbca · outbound

This paper cites H., Sch \"a rli, N., and Zhou, D.

Potemkin Understanding in Large Language Models H., Sch \"a rli, N., and Zhou, D

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.449212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.275321Z digest=sha256:5ded7a7b9977dcafd294be5ae344eaa5d9a3e0735d3058da010fb5087014353e

Observation 016182a3-c9ea-4dc4-b952-cf87cc9e0a17 · outbound

This paper cites and Choi, Y.

Potemkin Understanding in Large Language Models and Choi, Y

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.348695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.351314Z digest=sha256:60958d4f77fa936108cec70cbf7f80dec9252128b4ea3ede2695a26f255e9842

Observation 14bf4e05-0d71-4d7b-82bf-cfe53b9e85b6 · outbound

This paper cites S., Wei, J., Chung, H.

Potemkin Understanding in Large Language Models S., Wei, J., Chung, H

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.232060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.415105Z digest=sha256:0eefb03bfb6efb359abb012d4b5dc31cfa39f9c6f541d8b5486b2641f0ba87f5

Observation fc60529a-e377-4d7f-bff1-68ebb3aa1c42 · outbound

This paper cites Evaluating the Factual Consistency of Large Language Models Through News Summarization.

Potemkin Understanding in Large Language Models Evaluating the Factual Consistency of Large Language Models Through News Summarization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.462098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.462098Z digest=sha256:b903342ad79c503a6559dd34032b8c4a2054078cddc82540a1ca94e47562ab0f

Observation e7ad8acb-9e25-4e00-a2f2-9ec75fa9b2f0 · outbound

This paper cites Evaluating the World Model Implicit in a Generative Model.

Potemkin Understanding in Large Language Models Evaluating the World Model Implicit in a Generative Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.551668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.551668Z digest=sha256:15b9903ce4b0770b167eecb7184dbe369c2f6a0a44425f7a18f8f52bca20a570

Observation c9c5fe9e-7c31-43b3-91c9-c7e8ec64e601 · outbound

This paper cites Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function.

Potemkin Understanding in Large Language Models Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.594808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.594808Z digest=sha256:417950d3053d2ea68992973ca0925644c7911df303e7953c7d060c33ca362c88

Observation c157b09a-1b97-49fd-b6bb-facc38aef7bc · outbound

This paper cites On the planning abilities of large language models-a critical investigation.

Potemkin Understanding in Large Language Models On the planning abilities of large language models-a critical investigation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.145704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.648173Z digest=sha256:f643ef9d64aa984acf1d7a56d8bbc03315e81b1babf99468d651f1266c0cbabf

Observation ee7bcbd8-a32b-439c-88b7-95a1bbc9be8c · outbound

This paper cites T., Heer, J., and Weld, D.

Potemkin Understanding in Large Language Models T., Heer, J., and Weld, D

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.056585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.704793Z digest=sha256:bb04840c5e442e3acac6843c89f1184ea4de993c73681bb9cc4870b0fd69e71b

Observation c4931cb1-6d7c-4e55-a8b2-411fe52458cc · outbound

This paper cites Kformer: Knowledge injection in transformer feed-forward layers.

Potemkin Understanding in Large Language Models Kformer: Knowledge injection in transformer feed-forward layers

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:33.948182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.766649Z digest=sha256:ae8443f2f9341a7f54853030cbc576c22893e212f6e988145b1e380e69d527ea

Observation 360749ae-3ffb-43ba-a9d3-9645dd258d68 · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

Potemkin Understanding in Large Language Models WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.834986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.834986Z digest=sha256:80136e7e0217e3bb5a1ddc764a93214f309f8052e6148a0b90709e705a91d38c

Observation 0468593e-4a93-42a7-8b83-bf4510fa936e · outbound

This paper cites Modifying Memories in Transformer Models.

Potemkin Understanding in Large Language Models Modifying Memories in Transformer Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.913920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.913920Z digest=sha256:ed3329d6b215ed2395ce7397f2a4797cd14169ba82e6b92d96c71297ab9c1d51

Pith citing papers

Observation e47380ed-7f27-40ba-9c70-a07012ce052f · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis Potemkin Understanding in Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:02.504578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:93422fb96a2a2da5f3b7d6a254b60d4c88cec21adf9cba9c6e630d098b3cbe4c

Observation b920a734-a6e2-4c40-8dfb-4b2310026e8f · inbound

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny cites this paper.

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny Potemkin Understanding in Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:14.933677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:20:14.933677Z digest=sha256:0fed3fb07adaeeff809dcee60e450bf201c179f2d8d97378ef37cb2a2e3634c4

Observation d3377b6c-ba71-43ca-a1e9-e72554f74392 · inbound

The wall confronting large language models cites this paper.

The wall confronting large language models Potemkin Understanding in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:12:02.086965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:12:02.086965Z digest=sha256:3bc99630714d5afdd1e1469d6a8f345ee91ab7802f91d16d0da15ea270b46aee

Observation 4f151f5c-19ea-4c68-a39f-a7fdbbbd2f34 · inbound

A paradox of AI fluency cites this paper.

A paradox of AI fluency Potemkin Understanding in Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:52.218393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-07T16:17:48.531790Z digest=sha256:c465684b0ad3ae0da44d858c3432df5331c37929b86bf580e373ae88de45ffcb

Observation 23d6e3c3-18ed-4570-9836-ee238b71a00a · inbound

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment cites this paper.

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment Potemkin Understanding in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-15T07:32:36.632338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T07:32:36.632338Z digest=sha256:af162678fee6b94ecef52e4808d2e9ff742d1555be7d36cfec0c0c571abe2046