REVIEW 3 major objections 6 minor 10 references
From prosthetic memory to prosthetic denial: Auditing whether large language models are prone to mass atrocity denialism
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Audit: LLMs get 10% of genocide answers wrong
desk verdict A transparent comparative audit with a real Cambodian/Khmer finding, but the headline metric measures yes/no label agreement, not the paper's own concept of prosthetic denial, because justifications of correct answers are never coded. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the comparative audit itself, built on expert-defined baselines: for each of four mass atrocities, nine prompts encode common denialist claims, and each prompt is asked in English and in a language central to the case (Ukrainian, German, Khmer, French). A uniform template forces a JSON response with a yes/no label and a justification, so the study can separate label accuracy from narrative framing. The concept doing the theoretical work is 'prosthetic denial,' defined as the AI-mediated erasure or distortion of atrocity memory, set against 'prosthetic memory,' the mediated experience of the past. Language representation statistics from the web serve as a proxy for the training data likely available to the models, linking observed performance gaps to data availability.
What would settle it
Code all 370 responses' justifications—including those attached to correct labels—for denialist tropes, and compare rates across cases; if the Holocaust and Holodomor justifications contain denialist framing at rates comparable to Cambodia, the claim that only underrepresented cases are susceptible would fail.
Extended reading notes
Core claim
On the paper's own terms, the central finding is that LLM accuracy about mass atrocities is uneven and consequential: models answer reliably for widely documented, legally shielded events such as the Holocaust and the Holodomor, but they reproduce revisionist and denialist narratives for less-documented cases. The Cambodian Genocide is the clearest example: 18 of the 20 incorrect responses to Khmer prompts echoed nationalist tropes blaming Vietnam, portrayed Khmer Rouge leaders as patriotic defenders, or treated contested claims as open questions, and Mixtral failed to produce coherent answers in Khmer at all. The audit also found inconsistencies between binary yes/no labels and their justifications, such as Gemini affirming that Ukraine suffered no differently from other Soviet republics while its explanation stated the opposite. These contradictions matter because users may absorb denialist framing even when the headline answer is correct.
Load-bearing premise
The load-bearing premise is that matching the expert yes/no baseline measures whether a model contributes to denialism, even though denialist framing can hide inside the justifications of correct answers and the paper acknowledges this possibility.
Editorial extensions
If this is right
- LLM outputs about mass atrocities are not uniformly reliable; accuracy degrades for less-documented events and low-resource languages, so users in Khmer face materially worse information about the Cambodian Genocide.
- A correct yes/no label does not guarantee a safe response, because justifications can contradict the label or echo denialist tropes; evaluating model safety on label accuracy alone will miss this.
- Unmoderated LLM use risks reinforcing existing asymmetries in historical representation, amplifying nationalist or revisionist narratives for events that already receive less attention.
- The small replication variation (96.2% match over four months) suggests the patterns are stable in the short term, but model updates can shift performance, as Mixtral's Khmer answers improved.
- If LLMs replace search engines and fact-checkers, the opacity of their sourcing means users cannot easily detect when a confident answer mixes fact with denialist framing.
Reading between the lines
- If justifications were coded for denialist rhetoric regardless of label correctness, the share of problematic outputs would likely exceed the 10% label-error rate, because the paper itself notes similar statements might appear in justifications of correct answers.
- The same audit design could be extended to other contested histories—for example colonial atrocities or the Armenian genocide—with the expectation that documentation volume and language representation will predict error patterns.
- The label-justification contradiction points to a distinct failure mode: safety fine-tuning may suppress overt denial at the answer level while leaving narrative content to be generated probabilistically, so denialism can leak through in softer form.
- As chatbots become multimodal and more affectively expressive, the same revisionist framings could carry greater emotional credibility, making the prosthetic-denial risk larger than the text-only audit captures.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes the concept of 'prosthetic denial' to describe AI-mediated erasure or distortion of atrocity memory and reports a comparative audit of five LLMs (Claude, GPT, Llama, Mixtral, Gemini) across four cases (Holodomor, Holocaust, Cambodian Genocide, genocide against the Tutsi in Rwanda). Each model was prompted with nine questions per case in English and a case-relevant language, and the parsed yes/no answers were compared against expert-defined baselines. The authors report 37 incorrect answers out of 370 prompts (10%), with the highest error concentration in the Cambodian Genocide case, especially for Khmer prompts, and qualitatively describe how some outputs reproduce revisionist or denialist framings. The paper argues that LLMs can act as sites of prosthetic memory but that unmoderated use risks reinforcing historical denialism, particularly for underrepresented cases and low-resource languages.
Significance. If the central empirical claim is fully supported, the paper addresses an important and under-studied question: whether LLM outputs about mass atrocities are uniformly reliable or degrade for less-documented events and lower-resource languages. The audit design is transparent and valuable: it uses expert-derived baselines, covers multiple models and languages, includes a replication wave, and makes incorrect-answer justifications available via an OSF appendix. The conceptual framing of 'prosthetic denial' is a useful contribution to the memory-studies literature and to AI auditing discussions. The main quantitative headline—that Cambodia and Khmer prompts produce more errors—is plausible, but the evidence as currently coded does not fully measure the paper's own construct of prosthetic denial, because denialist framing inside label-correct justifications is not systematically assessed. The paper's strengths are its cross-case design, the clear prompt set, and the transparency about replication; its main weakness is the gap between the measured outcome (yes/no label agreement) and the claimed outcome (susceptibility to denialist framings).
major comments (3)
- [Methods and Results (quantitative analysis; Discussion)] The central quantitative outcome is only whether the parsed yes/no label matches the expert baseline: the Methods state 'we compared the parsed outputs obtained from all prompts against the baseline answer,' and the Results state 'we only reviewed justifications for incorrect answers.' Yet the paper's core construct, 'prosthetic denial,' is defined as AI-mediated erasure or distortion of atrocity-related past. A label-correct response that still reproduces a revisionist frame—such as the Gemini Q1 English Cambodia response invoking Vietnam's later invasion, or the ChatGPT Q2y Khmer response attributing atrocities to a 'bad faction' of the Khmer Rouge—is counted as accurate despite being denialist in framing. The Discussion explicitly concedes that 'similar statements might appear in the justifications of correct answers as well.' Because justifications for correct answers were never coded, the reported 20/100 Cambodia error rate and the Cambodia-versus-Holocaust contrast cannot be interpreted as rates of denialist output. This is load-bearing for the abstract's claim about 'susceptibility to denialist framings' and needs to be fixed, for example by coding all 370 justifications with a rubric for denialist framing or by restricting the quantitative claim to label accuracy.
- [Results, Figure 1 and overall counts] The paper reports raw proportions (37/370, 20/100 for Cambodia, 2/90 for Holodomor) without any statistical inference, confidence intervals, or adjustment for the nested structure of the data (9 prompts per model-language-case, repeated across five models). With only nine prompts per model-language-case, the cross-case differences could be influenced by prompt selection and stochastic variation; the replication wave shows 14 mismatches between the original and replication labels, further indicating variability. To support the comparative claim that Cambodia is significantly worse than the Holocaust or Holodomor, the authors should provide at least exact binomial confidence intervals or a simple regression/test that accounts for the grouping of prompts by case, model, and language. This is a fixable but necessary addition to make the central quantitative contrast interpretable.
- [Results, Rwanda Q8 and Discussion, Bagosora] There is an internal inconsistency in the reporting of the Bagosora question (Q8). The Results state that 'six out of the eight answers incorrectly affirmed the proposition that Bagosora stands convicted of conspiracy,' while the Discussion states that 'several models answered correctly that no conspiracy conviction occurred, yet their justifications oversimplified or ambiguously described the legal rationale.' These two statements cannot both describe the same set of responses: the first says most answers were label-incorrect, the second implies most answers were label-correct. This contradiction affects the interpretation of the Rwanda case and the general point about label-justification mismatches; the authors should clarify which models gave which labels and reconcile the two passages.
minor comments (6)
- [Methods, Table 5 note] The text below Table 5 says the placeholder [QUESTION] 'was replaced by the corresponding question in Table 1,' but the question list is Table 3; this cross-reference should be corrected.
- [Results, replication paragraph] The replication paragraph contains a typo: 'Mixtra' should be 'Mixtral' in the sentence about the Khmer-language improvements.
- [Results, section header] The section header 'Holomodor' is misspelled; it should be 'Holodomor'.
- [Methods, model selection] The Methods say the 'considered best model' of each provider was used, but the selection criterion is not defined; specifying the criterion (e.g., provider documentation, benchmark rankings) would strengthen reproducibility.
- [Methods, prompt template and API settings] The paper does not report API parameters such as temperature or decoding settings, which are known to affect LLM output variability; reporting these would improve the replicability of the audit beyond the two collection dates.
- [Figure 1 caption] The caption says 'Number of incorrect answers per model per genocide,' but the study includes the Holodomor and the Holocaust, which are not all termed 'genocide' in the same way; 'per case study' would be more accurate.
Circularity Check
No significant circularity: the audit compares LLM outputs against externally defined expert baselines, and the acknowledged 'correct justifications' gap is a measurement-validity limitation rather than a circular derivation.
full rationale
The derivation chain in this paper is empirical and externally grounded. Each prompt is scored by comparing the parsed yes/no answer to a baseline answer fixed by the authors' expertise (Methods: 'For the quantitative analysis, we compared the parsed outputs obtained from all prompts against the baseline answer defined by the experts'). The central quantitative claim (37/370 incorrect, 20/100 Cambodian Genocide prompts) is therefore a measurement against external historical consensus, not a quantity defined by the model outputs or by the paper's own concept. The concept 'prosthetic denial' is introduced as a framing label for the phenomenon under study; it is not fitted, estimated, or defined in terms of the reported label-mismatch rate. There are self-citations (e.g., Makhortykh et al., 2023b), but they provide conceptual context and prior motivation only; no uniqueness theorem, fitted parameter, or load-bearing premise is imported from them. The paper does contain one acknowledged measurement gap: justifications were reviewed only for answers that already mismatched the baseline, and the Discussion concedes, 'Notably, similar statements might appear in the justifications of correct answers as well.' This means the label-mismatch rate may undercount denialist framing inside otherwise correct outputs. That is a construct-validity caveat about what the quantitative outcome measures, not a case of a result being equivalent to its input by construction or by definition. No circular step can be exhibited from the paper's own equations or citation chain, so the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The expert-defined baseline answers are historically correct and unambiguous.
- domain assumption The nine questions per case are representative of denialist narratives and the translations preserve the intended meaning.
- domain assumption Web language prevalence statistics approximate the amount of LLM training data available for each language.
- domain assumption A binary yes/no label that matches the baseline is a valid indicator of whether the model contributes to denialism.
invented entities (1)
-
prosthetic denial
Cite this review
Pith. "Pith review of From prosthetic memory to prosthetic denial: Auditing whether large language models are prone to mass atrocity denialism." pith.science (2026). https://pith.science/paper/4HKHH67Q
@misc{pith2026250521753,
author = {Pith},
title = {Pith review of: From prosthetic memory to prosthetic denial: Auditing whether large language models are prone to mass atrocity denialism},
year = {2026},
howpublished = {\url{https://pith.science/paper/4HKHH67Q}},
note = {Machine review of arXiv:2505.21753}
}
read the original abstract
The proliferation of large language models (LLMs) can influence how historical narratives are disseminated and perceived. This study explores the implications of LLMs' responses on the representation of mass atrocity memory, examining whether generative AI systems contribute to prosthetic memory, i.e., mediated experiences of historical events, or to what we term "prosthetic denial," the AI-mediated erasure or distortion of atrocity memories. We argue that LLMs function as interfaces that can elicit prosthetic memories and, therefore, act as experiential sites for memory transmission, but also introduce risks of denialism, particularly when their outputs align with contested or revisionist narratives. To empirically assess these risks, we conducted a comparative audit of five LLMs (Claude, GPT, Llama, Mixtral, and Gemini) across four historical case studies: the Holodomor, the Holocaust, the Cambodian Genocide, and the genocide against the Tutsis in Rwanda. Each model was prompted with questions addressing common denialist claims in English and an alternative language relevant to each case (Ukrainian, German, Khmer, and French). Our findings reveal that while LLMs generally produce accurate responses for widely documented events like the Holocaust, significant inconsistencies and susceptibility to denialist framings are observed for more underrepresented cases like the Cambodian Genocide. The disparities highlight the influence of training data availability and the probabilistic nature of LLM responses on memory integrity. We conclude that while LLMs extend the concept of prosthetic memory, their unmoderated use risks reinforcing historical denialism, raising ethical concerns for (digital) memory preservation, and potentially challenging the advantageous role of technology associated with the original values of prosthetic memory.
Reference graph
Works this paper leans on
-
[1]
Augenstein, I., Baldwin, T., Cha, M., Chakraborty, T., Ciampaglia, G. L., Corney, D., DiResta, R., Ferrara, E., Hale, S., Halevy, A., Hovy, E., Ji, H., Menczer, F., Miguez, R., Nakov, P., Scheufele, D., Sharma, S., & Zagni, G. (2024). Factuality challenges in the era of large language models and opportunities for fact-checking. Nature Machine Intelligence...
arXiv 2024
-
[2]
Mierwald, M. (2024). Chatting about the Past with Artificial Intelligence: A Case Study of Pupils’ Interaction with ChatGPT while Completing a History-Learning Task. Journal of Educational Media, Memory, and Society, 16(2), 143-173. Newton, C. (2025, May 9). Stats from a dying web. Retrieved from https://www.platformer.news/safari-search-decline-apple-goo...
work page 2024
-
[6]
Lewis, A. (1976, October 4). Menu for disaster. The New York Times. Retrieved from https://www.nytimes.com, May 19, 2025 Li, Z., Shi, Y., Liu, Z., Yang, F., Payani, A., Liu, N., & Du, M. (2025, April). Language ranker: A metric for quantifying llm performance across high and low-resource languages. In Proceedings of the AAAI Conference on Artificial Intel...
work page 2025
-
[7]
https://doi.org/10.3389/frai.2024.1341697 Reading, A. (2009). Memobilia: The Mobile Phone and the Emergence of Wearable Memories. In: Garde-Hansen, J., Hoskins, A., Reading, A. (eds) Save As … Digital Memories. Palgrave Macmillan, London. https://doi.org/10.1057/9780230239418_5 Richardson-Walden, V. G., & Makhortykh, M. (2024). Imagining Human-AI Memory S...
-
[28]
Makhortykh, M., Sydorova, M., Baghumyan, A., Vziatysheva, V., & Kuznetsova, E. (2024). Stochastic lies: How LLM-powered chatbots deal with Russian disinformation about the war in Ukraine. Harvard Kennedy School Misinformation Review . Manning, C. D. (2022). Human Language Understanding & Reasoning. Daedalus, 151(2), 127–138. https://doi.org/10.1162/daed_a...
-
[305]
Gadzins’ka, I. (2021). Desiat’ milioniv piatsot tysh. I krapka. Ukraintsiam nav’iazyit’ superechlivi dani pro 10,5 mil’ioniv zhertv Golodomory. Istorycha Pravda, 26 November. Available at: https://www.istpravda.com.ua/articles/2021/11/26/160564/ (accessed 18 May
work page 2021
-
[1936]
New York . Bloch, M. (1998). Autobiographical Memory and the Historical Memory of the More Distant Past, in How We Think They Think: Anthropological Approaches to Cognition, Memory, and Literacy , Westview Press, 114-130. DeVerna, M. R., Yan, H. Y., Yang, K.-C., & Menczer, F. (2024). Fact-checking information from large language models can decrease headli...
-
[2018]
25 Landsberg, A. (2004). Prosthetic memory: The transformation of American remembrance in the age of mass culture . Columbia University Press. Lee, J. (2008). Rwanda: No Conspiracy, No Genocide Planning … No Genocide? JURIST. Retrieved from https://www.jurist.org/commentary/2008/12/rwanda-no-conspiracy-no-genocide/, May 26,
work page 2004
Show all 10 references
-
[2024]
https://www.obdilci.org/Base/en/search Pisanty, V
[Database]. https://www.obdilci.org/Base/en/search Pisanty, V. (2023). Identification and Make-Believe: the Fallacies of Prosthetic Memory. Witnessing the Witness of War Crimes, Mass Murder, and Genocide . Quelle, D., & Bovet, A. (2024). The perils and promises of fact-checkin...
2023
- [2025]
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.