REVIEW 3 major objections 3 minor 11 references
Can AI Rely on the Systematicity of Truth? The Challenge of Modelling Normative Domains
T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper argues that LLMs can only fill gaps in their data by exploiting the systematicity of truth, and that normative domains are too asystematic for that to work.
desk verdict A serious value-pluralist challenge to LLM progress optimism, with a soft empirical bridge but a robust agency thesis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the distinction between systematic and asystematic truth. Systematicity names the property of a body of truths being not merely consistent (free of contradiction) but coherent: truths stand in relations of rational support, so each truth can be recovered from others—an inferential redundancy that LLMs can in principle exploit to fill gaps and correct errors. The paper then uses value pluralism—the thesis that values are irreducibly multiple, often incompatible, and incommensurable—to establish that normative truths lack this redundancy, and uses the distinction between an impersonal 'What is to be done?' and a first-personal 'What should I do?' to show why the resulting judgements of importance remain the agent's own.
What would settle it
A controlled benchmark would settle it: take matched sets of facts from a systematic domain (e.g., geography) and a normative domain (e.g., what values bear on novel moral situations), withhold or corrupt the same fraction of each, and test whether an LLM recovers the withheld truths at comparable rates after training on the rest. If the model extrapolates normative truths as successfully as geographical ones despite the lack of inferential redundancy, the central claim would be false.
Extended reading notes
Core claim
The central claim is that LLMs' capacity to progress beyond incomplete and inaccurate training data depends on the systematicity of truth—the consistency and inferential coherence of true statements—and that in normative domains this systematicity is largely absent. On the paper's account, value pluralism shows that values are irreducibly diverse, incompatible, and incommensurable, so normative truths form a fragmented and tension-ridden landscape rather than a web. Consequently, the very mechanism that promises self-completion and self-correction in systematic domains—deriving missing truths from surrounding truths—cannot be relied on in ethics and politics. The paper further claims that this asystematicity intensifies the first-personal and personal character of practical deliberation: deciding what matters in a hard choice cannot be outsourced to an algorithm because it requires the agent's own judgement of importance and authenticity.
Load-bearing premise
The argument rests on the premise that LLMs can only move beyond their training data by exploiting a domain's inferential redundancy, so that where that redundancy is absent, no other mechanism—retrieval, symbolic reasoning, memory, or human-in-the-loop scaffolding—can make coverage comprehensive.
Editorial extensions
If this is right
- LLM self-completion and self-correction from training-data gaps should work in systematic empirical domains such as geography, but should falter precisely on normative questions where values conflict.
- Value-pluralist training data that explicitly represents conflicts can help surface considerations, but cannot supply the inferential redundancy needed to extrapolate beyond the data.
- The less systematic a domain is, the more first-personal judgement is required, so even a highly capable moral advisor leaves the final 'should I really do this?' question to the agent.
- Fine-tuning methods aimed at internal consistency and coherence cannot substitute for the systematicity missing from the normative landscape itself.
Reading between the lines
- Editorial inference: this argument predicts an empirically testable asymmetry—a model trained on sparse or corrupted data should recover withheld facts more reliably in systematic domains than in normative domains, once difficulty is matched.
- Editorial inference: if the paper is right, current alignment techniques that aggregate human preferences may not merely risk 'washing out' value conflicts; they may also create an illusion of comprehensiveness precisely where the model is depending on training data rather than inference.
- Editorial inference: the same asystematicity constraint would apply to future models with retrieval or memory scaffolding only if those mechanisms import normatively resolved answers from outside the model; otherwise they just shift the gap rather than closing it.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that the optimism that LLMs can overcome gaps and inaccuracies in training data by exploiting the systematicity of truth fails in normative domains. Drawing on the value-pluralist tradition (Berlin, Williams, Nagel, Chappell), it claims that normative truths are at least partly asystematic: values are irreducibly plural, incompatible, and often incommensurable, so truths about values do not form a consistent, coherent, inferentially interlocking web. It then infers that because LLMs cannot leverage systematicity in these domains, their self-completion and self-correction will be correspondingly harder, and comprehensive modelling of normative domains will be hampered. The paper further argues that asystematicity amplifies the first-personal and personal dimensions of practical deliberation, so the final practical judgment cannot be outsourced to an AI. The argument is explicitly conditional and carefully qualified throughout.
Significance. If the argument succeeds, it identifies a principled structural limit on LLM performance in ethics and politics, and it grounds a positive role for human agency in AI-assisted practical deliberation. The paper is significant for both AI ethics and the philosophy of AI: it connects a long-standing value-pluralist literature to a concrete capability claim, engages with existing systems such as ValuePrism and Kaleido, and carefully distinguishes output-level consistency and coherence from generation that actually relies on the wider systematicity of truth. Its conclusions are presented conditionally, which is methodologically honest and makes the paper a useful starting point for further empirical work. The main limitation is that the inference from truth-level asystematicity to corpus-level unlearnability is asserted rather than demonstrated, and the paper's own concessions in the Conclusion leave this bridge under-protected.
major comments (3)
- [§4] The central inference from asystematicity of ground truth to reduced LLM capability is under-argued in the paragraph beginning "The asystematicity of normative truths in turn has implications for the prospects of LLMs." LLMs are trained on corpora, not directly on the fabric of facts, and the paper itself notes in §3 that they internalise statistical patterns in text. A normative domain whose true propositions do not form a systematic web may nevertheless exhibit stable regularities in discourse about conflicts, trade-offs, and incommensurability; the paper does not explain why these corpus-level patterns cannot support learning and generalisation. Since the Abstract's claim that asystematicity renders progress "correspondingly harder" depends on this bridge, the argument needs either a defended premise that alternative mechanisms cannot compensate or a restriction of the conclusion to the claim that LLMs cannot rely on the specific route of systematicity.
- [§6] The Conclusion's concession that "alternative ways for LLMs to move beyond their training data may yet emerge" undermines the inference from "LLMs cannot leverage the systematic harmony of these domains" to "it should to that extent be harder for LLMs to comprehensively model normative domains." If retrieval-augmented generation, symbolic reasoning, external memory, or human-in-the-loop scaffolding could supply the needed redundancy, the predicted difficulty would not follow. The paper should either address these mechanisms and explain why they cannot restore comprehensiveness, or state more modestly that the systematicity-based route is unavailable without thereby asserting an overall increase in difficulty.
- [§4] The claim that a pluralist model such as Kaleido "cannot overcome the limitation imposed by the asystematicity of normative truth on its capacity to move beyond its training data" is presented as a direct consequence of asystematicity, but the example only shows that one system trained on synthetic GPT-4 data does not extrapolate via systematicity. It does not test whether other statistical or non-statistical mechanisms could improve coverage. This is an empirical premise that needs support or explicit identification as an open empirical question.
minor comments (3)
- [§3] The text refers to "autoregressive LMMs" where the intended term is presumably "LLMs"; this should be corrected.
- [§4] In the discussion of Sorensen et al., the citation appears as "Sorensen et al,. 2024" with a stray comma before the period, and the model name "Value Kaleidoskope" is spelled differently from the standard "Kaleidoscope" in the cited paper.
- [§5] The claim that Kaleido's relevance scores "measure the statistical relevance of a type of consideration to a type of situation" while importance "goes significantly beyond such merely statistical relevance" is asserted rather than argued; a sentence explaining why importance cannot be approximated by richer contextual features would help the reader assess this step.
Circularity Check
No significant circularity: the paper's argument is a conditional application of an independently sourced pluralist premise, with no fitted parameters or self-citation chain doing the work.
full rationale
The paper's central inference is conditional: if normative truths are asystematic, then LLMs cannot leverage the systematicity of truth for self-completion and self-correction, and so comprehensive modelling is to that extent harder. This is an application of the definition of systematicity to a stated mechanism (inferential redundancy), not a disguised restatement of the conclusion. The asystematicity premise is drawn from an independent value-pluralist literature (Berlin, Williams, Nagel, Chang, Wiggins and Williams, Chappell), not from the author's own prior work. The author's self-citations appear only as elaborations of Williamsian themes and are not load-bearing for the main inference. The paper also explicitly concedes that 'alternative ways for LLMs to move beyond their training data may yet emerge,' showing that the conclusion is not a fixed-point or fitted claim but a limited conditional. The weakest point is the bridging premise that LLM progress beyond training data must exploit inferential redundancy; this is asserted rather than empirically established, but that is an under-argument risk, not circularity. No equations, fitted parameters, or benchmark 'predictions' are present whose outputs are pre-encoded by their inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption Value pluralism is correct: there is a plurality of irreducibly distinct, incompatible, and incommensurable values, so normative truths are at least partly asystematic.
- domain assumption LLMs can only leverage systematicity indirectly, and their primary route to self-completion and self-correction is through consistency and coherence of the target domain.
- domain assumption Practical deliberation has an irreducibly first-personal dimension, so even complete impersonal advice leaves a first-personal question for the agent.
Cite this review
Pith. "Pith review of Can AI Rely on the Systematicity of Truth? The Challenge of Modelling Normative Domains." pith.science (2026). https://pith.science/paper/PHYDC2JF
@misc{pith2026250709676,
author = {Pith},
title = {Pith review of: Can AI Rely on the Systematicity of Truth? The Challenge of Modelling Normative Domains},
year = {2026},
howpublished = {\url{https://pith.science/paper/PHYDC2JF}},
note = {Machine review of arXiv:2507.09676}
}
read the original abstract
A key assumption fuelling optimism about the progress of large language models (LLMs) in accurately and comprehensively modelling the world is that the truth is systematic: true statements about the world form a whole that is not just consistent, in that it contains no contradictions, but coherent, in that the truths are inferentially interlinked. This holds out the prospect that LLMs might in principle rely on that systematicity to fill in gaps and correct inaccuracies in the training data: consistency and coherence promise to facilitate progress towards comprehensiveness in an LLM's representation of the world. However, philosophers have identified compelling reasons to doubt that the truth is systematic across all domains of thought, arguing that in normative domains, in particular, the truth is largely asystematic. I argue that insofar as the truth in normative domains is asystematic, this renders it correspondingly harder for LLMs to make progress, because they cannot then leverage the systematicity of truth. And the less LLMs can rely on the systematicity of truth, the less we can rely on them to do our practical deliberation for us, because the very asystematicity of normative domains requires human agency to play a greater role in practical thought.
Reference graph
Works this paper leans on
-
[1]
What if Dario Amodei Is Right About A.I.?
Abela, P. (2006). The Demands of Systematicity: Rational Judgment and the Structure of Nature. In G. Bird (Ed.), A Companion to Kant (pp. 408–422). Blackwell. Amodei, D. (2024). “What if Dario Amodei Is Right About A.I.?.” Interview by Ezra Klein. The Ezra Klein Show, New York Times Opinion, April 12,
work page 2006
-
[8]
16995/ pp. 17532 Queloz, M. (2025). The ethics of conceptualization: Tailoring thought and language to need. Oxford Uni- versity Press. https:// doi. org/
work page 2025
-
[9]
Leibniz and the Concept of a System
0001 Rescher, N. (1981). “Leibniz and the Concept of a System.” In Leibniz’s Metaphysics of Nature: A Group of Essays, 29–41. Dordrecht: Springer. Rescher, N. (1979). Cognitive Systematization: A Systems Theoretic Approach to a Coherentist Theory of Knowledge. Blackwell. Rescher, N. (2000). Kant and the Reach of Reason: Studies in Kant’s Theory of Rationa...
work page 1981
-
[10]
LLMs are Not Just Next Token Predictors
1038/ s41598- 025- 86510-0 Downes, S. M., Forber, P., & Grzankowski, Alex. (2024). “LLMs are Not Just Next Token Predictors.” arXiv arXiv:
work page 2024
-
[11]
Le concept de système de Leibniz à Condillac
Jahrhundert. Edited by Jürgen Blühdorn and Joachim Ritter, 63–88. Frankfurt am Main: Klostermann. Vickers, P. (2013). Understanding Inconsistent Science. Oxford University Press. Vieillard-Baron, J.-L. (1975). “Le concept de système de Leibniz à Condillac.” In Akten des II. Interna- tionalen Leibniz-Kongresses Hannover, 17.-22. Juli
work page 2013
-
[19]
Jahrhundert. Edited by Jürgen Blühdorn and Joachim Ritter, 99–122. Frankfurt am Main: Klostermann. Kekes, J. (1993). The Morality of Pluralism. Princeton University Press. Kitcher, P. (1986). Projecting the Order of Nature. In R. Butts (Ed.), Kant’s Philosophy of Material Nature (pp. 201–235). D. Reidel. Kretzmann, Norman, & Eleonore Stump. (1989). The Ca...
work page 1993
-
[1972]
Edited by Kurt Müller, Heinrich Schep- ers and Wilhelm Totok, 97–103. Wiesbaden: F. Steiner. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., ... & Fedus, W. (2022). Emergent abilities of large language models. arXiv preprint arXiv:2206.07682. Wiggins, D., & Williams, B. (1978). “Aurel Thomas Kolnai.” In Ethics, Value and Reality: Aure...
arXiv 2022
-
[2024]
A general language assistant as a laboratory for alignment
https:// www. nytim es. com/ 2024/ 04/ 12/ opini on/ ezra- klein- podca st- dario- amodei. html Askell, A., Bai, Y, Chen, A., Drain, D, Ganguli, D, Henighan, T., Jones, A., Joseph, N., Mann, B, & Das- Sarma, Nova. (2021). “A general language assistant as a laboratory for alignment.” arXiv preprint arXiv:
work page 2021
Show all 11 references
-
[2112]
Epistemology of Artificial Intelligence
00861. Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., ... & Kaplan, J. (2022). Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073. Beisbart, C. (forthcoming-a). “Epistemology of Artificial Intelligence.” In The Stanford Enc...
2022 arXiv
-
[2408]
Ideal Observer
04666. Dupré, J. (1995). The Disorder of Things: Metaphysical Foundations of the Disunity of Science. Harvard University Press. Elazar, Y., Kassner, N., Ravfogel, S., Ravichander, A., Hovy, E., Schütze, H., & Goldberg, Y. (2021). Measuring and Improving Consistency in Pretrain...
1995 arXiv
-
[2410]
Value Pluralism
02205. Losano, M. G. (1968). Sistema e struttura nel diritto, vol. 1: Dalle origini alla scuola storica. Turin: Giuffrè. MacIntyre, A. C. (2007). After Virtue: A Study in Moral Theory (3rd ed.). University of Notre Dame Press. Mason, E. (2023). “Value Pluralism.” In The Stanfo...
1968
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.