Pith. sign in

REVIEW 3 major objections 3 minor 11 references

Can AI Rely on the Systematicity of Truth? The Challenge of Modelling Normative Domains

T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that LLMs can only fill gaps in their data by exploiting the systematicity of truth, and that normative domains are too asystematic for that to work.

desk verdict A serious value-pluralist challenge to LLM progress optimism, with a soft empirical bridge but a robust agency thesis. read the letter →

arxiv 2507.09676 v1 pith:PHYDC2JF submitted 2025-07-13 cs.CY

classification cs.CY
keywords largelanguagemodelssystematicityoftruthvaluepluralismnormativedomainsself-completionself-correctionpracticaldeliberationAIethics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the optimistic assumption behind large language models' ability to comprehensively model the world—that true statements hang together in a systematically interlinked web—does not hold in normative domains. Drawing on value pluralism, it holds that truths about ethics and politics are often irreducibly conflicting, incommensurable, and disconnected, so LLMs cannot exploit inferential redundancy to fill gaps or correct errors there. As a result, the paper concludes, comprehensive modelling of normative domains should be harder for LLMs, and the final practical decision must remain the agent's own. A sympathetic reader should care because this gives a principled reason to expect AI moral advice to remain advisory no matter how much training data accumulates.

What carries the argument

The central machinery is the distinction between systematic and asystematic truth. Systematicity names the property of a body of truths being not merely consistent (free of contradiction) but coherent: truths stand in relations of rational support, so each truth can be recovered from others—an inferential redundancy that LLMs can in principle exploit to fill gaps and correct errors. The paper then uses value pluralism—the thesis that values are irreducibly multiple, often incompatible, and incommensurable—to establish that normative truths lack this redundancy, and uses the distinction between an impersonal 'What is to be done?' and a first-personal 'What should I do?' to show why the resulting judgements of importance remain the agent's own.

What would settle it

A controlled benchmark would settle it: take matched sets of facts from a systematic domain (e.g., geography) and a normative domain (e.g., what values bear on novel moral situations), withhold or corrupt the same fraction of each, and test whether an LLM recovers the withheld truths at comparable rates after training on the rest. If the model extrapolates normative truths as successfully as geographical ones despite the lack of inferential redundancy, the central claim would be false.

Watch

Extended reading notes

Core claim

The central claim is that LLMs' capacity to progress beyond incomplete and inaccurate training data depends on the systematicity of truth—the consistency and inferential coherence of true statements—and that in normative domains this systematicity is largely absent. On the paper's account, value pluralism shows that values are irreducibly diverse, incompatible, and incommensurable, so normative truths form a fragmented and tension-ridden landscape rather than a web. Consequently, the very mechanism that promises self-completion and self-correction in systematic domains—deriving missing truths from surrounding truths—cannot be relied on in ethics and politics. The paper further claims that this asystematicity intensifies the first-personal and personal character of practical deliberation: deciding what matters in a hard choice cannot be outsourced to an algorithm because it requires the agent's own judgement of importance and authenticity.

Load-bearing premise

The argument rests on the premise that LLMs can only move beyond their training data by exploiting a domain's inferential redundancy, so that where that redundancy is absent, no other mechanism—retrieval, symbolic reasoning, memory, or human-in-the-loop scaffolding—can make coverage comprehensive.

Editorial extensions

If this is right

  • LLM self-completion and self-correction from training-data gaps should work in systematic empirical domains such as geography, but should falter precisely on normative questions where values conflict.
  • Value-pluralist training data that explicitly represents conflicts can help surface considerations, but cannot supply the inferential redundancy needed to extrapolate beyond the data.
  • The less systematic a domain is, the more first-personal judgement is required, so even a highly capable moral advisor leaves the final 'should I really do this?' question to the agent.
  • Fine-tuning methods aimed at internal consistency and coherence cannot substitute for the systematicity missing from the normative landscape itself.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: this argument predicts an empirically testable asymmetry—a model trained on sparse or corrupted data should recover withheld facts more reliably in systematic domains than in normative domains, once difficulty is matched.
  • Editorial inference: if the paper is right, current alignment techniques that aggregate human preferences may not merely risk 'washing out' value conflicts; they may also create an illusion of comprehensiveness precisely where the model is depending on training data rather than inference.
  • Editorial inference: the same asystematicity constraint would apply to future models with retrieval or memory scaffolding only if those mechanisms import normatively resolved answers from outside the model; otherwise they just shift the gap rather than closing it.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper argues that the optimism that LLMs can overcome gaps and inaccuracies in training data by exploiting the systematicity of truth fails in normative domains. Drawing on the value-pluralist tradition (Berlin, Williams, Nagel, Chappell), it claims that normative truths are at least partly asystematic: values are irreducibly plural, incompatible, and often incommensurable, so truths about values do not form a consistent, coherent, inferentially interlocking web. It then infers that because LLMs cannot leverage systematicity in these domains, their self-completion and self-correction will be correspondingly harder, and comprehensive modelling of normative domains will be hampered. The paper further argues that asystematicity amplifies the first-personal and personal dimensions of practical deliberation, so the final practical judgment cannot be outsourced to an AI. The argument is explicitly conditional and carefully qualified throughout.

Significance. If the argument succeeds, it identifies a principled structural limit on LLM performance in ethics and politics, and it grounds a positive role for human agency in AI-assisted practical deliberation. The paper is significant for both AI ethics and the philosophy of AI: it connects a long-standing value-pluralist literature to a concrete capability claim, engages with existing systems such as ValuePrism and Kaleido, and carefully distinguishes output-level consistency and coherence from generation that actually relies on the wider systematicity of truth. Its conclusions are presented conditionally, which is methodologically honest and makes the paper a useful starting point for further empirical work. The main limitation is that the inference from truth-level asystematicity to corpus-level unlearnability is asserted rather than demonstrated, and the paper's own concessions in the Conclusion leave this bridge under-protected.

major comments (3)
  1. [§4] The central inference from asystematicity of ground truth to reduced LLM capability is under-argued in the paragraph beginning "The asystematicity of normative truths in turn has implications for the prospects of LLMs." LLMs are trained on corpora, not directly on the fabric of facts, and the paper itself notes in §3 that they internalise statistical patterns in text. A normative domain whose true propositions do not form a systematic web may nevertheless exhibit stable regularities in discourse about conflicts, trade-offs, and incommensurability; the paper does not explain why these corpus-level patterns cannot support learning and generalisation. Since the Abstract's claim that asystematicity renders progress "correspondingly harder" depends on this bridge, the argument needs either a defended premise that alternative mechanisms cannot compensate or a restriction of the conclusion to the claim that LLMs cannot rely on the specific route of systematicity.
  2. [§6] The Conclusion's concession that "alternative ways for LLMs to move beyond their training data may yet emerge" undermines the inference from "LLMs cannot leverage the systematic harmony of these domains" to "it should to that extent be harder for LLMs to comprehensively model normative domains." If retrieval-augmented generation, symbolic reasoning, external memory, or human-in-the-loop scaffolding could supply the needed redundancy, the predicted difficulty would not follow. The paper should either address these mechanisms and explain why they cannot restore comprehensiveness, or state more modestly that the systematicity-based route is unavailable without thereby asserting an overall increase in difficulty.
  3. [§4] The claim that a pluralist model such as Kaleido "cannot overcome the limitation imposed by the asystematicity of normative truth on its capacity to move beyond its training data" is presented as a direct consequence of asystematicity, but the example only shows that one system trained on synthetic GPT-4 data does not extrapolate via systematicity. It does not test whether other statistical or non-statistical mechanisms could improve coverage. This is an empirical premise that needs support or explicit identification as an open empirical question.
minor comments (3)
  1. [§3] The text refers to "autoregressive LMMs" where the intended term is presumably "LLMs"; this should be corrected.
  2. [§4] In the discussion of Sorensen et al., the citation appears as "Sorensen et al,. 2024" with a stray comma before the period, and the model name "Value Kaleidoskope" is spelled differently from the standard "Kaleidoscope" in the cited paper.
  3. [§5] The claim that Kaleido's relevance scores "measure the statistical relevance of a type of consideration to a type of situation" while importance "goes significantly beyond such merely statistical relevance" is asserted rather than argued; a sentence explaining why importance cannot be approximated by richer contextual features would help the reader assess this step.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the paper's argument is a conditional application of an independently sourced pluralist premise, with no fitted parameters or self-citation chain doing the work.

full rationale

The paper's central inference is conditional: if normative truths are asystematic, then LLMs cannot leverage the systematicity of truth for self-completion and self-correction, and so comprehensive modelling is to that extent harder. This is an application of the definition of systematicity to a stated mechanism (inferential redundancy), not a disguised restatement of the conclusion. The asystematicity premise is drawn from an independent value-pluralist literature (Berlin, Williams, Nagel, Chang, Wiggins and Williams, Chappell), not from the author's own prior work. The author's self-citations appear only as elaborations of Williamsian themes and are not load-bearing for the main inference. The paper also explicitly concedes that 'alternative ways for LLMs to move beyond their training data may yet emerge,' showing that the conclusion is not a fixed-point or fitted claim but a limited conditional. The weakest point is the bridging premise that LLM progress beyond training data must exploit inferential redundancy; this is asserted rather than empirically established, but that is an under-argument risk, not circularity. No equations, fitted parameters, or benchmark 'predictions' are present whose outputs are pre-encoded by their inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No numbers are fitted and no new entities are postulated. The argument is built from domain assumptions imported from value theory and from claims about LLM training, listed above.

assumptions (3)
  • domain assumption Value pluralism is correct: there is a plurality of irreducibly distinct, incompatible, and incommensurable values, so normative truths are at least partly asystematic.
    Imported in Section 4 from Berlin, Williams, Nagel, Chang, and others. The paper's normative conclusion depends on this contested thesis; the author is careful to argue conditionally ('if pluralists are right').
  • domain assumption LLMs can only leverage systematicity indirectly, and their primary route to self-completion and self-correction is through consistency and coherence of the target domain.
    Section 3 argues that SSL, SFT, and RLHF do not directly train for consistency and coherence, and Section 4 assumes that a lack of systematicity therefore hampers progress. This is the mechanism that turns asystematicity into difficulty.
  • domain assumption Practical deliberation has an irreducibly first-personal dimension, so even complete impersonal advice leaves a first-personal question for the agent.
    Section 5, drawing on Williams and Chang, distinguishes the impersonal 'What is to be done?' from the first-personal 'What should I do?' and argues the latter cannot be outsourced. This supports the human-agency conclusion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can AI Rely on the Systematicity of Truth? The Challenge of Modelling Normative Domains." pith.science (2026). https://pith.science/paper/PHYDC2JF

@misc{pith2026250709676,
  author       = {Pith},
  title        = {Pith review of: Can AI Rely on the Systematicity of Truth? The Challenge of Modelling Normative Domains},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PHYDC2JF}},
  note         = {Machine review of arXiv:2507.09676}
}
read the original abstract

A key assumption fuelling optimism about the progress of large language models (LLMs) in accurately and comprehensively modelling the world is that the truth is systematic: true statements about the world form a whole that is not just consistent, in that it contains no contradictions, but coherent, in that the truths are inferentially interlinked. This holds out the prospect that LLMs might in principle rely on that systematicity to fill in gaps and correct inaccuracies in the training data: consistency and coherence promise to facilitate progress towards comprehensiveness in an LLM's representation of the world. However, philosophers have identified compelling reasons to doubt that the truth is systematic across all domains of thought, arguing that in normative domains, in particular, the truth is largely asystematic. I argue that insofar as the truth in normative domains is asystematic, this renders it correspondingly harder for LLMs to make progress, because they cannot then leverage the systematicity of truth. And the less LLMs can rely on the systematicity of truth, the less we can rely on them to do our practical deliberation for us, because the very asystematicity of normative domains requires human agency to play a greater role in practical thought.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

11 extracted references · 8 canonical work pages

  1. [1]

    What if Dario Amodei Is Right About A.I.?

    Abela, P. (2006). The Demands of Systematicity: Rational Judgment and the Structure of Nature. In G. Bird (Ed.), A Companion to Kant (pp. 408–422). Blackwell. Amodei, D. (2024). “What if Dario Amodei Is Right About A.I.?.” Interview by Ezra Klein. The Ezra Klein Show, New York Times Opinion, April 12,

  2. [8]

    17532 Queloz, M

    16995/ pp. 17532 Queloz, M. (2025). The ethics of conceptualization: Tailoring thought and language to need. Oxford Uni- versity Press. https:// doi. org/

  3. [9]

    Leibniz and the Concept of a System

    0001 Rescher, N. (1981). “Leibniz and the Concept of a System.” In Leibniz’s Metaphysics of Nature: A Group of Essays, 29–41. Dordrecht: Springer. Rescher, N. (1979). Cognitive Systematization: A Systems Theoretic Approach to a Coherentist Theory of Knowledge. Blackwell. Rescher, N. (2000). Kant and the Reach of Reason: Studies in Kant’s Theory of Rationa...

  4. [10]

    LLMs are Not Just Next Token Predictors

    1038/ s41598- 025- 86510-0 Downes, S. M., Forber, P., & Grzankowski, Alex. (2024). “LLMs are Not Just Next Token Predictors.” arXiv arXiv:

  5. [11]

    Le concept de système de Leibniz à Condillac

    Jahrhundert. Edited by Jürgen Blühdorn and Joachim Ritter, 63–88. Frankfurt am Main: Klostermann. Vickers, P. (2013). Understanding Inconsistent Science. Oxford University Press. Vieillard-Baron, J.-L. (1975). “Le concept de système de Leibniz à Condillac.” In Akten des II. Interna- tionalen Leibniz-Kongresses Hannover, 17.-22. Juli

  6. [19]

    Striking a Balance: Alleviating Inconsistency in Pre-trained Models for Symmetric Classification Tasks

    Jahrhundert. Edited by Jürgen Blühdorn and Joachim Ritter, 99–122. Frankfurt am Main: Klostermann. Kekes, J. (1993). The Morality of Pluralism. Princeton University Press. Kitcher, P. (1986). Projecting the Order of Nature. In R. Butts (Ed.), Kant’s Philosophy of Material Nature (pp. 201–235). D. Reidel. Kretzmann, Norman, & Eleonore Stump. (1989). The Ca...

  7. [1972]

    Aurel Thomas Kolnai

    Edited by Kurt Müller, Heinrich Schep- ers and Wilhelm Totok, 97–103. Wiesbaden: F. Steiner. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., ... & Fedus, W. (2022). Emergent abilities of large language models. arXiv preprint arXiv:2206.07682. Wiggins, D., & Williams, B. (1978). “Aurel Thomas Kolnai.” In Ethics, Value and Reality: Aure...

  8. [2024]

    A general language assistant as a laboratory for alignment

    https:// www. nytim es. com/ 2024/ 04/ 12/ opini on/ ezra- klein- podca st- dario- amodei. html Askell, A., Bai, Y, Chen, A., Drain, D, Ganguli, D, Henighan, T., Jones, A., Joseph, N., Mann, B, & Das- Sarma, Nova. (2021). “A general language assistant as a laboratory for alignment.” arXiv preprint arXiv:

Show all 11 references
  1. [2112]

    Epistemology of Artificial Intelligence

    00861. Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., ... & Kaplan, J. (2022). Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073. Beisbart, C. (forthcoming-a). “Epistemology of Artificial Intelligence.” In The Stanford Enc...

  2. [2408]

    Ideal Observer

    04666. Dupré, J. (1995). The Disorder of Things: Metaphysical Foundations of the Disunity of Science. Harvard University Press. Elazar, Y., Kassner, N., Ravfogel, S., Ravichander, A., Hovy, E., Schütze, H., & Goldberg, Y. (2021). Measuring and Improving Consistency in Pretrain...

  3. [2410]

    Value Pluralism

    02205. Losano, M. G. (1968). Sistema e struttura nel diritto, vol. 1: Dalle origini alla scuola storica. Turin: Giuffrè. MacIntyre, A. C. (2007). After Virtue: A Study in Moral Theory (3rd ed.). University of Notre Dame Press. Mason, E. (2023). “Value Pluralism.” In The Stanfo...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.