REVIEW 2 major objections 4 minor 49 references
The essay claims that alignment-as-optimization cannot distinguish error from invention, because significance is not a scalar, yet this apparatus has inherited the authority to set legitimate language.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 07:00 UTC pith:IH5Q6ROC
load-bearing objection A serious critical essay on LLM alignment as 'optimization culture' — the central limit-claim is asserted rather than derived, but on its own humanistic terms the argument is strong enough to merit real peer review. the 2 major comments →
Optimization Is Not All You Need
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that an optimization procedure can measure how improbable a piece of generated text is, but it cannot tell whether that unlikelihood is error or invention. Significance is not a scalar property: surprisal is a computable property of tokens, while meaning arises only in an event of interpretation. Because reward models, loss functions, benchmarks, and system prompts can only register deviation from a statistical baseline, they must classify metaphor and misinformation, confusion and invention, slop and the start of a new form as the same thing—and they suppress both. Within half a decade, this apparatus has taken over the authority once held by academies, schoolrooms, gra
What carries the argument
The carrying mechanism is the language-model stack—pre-training objective, decoding and sampling, preference tuning and reward modeling, benchmarking, and interface design. Its unifying operation is scalarization: the collapse of multidimensional judgment into a single number such as loss, reward, or benchmark score. The paper's load-bearing distinction is surprisal versus significance: surprisal is a token-level, computable quantity; significance is an event of interpretation that no scalar indexes. The stack can compute the former and therefore must treat the latter as if it did not exist.
Load-bearing premise
The load-bearing premise is the in-principle claim, not an empirical result, that no scalar can represent whether an improbable continuation is error or invention; the paper itself concedes (Section 6, note 19) that it offers no empirical study of the distributional tail in current models, so if a reward model or learned judge could be trained to make that distinction, the argument would collapse.
What would settle it
A controlled experiment in which a reward model is trained on a corpus of deviations—fabricated citations, metaphors, genre-shifts, confabulations—and then asked to classify new deviations as error or invention; if any model or learned judge reliably outperforms a baseline that simply flags all anomalies, across unseen contexts, the paper's limit-claim is false.
If this is right
- Hallucination and invention are formally indistinguishable to the reward model, so suppressing misinformation also suppresses metaphor, style, and the improbable-but-useful connection.
- Because the authority to define legitimate language has moved from contestable institutions to scalar metrics, there is no appeals process: the regime cannot be argued with.
- Training a model on 'weird' or creative text does not restore variance; it teaches the model to simulate the surface of a swerve, turning surprise into a genre and a category to optimize.
- Chain-of-thought reasoning makes the problem worse: the linear trace is a script that pre-empts the uncaused deviation, and because traces are often unfaithful, transparency becomes an alibi for opacity.
- Optimization is sufficient for tasks with verifiable targets such as mathematics and code, but mismatched to humanistic inquiry, where variance is a medium of thought rather than a defect.
Where Pith is reading between the lines
- If the limit-claim holds, the relevant design target is not more capable reward models but contestable evaluation: interfaces, review processes, or provenance regimes that let readers argue with the standard instead of merely complying with it.
- The argument predicts a testable asymmetry: as preference tuning scales, the population-level diversity of model outputs should decline even when individual outputs are rated as creative—current studies hint at this, but the paper does not itself perform the measurement.
- An 'opt-out' variant of the stack—where users control sampling, inspect system prompts, and can appeal output classifications—would test whether a less foreclosing ecology can survive economic incentives; the paper leaves that question open.
- The essay implicitly offers a negative definition of alignment: a model whose outputs sustain reinterpretation and disagreement would be judged misaligned, meaning alignment success and hermeneutic value may be inverse.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This essay argues that the alignment of large language models should be understood not (only) as an engineering achievement but as an expression of 'optimization culture'—a late-modern conviction that measurable improvement along predefined axes exhausts the question of value. Tracing the logic of optimization through pretraining, decoding, preference tuning, benchmarking, interfaces, and reasoning traces, the authors contend that the resulting 'stack' imposes a normative linguistic standard that suppresses variance, treating improbable output as defect rather than as potential invention. The central limit-claim is that an optimization procedure can measure how improbable a continuation is but cannot tell whether that unlikelihood is error or invention, because 'significance is an event of interpretation, and no scalar indexes it.' The paper concludes that the apparatus of linguistic authority has been handed to reward models and benchmarks—an apparatus that 'executes the office of judgment with no capacity for judging.' It proposes 'controlled variance' as a critical orientation and draws on examples from GPT-2, contemporary aligned assistants, and empirical studies of creative homogenization.
Significance. If the central claim holds, the essay makes an important intervention: it reframes alignment as a normative and disciplinary project rather than a neutral safety measure, and it gives humanistic study a concrete stake in the evaluation of generative systems. The paper is well-written, historically informed, and productively interdisciplinary. It cites relevant empirical work—Wenger and Kenett on creative homogeneity, Koch et al. on benchmark concentration, Kirk et al. on RLHF diversity loss—and its close readings of GPT-2 samples are genuinely illuminating. The manuscript is also transparent about some of its own limits (e.g., n.19). However, the central in-principle impossibility claim is asserted rather than demonstrated, and the essay's own account of humanistic judgment is too thin to sustain the asymmetry it posits between optimization procedures and interpretive communities. The essay could be made rigorous by either defending the limit-claim philosophically or weakening it to a claim about current systems and practices.
major comments (2)
- [§1, §5 ('ouroboros'), §6] The central claim that no optimization procedure can distinguish error from invention is load-bearing and is not argued for. The paper states 'significance is an event of interpretation, and no scalar indexes it' (§1) and 'A learned judge does not escape this. Trained on precedent, it can only rate a deviation against what it knows; the deviation that revalues precedent is the one it cannot recognize' (§6). These are in-principle limit-claims. The evidence offered—the unicorn example, the Claude Opus 4.7 response, the Lisbeth sample, Wenger and Kenett's findings—establishes that current models and reward models fail in these ways, not that no learned judge could condition on genre, user intent, and a sufficiently rich history of judgments. Similarly, the 'ouroboros argument' in §5 ('an LLM trained on weirdness... learns to simulate the surface of a break') is another unargued impossibili
- [§5, §6] The paper's own alternative—humanistic interpretation and 'controlled variance'—does not explain how human critics escape the problem it attributes to learned judges. The text asserts that 'a poetics can' determine whether deviation functions as error or invention (§5) and that meaning is 'an event of reading, structured by delay and return' (§6), but it never gives an account of how human readers recognize a deviation that revalues precedent, nor why a learned judge trained on human judgments could not in principle do the same. If the intended asymmetry is about temporality and sociality—interpretation happening over time, across communities—rather than about scalar representation per se, the argument should say so explicitly. As written, the essay risks proving too much: if no procedure can recognize the deviation that revalues precedent, then humanistic judgment is either left unexpla
minor comments (4)
- [§3] The historical contrast between the old regime of 'judges who could be argued with' and the new unappealable stack is idealized. For most speakers, academies, examiners, and schoolrooms were not open to appeal; the claim needs qualification or historical evidence. This does not affect the central argument, but it weakens a supporting contrast.
- [§4, fn 15] The single Claude Opus 4.7 output is presented as representative without systematic sampling. A short acknowledgment of the example's illustrative status, or a small set of comparable outputs, would strengthen the point.
- [References] The reference to Milička et al. appears with a mis-encoded character, and the URL is given as an HTML page rather than an abstract or PDF link. Please fix the encoding and provide stable URLs.
- [Figure 2] The close reading of Figure 2 is detailed and depends on the reader seeing the sample. Ensure the figure is legible in the final version and ideally include the full output text in the caption or an appendix.
Circularity Check
No significant circularity: the central claim is a stated philosophical limit-claim, not an output forced by fitted inputs or self-citation.
full rationale
This paper does not contain a derivation chain of the kind that could be circular. It is an interpretive essay: it offers no equations, fits no parameters, and makes no quantitative prediction that is then 'confirmed' from its own inputs. The central thesis—'an optimization procedure can measure how improbable a continuation is; it cannot measure whether the improbability signifies'—is asserted as a claim about the nature of significance ('significance is an event of interpretation, and no scalar indexes it'), supported by external empirical studies (Anderson et al. 2024; Wenger and Kenett 2025; Sui et al. 2024), close readings of GPT-2 outputs, and genealogical arguments about audit culture and legitimate language. The authors' own prior papers are cited only as background: Hua and Raley (2023) supports a claim about GPT-2 ancillary code enabling creative interfaces, and Hua and Raley (2020) is referenced as an earlier call for humanists to engage evaluation. Neither supplies the load-bearing premise. Footnote 19 explicitly concedes that no empirical tail study is provided and says the argument would not be exhausted by one; this is a limitation in evidential support, not circular reasoning. The skeptical worry that the 'cannot' claim is asserted rather than proven is a substantive objection to the argument's soundness, but an unsupported premise is not a circular step: nothing in the paper reduces, by construction or by self-citation, to its own input. Accordingly, the circularity score is 1 (at most a minor, non-load-bearing self-citation; no circularity in the central argument).
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption No scalar can index significance; significance is an event of interpretation.
- domain assumption Optimization procedures cannot in principle distinguish error from invention.
- domain assumption Pre-optimization linguistic authority was administered by contestable judges; the new regime has no audience for appeals.
- domain assumption Humanistic inquiry depends on variance as a medium of thought.
invented entities (3)
-
optimization culture
no independent evidence
-
controlled variance
no independent evidence
-
reward model prosody
no independent evidence
read the original abstract
In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering achievement; we read it instead as the newest expression of optimization culture: the conviction, older than the technology, that measurable improvement along predefined axes exhausts the question of value. Tracing that conviction through the stack-pretraining, decoding, preference tuning, benchmarking, interface-and back through its genealogy in the audit society, we arrive at the limit: an optimization procedure can measure how improbable a piece of generated text is; it cannot tell whether that unlikelihood is error or invention. A procedure that cannot make that distinction has nonetheless, within half a decade, assumed the authority to set the protocols of legitimate language. Held for centuries by academies and schoolrooms, grammars and examiners, this authority has been given over to loss functions, reward models, benchmarks, and system prompts: an apparatus that executes the office of judgment with no capacity for judging.
Figures
Reference graph
Works this paper leans on
-
[1]
Frontiers in Handwriting Recognition (ICFHR), 2014 14th International Conference on , pages=
Real-time segmentation of on-line handwritten arabic script , author=. Frontiers in Handwriting Recognition (ICFHR), 2014 14th International Conference on , pages=. 2014 , organization=
2014
-
[2]
Soft Computing and Pattern Recognition (SoCPaR), 2014 6th International Conference of , pages=
Fast classification of handwritten on-line Arabic characters , author=. Soft Computing and Pattern Recognition (SoCPaR), 2014 6th International Conference of , pages=. 2014 , organization=
2014
-
[3]
arXiv preprint arXiv:1804.09028 , year=
Estimate and Replace: A Novel Approach to Integrating Deep Neural Networks with Existing Applications , author=. arXiv preprint arXiv:1804.09028 , year=
-
[4]
Proceedings of the 16th Conference on Creativity and Cognition , year=
Homogenization Effects of Large Language Models on Human Creative Ideation , author=. Proceedings of the 16th Conference on Creativity and Cognition , year=
-
[5]
Claude's Constitution: Our Vision for Claude's Character , author=
-
[6]
arXiv preprint arXiv:2212.08073 , year=
Constitutional AI: Harmlessness from AI Feedback , author=. arXiv preprint arXiv:2212.08073 , year=
-
[7]
Mythologies , author=
-
[8]
Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages=
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? , author=. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages=. 2021 , organization=
2021
-
[9]
The Atlantic , year=
AI Has Lost Its Magic , author=. The Atlantic , year=
-
[10]
Language and Symbolic Power , author=
-
[11]
The Stack: On Software and Sovereignty , author=
-
[12]
From Memory to Written Record: England 1066--1307 , author=
-
[13]
arXiv preprint arXiv:2501.12948 , year=
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning , author=. arXiv preprint arXiv:2501.12948 , year=
-
[14]
Four Quartets , author=
-
[15]
Communications of the ACM , volume=
Datasheets for Datasets , author=. Communications of the ACM , volume=
-
[16]
Ghost Work: How Amazon, Google, and Uber Are Creating a New Global Underclass , author=
-
[17]
AI & Society , volume=
Prompting Meaning: A Hermeneutic Approach to Optimising Prompt Engineering with ChatGPT , author=. AI & Society , volume=
-
[18]
International Conference on Learning Representations , year=
The Curious Case of Neural Text Degeneration , author=. International Conference on Learning Representations , year=
-
[19]
Digital Humanities Quarterly , volume=
Playing With Unicorns: AI Dungeon and Citizen NLP , author=. Digital Humanities Quarterly , volume=. 2020 , url=
2020
-
[20]
Digital Humanities Quarterly , volume=
How to Do Things with Deep Learning Code , author=. Digital Humanities Quarterly , volume=. 2023 , url=
2023
-
[21]
International Conference on Learning Representations , year=
Understanding the Effects of RLHF on LLM Generalisation and Diversity , author=. International Conference on Learning Representations , year=
-
[22]
Discourse Networks 1800/1900 , author=
1900
-
[23]
ACM AI Letters , volume=
Why Slop Matters , author=. ACM AI Letters , volume=
-
[24]
Proceedings of the NeurIPS Datasets and Benchmarks Track , year=
Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research , author=. Proceedings of the NeurIPS Datasets and Benchmarks Track , year=
-
[25]
arXiv preprint arXiv:2307.13702 , year=
Measuring Faithfulness in Chain-of-Thought Reasoning , author=. arXiv preprint arXiv:2307.13702 , year=
-
[26]
Why Greatness Cannot Be Planned: The Myth of the Objective , author=
-
[27]
Artificial Intelligence: A Paper Symposium , publisher=
Artificial Intelligence: A General Survey , author=. Artificial Intelligence: A Paper Symposium , publisher=
-
[28]
PMLA , volume=
Toward a Diversity Stack: Digital Humanities and Diversity as Technical Problem , author=. PMLA , volume=
-
[29]
The Postmodern Condition: A Report on Knowledge , author=
-
[30]
arXiv preprint arXiv:2509.10179 , year=
Benchmark of Stylistic Variation in LLM-Generated Texts , author=. arXiv preprint arXiv:2509.10179 , year=
-
[31]
2019 , month = feb, url =
Better Language Models and Their Implications , author =. 2019 , month = feb, url =
2019
-
[32]
2019 , month = may, url =
GPT-2 Output Dataset , author =. 2019 , month = may, url =
2019
-
[33]
Advances in Neural Information Processing Systems , volume=
Training Language Models to Follow Instructions with Human Feedback , author=. Advances in Neural Information Processing Systems , volume=
-
[34]
Trust in Numbers: The Pursuit of Objectivity in Science and Public Life , author=
-
[35]
The Audit Society: Rituals of Verification , author=
-
[36]
Language Models Are Unsupervised Multitask Learners , author=
-
[37]
Advances in Neural Information Processing Systems , volume=
Direct Preference Optimization: Your Language Model Is Secretly a Reward Model , author=. Advances in Neural Information Processing Systems , volume=
-
[38]
Nature , volume=
Mathematical Discoveries from Program Search with Large Language Models , author=. Nature , volume=
-
[39]
Proceedings of the First Conference on Language Modeling , year=
Benchmarks as Microscopes: A Call for Model Metrology , author=. Proceedings of the First Conference on Language Modeling , year=
-
[40]
The Birth of Physics , author=
-
[41]
arXiv preprint arXiv:2310.13548 , year=
Towards Understanding Sycophancy in Language Models , author=. arXiv preprint arXiv:2310.13548 , year=
-
[42]
Strathern, Marilyn , journal=
-
[43]
arXiv preprint arXiv:2406.04175 , year=
Confabulation: The Surprising Value of Large Language Model Hallucinations , author=. arXiv preprint arXiv:2406.04175 , year=
-
[44]
Advances in Neural Information Processing Systems , volume=
Attention Is All You Need , author=. Advances in Neural Information Processing Systems , volume=
-
[45]
Advances in Neural Information Processing Systems , volume=
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , author=. Advances in Neural Information Processing Systems , volume=. 2022 , url=
2022
-
[46]
Computer Power and Human Reason: From Judgment to Calculation , author=
-
[47]
arXiv preprint arXiv:2501.19361 , year=
We're Different, We're the Same: Creative Homogeneity Across LLMs , author=. arXiv preprint arXiv:2501.19361 , year=
-
[48]
arXiv preprint arXiv:2309.03409 , year=
Large Language Models as Optimizers , author=. arXiv preprint arXiv:2309.03409 , year=
-
[49]
arXiv preprint arXiv:1909.08593 , year=
Fine-Tuning Language Models from Human Preferences , author=. arXiv preprint arXiv:1909.08593 , year=
Pith/arXiv arXiv 1909
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.