REVIEW 3 major objections 5 minor 22 references
Justifications for Democratizing AI Alignment and Their Prospects
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Neither expert rule nor public rule can legitimate AI alignment by itself; the paper argues for hybrid institutions combining expert judgment, participatory input, and safeguards against AI monopolies.
desk verdict A clean conceptual map of democratic vs. epistocratic justifications that earns its place, but its hybrid conclusion leans on an exit-based coercion account that needs explicit defense. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by the 'justificatory gap': normative and metanormative uncertainty — reasonable disagreement about which reasons matter, how strongly, and whether any unique normative truth exists — eliminates the theoretical justification that would legitimize expert-chosen alignment constraints, leaving a space that democratic approaches try to fill with political justification. Within that frame, the load-bearing mechanism is the coercion analysis with its background conditions (the account the paper draws on calls them 'tempering factors'): an aligned AI only actually coerces a user when the user cannot freely act on their own desires, for instance because the AI is the de facto or de jure standard for some purpose and switching is unavailable or too costly. This mechanism is what lets the paper defuse the democratic justification, since background conditions can be controlled — multiple differently aligned AIs can be kept available — and it is also what lets epistocratic approaches answer the coercion worry. A secondary piece of machinery is the three-step decomposition of producing normative constraints (identify the relevant reasons, measure their relative strength, aggregate them into overall verdicts), which shows how hybrid approaches can partition the labor between experts and the affected public scenario by scenario, step by step, or within a single step.
What would settle it
Examine a real deployment in which one AI is the only reasonable gateway to an essential function — for example, a government service whose sole interface is an aligned chatbot — and check whether users can realistically switch to an alternative or act on their own. Under the paper's own criteria, users without a free alternative are genuinely coerced, so finding that such coercion occurs routinely even in markets with nominally competing AI products would falsify the claim that keeping alternatives available neutralizes the coercion-based justification for democratic input. Conversely, deployments where alternatives genuinely exist and users report no restraint would support the paper's conclusion.
Extended reading notes
Core claim
The paper's central claim is that neither purely epistocratic nor purely democratic approaches to the normative problem of AI alignment are sufficient on their own, and that the coercion-based justification for democratic approaches is weaker than its proponents assume. First, because people reasonably disagree about what is right and even about whether a unique normative truth exists, normative and metanormative uncertainty removes the sure theoretical justification that would legitimize any imposed set of constraints; this justificatory gap is what democratic participation aims to fill with political justification. Second, the most promising non-instrumental justification for democracy — that expert-driven alignment illegitimately coerces users — must establish four propositions: that users can be coerced through an AI's alignment, that such coercion would be unjustified, that democratic procedures can produce a justification legitimizing it, and that epistocratic approaches cannot prevent it. The authors argue none of these comes for free: coercion only occurs under background conditions (a user who can freely use another AI or act on their own is not coerced), democratic procedures face a bootstrapping problem and risk settling on a minimal normative denominator, and epistocratic approaches can invoke decision rules such as maximise expected choiceworthiness to justify their choices under uncertainty while also preventing coercion by keeping any single AI from becoming the de facto or de jure standard. The conclusion is not that democratic participation is worthless but that context-sensitive hybrid frameworks, combining expert judgment with participatory input and institutional safeguards against AI monopolization, are the most suitable path.
Load-bearing premise
The entire argument leans on a specific, contested account of coercion, namely that a user is not coerced by an AI's noncompliance if they can freely use a different AI or act on their own; if one instead accepts a broader account of coercion (structural coercion, or mere subjection to another's will), the background-conditions response fails and the democratic justification for alignment retains its full force.
Editorial extensions
If this is right
- Democratic proposals for AI alignment must pass a four-part test — users can be coerced through alignment, such coercion would be illegitimate, democratic procedures can legitimize it, and epistocratic approaches cannot prevent it — and none of the four parts comes for free.
- Crowdsourcing, RLHF, and constitutional AI should be treated as technical implementation techniques: they are silent on who determines the normative constraints, so debates that treat them as inherently democratic settle nothing about the normative problem.
- Ensuring that no single AI becomes the de facto or de jure standard is not merely a market-structure concern; it is the concrete mechanism that prevents alignment from being coercive in the first place.
- Decision rules designed for normative uncertainty, such as maximise expected choiceworthiness, give epistocratic approaches a practical justification for their chosen constraints, so democratic proponents must engage those rules rather than assume expertise cannot justify.
- The question the paper leaves open is contextual: which aspects of the normative problem should be settled by expert knowledge, which by democratic input, and under what institutional conditions — the answer is expected to vary by application context.
Reading between the lines
- The analysis implies an empirical prediction: the force of the democratic, anti-coercion case should track market structure, being strong where AI use is effectively mandatory (government services, default preinstalled assistants, employer-mandated tools) and weak where genuine exit options exist; measuring user exit options across deployment contexts and correlating them with reported value impos
- The 'lowest common normative denominator' objection suggests a testable failure mode: democratic processes among diverse stakeholders may systematically produce alignment constraints too thin to regulate AI behavior, forcing a return to expert content; a deliberation experiment on concrete alignment scenarios could measure how much common ground actually emerges.
- The paper analyzes coercion of users, but its own definition of affected stakeholders includes people only indirectly touched by an AI's behavior; extending the coercion analysis to bystanders and third parties would clarify when democratic input matters beyond the immediate user-AI interaction.
- Because the paper notes that normative uncertainty is not evenly distributed, a design principle for the hybrid frameworks it endorses suggests itself: use democratic input where disagreement runs deep and expert deference where near-consensus exists, and test whether such a division of labor is stable in concrete alignment case studies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses the normative problem of AI alignment, defined as the question of what normative constraints an AI system should satisfy, as opposed to the technical problem of implementing them. It contrasts democratic approaches, in which affected stakeholders determine the constraints, with epistocratic approaches, which defer to normative experts. The authors distinguish instrumental justifications (better outcomes, wisdom of crowds) from non-instrumental justifications (preventing illegitimate authority or coercion). They argue that normative and metanormative uncertainty create a justificatory gap that democratic approaches try to fill with political justification, but they identify unresolved burdens for the coercion-based justification, especially the claim that coercion only occurs when an AI is the de facto or de jure standard. They conclude that neither purely democratic nor purely epistocratic approaches may suffice on their own, pointing toward hybrid frameworks with anti-monopoly safeguards.
Significance. If the analysis holds, the paper makes a useful contribution by clearing away a common conflation of technical and normative alignment and by laying out the exact argumentative burdens that a coercion-based democratic justification must meet. Its taxonomy of instrumental versus non-instrumental justifications, and its list of four necessary propositions for the coercion argument, are genuinely helpful for future work. The paper engages fairly with opposing views, including epistemic decision rules and anti-monopoly safeguards, and it explicitly acknowledges tensions it does not resolve. The main value is as a map of the dialectical terrain rather than as a positive theory; its conclusions are conditional and forward-looking. It would be more significant if the coercion account were defended and if the transition from identified burdens to the hybrid conclusion were made explicit.
major comments (3)
- [4 (coercion, vegan-supermarket example)] The background-conditions response to coercion is load-bearing for proposition (iv) and for the Section 5 conclusion, but it relies on an exit-based conception of coercion that is asserted rather than defended. The claim that a user is not coerced when she can 'freely go elsewhere' works only on a narrow, negative-liberty account of coercion; on broader accounts, such as structural coercion or domination as subjection to another's will, the availability of a differently aligned AI does not remove the coercive character of the noncompliant AI. The paper itself concedes the 'great opportunity cost' caveat but never specifies when opportunity costs become coercion. Since the conclusion that epistocratic approaches can 'take the sting out of' the non-instrumental justification depends on this account, the authors need to defend it against alternative theories of coercion or explicitly restrict the conclusion to cases of genuinely costless exit.
- [4 (epistocratic response, decision rules)] The suggestion that epistocratic approaches can justify coercion through decision rules such as maximise expected choiceworthiness assumes that a practical justification for choosing certain normative constraints transfers directly to the coercion of users who do not endorse those constraints. That transfer is not trivial: a decision rule may give designers a reason to implement a constraint without giving the affected user a justification for being subjected to it, particularly under the normative and metanormative uncertainty the paper itself emphasizes. Because this undercutting move contributes to the conclusion that epistocratic approaches can handle the coercion objection, the transfer claim needs explicit defense.
- [5 (Conclusion)] The central conclusion that 'neither purely epistocratic nor purely democratic approaches ... may be sufficient on their own' is not established by the preceding analysis. Sections 3 and 4 show at most that instrumental justifications require further empirical work and that the coercion-based non-instrumental justification faces unanswered objections. That is a directed challenge to democratic proponents; it does not by itself show that pure epistocratic or pure democratic solutions are insufficient, nor does it show that hybrid frameworks are superior. As it stands, the hybrid conclusion is a research prediction rather than an implication of the argument. Please either mark it explicitly as a forward-looking conjecture or supply an argument that the identified burdens cannot be met by either pure approach.
minor comments (5)
- [Abstract] There is a typographical artifact in the abstract: 'aimtofillthroughpoliticalratherthantheoreticaljustification' should be 'aim to fill through political rather than theoretical justification'. The author affiliation also contains a typo: 'Saarbücken' should be 'Saarbrücken'.
- [1] The discussion of Schuster and Kilov would benefit from a brief statement of the disagreement: the authors claim that Schuster and Kilov conflate the technical and normative problems, but the paragraph asserts this rather than showing where exactly the conflation occurs.
- [3] The examples 'Brexit, Trump, the climate crisis' are too compressed to carry the argument. If these are meant as cases where people vote against their own interests, at least one example should be spelled out, or they should be replaced with less contested cases.
- [4] The clarification that 'the primary coercer is not the AI itself but the person or organisation that defines the normative constraints' introduces an agency question that is never revisited: in the later examples (the AI assistant refusing to buy meat), the corporate and technical chain of responsibility is not identified. A sentence about who, in the authors' view, defines constraints in current deployment contexts would help.
- [2] The paper moves from observed 'reasonable disagreement' to 'normative uncertainty' as the rational response. This inference would benefit from a footnote distinguishing epistemic uncertainty from the possibility that disagreement merely reflects different evaluative perspectives; not all readers will accept that persistent disagreement entails uncertainty about the normative facts.
Circularity Check
No significant circularity: the argument is a critical philosophical analysis whose conclusion follows from the stated objections and conditions, not from definitional or self-citational construction.
full rationale
The paper's central conclusion—that neither purely epistocratic nor purely democratic approaches may be sufficient and that hybrid frameworks merit consideration—is reached through an open-ended critique of four required propositions for the coercion-based non-instrumental justification of democratic alignment. Each proposition is treated as a burden of proof for proponents of democratic approaches, not as an assumption built into the definitions. The paper explicitly leaves open whether the propositions can be established and identifies objections and alternative resources for epistocratic approaches, such as background-condition control and decision rules under normative uncertainty. No equation or formal derivation is present, and no fitted parameter is renamed as a prediction. The self-citations, notably Baum (2025) for a taxonomy of AI alignment, serve as background classification rather than load-bearing premises, and the argument does not depend on an unverified uniqueness theorem or on the authors' prior work to force its conclusion. The coercion discussion invokes a contested exit-based account of coercion, but this is an ordinary substantive philosophical assumption subject to external criticism, not a case of the paper defining its conclusion into existence. Accordingly, no circular step meeting the required evidentiary standard can be identified, and the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Normative and metanormative uncertainty is a necessary condition for the possibility of illegitimate authority or coercion.
- domain assumption Coercion is prima facie wrong.
- ad hoc to paper A user is coerced by an AI's noncompliance only if the AI is the de facto or de jure standard, i.e., the user cannot freely go elsewhere.
- domain assumption Decision rules like maximize expected choiceworthiness can provide a practical justification for normative constraints under uncertainty.
Cite this review
Pith. "Pith review of Justifications for Democratizing AI Alignment and Their Prospects." pith.science (2026). https://pith.science/paper/MDCJWC4B
@misc{pith2026250719548,
author = {Pith},
title = {Pith review of: Justifications for Democratizing AI Alignment and Their Prospects},
year = {2026},
howpublished = {\url{https://pith.science/paper/MDCJWC4B}},
note = {Machine review of arXiv:2507.19548}
}
read the original abstract
The AI alignment problem comprises both technical and normative dimensions. While technical solutions focus on implementing normative constraints in AI systems, the normative problem concerns determining what these constraints should be. This paper examines justifications for democratic approaches to the normative problem -- where affected stakeholders determine AI alignment -- as opposed to epistocratic approaches that defer to normative experts. We analyze both instrumental justifications (democratic approaches produce better outcomes) and non-instrumental justifications (democratic approaches prevent illegitimate authority or coercion). We argue that normative and metanormative uncertainty create a justificatory gap that democratic approaches aim to fill through political rather than theoretical justification. However, we identify significant challenges for democratic approaches, particularly regarding the prevention of illegitimate coercion through AI alignment. Our analysis suggests that neither purely epistocratic nor purely democratic approaches may be sufficient on their own, pointing toward hybrid frameworks that combine expert judgment with participatory input alongside institutional safeguards against AI monopolization.
Reference graph
Works this paper leans on
-
[1]
et al.: A General Language Assistant as a Laboratory for Alignment
Askell, A. et al.: A General Language Assistant as a Laboratory for Alignment. (2021), arXiv:2112.00861
arXiv 2021
-
[2]
et al.: The Moral Machine experiment
Awad, E. et al.: The Moral Machine experiment. Nature, 563, 59–64. (2018)
work page 2018
-
[3]
et al.: Training a Helpful and Harmless Assistant with Reinforcement Learn- ing from Human Feedback
Bai, Y. et al.: Training a Helpful and Harmless Assistant with Reinforcement Learn- ing from Human Feedback. (2022), arXiv:2204.05862
arXiv 2022
-
[4]
et al.: Constitutional AI: Harmlessness from AI Feedback
Bai, Y. et al.: Constitutional AI: Harmlessness from AI Feedback. (2022), arXiv:2212.08073
arXiv 2022
-
[5]
Baum, K.: Disentangling AI Alignment: A Structured Taxonomy Beyond Safety and Ethics. In: Steffen, B. (ed.) AISoLA 2024 Post-Proceedings. Lecture Notes in Computer Science. Springer, Cham (forthcoming). arXiv preprint arXiv:2506.06286 (2025)
arXiv 2025
-
[6]
AI & Society, 35(1), 165– 176 (2020)
Baum, S.: Social Choice Ethics in Artificial Intelligence. AI & Society, 35(1), 165– 176 (2020). https://doi.org/10.1007/s00146-017-0760-1 14 A. Steingrüber and K. Baum
-
[7]
Christiano, P. F. et al.: Deep reinforcement learning from human preferences. Ad- vances in neural information processing systems, 30. (2017)
work page 2017
-
[8]
The Stanford Encyclopedia of Philosophy
Christiano, T., Bajaj, S.: Democracy. The Stanford Encyclopedia of Philosophy. (2024), https://plato.stanford.edu/archives/sum2024/entries/democracy/
work page 2024
Show all 22 references
-
[9]
Minds and Machines, 2(5), 411–437
Gabriel, I.: Artificial intelligence, values, and alignment. Minds and Machines, 2(5), 411–437. (2020)
2020
-
[10]
Philosophical Studies, (2025)
Gabriel, I., Keeling, G.: A matter of principle? AI alignment as the fair treatment of claims. Philosophical Studies, (2025)
2025
-
[11]
E., Spiekermann, K.: An Epistemic Theory of Democracy, Oxford Uni- versity Press, Oxford
Goodin, R. E., Spiekermann, K.: An Epistemic Theory of Democracy, Oxford Uni- versity Press, Oxford. (2018)
2018
-
[12]
T., Papyshev, G., Wong, J
Huang, L. T., Papyshev, G., Wong, J. K.: Democratizing AI alignment: from au- thoritarian to democratic AI ethics. AI and Ethics, 5, 11–18. (2025)
2025
-
[13]
et al.: Can Machines Learn Morality? The Delphi Experiment
Jiang, L. et al.: Can Machines Learn Morality? The Delphi Experiment. (2022), arXiv:2110.07574
2022 arXiv
-
[14]
W., Hooker, J., Donaldson, T.: Taking Principles Seriously: A Hybrid Ap- proach to Value Alignment in Artificial Intelligence
Kim, T. W., Hooker, J., Donaldson, T.: Taking Principles Seriously: A Hybrid Ap- proach to Value Alignment in Artificial Intelligence. Journal of Artificial Intelligence Research, 70, 871–890. (2021)
2021
-
[15]
FAccT ’25: Proceedings of the 2025 ACM Conference on Fairness, Ac- countability, and Transparency, 2671 - 2681
Kneer, M., Viehoff, J.: The Hard Problem of AI Alignment: Value Forks in Moral Judgment. FAccT ’25: Proceedings of the 2025 ACM Conference on Fairness, Ac- countability, and Transparency, 2671 - 2681. (2025)
2025
-
[16]
Kolodny,N.:ThePeckingOrder,HarvardUniversityPress,Cambridge,MA.(2023)
2023
-
[17]
MacAskill, W., Bykvist, K., Ord, T.: Moral Uncertainty, Oxford University Press, Oxford. (2020)
2020
-
[18]
Oxford University Press, Oxford
Raz, J.: The Morality of Freedom. Oxford University Press, Oxford. (1986)
1986
-
[19]
AI and Ethics, 5, 3727–3741
Riesen, E., Boespflug, M.: Aligning with ideal values: a proposal for anchoring AI in moral expertise. AI and Ethics, 5, 3727–3741. (2025)
2025
-
[20]
Philosophy and Public Affairs, 32(1), 2–35
Ripstein, A.: Authority and Coercion. Philosophy and Public Affairs, 32(1), 2–35. (2004)
2004
-
[21]
AI & Society
Schuster, N., Kilov, D.: Moral disagreement and the limits of AI value alignment: a dual challenge of epistemic justification and political legitimacy. AI & Society. (2025)
2025
-
[22]
et al.: Fine-Tuning Language Models from Human Preferences
Ziegler, D. et al.: Fine-Tuning Language Models from Human Preferences. (2019), arXiv:1909.08593
2019 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.