Pith. sign in

REVIEW 3 major objections 5 minor 22 references

Justifications for Democratizing AI Alignment and Their Prospects

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Neither expert rule nor public rule can legitimate AI alignment by itself; the paper argues for hybrid institutions combining expert judgment, participatory input, and safeguards against AI monopolies.

desk verdict A clean conceptual map of democratic vs. epistocratic justifications that earns its place, but its hybrid conclusion leans on an exit-based coercion account that needs explicit defense. read the letter →

arxiv 2507.19548 v1 pith:MDCJWC4B submitted 2025-07-24 cs.CY cs.AI

classification cs.CYcs.AI
keywords AIalignmentlegitimacydemocraticjustificationpublicreasonvalueimpositionepistocracynormativeuncertaintycoercion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks who should decide what an AI system may and may not do — the normative half of the alignment problem, as opposed to the technical half of implementing constraints. The authors weigh two answers: deferring to normative experts (epistocratic approaches) or letting everyone affected decide (democratic approaches), and they clarify that techniques like crowdsourcing, RLHF, and constitutional AI are implementation methods, compatible with either answer. Their central argument is that deep, reasonable disagreement about what is right leaves a justificatory gap: no one can prove their values are the correct ones, so expert-imposed alignment stands on weak ground, and democratic participation is meant to supply political legitimacy instead. Working through the strongest non-instrumental justification for democracy — that expert-driven alignment illegitimately coerces users — they conclude it is not decisive, because coercion depends on background conditions such as whether users can freely switch to a differently aligned AI, and because experts have their own resources, including decision rules for normative uncertainty and the ability to prevent any AI from becoming a de facto standard. The paper concludes that neither pure approach suffices on its own, pointing toward hybrid frameworks that combine expert judgment with targeted participatory input and institutional safeguards against uniform alignment and AI monopolies.

What carries the argument

The argument is carried by the 'justificatory gap': normative and metanormative uncertainty — reasonable disagreement about which reasons matter, how strongly, and whether any unique normative truth exists — eliminates the theoretical justification that would legitimize expert-chosen alignment constraints, leaving a space that democratic approaches try to fill with political justification. Within that frame, the load-bearing mechanism is the coercion analysis with its background conditions (the account the paper draws on calls them 'tempering factors'): an aligned AI only actually coerces a user when the user cannot freely act on their own desires, for instance because the AI is the de facto or de jure standard for some purpose and switching is unavailable or too costly. This mechanism is what lets the paper defuse the democratic justification, since background conditions can be controlled — multiple differently aligned AIs can be kept available — and it is also what lets epistocratic approaches answer the coercion worry. A secondary piece of machinery is the three-step decomposition of producing normative constraints (identify the relevant reasons, measure their relative strength, aggregate them into overall verdicts), which shows how hybrid approaches can partition the labor between experts and the affected public scenario by scenario, step by step, or within a single step.

What would settle it

Examine a real deployment in which one AI is the only reasonable gateway to an essential function — for example, a government service whose sole interface is an aligned chatbot — and check whether users can realistically switch to an alternative or act on their own. Under the paper's own criteria, users without a free alternative are genuinely coerced, so finding that such coercion occurs routinely even in markets with nominally competing AI products would falsify the claim that keeping alternatives available neutralizes the coercion-based justification for democratic input. Conversely, deployments where alternatives genuinely exist and users report no restraint would support the paper's conclusion.

Watch

Extended reading notes

Core claim

The paper's central claim is that neither purely epistocratic nor purely democratic approaches to the normative problem of AI alignment are sufficient on their own, and that the coercion-based justification for democratic approaches is weaker than its proponents assume. First, because people reasonably disagree about what is right and even about whether a unique normative truth exists, normative and metanormative uncertainty removes the sure theoretical justification that would legitimize any imposed set of constraints; this justificatory gap is what democratic participation aims to fill with political justification. Second, the most promising non-instrumental justification for democracy — that expert-driven alignment illegitimately coerces users — must establish four propositions: that users can be coerced through an AI's alignment, that such coercion would be unjustified, that democratic procedures can produce a justification legitimizing it, and that epistocratic approaches cannot prevent it. The authors argue none of these comes for free: coercion only occurs under background conditions (a user who can freely use another AI or act on their own is not coerced), democratic procedures face a bootstrapping problem and risk settling on a minimal normative denominator, and epistocratic approaches can invoke decision rules such as maximise expected choiceworthiness to justify their choices under uncertainty while also preventing coercion by keeping any single AI from becoming the de facto or de jure standard. The conclusion is not that democratic participation is worthless but that context-sensitive hybrid frameworks, combining expert judgment with participatory input and institutional safeguards against AI monopolization, are the most suitable path.

Load-bearing premise

The entire argument leans on a specific, contested account of coercion, namely that a user is not coerced by an AI's noncompliance if they can freely use a different AI or act on their own; if one instead accepts a broader account of coercion (structural coercion, or mere subjection to another's will), the background-conditions response fails and the democratic justification for alignment retains its full force.

Editorial extensions

If this is right

  • Democratic proposals for AI alignment must pass a four-part test — users can be coerced through alignment, such coercion would be illegitimate, democratic procedures can legitimize it, and epistocratic approaches cannot prevent it — and none of the four parts comes for free.
  • Crowdsourcing, RLHF, and constitutional AI should be treated as technical implementation techniques: they are silent on who determines the normative constraints, so debates that treat them as inherently democratic settle nothing about the normative problem.
  • Ensuring that no single AI becomes the de facto or de jure standard is not merely a market-structure concern; it is the concrete mechanism that prevents alignment from being coercive in the first place.
  • Decision rules designed for normative uncertainty, such as maximise expected choiceworthiness, give epistocratic approaches a practical justification for their chosen constraints, so democratic proponents must engage those rules rather than assume expertise cannot justify.
  • The question the paper leaves open is contextual: which aspects of the normative problem should be settled by expert knowledge, which by democratic input, and under what institutional conditions — the answer is expected to vary by application context.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The analysis implies an empirical prediction: the force of the democratic, anti-coercion case should track market structure, being strong where AI use is effectively mandatory (government services, default preinstalled assistants, employer-mandated tools) and weak where genuine exit options exist; measuring user exit options across deployment contexts and correlating them with reported value impos
  • The 'lowest common normative denominator' objection suggests a testable failure mode: democratic processes among diverse stakeholders may systematically produce alignment constraints too thin to regulate AI behavior, forcing a return to expert content; a deliberation experiment on concrete alignment scenarios could measure how much common ground actually emerges.
  • The paper analyzes coercion of users, but its own definition of affected stakeholders includes people only indirectly touched by an AI's behavior; extending the coercion analysis to bystanders and third parties would clarify when democratic input matters beyond the immediate user-AI interaction.
  • Because the paper notes that normative uncertainty is not evenly distributed, a design principle for the hybrid frameworks it endorses suggests itself: use democratic input where disagreement runs deep and expert deference where near-consensus exists, and test whether such a division of labor is stable in concrete alignment case studies.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper addresses the normative problem of AI alignment, defined as the question of what normative constraints an AI system should satisfy, as opposed to the technical problem of implementing them. It contrasts democratic approaches, in which affected stakeholders determine the constraints, with epistocratic approaches, which defer to normative experts. The authors distinguish instrumental justifications (better outcomes, wisdom of crowds) from non-instrumental justifications (preventing illegitimate authority or coercion). They argue that normative and metanormative uncertainty create a justificatory gap that democratic approaches try to fill with political justification, but they identify unresolved burdens for the coercion-based justification, especially the claim that coercion only occurs when an AI is the de facto or de jure standard. They conclude that neither purely democratic nor purely epistocratic approaches may suffice on their own, pointing toward hybrid frameworks with anti-monopoly safeguards.

Significance. If the analysis holds, the paper makes a useful contribution by clearing away a common conflation of technical and normative alignment and by laying out the exact argumentative burdens that a coercion-based democratic justification must meet. Its taxonomy of instrumental versus non-instrumental justifications, and its list of four necessary propositions for the coercion argument, are genuinely helpful for future work. The paper engages fairly with opposing views, including epistemic decision rules and anti-monopoly safeguards, and it explicitly acknowledges tensions it does not resolve. The main value is as a map of the dialectical terrain rather than as a positive theory; its conclusions are conditional and forward-looking. It would be more significant if the coercion account were defended and if the transition from identified burdens to the hybrid conclusion were made explicit.

major comments (3)
  1. [4 (coercion, vegan-supermarket example)] The background-conditions response to coercion is load-bearing for proposition (iv) and for the Section 5 conclusion, but it relies on an exit-based conception of coercion that is asserted rather than defended. The claim that a user is not coerced when she can 'freely go elsewhere' works only on a narrow, negative-liberty account of coercion; on broader accounts, such as structural coercion or domination as subjection to another's will, the availability of a differently aligned AI does not remove the coercive character of the noncompliant AI. The paper itself concedes the 'great opportunity cost' caveat but never specifies when opportunity costs become coercion. Since the conclusion that epistocratic approaches can 'take the sting out of' the non-instrumental justification depends on this account, the authors need to defend it against alternative theories of coercion or explicitly restrict the conclusion to cases of genuinely costless exit.
  2. [4 (epistocratic response, decision rules)] The suggestion that epistocratic approaches can justify coercion through decision rules such as maximise expected choiceworthiness assumes that a practical justification for choosing certain normative constraints transfers directly to the coercion of users who do not endorse those constraints. That transfer is not trivial: a decision rule may give designers a reason to implement a constraint without giving the affected user a justification for being subjected to it, particularly under the normative and metanormative uncertainty the paper itself emphasizes. Because this undercutting move contributes to the conclusion that epistocratic approaches can handle the coercion objection, the transfer claim needs explicit defense.
  3. [5 (Conclusion)] The central conclusion that 'neither purely epistocratic nor purely democratic approaches ... may be sufficient on their own' is not established by the preceding analysis. Sections 3 and 4 show at most that instrumental justifications require further empirical work and that the coercion-based non-instrumental justification faces unanswered objections. That is a directed challenge to democratic proponents; it does not by itself show that pure epistocratic or pure democratic solutions are insufficient, nor does it show that hybrid frameworks are superior. As it stands, the hybrid conclusion is a research prediction rather than an implication of the argument. Please either mark it explicitly as a forward-looking conjecture or supply an argument that the identified burdens cannot be met by either pure approach.
minor comments (5)
  1. [Abstract] There is a typographical artifact in the abstract: 'aimtofillthroughpoliticalratherthantheoreticaljustification' should be 'aim to fill through political rather than theoretical justification'. The author affiliation also contains a typo: 'Saarbücken' should be 'Saarbrücken'.
  2. [1] The discussion of Schuster and Kilov would benefit from a brief statement of the disagreement: the authors claim that Schuster and Kilov conflate the technical and normative problems, but the paragraph asserts this rather than showing where exactly the conflation occurs.
  3. [3] The examples 'Brexit, Trump, the climate crisis' are too compressed to carry the argument. If these are meant as cases where people vote against their own interests, at least one example should be spelled out, or they should be replaced with less contested cases.
  4. [4] The clarification that 'the primary coercer is not the AI itself but the person or organisation that defines the normative constraints' introduces an agency question that is never revisited: in the later examples (the AI assistant refusing to buy meat), the corporate and technical chain of responsibility is not identified. A sentence about who, in the authors' view, defines constraints in current deployment contexts would help.
  5. [2] The paper moves from observed 'reasonable disagreement' to 'normative uncertainty' as the rational response. This inference would benefit from a footnote distinguishing epistemic uncertainty from the possibility that disagreement merely reflects different evaluative perspectives; not all readers will accept that persistent disagreement entails uncertainty about the normative facts.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the argument is a critical philosophical analysis whose conclusion follows from the stated objections and conditions, not from definitional or self-citational construction.

full rationale

The paper's central conclusion—that neither purely epistocratic nor purely democratic approaches may be sufficient and that hybrid frameworks merit consideration—is reached through an open-ended critique of four required propositions for the coercion-based non-instrumental justification of democratic alignment. Each proposition is treated as a burden of proof for proponents of democratic approaches, not as an assumption built into the definitions. The paper explicitly leaves open whether the propositions can be established and identifies objections and alternative resources for epistocratic approaches, such as background-condition control and decision rules under normative uncertainty. No equation or formal derivation is present, and no fitted parameter is renamed as a prediction. The self-citations, notably Baum (2025) for a taxonomy of AI alignment, serve as background classification rather than load-bearing premises, and the argument does not depend on an unverified uniqueness theorem or on the authors' prior work to force its conclusion. The coercion discussion invokes a contested exit-based account of coercion, but this is an ordinary substantive philosophical assumption subject to external criticism, not a case of the paper defining its conclusion into existence. Accordingly, no circular step meeting the required evidentiary standard can be identified, and the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper relies on several domain assumptions from political philosophy and metaethics. The most fragile is the specific account of coercion (going-elsewhere condition) because the paper's challenge to democratic justification depends on it. No free parameters or invented entities are involved in this conceptual analysis.

assumptions (4)
  • domain assumption Normative and metanormative uncertainty is a necessary condition for the possibility of illegitimate authority or coercion.
    Section 2 argues that uncertainty eliminates a 'sure theoretical justification' that could legitimise coercion; this is a contested metaethical premise about legitimacy.
  • domain assumption Coercion is prima facie wrong.
    Section 4 lists this as the first ingredient in establishing that AI coercion would be unjustified; the paper notes a tension with normative uncertainty but does not resolve it.
  • ad hoc to paper A user is coerced by an AI's noncompliance only if the AI is the de facto or de jure standard, i.e., the user cannot freely go elsewhere.
    Section 4 uses the supermarket and bowling club examples to support this; this account of coercion drives the response to proposition (i) but is not defended against alternative views.
  • domain assumption Decision rules like maximize expected choiceworthiness can provide a practical justification for normative constraints under uncertainty.
    Section 4 grants that epistocrats could invoke such rules to justify coercion, thereby undercutting the democratic advantage; the paper does not critically assess the validity of those rules.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Justifications for Democratizing AI Alignment and Their Prospects." pith.science (2026). https://pith.science/paper/MDCJWC4B

@misc{pith2026250719548,
  author       = {Pith},
  title        = {Pith review of: Justifications for Democratizing AI Alignment and Their Prospects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MDCJWC4B}},
  note         = {Machine review of arXiv:2507.19548}
}
read the original abstract

The AI alignment problem comprises both technical and normative dimensions. While technical solutions focus on implementing normative constraints in AI systems, the normative problem concerns determining what these constraints should be. This paper examines justifications for democratic approaches to the normative problem -- where affected stakeholders determine AI alignment -- as opposed to epistocratic approaches that defer to normative experts. We analyze both instrumental justifications (democratic approaches produce better outcomes) and non-instrumental justifications (democratic approaches prevent illegitimate authority or coercion). We argue that normative and metanormative uncertainty create a justificatory gap that democratic approaches aim to fill through political rather than theoretical justification. However, we identify significant challenges for democratic approaches, particularly regarding the prevention of illegitimate coercion through AI alignment. Our analysis suggests that neither purely epistocratic nor purely democratic approaches may be sufficient on their own, pointing toward hybrid frameworks that combine expert judgment with participatory input alongside institutional safeguards against AI monopolization.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

22 extracted references · 15 canonical work pages

  1. [1]

    et al.: A General Language Assistant as a Laboratory for Alignment

    Askell, A. et al.: A General Language Assistant as a Laboratory for Alignment. (2021), arXiv:2112.00861

  2. [2]

    et al.: The Moral Machine experiment

    Awad, E. et al.: The Moral Machine experiment. Nature, 563, 59–64. (2018)

  3. [3]

    et al.: Training a Helpful and Harmless Assistant with Reinforcement Learn- ing from Human Feedback

    Bai, Y. et al.: Training a Helpful and Harmless Assistant with Reinforcement Learn- ing from Human Feedback. (2022), arXiv:2204.05862

  4. [4]

    et al.: Constitutional AI: Harmlessness from AI Feedback

    Bai, Y. et al.: Constitutional AI: Harmlessness from AI Feedback. (2022), arXiv:2212.08073

  5. [5]

    In: Steffen, B

    Baum, K.: Disentangling AI Alignment: A Structured Taxonomy Beyond Safety and Ethics. In: Steffen, B. (ed.) AISoLA 2024 Post-Proceedings. Lecture Notes in Computer Science. Springer, Cham (forthcoming). arXiv preprint arXiv:2506.06286 (2025)

  6. [6]

    AI & Society, 35(1), 165– 176 (2020)

    Baum, S.: Social Choice Ethics in Artificial Intelligence. AI & Society, 35(1), 165– 176 (2020). https://doi.org/10.1007/s00146-017-0760-1 14 A. Steingrüber and K. Baum

  7. [7]

    Christiano, P. F. et al.: Deep reinforcement learning from human preferences. Ad- vances in neural information processing systems, 30. (2017)

  8. [8]

    The Stanford Encyclopedia of Philosophy

    Christiano, T., Bajaj, S.: Democracy. The Stanford Encyclopedia of Philosophy. (2024), https://plato.stanford.edu/archives/sum2024/entries/democracy/

Show all 22 references
  1. [9]

    Minds and Machines, 2(5), 411–437

    Gabriel, I.: Artificial intelligence, values, and alignment. Minds and Machines, 2(5), 411–437. (2020)

  2. [10]

    Philosophical Studies, (2025)

    Gabriel, I., Keeling, G.: A matter of principle? AI alignment as the fair treatment of claims. Philosophical Studies, (2025)

  3. [11]

    E., Spiekermann, K.: An Epistemic Theory of Democracy, Oxford Uni- versity Press, Oxford

    Goodin, R. E., Spiekermann, K.: An Epistemic Theory of Democracy, Oxford Uni- versity Press, Oxford. (2018)

  4. [12]

    T., Papyshev, G., Wong, J

    Huang, L. T., Papyshev, G., Wong, J. K.: Democratizing AI alignment: from au- thoritarian to democratic AI ethics. AI and Ethics, 5, 11–18. (2025)

  5. [13]

    et al.: Can Machines Learn Morality? The Delphi Experiment

    Jiang, L. et al.: Can Machines Learn Morality? The Delphi Experiment. (2022), arXiv:2110.07574

  6. [14]

    W., Hooker, J., Donaldson, T.: Taking Principles Seriously: A Hybrid Ap- proach to Value Alignment in Artificial Intelligence

    Kim, T. W., Hooker, J., Donaldson, T.: Taking Principles Seriously: A Hybrid Ap- proach to Value Alignment in Artificial Intelligence. Journal of Artificial Intelligence Research, 70, 871–890. (2021)

  7. [15]

    FAccT ’25: Proceedings of the 2025 ACM Conference on Fairness, Ac- countability, and Transparency, 2671 - 2681

    Kneer, M., Viehoff, J.: The Hard Problem of AI Alignment: Value Forks in Moral Judgment. FAccT ’25: Proceedings of the 2025 ACM Conference on Fairness, Ac- countability, and Transparency, 2671 - 2681. (2025)

  8. [16]

    Kolodny,N.:ThePeckingOrder,HarvardUniversityPress,Cambridge,MA.(2023)

  9. [17]

    MacAskill, W., Bykvist, K., Ord, T.: Moral Uncertainty, Oxford University Press, Oxford. (2020)

  10. [18]

    Oxford University Press, Oxford

    Raz, J.: The Morality of Freedom. Oxford University Press, Oxford. (1986)

  11. [19]

    AI and Ethics, 5, 3727–3741

    Riesen, E., Boespflug, M.: Aligning with ideal values: a proposal for anchoring AI in moral expertise. AI and Ethics, 5, 3727–3741. (2025)

  12. [20]

    Philosophy and Public Affairs, 32(1), 2–35

    Ripstein, A.: Authority and Coercion. Philosophy and Public Affairs, 32(1), 2–35. (2004)

  13. [21]

    AI & Society

    Schuster, N., Kilov, D.: Moral disagreement and the limits of AI value alignment: a dual challenge of epistemic justification and political legitimacy. AI & Society. (2025)

  14. [22]

    et al.: Fine-Tuning Language Models from Human Preferences

    Ziegler, D. et al.: Fine-Tuning Language Models from Human Preferences. (2019), arXiv:1909.08593

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.