Pith. sign in

REVIEW 23 cited by

Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.10271 v2 pith:DMLCBG67 submitted 2024-04-16 cs.LG cs.AIcs.CLcs.CYcs.GT

classification cs.LGcs.AIcs.CLcs.CYcs.GT
keywords choicehumansinputsocialapproachbehaviorcollectivefeedback
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tuning, called reinforcement learning from human feedback, learns from humans' expressed preferences over multiple outputs. Another approach is constitutional AI, in which the input from humans is a list of high-level principles. But how do we deal with potentially diverging input from humans? How can we aggregate the input into consistent data about "collective" preferences or otherwise use it to make collective choices about model behavior? In this paper, we argue that the field of social choice is well positioned to address these questions, and we discuss ways forward for this agenda, drawing on discussions in a recent workshop on Social Choice for AI Ethics and Safety held in Berkeley, CA, USA in December 2023.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Internal Pluralism and the Limits of Pairwise Comparisons

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Under internal pluralism, forced local pairwise comparisons erase inseparable priorities and distort conflicted answers, while allowing indecision reports can sharply reduce queries needed to learn preference weights.

  2. Do Large Language Model Voters Strategize? An Oracle-Based Benchmark for Manipulation under Voting Rules

    cs.GT 2026-06 unverdicted novelty 7.0 of 10

    Introduces an oracle benchmark supplying exact ground truth on LLM strategic manipulation rates across five voting rules using 600 election instances.

  3. Power and Limitations of Aggregation in Compound AI Systems

    cs.AI 2026-02 conditional novelty 7.0 of 10

    In a principal-agent model of compound AI, aggregation expands the set of outputs a designer can elicit exactly when one of three mechanisms — feasibility expansion, support expansion, or binding set contraction — hol...

  4. Selective Response Strategies for GenAI

    cs.AI 2025-02 reject novelty 7.0 of 10

    Selective response, withholding answers to drive users to human forums, can in a stylized model increase both GenAI revenue and user welfare, and near-optimal policies can be computed approximately.

  5. JuStRank: Benchmarking LLM Judges for System Ranking

    cs.CL 2024-12 conditional novelty 7.0 of 10

    JuStRank ranks AI judges by how well their aggregated scores reproduce the Chatbot Arena human system ranking, revealing that judge realization and bias, not just model size, determine ranking quality.

  6. Metanormative Theory for RL-Based Moral Agents

    cs.AI 2026-08 conditional novelty 6.0 of 10

    The authors argue that RL agents should be classified as moral only relative to a distinct moral reward function, and use this criterion to assess three RL-based value alignment proposals.

  7. Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Pluralistic alignment should be reframed as socially grounded coordination, using roles, deliberative interaction, field-aware weighting, and trajectory-level audit rather than output diversification.

  8. Bounded Morality: Defining the Space of Moral Computation

    cs.AI 2026-04 conditional novelty 6.0 of 10

    Finite agents face an unavoidable breadth-depth tradeoff in moral computation, so ethical theories are resource strategies and AI alignment requires capacity scaling, not human imitation.

  9. Collaborating with GenAI: Incentives and Replacements

    cs.GT 2025-08 conditional novelty 6.0 of 10

    Generative AI can collapse worker effort in a stylized team game, and selecting the optimal team is NP-complete.

  10. Quantitative Relaxations of Arrow's Axioms

    cs.GT 2025-06 conditional novelty 6.0 of 10

    A new quantitative framework measures the degree to which voting rules violate Arrow's independence and unanimity axioms, and an empirical study finds Borda performs best on Scottish and synthetic elections.

  11. The Battling Influencers Game: Nash Equilibria Structure of a Potential Game and Implications to Value Alignment

    cs.GT 2025-02 conditional novelty 6.0 of 10

    A new potential game shows that when influencers compete to shape a receiver's aggregate opinion, any pure Nash equilibrium forces all but at most one influencer to the most extreme allowed action.

  12. Clone-Robust AI Alignment

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A Voronoi-weighted maximum likelihood estimator for RLHF is robust to adding approximate clone responses, unlike the standard regularized MLE.

  13. Evaluating the Prompt Steerability of Large Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A formal benchmark with steerability indices shows that six open-weight LLMs are only partially steerable by prompting, with strong baseline skew and directional asymmetry.

  14. AI Value Alignment for Evolving Social Norms

    cs.CY 2026-07 conditional novelty 5.5 of 10

    Static AI alignment to historical user values produces value lock-in and can collapse distinct social norms into a maladaptive consensus, so alignment should be dynamic and adaptive.

  15. Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction

    cs.AI 2026-04 conditional novelty 5.5 of 10

    Alignment should be a control problem over layered, dynamic, interaction-constructed preference trajectories, constrained by coherence, reflective endorsement, bounded influence, epistemic integrity, and empowerment.

  16. What Voting Rules Actually Do: A Data-Driven Analysis of Multi-Winner Voting

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    A data-driven framework for counting axiom violations across preference distributions, with the claim that trained neural-network rules minimize violations better than traditional multi-winner rules.

  17. Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory

    stat.ML 2025-06 conditional novelty 5.0 of 10

    RLHF reward modeling satisfies pairwise majority and Condorcet consistency when each response pair is labeled once, because the maximum likelihood ranking then matches the Copeland rule.

  18. Configurable Preference Tuning with Rubric-Guided Synthetic Data

    cs.CL 2025-06 conditional novelty 5.0 of 10

    CPT fine-tunes LLMs with DPO on rubric-guided synthetic preferences so that a system prompt can reconfigure output style at inference, with in-distribution accuracy gains over baselines.

  19. Data Sharing with a Generative AI Competitor

    cs.GT 2025-05 conditional novelty 5.0 of 10

    In a two-stage data-sharing game, the unique equilibrium is either that the firm shares just enough data to stop the platform buying expert data, or that the firm shares an amount that maximizes its payoff while the p...

  20. Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models

    cs.LG 2025-01 conditional novelty 5.0 of 10

    Applying a utility-inspired, threshold-based transformation to individual rewards before summing them improved the harmlessness of an RLHF-trained 2B language model without reducing helpfulness.

  21. Online Learning from Strategic Human Feedback in LLM Fine-Tuning

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A multiplicative-weight aggregation rule with the Brier score makes truthful human feedback a dominant strategy in online RLHF and yields O(sqrt(T)) regret.

  22. Beyond the Binary: Capturing Diverse Preferences With Reward Regularization

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Reward models trained with a margin derived from synthetic LLM judgments better match aggregate human preferences than standard binary-trained reward models, mainly on subjective prompts.

  23. The Problem of Social Cost in Multi-Agent General Reinforcement Learning: Survey and Synthesis

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A market mechanism based on VCG payments is defined for general reinforcement learning agents, with proofs of Bayes-Nash incentive compatibility and individual rationality, plus illustrative applications.

Pith tools