REVIEW 23 cited by
Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tuning, called reinforcement learning from human feedback, learns from humans' expressed preferences over multiple outputs. Another approach is constitutional AI, in which the input from humans is a list of high-level principles. But how do we deal with potentially diverging input from humans? How can we aggregate the input into consistent data about "collective" preferences or otherwise use it to make collective choices about model behavior? In this paper, we argue that the field of social choice is well positioned to address these questions, and we discuss ways forward for this agenda, drawing on discussions in a recent workshop on Social Choice for AI Ethics and Safety held in Berkeley, CA, USA in December 2023.
Forward citations
Cited by 23 Pith papers
-
Internal Pluralism and the Limits of Pairwise Comparisons
Under internal pluralism, forced local pairwise comparisons erase inseparable priorities and distort conflicted answers, while allowing indecision reports can sharply reduce queries needed to learn preference weights.
-
Do Large Language Model Voters Strategize? An Oracle-Based Benchmark for Manipulation under Voting Rules
Introduces an oracle benchmark supplying exact ground truth on LLM strategic manipulation rates across five voting rules using 600 election instances.
-
Power and Limitations of Aggregation in Compound AI Systems
In a principal-agent model of compound AI, aggregation expands the set of outputs a designer can elicit exactly when one of three mechanisms — feasibility expansion, support expansion, or binding set contraction — hol...
-
Selective Response Strategies for GenAI
Selective response, withholding answers to drive users to human forums, can in a stylized model increase both GenAI revenue and user welfare, and near-optimal policies can be computed approximately.
-
JuStRank: Benchmarking LLM Judges for System Ranking
JuStRank ranks AI judges by how well their aggregated scores reproduce the Chatbot Arena human system ranking, revealing that judge realization and bias, not just model size, determine ranking quality.
-
Metanormative Theory for RL-Based Moral Agents
The authors argue that RL agents should be classified as moral only relative to a distinct moral reward function, and use this criterion to assess three RL-based value alignment proposals.
-
Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory
Pluralistic alignment should be reframed as socially grounded coordination, using roles, deliberative interaction, field-aware weighting, and trajectory-level audit rather than output diversification.
-
Bounded Morality: Defining the Space of Moral Computation
Finite agents face an unavoidable breadth-depth tradeoff in moral computation, so ethical theories are resource strategies and AI alignment requires capacity scaling, not human imitation.
-
Collaborating with GenAI: Incentives and Replacements
Generative AI can collapse worker effort in a stylized team game, and selecting the optimal team is NP-complete.
-
Quantitative Relaxations of Arrow's Axioms
A new quantitative framework measures the degree to which voting rules violate Arrow's independence and unanimity axioms, and an empirical study finds Borda performs best on Scottish and synthetic elections.
-
The Battling Influencers Game: Nash Equilibria Structure of a Potential Game and Implications to Value Alignment
A new potential game shows that when influencers compete to shape a receiver's aggregate opinion, any pure Nash equilibrium forces all but at most one influencer to the most extreme allowed action.
-
Clone-Robust AI Alignment
A Voronoi-weighted maximum likelihood estimator for RLHF is robust to adding approximate clone responses, unlike the standard regularized MLE.
-
Evaluating the Prompt Steerability of Large Language Models
A formal benchmark with steerability indices shows that six open-weight LLMs are only partially steerable by prompting, with strong baseline skew and directional asymmetry.
-
AI Value Alignment for Evolving Social Norms
Static AI alignment to historical user values produces value lock-in and can collapse distinct social norms into a maladaptive consensus, so alignment should be dynamic and adaptive.
-
Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
Alignment should be a control problem over layered, dynamic, interaction-constructed preference trajectories, constrained by coherence, reflective endorsement, bounded influence, epistemic integrity, and empowerment.
-
What Voting Rules Actually Do: A Data-Driven Analysis of Multi-Winner Voting
A data-driven framework for counting axiom violations across preference distributions, with the claim that trained neural-network rules minimize violations better than traditional multi-winner rules.
-
Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
RLHF reward modeling satisfies pairwise majority and Condorcet consistency when each response pair is labeled once, because the maximum likelihood ranking then matches the Copeland rule.
-
Configurable Preference Tuning with Rubric-Guided Synthetic Data
CPT fine-tunes LLMs with DPO on rubric-guided synthetic preferences so that a system prompt can reconfigure output style at inference, with in-distribution accuracy gains over baselines.
-
Data Sharing with a Generative AI Competitor
In a two-stage data-sharing game, the unique equilibrium is either that the firm shares just enough data to stop the platform buying expert data, or that the firm shares an amount that maximizes its payoff while the p...
-
Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models
Applying a utility-inspired, threshold-based transformation to individual rewards before summing them improved the harmlessness of an RLHF-trained 2B language model without reducing helpfulness.
-
Online Learning from Strategic Human Feedback in LLM Fine-Tuning
A multiplicative-weight aggregation rule with the Brier score makes truthful human feedback a dominant strategy in online RLHF and yields O(sqrt(T)) regret.
-
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization
Reward models trained with a margin derived from synthetic LLM judgments better match aggregate human preferences than standard binary-trained reward models, mainly on subjective prompts.
-
The Problem of Social Cost in Multi-Agent General Reinforcement Learning: Survey and Synthesis
A market mechanism based on VCG payments is defined for general reinforcement learning agents, with proofs of Bayes-Nash incentive compatibility and individual rationality, plus illustrative applications.
Discussion (0). Continue with ORCID to comment.