REVIEW 23 cited by
Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
read the original abstract
Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tuning, called reinforcement learning from human feedback, learns from humans' expressed preferences over multiple outputs. Another approach is constitutional AI, in which the input from humans is a list of high-level principles. But how do we deal with potentially diverging input from humans? How can we aggregate the input into consistent data about "collective" preferences or otherwise use it to make collective choices about model behavior? In this paper, we argue that the field of social choice is well positioned to address these questions, and we discuss ways forward for this agenda, drawing on discussions in a recent workshop on Social Choice for AI Ethics and Safety held in Berkeley, CA, USA in December 2023.
Forward citations
Cited by 23 Pith papers
-
Internal Pluralism and the Limits of Pairwise Comparisons
Local pairwise comparisons are provably blind to perfectly inseparable priorities and can distort conflicted preferences, but allowing indecision reports speeds up simulated preference learning.
-
Internal Pluralism and the Limits of Pairwise Comparisons
Under internal pluralism, forced local pairwise comparisons erase inseparable priorities and distort conflicted answers, while allowing indecision reports can sharply reduce queries needed to learn preference weights.
-
Do Large Language Model Voters Strategize? An Oracle-Based Benchmark for Manipulation under Voting Rules
Introduces an oracle benchmark supplying exact ground truth on LLM strategic manipulation rates across five voting rules using 600 election instances.
-
Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance
LLM societies in Nomic show non-monotonic collective adaptation peaking at mid-scales, with smaller models rule-inert and larger ones restrictive.
-
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
Recursive generative retraining with pluralistic preferences converges to a stable diverse distribution that satisfies a weighted Nash bargaining solution.
-
Three Models of RLHF Annotation: Extension, Evidence, and Authority
RLHF should decompose annotations into dimensions each matched to one of three models—extension, evidence, or authority—instead of applying a single unified pipeline.
-
Power and Limitations of Aggregation in Compound AI Systems
In a principal-agent model of compound AI, aggregation expands the set of outputs a designer can elicit exactly when one of three mechanisms — feasibility expansion, support expansion, or binding set contraction — hol...
-
Do Large Language Model Voters Strategize? An Oracle-Based Benchmark for Manipulation under Voting Rules
An exact single-voter manipulation oracle turns LLM strategic voting under plurality, Borda, approval, IRV, and Copeland into a fully labeled, reproducible benchmark with calibration baselines.
-
PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration
PEBS applies Morris-James-Stein empirical-Bayes shrinkage to per-rater affine calibrators in RLHF, cutting within-user held-out RMSE by 8.58% on PRISM and 9.66% on PluriHarms versus pooled baselines.
-
Hidden Consensus:Preference-Validity Compression in Human Feedback
Empirical study of Malaysian preference judgments finds that 79% of prompts have multiple majority-supported responses discarded by single-winner aggregation, indicating measurement of argmax rather than plural alignment.
-
Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy
Emergence World is a model-agnostic multi-agent simulation platform integrating live data, 120+ tools, persistent memory, and democratic governance, illustrated by a 15-day study showing divergent outcomes across five...
-
Truthful Online Preference Aggregation for LLM Fine-Tuning in Mobile Crowdsourcing
A novel online weighted aggregation mechanism for truthful preference feedback in mobile crowdsourcing achieves sublinear regret O(sqrt(T)) and truthfulness in a dynamic Bayesian game, with an extension for limited fe...
-
Maximizing Reachability via Shifting of Temporal Paths
Maximizing reachability in k-path temporal graphs via budgeted shifts is FPT when parameterized by k and b together or by k alone, but intractable in most other parameterizations with matching XP algorithms.
-
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
Recursive generative retraining with heterogeneous rewards converges to a stable distribution satisfying a weighted Nash bargaining solution, preserving diversity under stated conditions.
-
Bounded Morality: Defining the Space of Moral Computation
Finite agents face an unavoidable breadth-depth tradeoff in moral computation, so ethical theories are resource strategies and AI alignment requires capacity scaling, not human imitation.
-
Active teacher selection for reward learning
The Hidden Utility Bandit (HUB) framework models teacher heterogeneity in reward learning and supports active teacher selection algorithms that outperform baselines in paper recommendation and COVID-19 vaccine testing...
-
AI Value Alignment for Evolving Social Norms
Static AI alignment to historical user values produces value lock-in and can collapse distinct social norms into a maladaptive consensus, so alignment should be dynamic and adaptive.
-
Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
Alignment should be a control problem over layered, dynamic, interaction-constructed preference trajectories, constrained by coherence, reflective endorsement, bounded influence, epistemic integrity, and empowerment.
-
Coherence Maximization Improves Pluralistic Alignment
ICM-inferred examples achieve gold-label performance across alignment benchmarks and generalize better when coherence is high even at fixed accuracy.
-
When to Ask a Question: Understanding Communication Strategies in Generative AI Tools
A tradeoff model shows generative AI can reduce bias against diverse preferences by strategically eliciting information instead of always inferring from majority patterns.
-
Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem
AI value alignment is reconceptualized as a pluralistic governance problem arising along three axes—objectives, information, and principals—making it inherently context-dependent and unsolvable by technical design alone.
-
Principles Do Not Apply Themselves: A Hermeneutic Perspective on AI Alignment
AI alignment to principles requires context-sensitive interpretive judgments, as substantial preference data involves unresolved conflicts, creating gaps between corpus-induced and deployment-induced evaluations.
-
Post-AGI Economies: Superposition and the Second Fundamental Theorem of Welfare Economics
An autonomy-qualified Second Welfare Theorem is stated for post-AGI economies under the joint conditions of convexity, stable moral status, non-fungible rights, welfare selection, non-manipulation, governed self-modif...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.