Pith. sign in

Ai alignment and social choice: Fundamental limitations and policy implications

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

years

2026 5 2025 1

representative citing papers

Variance-aware Reward Modeling with Anchor Guidance

stat.ML · 2026-05-12 · unverdicted · novelty 7.0

Anchor-guided variance-aware reward modeling uses two response-level anchors to resolve non-identifiability in Gaussian models of pluralistic preferences, yielding provable identification, a joint training objective, and improved RLHF performance.

AI Alignment From Social Choice Perspectives

cs.AI · 2026-06-19 · unverdicted · novelty 3.0

This survey examines applications of social choice theory to aggregating human feedback in AI alignment, identifying failure modes and expanding design options for disagreement.

citing papers explorer

Showing 6 of 6 citing papers.