REVIEW 6 cited by
Aligning Language Models to User Opinions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
An important aspect of developing LLMs that interact with humans is to align models' behavior to their users. It is possible to prompt an LLM into behaving as a certain persona, especially a user group or ideological persona the model captured during its pertaining stage. But, how to best align an LLM with a specific user and not a demographic or ideological group remains an open question. Mining public opinion surveys (by Pew Research), we find that the opinions of a user and their demographics and ideologies are not mutual predictors. We use this insight to align LLMs by modeling both user opinions as well as user demographics and ideology, achieving up to 7 points accuracy gains in predicting public opinions from survey questions across a broad set of topics. In addition to the typical approach of prompting LLMs with demographics and ideology, we discover that utilizing the most relevant past opinions from individual users enables the model to predict user opinions more accurately.
Forward citations
Cited by 6 Pith papers
-
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models
LLM probability estimates violate the law of total probability across partitions, and subgroup-aggregated estimates often beat direct population-level estimates (the macro fallacy).
-
People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe
Across 10 LLMs and the European Social Survey, AI alignment favors wealthier, more educated, less religious, and more politically interested groups, with country of residence explaining as much variance as all sociode...
-
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
A new dataset and benchmark maps movie clips to distributions of audience emotional reactions derived from YouTube comments, showing that finetuned vision-language models can predict these distributions from video alone.
-
Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment
Persona-judge applies speculative decoding between two preference-prompted copies of the same LLM to achieve training-free personalized alignment.
-
Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods
LLMs extend, rather than replace, classical social science methods, with a proposed three-tier bias framework for LLM-augmented surveys.
-
Embedding-to-Prefix: Parameter-Efficient Personalization for Pre-Trained Large Language Models
E2P projects pre-computed user embeddings into a single soft prefix token for frozen LLMs, reporting gains on four personalization tasks, though its reproduction scripts write zero embeddings.
Discussion (0). Continue with ORCID to comment.