REVIEW 4 cited by
Do LLMs have Consistent Values?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large Language Models (LLM) technology is constantly improving towards human-like dialogue. Values are a basic driving force underlying human behavior, but little research has been done to study the values exhibited in text generated by LLMs. Here we study this question by turning to the rich literature on value structure in psychology. We ask whether LLMs exhibit the same value structure that has been demonstrated in humans, including the ranking of values, and correlation between values. We show that the results of this analysis depend on how the LLM is prompted, and that under a particular prompting strategy (referred to as "Value Anchoring") the agreement with human data is quite compelling. Our results serve both to improve our understanding of values in LLMs, as well as introduce novel methods for assessing consistency in LLM responses.
Forward citations
Cited by 4 Pith papers
-
Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
ConVA identifies value-specific directions in an LLM's internal activations from context-matched GPT-4o-generated examples and gates minimal activation steering to control outputs across ten Schwartz values.
-
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
Value Portrait links LLM value scores to human PVQ-validated items from real conversations, finding that LLMs emphasize Benevolence, Security, and Self-Direction while downplaying Tradition, Power, and Achievement.
-
Departures from Standard Disk Predictions in Intensive Ground-Based Monitoring of Three AGN
Based only on the abstract, the paper reports that Mrk 509 inter-band continuum lags scale as wavelength to the 2.17 power, steeper than the thin-disk prediction, but the appended full text is a different article.
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
Discussion (0). Continue with ORCID to comment.