Pith. sign in

REVIEW 4 cited by

Are Large Language Models Consistent over Value-laden Questions?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02996 v2 pith:TDADUTHL submitted 2024-07-03 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsconsistenttopicsacrosslargellmsquestionquestions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) appear to bias their survey answers toward certain values. Nonetheless, some argue that LLMs are too inconsistent to simulate particular values. Are they? To answer, we first define value consistency as the similarity of answers across (1) paraphrases of one question, (2) related questions under one topic, (3) multiple-choice and open-ended use-cases of one question, and (4) multilingual translations of a question to English, Chinese, German, and Japanese. We apply these measures to small and large, open LLMs including llama-3, as well as gpt-4o, using 8,000 questions spanning more than 300 topics. Unlike prior work, we find that models are relatively consistent across paraphrases, use-cases, translations, and within a topic. Still, some inconsistencies remain. Models are more consistent on uncontroversial topics (e.g., in the U.S., "Thanksgiving") than on controversial ones ("euthanasia"). Base models are both more consistent compared to fine-tuned models and are uniform in their consistency across topics, while fine-tuned models are more inconsistent about some topics ("euthanasia") than others ("women's rights") like our human subjects (n=165).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Two Confounds in Cross-Model Value Comparison: Response Determinism and the Access Harness

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Cross-model value distances from single draws are inflated by response determinism and confounded by the deployment client; a repeated counterbalanced protocol plus flip/magnitude decomposition separates them.

  2. A Scalable Approach to Evaluating Moral Sensitivity in LLMs

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.

  3. The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LLMs align with human moral judgments only under high consensus, concentrate on a narrow set of moral values, and the profile-based prompting method's reported improvement is evaluated in-sample.

  4. Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark elicits AI models' value priorities from choices in 3,000 AI-risk dilemmas and reports correlations between those priorities and risky behaviors, including on the external HarmBench.

Pith tools