LLMs match condition-level patterns in a noodle purchase survey but fail to replicate distributional structure, with no model beating a pooled human baseline for purchase quantities.
arXiv preprint arXiv:2408.06929 , eprint =
5 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Frontier LLMs approximate human story morals but show markedly less cross-linguistic variation and narrower value focus than human responses across 14 language-culture pairs.
Centralized matching mechanisms outperform free negotiation in stability and efficiency with LLM agents, who also report preferences truthfully more often than humans, though not always in line with strategy-proofness predictions.
A large synthetic dataset of police personas and SJTs is used to claim LLMs display measurable trait-consistent behavior, but the validation is internal and several headline analyses are absent from the text.
Mod-Guide uses RAG with a community co-created corpus to make LLM moderation responses more contextually accurate for insensitive speech toward Bangladesh's Hindu and Chakma minorities, with mixed-method evaluation showing differences by ethnic background.
citing papers explorer
-
Beyond Averages: Evaluating LLMs on Human Survey Replication at the Distributional Level
LLMs match condition-level patterns in a noodle purchase survey but fail to replicate distributional structure, with no model beating a pooled human baseline for purchase quantities.
-
Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation
Frontier LLMs approximate human story morals but show markedly less cross-linguistic variation and narrower value focus than human responses across 14 language-culture pairs.
-
Do Matching Mechanisms Work with LLM Agents?
Centralized matching mechanisms outperform free negotiation in stability and efficiency with LLM agents, who also report preferences truthfully more often than humans, though not always in line with strategy-proofness predictions.
-
Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
A large synthetic dataset of police personas and SJTs is used to claim LLMs display measurable trait-consistent behavior, but the validation is internal and several headline analyses are absent from the text.
-
Mod-Guide: An LLM-based Content Moderation Feedback System to Address Insensitive Speech toward Indigenous Ethnic and Religious Minority Communities
Mod-Guide uses RAG with a community co-created corpus to make LLM moderation responses more contextually accurate for insensitive speech toward Bangladesh's Hindu and Chakma minorities, with mixed-method evaluation showing differences by ethnic background.