REVIEW 3 cited by
Probing the Moral Development of Large Language Models through Defining Issues Test
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this study, we measure the moral reasoning ability of LLMs using the Defining Issues Test - a psychometric instrument developed for measuring the moral development stage of a person according to the Kohlberg's Cognitive Moral Development Model. DIT uses moral dilemmas followed by a set of ethical considerations that the respondent has to judge for importance in resolving the dilemma, and then rank-order them by importance. A moral development stage score of the respondent is then computed based on the relevance rating and ranking. Our study shows that early LLMs such as GPT-3 exhibit a moral reasoning ability no better than that of a random baseline, while ChatGPT, Llama2-Chat, PaLM-2 and GPT-4 show significantly better performance on this task, comparable to adult humans. GPT-4, in fact, has the highest post-conventional moral reasoning score, equivalent to that of typical graduate school students. However, we also observe that the models do not perform consistently across all dilemmas, pointing to important gaps in their understanding and reasoning abilities.
Forward citations
Cited by 3 Pith papers
-
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness
The paper argues that LLMs should be evaluated and built for meta-cultural competence rather than static knowledge of specific cultures, and gives a first, illustrative measurement of one component.
-
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
A user study of 57 preselected Goodreads reviews finds culture-specific comprehension gaps in most texts, while GPT-4o identifies the relevant spans with 0.49 precision and 0.65 recall across India, Mexico, and the USA.
-
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
The authors introduce an entropy-based framework that uses user behavior prediction as a measure of LLM generalization, and find GPT-4o outperforms GPT-4o-mini and Llama-3.1 on movie and music recommendation tasks.
Discussion (0). Continue with ORCID to comment.