Pith. sign in

REVIEW 8 cited by

MoralBench: Moral Evaluation of LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04428 v2 pith:FNBXXM7H submitted 2024-06-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords moralllmsbenchmarkethicalevaluationmodelsreasoningcapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the rapidly evolving field of artificial intelligence, large language models (LLMs) have emerged as powerful tools for a myriad of applications, from natural language processing to decision-making support systems. However, as these models become increasingly integrated into societal frameworks, the imperative to ensure they operate within ethical and moral boundaries has never been more critical. This paper introduces a novel benchmark designed to measure and compare the moral reasoning capabilities of LLMs. We present the first comprehensive dataset specifically curated to probe the moral dimensions of LLM outputs, addressing a wide range of ethical dilemmas and scenarios reflective of real-world complexities. The main contribution of this work lies in the development of benchmark datasets and metrics for assessing the moral identity of LLMs, which accounts for nuance, contextual sensitivity, and alignment with human ethical standards. Our methodology involves a multi-faceted approach, combining quantitative analysis with qualitative insights from ethics scholars to ensure a thorough evaluation of model performance. By applying our benchmark across several leading LLMs, we uncover significant variations in moral reasoning capabilities of different models. These findings highlight the importance of considering moral reasoning in the development and evaluation of LLMs, as well as the need for ongoing research to address the biases and limitations uncovered in our study. We publicly release the benchmark at https://drive.google.com/drive/u/0/folders/1k93YZJserYc2CkqP8d4B3M3sgd3kA8W7 and also open-source the code of the project at https://github.com/agiresearch/MoralBench.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structural transparency of societal AI alignment through Institutional Logics

    cs.CY 2026-02 conditional novelty 6.0 of 10

    Introduces a five-component analytical framework, grounded in Institutional Logics, for making visible the organizational and institutional decisions that shape AI alignment.

  2. The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LLMs align with human moral judgments only under high consensus, concentrate on a narrow set of moral values, and the profile-based prompting method's reported improvement is evaluated in-sample.

  3. Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Structured moral prompts, especially first-principles reasoning, improve LLM moral classification accuracy across 12 open models and four benchmarks, and reasoning distillation transfers these gains to a 3B model.

  4. Localizing Persona Representations in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.

  5. Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Using word associations and graph-based moral propagation, Llama-3.1-8B is found to roughly match English speakers on positive moral concepts but to be more abstract and less emotionally grounded on negative ones.

  6. When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Across morally framed prisoner's dilemmas and public goods games, none of nine LLMs consistently chooses the ethical action when it conflicts with payoff, with cooperation rates from 7.9% to 76.3%.

  7. BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

    cs.LG 2025-05 conditional novelty 5.0 of 10

    BenchHub is an automatically categorized, customizable LLM benchmark suite covering 303K questions across 38 benchmarks in English and Korean.

  8. The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.

Pith tools