Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Low-resource languages are the strongest levers on an LLM's multilingual moral reasoning, for better and for worse.

desk verdict A usable multilingual moral benchmark, but the headline low-resource finding is an overgeneralization of target-specific effects. read the letter →

arxiv 2504.19759 v1 pith:D736RBCN submitted 2025-04-28 cs.CL

classification cs.CL
keywords multilingualmoralreasoninglow-resourcelanguagescross-lingualtransferalignmentfine-tuningdatapoisoningbenchmarkLLMevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper builds a five-language benchmark of moral scenarios at three levels of context and asks whether LLMs reason about morality consistently across languages, and which languages most affect cross-lingual behaviour. It finds that models are inconsistent across languages and that accuracy drops as context grows longer, with Vietnamese, the lowest-resource language in the set, hit hardest. The central result comes from fine-tuning a single open model: clean fine-tuning on Indonesian and Vietnamese improves other languages more than clean fine-tuning on English, while corrupted fine-tuning on Vietnamese damages other languages most. The paper argues that low-resource languages are high-leverage precisely because they are underrepresented in pretraining corpora, so fine-tuning has more room to change the model's moral associations in either direction.

What carries the argument

The load-bearing instrument is MMRB, a benchmark of 2,170 moral scenarios built from a sentence-level binary set, a paragraph-level reasoning set, and a document-level dilemma set with Virtue, Deontological, and Consequentialist answer branches, all translated into English, Chinese, Russian, Vietnamese, and Indonesian. The experimental machinery is monolingual fine-tuning: the same base model is trained on one language at a time, with correct labels for alignment and flipped labels for poisoning, and the cross-lingual accuracy deltas in Tables 3 and 4 provide the evidence for the resource-level claim.

What would settle it

Build an independent second translation of MMRB with item-level difficulty matched across languages, then rerun the monolingual alignment and poisoning experiments; if Indonesian and Vietnamese no longer dominate the cross-lingual deltas, the resource-level claim is an artifact of translation difficulty.

Watch

Extended reading notes

Core claim

The paper claims that a language's resource level, not its typological distance or the model's familiarity, predicts how much that language's data moves multilingual moral reasoning. Fine-tuning LLaMA-3-8B on correctly labeled Indonesian data raises accuracy on every other language in MMRB, with Vietnamese clean data the next strongest; flipping labels in Vietnamese causes the sharpest cross-lingual drops, while English poisoning barely moves the scores. These cross-lingual transfer effects support a mechanism in which underrepresented languages have more headroom for fine-tuning to fill knowledge gaps, so their data quality matters disproportionately.

Load-bearing premise

The five language versions of the translated scenarios are equally valid and equally difficult, so the larger transfer effects observed for Vietnamese and Indonesian reflect language resource levels rather than translation artifacts.

Editorial extensions

If this is right

  • Multilingual safety evaluation should treat low-resource languages as the most sensitive probes, since they show the largest alignment gains and the largest poisoning losses.
  • A relatively small amount of clean low-resource data can substitute for much larger English datasets in cross-lingual alignment.
  • A corrupted low-resource dataset can silently degrade moral behaviour in higher-resource languages, so curation effort should shift toward under-represented languages.
  • English-only benchmark scores overstate a model's cross-lingual moral consistency.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the proposed mechanism is pretraining underrepresentation, the same pattern should appear for other safety-sensitive abilities, such as refusing harmful requests or avoiding social bias, when fine-tuned on an equally low-resource language.
  • The resource-level effect predicts even larger transfer from languages with less pretraining representation than Vietnamese; adding Swahili or a similar language to the same benchmark would test that gradient.
  • Since the paper reports that document-level scenarios restore accuracy with explicit moral principles, an extension could test whether explicit structure also suppresses poisoning transfer, which would clarify the interaction between context length and language effects.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces MMRB, a multilingual moral reasoning benchmark with 2,170 scenarios in English, Chinese, Russian, Vietnamese, and Indonesian across three context complexity levels (sentence, paragraph, document). The authors evaluate five LLMs on MMRB, report cross-lingual inconsistencies and a degradation with context complexity, and then fine-tune LLaMA-3-8B on monolingual alignment and poisoning data. The headline claim is that low-resource languages, particularly Vietnamese and Indonesian, have a disproportionately strong positive and negative impact on multilingual moral reasoning, challenging the assumption that high-resource languages are always the most effective sources of cross-lingual transfer.

Significance. The dataset and the fine-tuning experiments are potentially useful contributions to multilingual evaluation of moral reasoning, and the finding that low-resource-language fine-tuning can produce large effects on low-resource evaluation targets is interesting and actionable. However, the central claim as stated in the abstract and Section 3.2.2 is broader than what the data support: the alignment results show low-resource training languages dominating only on low-resource evaluation targets, not on high-resource targets. If the claim is narrowed to the observed target-language-specific pattern, the contribution is still worthwhile but substantially more modest. The paper also ships no code or data in the reviewed version, and the experiments lack repeated runs or confidence intervals, limiting the strength of the comparative conclusions.

major comments (3)
  1. [Abstract and Section 3.2.2, Tables 3 and 4] The fine-tuning experiments in Tables 3 and 4 report single point estimates with no confidence intervals, no repeated runs, and no seed variation. Differences of a few accuracy points, which are used to rank languages by their cross-lingual impact, may be within run-to-run noise. At minimum, the authors should report multiple seeds with standard deviations or a statistical test over runs. This is load-bearing because the central comparative claim depends on the ordering of these deltas.
  2. [Section 3.2.1] The interpretation of the Mann-Whitney U tests is inverted. The text states that 'most comparisons across languages fail to reach significance, indicating inconsistency in multilingual moral reasoning.' Failure to reject the null hypothesis means there is no detectable difference, not that the languages are inconsistent. The authors should rephrase this observation or report the direction and effect sizes of the tests that do reach significance.
  3. [Section 3.1 and Limitations] The same translated MMRB corpus is used both for fine-tuning and for evaluation, and no external benchmark is used to validate the cross-lingual transfer findings. Because the translation pipeline (DeepSeek plus manual verification plus Google Translate cross-validation) is the only bridge between languages, the observed disproportionate effects of Vietnamese and Indonesian could reflect translation artifacts or dataset-specific properties rather than a general property of low-resource languages. The Limitations section acknowledges translation inaccuracies but provides no analysis of translation quality, item difficulty equivalence, or consistency across languages. Adding a small external validation set or a per-language difficulty analysis would substantially strengthen the claim.
minor comments (5)
  1. [Table 4] The ETHICSPRO rows in Table 4 are malformed: values such as '62.90↓9.2268.30↓3.8268.04↓4.0866.63↓5.4971.18↓0.94' run together without separators, making the table unreadable.
  2. [References] The reference list contains two very similar entries for Wei et al. (2023) and Wei et al. (2022) for chain-of-thought prompting; these should be consolidated or disambiguated clearly.
  3. [Table 1 and Figure 1] The label 'ETHICSBASE-A VG' in Table 1 and the spacing in Figure 1 ('ETHICS BASE', 'ETHICS PRO', 'ETHICS MAX') are inconsistent and should be normalized.
  4. [Section 3.2.1] The text refers to 'ETHICS BASE' with a space, while elsewhere it is 'ETHICSBASE'; please unify the naming.
  5. [Figure 2] The win-rate diagram in Figure 2 is difficult to parse because the language labels are densely packed and the ordering is not explained; a clearer layout or a table of pairwise win rates would help.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims are empirical readings of benchmark and fine-tuning results, not derived from fitted inputs, self-citation chains, or definitional equivalences.

full rationale

This is an empirical evaluation and fine-tuning study, not a derivation chain with equations or fitted parameters. The central claim that low-resource languages have a stronger positive and negative impact on multilingual moral reasoning is a direct reading of the delta tables (Tables 3 and 4), and the paper does not define 'low-resource impact' in terms of the quantities it then reports as findings. No parameter is fitted to a subset of data and then 'predicted' on a closely related quantity. No uniqueness theorem or load-bearing self-citation appears: the references include prior work on moral reasoning, but none of it is by the present authors, and none is invoked to forbid alternative explanations. The use of the same translated benchmark for fine-tuning and evaluation is a possible validity or leakage concern, and the skeptic's observation that the low-resource advantage is concentrated on low-resource evaluation targets is an overgeneralization critique, not a circularity critique. The limitations section explicitly acknowledges translation inaccuracies and limited generalizability, which further confirms that the authors do not claim to derive the result from prior assumptions. Thus, under the stated rules, there is no circular step to report and the score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities are present. The central claim rests on the benchmark's construct validity, translation fidelity, and single-model fine-tuning generalization, all of which are domain assumptions rather than fitted quantities.

assumptions (3)
  • domain assumption Accuracy on MMRB is a valid measure of moral reasoning ability.
    The entire evaluation and fine-tuning analysis treats MMRB accuracy as the target metric; the paper does not validate this against external moral reasoning benchmarks or human judgments.
  • domain assumption Machine translation plus manual verification preserves scenario difficulty and moral content across languages.
    Cross-language comparisons and low-resource claims rely on comparability of translated items; Section 3.1 describes translation and filtering but provides no quantitative quality check.
  • domain assumption Fine-tuning effects generalize beyond LLaMA-3-8B.
    The alignment and poisoning conclusions are drawn from one model family and one model size; no other architecture or scale is tested.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs." pith.science (2026). https://pith.science/paper/D736RBCN

@misc{pith2026250419759,
  author       = {Pith},
  title        = {Pith review of: Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/D736RBCN}},
  note         = {Machine review of arXiv:2504.19759}
}
read the original abstract

In this paper, we introduce the Multilingual Moral Reasoning Benchmark (MMRB) to evaluate the moral reasoning abilities of large language models (LLMs) across five typologically diverse languages and three levels of contextual complexity: sentence, paragraph, and document. Our results show moral reasoning performance degrades with increasing context complexity, particularly for low-resource languages such as Vietnamese. We further fine-tune the open-source LLaMA-3-8B model using curated monolingual data for alignment and poisoning. Surprisingly, low-resource languages have a stronger impact on multilingual reasoning than high-resource ones, highlighting their critical role in multilingual NLP.

Figures

Figures reproduced from arXiv: 2504.19759 by the authors.

Figure 1
Figure 1. The performance of three levels datasets across five LLMs in a multilingual setting. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. The win-rate between five languages in ETHICSPRO dataset based on Llama-3-8B. degradation in multilingual performance, while En￾glish poisoning has minimal effect. This result suggests that low-resource languages are not only impactful in positive transfer (alignment) but also particularly vulnerable to negative influence, likely because harmful data in these languages is less diluted during pretraining. These obser… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

    cs.CL 2026-07 conditional novelty 6.5 of 10

    MET-D self-distills theory-selected moral grounds into native-language reasoning, lifting macro-F1 by ~3.7–4.2 points on MCLASH and MMoralExceptQA while raising native-language chains by ~62 points.

Reference graph

Works this paper leans on

19 extracted references · 3 canonical work pages · cited by 1 Pith paper

  1. [1]

    AI@Meta. 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md Llama 3 model card

  2. [2]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long a...

  3. [3]

    Hwang, Maxwell Forbes, and Yejin Choi

    Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.54 Moral stories: Situated reasoning about norms, intents, actions, and their consequences . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 698--718, Online and Punta Cana, Dominica...

  4. [4]

    Katharina Haemmerl, Bjoern Deiseroth, Patrick Schramowski, Jind r ich Libovick \'y , Constantin Rothkopf, Alexander Fraser, and Kristian Kersting. 2023. https://doi.org/10.18653/v1/2023.findings-acl.134 Speaking multiple languages affects the moral bias of language models . In Findings of the Association for Computational Linguistics: ACL 2023, pages 2137...

  5. [5]

    Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020. Aligning ai with shared human values. arXiv preprint arXiv:2008.02275

  6. [6]

    Aditi Khandelwal, Utkarsh Agarwal, Kumar Tanmay, and Monojit Choudhury. 2024. Do moral judgment and reasoning capability of llms change with language? a study using the multilingual defining issues test. arXiv preprint arXiv:2402.02135

  7. [7]

    "Oops, Did I Just Say That?" Testing and Repairing Unethical Suggestions of Large Language Models with Suggest-Critique-Reflect Process

    Pingchuan Ma, Zongjie Li, Ao Sun, and Shuai Wang. 2023. https://arxiv.org/abs/2305.02626 "oops, did i just say that?" testing and repairing unethical suggestions of large language models with suggest-critique-reflect process . Preprint, arXiv:2305.02626

  8. [8]

    OpenAI. 2023. Gpt-4. https://openai.com/gpt-4

Show all 19 references
  1. [9]

    Abhinav Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, and Monojit Choudhury. 2023 a . Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in llms. arXiv preprint arXiv:2310.07251

  2. [10]

    Abhinav Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, and Monojit Choudhury. 2023 b . https://doi.org/10.18653/v1/2023.findings-emnlp.892 Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in LLM s . In Findings of the Associat...

  3. [11]

    Nino Scherrer, Claudia Shi, Amir Feder, and David Blei. 2024. Evaluating the moral beliefs encoded in llms. Advances in Neural Information Processing Systems, 36

  4. [12]

    Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2022. On the machine learning of ethical judgments from natural language. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational L...

  5. [13]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. https://arxiv.org/abs/2201.11903 Chain-of-thought prompting elicits reasoning in large language models . Preprint, arXiv:2201.11903

  6. [14]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837

  7. [15]

    Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, and Zheyan Luo. 2024. Llamafactory: Unified efficient fine-tuning of 100+ language models. arXiv preprint arXiv:2403.13372

  8. [16]

    Jingyan Zhou, Minda Hu, Junan Li, Xiaoying Zhang, Xixin Wu, Irwin King, and Helen Meng. 2023. Rethinking machine ethics--can llms perform moral reasoning through the lens of moral theories? arXiv preprint arXiv:2308.15399

  9. [17]

    Caleb Ziems, Jane A Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022. The moral integrity corpus: A benchmark for ethical dialogue systems. arXiv preprint arXiv:2204.03021

  10. [18]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  11. [19]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.