Pith. sign in

REVIEW 3 major objections 4 minor 41 references

The paper introduces Pancasila-Dilemmas, a 1,834-question benchmark of Indonesian news dilemmas, and shows that all 50 evaluated LLMs land at a Probability Match Score around 0.50 and a Max-Vote Agreement Score around 0.72, far below strong

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 16:12 UTC pith:ZKQT3QYN

load-bearing objection Useful first Pancasila-grounded Indonesian dilemma benchmark, but the abstract's headline numbers are contradicted by its own Table 2 and the 5-annotator ground truth is thin. the 3 major comments →

arxiv 2607.18066 v1 pith:ZKQT3QYN submitted 2026-07-20 cs.CL

Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila

classification cs.CL
keywords Pancasilavalue alignmentLLM evaluationIndonesian culturemoral dilemmasmultiple-choice benchmarkprobability match scoremax-vote agreement
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Pancasila-Dilemmas is a new benchmark of 1,834 multiple-choice dilemmas taken from real Indonesian news and labelled with one of the five Pancasila values: Religion, Humanity, Unity, Democracy, and Social Justice. Each dilemma was answered by five Indonesian citizens, so instead of a single gold answer the benchmark records a distribution of human preferences. The paper evaluates 50 LLMs and reports that the best models reach only about 0.50 Probability Match Score and 0.72 Max-Vote Agreement Score, meaning they capture only part of the human preference spread and frequently miss the majority answer. The hardest value categories are Religion and Unity, and models over-generate cooperative 'Collaborating' responses compared to the wider range of human conflict styles. The authors conclude that current LLMs are substantially misaligned with Indonesian value preferences and that regional training reduces but does not close the gap.

Core claim

The central claim is that nation-specific value alignment is not captured by current LLM training even when models are frontier, Indonesian-tuned, or prompted to roleplay a Pancasila-guided citizen. On the new benchmark, best models reach PMS around 0.50 and MVAS around 0.72 under both prompting conditions, far from the human reference; Religion and Unity dilemmas produce the lowest agreement and the largest variance, while Democracy and Humanity are easier because they resemble values well represented in global pretraining. The paper also finds that regionally adapted models outperform their backbones consistently but modestly, and that an explicit Pancasila prompt helps only models that al

What carries the argument

Pancasila-Dilemmas, a benchmark of 1,834 dilemma questions derived from Indonesian news, each with a scenario, a question, four plausible action options, a Pancasila value label, and five human annotator answers per item. Evaluation uses two complementary metrics: Probability Match Score (PMS), which gives the fraction of annotators who chose the model's option and thus measures agreement with the full preference distribution, and Max-Vote Agreement Score (MVAS), which checks whether the model picks the most common human answer. The paper also tags answer options with a five-way conflict-mode taxonomy (Competing, Accommodating, Avoiding, Collaborating, Compromising) to compare how model resp

Load-bearing premise

The load-bearing premise is that five annotators per question provide a stable, representative picture of Indonesian public value preferences; if that reference set is unrepresentative or unstable, the reported LLM-human gap is an artifact of the reference rather than a genuine value misalignment.

What would settle it

Re-answer a random sample of the 1,834 dilemmas with a larger, demographically stratified Indonesian panel, for example 50 respondents per question, and recompute the two scores. If the human majority answers or preference spreads shift enough that the same LLMs re-scored cross the reported 0.50 PMS / 0.72 MVAS thresholds, the paper's blanket misalignment conclusion would depend on its small annotation pool rather than on stable LLM behavior.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Models deployed in Indonesian-language settings will frequently select options that only a minority of Indonesians would choose in everyday ethical conflicts.
  • Religion and Unity dilemmas are the most likely failure points, so safety and alignment work for Indonesian users should prioritize those categories.
  • Adding Indonesian and Southeast Asian data during continued pretraining yields real but small alignment gains, so data exposure alone is not sufficient.
  • Asking models to adopt a Pancasila perspective helps only models already familiar with Indonesian values; for other models it adds ambiguity rather than improving answers.
  • Value benchmarks for pluralistic societies should record a distribution of human judgments over several annotators rather than assume a single correct answer.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the five-annotator reference shows only slight inter-annotator agreement (Fleiss kappa around 0.17–0.19), the reported thresholds are estimates against a noisy reference; a larger, stratified Indonesian panel could shift the numeric headlines even if the qualitative gap remains.
  • The finding that LLMs overuse the 'Collaborating' conflict style while humans spread across all five styles suggests models may be optimizing for pleasant, cooperative output generally rather than for Indonesian conflict norms; comparing the same models on other non-Western value benchmarks would test whether this pattern is global.
  • Since the questions and answer options are LLM-generated before human proofreading, the option space may carry a generation bias; a fully human-authored version of the benchmark would reveal whether part of the measured misalignment is a construction artifact.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces Pancasila-Dilemmas, a benchmark of 1,834 multiple-choice dilemma questions derived from Indonesian news and labeled by five Pancasila values. The dataset was proofread by native speakers and answered by 185 Indonesian participants with five annotators per item. The authors evaluate 50 closed- and open-source LLMs under two prompting conditions and report Probabilistic Match Score (PMS), Max-Vote Agreement Score (MVAS), and a normalized PMS. The headline claim is that all evaluated LLMs remain below 0.5 PMS and 0.72 MVAS, that LLMs struggle most on Religion and Unity dilemmas, and that this reveals a significant gap in Indonesian value alignment. The paper also analyzes open-ended responses using the Thomas-Kilmann Conflict Mode Instrument, finding that LLMs are overly biased toward Collaborating responses relative to humans.

Significance. If the dataset and human response distributions are accepted as a valid reference, this is a useful contribution to non-Western, nation-specific value alignment evaluation. The paper's strengths include a locally grounded scenario source (Kompas news), a relatively large model coverage (50 models, base and instruction-tuned, international and regional), two prompt settings, and evaluation metrics that explicitly target the distributional nature of subjective moral judgments rather than a single gold label. The release of the dataset and the demographic documentation of annotators are also valuable. These features make the benchmark potentially reusable for future Indonesian value-alignment work and for comparative studies of culturally grounded evaluation.

major comments (3)
  1. [Abstract and Table 2] The Abstract states: 'all evaluated LLMs achieves less than 0.5 Probability Match Score (PMS) and 0.72 Max-Vote Agreement Score (MVAS).' This universal claim is directly contradicted by the paper's own Table 2. Under the Universal Prompt, GPT-5.4 obtains PMS 0.5154 and MVAS 0.7317; Gemma-4-31B-IT obtains 0.5133 and 0.7279; Qwen-3.6-Max obtains 0.5129 and 0.7268; Kimi-K2.5 obtains 0.5104 and 0.7263. Under the Pancasila Prompt, Qwen-3.5-27B obtains PMS 0.5082 and MVAS 0.7235. Several other models also exceed the stated thresholds. The more cautious phrasing in §4.4 ('around 0.50' and 'around 0.72') is accurate, but the Abstract's quantitative claim is false as written. Because this is the paper's headline finding, it must be corrected before the results can be used as evidence for a 'significant gap'.
  2. [§4.4, Impact of Parameter Sizes] The text states that within the Qwen3 family under the Pancasila Prompt, performance improves from 'Qwen3-0.6B (PMS: 0.3226, MVAS: 0.4067)' to 'Qwen3-32B (PMS: 0.4995, MVAS: 0.7017)'. However, Table 2 lists Qwen3-0.6B-Base with PMS 0.3744 and MVAS 0.5038; no row in Table 2 reports 0.3226/0.4067. This is an internal inconsistency in a result used to support the scaling trend. The authors should either correct the numbers or clarify which model and setting these values refer to.
  3. [§3.4, §4.3, and Limitations] The paper's central conclusion that LLMs are 'substantially misaligned' with Indonesian public preferences rests on treating the five-annotator response distribution c(y) as the reference. With only five annotators per item, each item's distribution is restricted to multiples of 0.2 and has high sampling variance; the reported Fleiss kappa of 0.168–0.194 is in the 'slight agreement' range. The Limitations section acknowledges the small respondent pool, but the main quantitative claims do not quantify the resulting uncertainty. The authors should add robustness analyses, such as bootstrap confidence intervals for the aggregate PMS/MVAS, a comparison against a random-choice baseline, or per-item score distributions, to support the strength of the 'significant gap' conclusion. This is not an alternative to the current analysis but a necessary check on whether the reported gaps are larger th
minor comments (4)
  1. [Table 2 and §4.4] The model name is inconsistent: Table 2 uses 'Gemini-3.1-Pro' while the text refers to 'Gemini-3.1-Pro-Preview'. Please unify.
  2. [Table 2 caption] The caption says 'PMS_Norm is the PMS score normalized by human baseline with score 0.616,' but the text does not explain how the human baseline 0.616 was computed. A sentence with the definition (e.g., average self-agreement of annotators) would aid reproducibility.
  3. [Abstract] Grammar: 'all evaluated LLMs achieves' should be 'all evaluated LLMs achieve'; similar issues occur elsewhere (e.g., 'comprises of' in §3). A copyedit pass is recommended.
  4. [Figure 3] The colorbar range is 0.38–0.48, but many reported PMS values exceed 0.50. The heatmap may be clipped or using a different scale; please clarify the mapping.

Circularity Check

0 steps flagged

No significant circularity: PMS and MVAS compare LLM choices to independently collected human annotations; no fitted parameter or self-citation chain forces the reported gap.

full rationale

The evaluation chain is externally anchored. Human response counts c(y) come from 185 independent Indonesian annotators (Section 3.4), and LLM outputs are produced with fixed greedy decoding (Section 4.1). PMS and MVAS are direct comparison metrics, not fitted quantities: PMS(y_hat)=c(y_hat)/sum_y' c(y') and MVAS(y_hat)=1[y_hat in argmax_y c(y)] (Sections 4.3.1-4.3.2). Nothing in the metric definitions is calibrated to make LLMs score low; the 'around 0.50 / around 0.72' statement is an empirical summary of Table 2. The dataset-generation pipeline uses GPT-4o and native-speaker proofreading (Sections 3.2-3.3); this is a data-construction limitation acknowledged in the Limitations section, not circularity, because the reference standard is human preference, not model output. Same-lab citations in Section 2 (e.g., Xu et al. 2024/2025, Yu et al. 2024, Zheng et al. 2026, Zhu et al. 2026) are contextual related work and are not load-bearing for the benchmark's measurements. The Abstract's universal quantifier claim ('all evaluated LLMs achieves less than 0.5 PMS and 0.72 MVAS') is internally inconsistent with Table 2 entries such as GPT-5.4 under the Universal Prompt (PMS 0.5154, MVAS 0.7317) and Qwen-3.5-27B under the Pancasila Prompt (PMS 0.5082, MVAS 0.7235); however, that is a correctness/consistency issue, not circularity. No equation reduces a prediction to its input, and no fitted parameter is renamed as a result.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 0 invented entities

No fitted mathematical parameters drive the result; the numbers listed are design choices and a normalization constant. The benchmark itself is a construct, not an invented entity. The main external assumptions are about representativeness of the news corpus, the GPT-4o generation/proofreading loop, and the 5-annotator reference distribution; all are flagged in the Limitations or §3.

free parameters (3)
  • MinHash deduplication threshold = 0.2
    Hand-chosen in §3.1.1 to balance size and diversity, reducing 279,362 titles to 9,155; affects which news enters the benchmark.
  • Annotators per item = 5
    Design choice (§3.4) that fixes the human reference distribution to increments of 0.2 and makes agreement statistics coarse; central to PMS/MVAS.
  • Human PMS baseline for PMS_Norm = 0.616
    Normalization constant used in Table 2 and §4.4; reported without derivation, computed from the same annotation data.
axioms (6)
  • domain assumption Pancasila's five values are a complete and adequate taxonomy for Indonesian value dilemmas.
    All 1,834 items are labeled into Religion/Humanity/Unity/Democracy/Social Justice (§3.2, Table 1); if the taxonomy omits conflicts, the benchmark's coverage claim fails.
  • domain assumption The 5-annotator human distribution is an adequate proxy for the Indonesian public's preferences.
    PMS/MVAS compare models to c(y) from five annotators (§4.3). Low Fleiss kappa (0.168-0.194) and a convenience sample of 185 paid participants make this assumption fragile.
  • domain assumption GPT-4o-generated scenarios and options, after proofreading, are unbiased representations of real dilemmas.
    Generation pipeline in §3.2 uses GPT-4o; Limitations admits model-generation bias; if options are artificial, results measure the generator rather than Indonesian values.
  • domain assumption Single-letter forced choice under temperature=0 measures value preference.
    The evaluation protocol (§4.1) assumes the selected option reflects alignment, not formatting behavior or social desirability; §5.2 shows LLMs overuse 'Collaborating'.
  • domain assumption Kompas 2024 headlines classified by GPT-4o are representative of daily Indonesian dilemmas.
    Single-news-source, single-year corpus (§3.1); classification by GPT-4o with only 300 manually annotated titles may introduce selection bias.
  • domain assumption TKI mode labels from GPT-4o-Mini are valid for both human and LLM responses.
    Used in §5.2 open-ended analysis; no validation or agreement statistics for the auto-labeling are provided.

pith-pipeline@v1.3.0-alltime-deepseek · 16576 in / 16976 out tokens · 162811 ms · 2026-08-01T16:12:32.144747+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila." pith.science (2026). https://pith.science/paper/ZKQT3QYN

@misc{pith2026260718066,
  author       = {Pith},
  title        = {Pith review of: Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZKQT3QYN}},
  note         = {Machine review of arXiv:2607.18066}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The value alignment of large language models (LLMs) is crucial for ensuring responses align with human intention and value preferences. However, most evaluations of value alignment focus on Western or universal values, while assessments grounded in the value systems of specific countries remain scarce. In this paper, we introduce Pancasila-Dilemmas, an evaluation dataset of 1,834 questions derived from Indonesian news, classified by 5 values of Pancasila: Religion, Humanity, Unity, Democracy, and Social Justice. This dataset reflects daily life in Indonesia, making it suitable for measuring the value alignment of LLMs deployed for Indonesia. To ensure a more rigorous evaluation, we choose scenarios containing dilemmas. The dataset is proofread by native speakers and answered by 5 diverse Indonesian citizens. We evaluate 50 closed- and open-source LLMs on our dataset. Results reveal that all evaluated LLMs achieves less than 0.5 Probability Match Score (PMS) and 0.72 Max-Vote Agreement Score (MVAS). Compared by each values, LLMs mostly struggle in Religion and Unity dilemma cases. This highlights a significant gap in capturing Indonesian values. The dataset is publicly available at https://github.com/tjunlp-lab/Pancasila-Dilemmas.

Figures

Figures reproduced from arXiv: 2607.18066 by Darren Keanly Martin, Deyi Xiong, Irfan, Jayvin Fernando, Julianti, Supryadi, Yuqi Ren.

Figure 1
Figure 1. Figure 1: An illustrated example (AI generated) of the [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The framework of our Pancasila-Dilemmas data curation and evaluation settings. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The heatmaps of PMS metrics for each LLM at different value categories. The darker colors are the best [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The results of TKI conflict resolution between [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Screenshot of the platform for dataset validation. [PITH_FULL_IMAGE:figures/full_fig_p014_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 9 canonical work pages

  1. [1]

    Gema Keadilan , volume =

    Marshandha Ardhani and Irma Utaminingsih and Izzati Ardana and Riska Fitriono , title =. Gema Keadilan , volume =. 2022 , keywords =. doi:10.14710/gk.2022.16167 , url =

  2. [2]

    Widyadari , author=

    IMPLEMENTASI NILAI NILAI PANCASILA DALAM PENGUATAN KARAKTER BANGSA , volume=. Widyadari , author=. 2020 , month=

  3. [3]

    K or NAT : LLM Alignment Benchmark for K orean Social Values and Common Knowledge

    Lee, Jiyoung and Kim, Minwoo and Kim, Seungho and Kim, Junghwan and Won, Seunghyun and Lee, Hwaran and Choi, Edward. K or NAT : LLM Alignment Benchmark for K orean Social Values and Common Knowledge. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1/2024.findings-acl.666

  4. [4]

    First Conference on Language Modeling , year=

    Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks? , author=. First Conference on Language Modeling , year=

  5. [5]

    ``Seeing the Big through the Small'': Can LLM s Approximate Human Judgment Distributions on NLI from a Few Explanations?

    Chen, Beiduo and Wang, Xinpeng and Peng, Siyao and Litschko, Robert and Korhonen, Anna and Plank, Barbara. ``Seeing the Big through the Small'': Can LLM s Approximate Human Judgment Distributions on NLI from a Few Explanations?. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024. doi:10.18653/v1/2024.findings-emnlp.842

  6. [6]

    A Multi-Labeled Dataset for I ndonesian Discourse: Examining Toxicity, Polarization, and Demographics Information

    Susanto, Lucky and Wijanarko, Musa Izzanardi and Pratama, Prasetia Anugrah and Tang, Zilu and Akyas, Fariz and Hong, Traci and Idris, Ika Karlina and Aji, Alham Fikri and Wijaya, Derry Tanti. A Multi-Labeled Dataset for I ndonesian Discourse: Examining Toxicity, Polarization, and Demographics Information. Findings of the Association for Computational Ling...

  7. [7]

    COPAL - ID : I ndonesian Language Reasoning with Local Culture and Nuances

    Wibowo, Haryo and Fuadi, Erland and Nityasya, Made and Prasojo, Radityo Eko and Aji, Alham. COPAL - ID : I ndonesian Language Reasoning with Local Culture and Nuances. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. doi:10.18653/v1...

  8. [8]

    Large Language Models Only Pass Primary School Exams in I ndonesia: A Comprehensive Test on I ndo MMLU

    Koto, Fajri and Aisyah, Nurul and Li, Haonan and Baldwin, Timothy. Large Language Models Only Pass Primary School Exams in I ndonesia: A Comprehensive Test on I ndo MMLU. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.760

  9. [9]

    I ndo C ulture: Exploring Geographically Influenced Cultural Commonsense Reasoning Across Eleven I ndonesian Provinces

    Koto, Fajri and Mahendra, Rahmad and Aisyah, Nurul and Baldwin, Timothy. I ndo C ulture: Exploring Geographically Influenced Cultural Commonsense Reasoning Across Eleven I ndonesian Provinces. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00726

  10. [10]

    Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in I ndonesia

    Koto, Fajri. Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in I ndonesia. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track). 2025. doi:10.18653/v1/2025.naacl-industry.69

  11. [11]

    I ndo S afety: Culturally Grounded Safety for LLM s in I ndonesian Languages

    Azmi, Muhammad Falensi and Al Kautsar, Muhammad Dehan and Wicaksono, Alfan Farizki and Koto, Fajri. I ndo S afety: Culturally Grounded Safety for LLM s in I ndonesian Languages. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.465

  12. [12]

    One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in I ndonesia

    Aji, Alham Fikri and Winata, Genta Indra and Koto, Fajri and Cahyawijaya, Samuel and Romadhony, Ade and Mahendra, Rahmad and Kurniawan, Kemal and Moeljadi, David and Prasojo, Radityo Eko and Baldwin, Timothy and Lau, Jey Han and Ruder, Sebastian. One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in I ndonesia. Proceed...

  13. [13]

    CM oral E val: A Moral Evaluation Benchmark for C hinese Large Language Models

    Yu, Linhao and Leng, Yongqi and Huang, Yufei and Wu, Shang and Liu, Haixin and Ji, Xinmeng and Zhao, Jiahui and Song, Jinwang and Cui, Tingting and Cheng, Xiaoqing and Liutao, Liutao and Xiong, Deyi. CM oral E val: A Moral Evaluation Benchmark for C hinese Large Language Models. Findings of the Association for Computational Linguistics: ACL 2024. 2024. do...

  14. [14]

    CoRR , volume =

    Ping Wu and Guobin Shen and Dongcheng Zhao and Yuwei Wang and Yiting Dong and Yu Shi and Enmeng Lu and Feifei Zhao and Yi Zeng , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2506.01495 , eprinttype =. 2506.01495 , timestamp =

  15. [15]

    Benchmarking Multi-National Value Alignment for Large Language Models

    Ju, Chengyi and Shi, Weijie and Liu, Chengzhong and Ji, Jiaming and Zhang, Jipeng and Zhang, Ruiyuan and Xu, Jiajie and Yang, Yaodong and Han, Sirui and Guo, Yike. Benchmarking Multi-National Value Alignment for Large Language Models. Findings of the Association for Computational Linguistics: ACL 2025. 2025. doi:10.18653/v1/2025.findings-acl.1028

  16. [16]

    Taylor Sorensen and Liwei Jiang and Jena D. Hwang and Sydney Levine and Valentina Pyatkin and Peter West and Nouha Dziri and Ximing Lu and Kavel Rao and Chandra Bhagavatula and Maarten Sap and John Tasioulas and Yejin Choi , editor =. Value Kaleidoscope: Engaging. Thirty-Eighth. 2024 , url =. doi:10.1609/AAAI.V38I18.29970 , timestamp =

  17. [17]

    Hannah Rose Kirk and Alexander Whitefield and Paul R. The. Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 , year =

  18. [18]

    Exploring Multilingual Concepts of Human Values in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages? , booktitle =

    Shaoyang Xu and Weilong Dong and Zishan Guo and Xinwei Wu and Deyi Xiong , editor =. Exploring Multilingual Concepts of Human Values in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages? , booktitle =. 2024 , url =. doi:10.18653/V1/2024.FINDINGS-EMNLP.96 , timestamp =

  19. [19]

    Learning Human-like Representations to Enable Learning Human Values , booktitle =

    Andrea Wynn and Ilia Sucholutsky and Tom Griffiths , editor =. Learning Human-like Representations to Enable Learning Human Values , booktitle =. 2024 , url =

  20. [20]

    Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations , booktitle =

    Chaoyi Xiang and Chunhua Liu and Simon De Deyne and Lea Frermann , editor =. Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations , booktitle =. 2025 , url =

  21. [21]

    Can Language Models Reason about Individualistic Human Values and Preferences? , booktitle =

    Liwei Jiang and Taylor Sorensen and Sydney Levine and Yejin Choi , editor =. Can Language Models Reason about Individualistic Human Values and Preferences? , booktitle =. 2025 , url =

  22. [22]

    Long Ouyang and Jeffrey Wu and Xu Jiang and Diogo Almeida and Carroll L. Wainwright and Pamela Mishkin and Chong Zhang and Sandhini Agarwal and Katarina Slama and Alex Ray and John Schulman and Jacob Hilton and Fraser Kelton and Luke Miller and Maddie Simens and Amanda Askell and Peter Welinder and Paul F. Christiano and Jan Leike and Ryan Lowe , editor =...

  23. [23]

    The Thirteenth International Conference on Learning Representations,

    Yu Ying Chiu and Liwei Jiang and Yejin Choi , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  24. [24]

    Gemma 2: Improving Open Language Models at a Practical Size , journal =

    Morgane Rivi. Gemma 2: Improving Open Language Models at a Practical Size , journal =. 2024 , url =. doi:10.48550/ARXIV.2408.00118 , eprinttype =. 2408.00118 , timestamp =

  25. [25]

    CoRR , volume =

    Gemma Team , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2503.19786 , eprinttype =. 2503.19786 , timestamp =

  26. [26]

    CoRR , volume =

    An Yang and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chengyuan Li and Dayiheng Liu and Fei Huang and Haoran Wei and Huan Lin and Jian Yang and Jianhong Tu and Jianwei Zhang and Jianxin Yang and Jiaxi Yang and Jingren Zhou and Junyang Lin and Kai Dang and Keming Lu and Keqin Bao and Kexin Yang and Le Yu and Mei Li and Mi...

  27. [27]

    CoRR , volume =

    An Yang and Anfeng Li and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chang Gao and Chengen Huang and Chenxu Lv and Chujie Zheng and Dayiheng Liu and Fan Zhou and Fei Huang and Feng Hu and Hao Ge and Haoran Wei and Huan Lin and Jialong Tang and Jian Yang and Jianhong Tu and Jianwei Zhang and Jian Yang and Jiaxi Yang and Ji...

  28. [28]

    CoRR , volume =

    Llama Team , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2407.21783 , eprinttype =. 2407.21783 , timestamp =

  29. [29]

    S ea LLM s 3: Open Foundation and Chat Multilingual Large Language Models for S outheast A sian Languages

    Zhang, Wenxuan and Chan, Hou Pong and Zhao, Yiran and Aljunied, Mahani and Wang, Jianyu and Liu, Chaoqun and Deng, Yue and Hu, Zhiqiang and Xu, Weiwen and Chia, Yew Ken and Li, Xin and Bing, Lidong. S ea LLM s 3: Open Foundation and Chat Multilingual Large Language Models for S outheast A sian Languages. Proceedings of the 2025 Conference of the Nations o...

  30. [30]

    CoRR , volume =

    Raymond Ng and Thanh Ngan Nguyen and Yuli Huang and Ngee Chia Tai and Wai Yi Leong and Wei Qi Leong and Xianbin Yong and Jian Gang Ngui and Yosephine Susanto and Nicholas Cheng and Hamsawardhini Rengarajan and Peerat Limkonchotiwat and Adithya Venkatadri Hulagadri and Kok Wai Teng and Yeo Yeow Tong and Bryan Siow and Wei Yi Teo and Wayne Lau and Choon Men...

  31. [31]

    Gomez and Phil Blunsom and Marzieh Fadaee and Ahmet

    Viraat Aryabumi and John Dang and Dwarak Talupuru and Saurabh Dash and David Cairuz and Hangyu Lin and Bharat Venkitesh and Madeline Smith and Jon Ander Campos and Yi Chern Tan and Kelly Marchisio and Max Bartolo and Sebastian Ruder and Acyr Locatelli and Julia Kreutzer and Nick Frosst and Aidan N. Gomez and Phil Blunsom and Marzieh Fadaee and Ahmet. Aya ...

  32. [32]

    Gomez and Ivan Zhang and Phil Blunsom and Nick Frosst and Joelle Pineau and Beyza Ermis and Ahmet

    Alejandro Salamanca and Diana Abagyan and Daniel D'souza and Ammar Khairi and David Mora and Saurabh Dash and Viraat Aryabumi and Sara Rajaee and Mehrnaz Mofakhami and Ananya Sahu and Thomas Euyang and Brittawnya Prince and Madeline Smith and Hangyu Lin and Acyr Locatelli and Sara Hooker and Tom Kocmi and Aidan N. Gomez and Ivan Zhang and Phil Blunsom and...

  33. [33]

    Kilmann and Kenneth W

    Ralph H. Kilmann and Kenneth W. Thomas , title =. Educational and Psychological Measurement , volume =. 1977 , doi =. https://doi.org/10.1177/001316447703700204 , abstract =

  34. [34]

    Zhao and Kelvin Guu and Adams Wei Yu and Brian Lester and Nan Du and Andrew M

    Jason Wei and Maarten Bosma and Vincent Y. Zhao and Kelvin Guu and Adams Wei Yu and Brian Lester and Nan Du and Andrew M. Dai and Quoc V. Le , title =. The Tenth International Conference on Learning Representations,. 2022 , url =

  35. [35]

    International Journal of Computer Science and Humanitarian AI , author=

    Design and Implementation of Chatbot Pancasila for Teaching Pancasila and Character Building for University’s Students , volume=. International Journal of Computer Science and Humanitarian AI , author=. 2025 , month=. doi:10.21512/ijcshai.v2i2.14418 , abstractNote=

  36. [36]

    Gema Keadilan , volume =

    Erlina Aryani and Nurhalisa Fadjrin and Tsania Azzahro’ and Riska Fitriono , title =. Gema Keadilan , volume =. 2022 , keywords =. doi:10.14710/gk.2022.16430 , url =

  37. [37]

    Manning and Stefano Ermon and Chelsea Finn , editor =

    Rafael Rafailov and Archit Sharma and Eric Mitchell and Christopher D. Manning and Stefano Ermon and Chelsea Finn , editor =. Direct Preference Optimization: Your Language Model is Secretly a Reward Model , booktitle =. 2023 , url =

  38. [38]

    Zhihong Shao and Peiyi Wang and Qihao Zhu and Runxin Xu and Junxiao Song and Mingchuan Zhang and Y. K. Li and Y. Wu and Daya Guo , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2402.03300 , eprinttype =. 2402.03300 , timestamp =

  39. [39]

    Self-Pluralising Culture Alignment for Large Language Models

    Xu, Shaoyang and Leng, Yongqi and Yu, Linhao and Xiong, Deyi. Self-Pluralising Culture Alignment for Large Language Models. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. doi:10.18653/v1/2025.naacl-long.350

  40. [40]

    Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q -Sorts

    Zheng, Jingting and Ren, Yuqi and Yu, Linhao and Leng, Yongqi and Xiong, Deyi. Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q -Sorts. Proceedings of the 64th Annual Meeting of the A ssociation for C omputational L inguistics (Volume 1: Long Papers). 2026. doi:10.18653/v1/2026.acl-long.830

  41. [41]

    DVM ap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping

    Zhu, Pengyun and Ren, Yuqi and Wang, Zhen and Yang, Lei and Xiong, Deyi. DVM ap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping. Proceedings of the 64th Annual Meeting of the A ssociation for C omputational L inguistics (Volume 1: Long Papers). 2026. doi:10.18653/v1/2026.acl-long.909