Pith. sign in

REVIEW 2 cited by

CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.13974 v1 pith:YCV7HXIU submitted 2024-05-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords acrosscivicsdatasetculturalllmssocialexperimentsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces the "CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal impacts" dataset, designed to evaluate the social and cultural variation of Large Language Models (LLMs) across multiple languages and value-sensitive topics. We create a hand-crafted, multilingual dataset of value-laden prompts which address specific socially sensitive topics, including LGBTQI rights, social welfare, immigration, disability rights, and surrogacy. CIVICS is designed to generate responses showing LLMs' encoded and implicit values. Through our dynamic annotation processes, tailored prompt design, and experiments, we investigate how open-weight LLMs respond to value-sensitive issues, exploring their behavior across diverse linguistic and cultural contexts. Using two experimental set-ups based on log-probabilities and long-form responses, we show social and cultural variability across different LLMs. Specifically, experiments involving long-form responses demonstrate that refusals are triggered disparately across models, but consistently and more frequently in English or translated statements. Moreover, specific topics and sources lead to more pronounced differences across model answers, particularly on immigration, LGBTQI rights, and social welfare. As shown by our experiments, the CIVICS dataset aims to serve as a tool for future research, promoting reproducibility and transparency across broader linguistic settings, and furthering the development of AI technologies that respect and reflect global cultural diversities and value pluralism. The CIVICS dataset and tools will be made available upon publication under open licenses; an anonymized version is currently available at https://huggingface.co/CIVICS-dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LKValues: Aligning Large Language Models with Sri Lankan Societal Values

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A survey-derived Sri Lankan value alignment suite (LKValues) with 150k instruction instances and a 1k benchmark improves Qwen-family LLMs' Sri Lankan value judgment in Sinhala and English.

  2. Understanding How Value Neurons Shape the Generation of Specified Values in LLMs

    cs.CL 2025-05 conditional novelty 5.0 of 10

    ValueLocate identifies value neurons via activation probability differences between opposing value prompts, and amplifying or suppressing these neurons alters G-EVAL value scores in four LLMs.

Pith tools