A cross-evaluation setup with native-speaker SME rubrics benchmarks LLMs on underrepresented Arabic dialects and quantifies automated judge bias and cultural-reasoning gaps.
Mind the gap in cultural alignment: Task-aware culture management for large language models
2 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CL 2years
2026 2roles
background 1polarities
background 1representative citing papers
DISCA converts within-country disagreement among World Values Survey personas into a bounded logit correction that reduces cultural misalignment by 10-24% on MultiTP for models 3.8B and larger across 20 countries, without any weight updates.
citing papers explorer
-
Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth
A cross-evaluation setup with native-speaker SME rubrics benchmarks LLMs on underrepresented Arabic dialects and quantifies automated judge bias and cultural-reasoning gaps.
-
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
DISCA converts within-country disagreement among World Values Survey personas into a bounded logit correction that reduces cultural misalignment by 10-24% on MultiTP for models 3.8B and larger across 20 countries, without any weight updates.