A cross-evaluation setup with native-speaker SME rubrics benchmarks LLMs on underrepresented Arabic dialects and quantifies automated judge bias and cultural-reasoning gaps.
arXiv preprint arXiv:2409.12623 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
Proposes a three-level taxonomy of Cultural Awareness, Cultural Sensitivity, and Cultural Competence for AI evaluation, grounded in intercultural communication scholarship to improve validity in multicultural contexts.
citing papers explorer
-
Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth
A cross-evaluation setup with native-speaker SME rubrics benchmarks LLMs on underrepresented Arabic dialects and quantifies automated judge bias and cultural-reasoning gaps.
-
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
Proposes a three-level taxonomy of Cultural Awareness, Cultural Sensitivity, and Cultural Competence for AI evaluation, grounded in intercultural communication scholarship to improve validity in multicultural contexts.