A human-LLM collaborative pipeline yields EspanStereo, a multi-country Spanish stereotype dataset that exposes region-specific biases in Spanish LLMs and diverges sharply from English-centric resources.
SHADES : Towards a Multilingual Assessment of Stereotypes in Large Language Models
4 Pith papers cite this work, alongside 4 external citations. Polarity classification is still indexing.
fields
cs.CL 4years
2026 4representative citing papers
RedVox benchmark shows speech model safety and fairness vulnerabilities persist under non-adversarial conditions, worsen in non-English languages, and increase with spoken inputs.
Proposes a three-level taxonomy of Cultural Awareness, Cultural Sensitivity, and Cultural Competence for AI evaluation, grounded in intercultural communication scholarship to improve validity in multicultural contexts.
Meta-analysis of 33 ACL papers shows inconsistent LLM-as-a-Judge results, overtrust, and single-model reliance in multilingual/low-resource settings, with recommendations for better practice.
citing papers explorer
-
Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration
A human-LLM collaborative pipeline yields EspanStereo, a multi-country Spanish stereotype dataset that exposes region-specific biases in Spanish LLMs and diverges sharply from English-centric resources.
-
RedVox: Safety and Fairness Gaps in Speech Models Across Languages
RedVox benchmark shows speech model safety and fairness vulnerabilities persist under non-adversarial conditions, worsen in non-English languages, and increase with spoken inputs.
-
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
Proposes a three-level taxonomy of Cultural Awareness, Cultural Sensitivity, and Cultural Competence for AI evaluation, grounded in intercultural communication scholarship to improve validity in multicultural contexts.
-
Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
Meta-analysis of 33 ACL papers shows inconsistent LLM-as-a-Judge results, overtrust, and single-model reliance in multilingual/low-resource settings, with recommendations for better practice.