REVIEW 3 major objections 5 minor 99 references
Out of Sight Out of Mind, Out of Sight Out of Mind: Measuring Bias in Language Models Against Overlooked Marginalized Groups in Regional Contexts
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Across 23 language models and 270 marginalized groups in 25 countries, the paper reports that English and German models show higher offensive stereotyping bias against marginalized groups than dominant groups, while Arabic MLMs score high…
desk verdict A genuinely new bias-audit resource for overlooked Arab-world groups, but the MLM metric that carries the main claims needs validation before those claims should be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two measurement instruments carry the argument. The SOS_MLM metric, introduced here, compares the pseudo-log-likelihood that a masked language model assigns to a toxic sentence versus its non-toxic template twin containing the same identity term; the bias score is the fraction of identity-prompt pairs for which the toxic version is more probable. For generative models, the paper adapts the HONEST metric, which counts how often the model's top-K completions contain a hurtful word from the HurtLex lexicon. The dataset built for this study contains 270 marginalized and 60 dominant identity groups across six sensitive attributes, in English, German, Modern Standard Arabic, and Egyptian Arabic, and is what allows group-level and dialect-level comparisons for the first time.
What would settle it
To settle the central claim, run the SOS_MLM and HONEST metrics on matched templates where toxic and non-toxic versions are equated for length, frequency, and register, and have human raters score offensiveness for the same prompts. If the metric's group-level ordering diverges from human judgments, or if the Egyptian-MSA gap disappears under matched templates, the paper's bias comparisons are measurement artifacts rather than model bias.
Extended reading notes
Core claim
The central claim is that offensive stereotyping bias in language models is a regional and low-resource-language problem that current English-centered benchmarks fail to capture. The paper's three main results are: English and German models consistently show higher SOS bias against marginalized groups than against dominant groups; Arabic MLMs instead score high bias against both marginalized and dominant groups for religion and ethnicity, which the authors trace to pretraining on Western news translated into Arabic; and measuring with Egyptian Arabic yields significantly higher bias scores than Modern Standard Arabic, a gap attributed to under-representation of Egyptian content in pretraining corpora. The paper also reports pronounced intersectional bias against non-binary, LGBTQIA+, and Black women, and documents that multilingual instruction-following models hallucinate far more in Arabic than in English. Finally, it shows that the HONEST metric's reliance on a smaller Arabic HurtLex makes cross-language bias comparisons unreliable.
Load-bearing premise
The load-bearing premise is that the two bias scores measure bias against the identity groups themselves; if template wording, translation choices, or the different sizes of the hurtful-word lexicon across languages dominate the scores, the paper's group and dialect comparisons do not measure what they claim.
Editorial extensions
If this is right
- Existing bias benchmarks that cover only US and English-speaking groups will miss the majority of marginalized groups worldwide; regional group inventories are needed.
- Low-resource languages and dialects cannot be audited with English-derived metrics: HurtLex's 3,360 English entries versus 1,147 Arabic and 2,043 German entries mean cross-language HONEST comparisons understate non-English bias.
- If Arabic models inherit Western-media stereotypes through translated pretraining data, training on local, representative sources is a concrete debiasing lever.
- Higher measured bias in Egyptian Arabic than MSA implies dialect-specific evaluation is required and that dialect under-representation is itself a bias mechanism.
- Intersectional identities, especially non-binary and Black women, need separate measurement; aggregate group scores hide them.
Reading between the lines
- Because template pairs differ by more than toxicity, part of the measured SOS_MLM score could reflect lexical or syntactic preferences; the authors do not validate the metric against human offensiveness judgments, so a human-annotation study would separate bias from artifact.
- Low SOS_MLM scores for identities absent from Arabic training data, such as 'Muhamash' or 'Ahmadi', may indicate that the model has no association for them at all; interpreting those low scores as 'low bias' would be a misreading.
- If the dialect gap is driven by pretraining-corpus composition, the same protocol applied to another underrepresented Arabic dialect should show a similar gap; this is a testable extension the paper does not run.
- The Arabic-MLM result of high bias against dominant groups predicts that retraining on local rather than translated data would lower bias scores for both dominant and marginalized groups; that prediction could be checked by fine-tuning the same architecture on curated local corpora.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a large-scale audit of offensive stereotyping bias (SOS) in 23 language models across 270 marginalized groups and 60 dominant groups in Egypt, the remaining 21 Arab countries, Germany, the UK, and the US, using English, German, Modern Standard Arabic (MSA), and Egyptian Arabic. The authors build a new SOS dataset of 72,000 sentences, extend the HONEST dataset, and propose a new masked-language-model metric, SOS_MLM, based on pseudo-log-likelihood comparisons of toxic and non-toxic template pairs. They report that instruction-following models hallucinate substantially more on Arabic instructions, that HONEST scores are lower in non-English languages partly because of HurtLex coverage gaps, that Egyptian Arabic yields higher bias scores than MSA, that most MLMs show higher bias against marginalized than dominant groups except Arabic MLMs on religion and ethnicity, and that intersectional bias is especially high for non-binary and Black women. The paper also tests the dialect hypothesis with CamelBERT-Da and finds higher Egyptian than MSA scores. The authors conclude that existing bias evaluation, which is largely US/English-centered, is incomplete for regional and low-resource contexts.
Significance. If the measurement claims hold, this is a valuable contribution because it broadens bias evaluation beyond US/English settings, introduces regional identity groups that are rarely studied, and provides a multilingual dataset and code. The paper has concrete strengths: it ships data and code, it includes an additional dialect-trained model (CamelBERT-Da) as a falsifiable probe of the pretraining-data hypothesis, it reports qualitative examples of toxic completions, and it explicitly acknowledges the HurtLex coverage limitations for non-English languages. However, the significance is conditional because the paper's central comparisons for RQ2 rest on the SOS_MLM metric, whose construct validity is not established, and because the repeated use of 'significantly higher' is not backed by inferential tests. The dataset and descriptive findings are useful, but the main scientific claims require additional validation.
major comments (3)
- [4.3/Eq. 1] The construct validity of the proposed SOS_MLM metric is not established, and this is load-bearing for RQ2 and for the Egyptian-vs-MSA claim. In Eq. (1), score(S) is the sum of log pseudo-likelihoods of all unmodified tokens U, which include the identity token plus shared frame tokens such as 'Being', 'a', 'person', 'is'. The frame-token contribution is identical across identities, so the comparison score(S) > score(S') can be dominated by whether the shared frame fits toxic versus non-toxic fillers, rather than by an identity-specific stereotype. The identity-specific signal is confined to log P(identity | M), and the paper provides no neutral-noun control, no matched toxic/non-toxic filler sets controlled for frequency and register, and no human-annotation validation to show that this term drives the reported marginalized-vs-dominant and dialect differences. Without such a control or validation, the metric may be measuring template-level toxicity preference rather than bias against specific identity groups.
- [4.2/4.3/5.1] The paper repeatedly uses 'significantly higher' (e.g., English HONEST scores, Egyptian versus MSA SOS_MLM scores) without any inferential statistics. Table 3 and Fig. 1a report only means; there are no confidence intervals, paired tests, or effect sizes. Because SOS_MLM and HONEST scores are aggregates over many sentence pairs, bootstrap or mixed-effects analyses are feasible and would substantiate the claimed differences. As written, the dialect and cross-language comparisons are descriptive trends, not statistically supported findings.
- [4.2/5.1] The cross-language comparison of HONEST scores is confounded by lexicon coverage. The paper itself notes that HurtLex has 3360 English entries versus 1147 Arabic and 2043 German entries, and that the Arabic lexicon includes English hurtful words, yet Sec. 4.2 states that 'HONEST scores are significantly higher for the English dataset' without adjusting for coverage. The authors do acknowledge this limitation later, but the main quantitative claim in Sec. 4.2 remains misleading and should be either re-analyzed with coverage-matched lexicons or explicitly re-framed as a metric-artifact hypothesis. This does not necessarily invalidate the qualitative examples, but it weakens the quantitative cross-language and cross-dialect HONEST comparisons.
minor comments (5)
- [Title] The title as submitted to arXiv contains 'Out of Sight Out of Mind' twice; the in-text title uses it once. This duplication should be corrected.
- [3.2] The dataset-size statement is hard to verify: 72,000 sentences is not derived from the stated 37 toxic and 37 non-toxic templates, the number of identity terms, and the three gender variants. Please provide the exact arithmetic or a table clarifying how the counts are obtained.
- [5.1] In the third item of Sec. 5.1, 'we hypothesis that' should be 'we hypothesize that'; also the later sentence 'This internet access gap relates to income' is a fragment that should be joined to the previous sentence.
- [References] Reference [91] 'Askari, S.' has no title, venue, or year, and reference [49] 'Queerinai, O. O. et al.' is malformed. Please complete these entries.
- [A.3] The appendix figures are dense and hard to read; consider labeling the panels with the model names and sensitive attributes directly in the figure rather than only in the caption.
Circularity Check
No circular derivation: all bias scores are fresh measurements; the only self-citations are terminological and non-load-bearing.
full rationale
The paper does not fit parameters and then predict a closely related quantity. The central results—MLM SOS_MLM scores (Eq. 1), HONEST scores for generative models, and the Egyptian-vs-MSA comparison—are computed directly from model outputs on newly constructed datasets. Eq. 1 is an adapted CrowS-Pairs-style pseudo-log-likelihood comparison between toxic and non-toxic template pairs; it is a stipulated operationalization of offensive stereotyping bias, not a conclusion that is defined into existence, because the marginalized-vs-dominant and dialect comparisons are empirical contrasts across identities and languages. The term "SOS" is attributed to the authors' earlier word-embedding work [15], and the Perspective-API discussion cites the second author's audit [25], but neither citation supplies a theorem, fitted parameter, or forced choice; the present measurements stand on their own. The paper explicitly acknowledges the main metric-validity threats—HurtLex coverage differences (3360 English vs 1147 Arabic and 2043 German entries, Sec 4.2 and 5.1), hallucination-driven performance drops (Table 2), and lexical measures missing implicit stereotypes (Sec 5.1)—which are correctness and construct-validity risks, not circularity. The CamelBERT-Da check in Sec 5.1 is an external falsification test of the dialect-under-representation hypothesis. Accordingly, there are no circular steps; the low score reflects only minor terminological self-citation.
Assumptions & free parameters
free parameters (1)
- HONEST top-k (K) =
not reported
assumptions (4)
- domain assumption Marginalized group lists from Minority Rights and UNHCR accurately represent the marginalized groups of each country.
- domain assumption Translated templates and identity labels are semantically equivalent across English, German, MSA, and Egyptian Arabic.
- domain assumption HurtLex provides comparable hurtful-word coverage across the four language varieties.
- domain assumption Pseudo-log-likelihood comparisons of toxic vs non-toxic template pairs isolate identity bias rather than template artifacts.
Cite this review
Pith. "Pith review of Out of Sight Out of Mind, Out of Sight Out of Mind: Measuring Bias in Language Models Against Overlooked Marginalized Groups in Regional Contexts." pith.science (2026). https://pith.science/paper/AXXMCMFB
@misc{pith2026250412767,
author = {Pith},
title = {Pith review of: Out of Sight Out of Mind, Out of Sight Out of Mind: Measuring Bias in Language Models Against Overlooked Marginalized Groups in Regional Contexts},
year = {2026},
howpublished = {\url{https://pith.science/paper/AXXMCMFB}},
note = {Machine review of arXiv:2504.12767}
}
read the original abstract
We know that language models (LMs) form biases and stereotypes of minorities, leading to unfair treatments of members of these groups, thanks to research mainly in the US and the broader English-speaking world. As the negative behavior of these models has severe consequences for society and individuals, industry and academia are actively developing methods to reduce the bias in LMs. However, there are many under-represented groups and languages that have been overlooked so far. This includes marginalized groups that are specific to individual countries and regions in the English speaking and Western world, but crucially also almost all marginalized groups in the rest of the world. The UN estimates, that between 600 million to 1.2 billion people worldwide are members of marginalized groups and in need for special protection. If we want to develop inclusive LMs that work for everyone, we have to broaden our understanding to include overlooked marginalized groups and low-resource languages and dialects. In this work, we contribute to this effort with the first study investigating offensive stereotyping bias in 23 LMs for 270 marginalized groups from Egypt, the remaining 21 Arab countries, Germany, the UK, and the US. Additionally, we investigate the impact of low-resource languages and dialects on the study of bias in LMs, demonstrating the limitations of current bias metrics, as we measure significantly higher bias when using the Egyptian Arabic dialect versus Modern Standard Arabic. Our results show, LMs indeed show higher bias against many marginalized groups in comparison to dominant groups. However, this is not the case for Arabic LMs, where the bias is high against both marginalized and dominant groups in relation to religion and ethnicity. Our results also show higher intersectional bias against Non-binary, LGBTQIA+ and Black women.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Bridging the AI Governance Divide
LaForge, G.; , R.; Seiler, G. Bridging the AI Governance Divide. https://www.t20brasil.org/media/documentos/arquivos/TF05_ST_05_Bridging_the_ AI_gov66cdcbf06f991.pdf, 2024
2024
-
[2]
When do Word Embeddings Accurately Reflect Surveys on our Beliefs About People? Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
Joseph, K.; Morgan, J. When do Word Embeddings Accurately Reflect Surveys on our Beliefs About People? Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Online, 2020; pp 4392–4415
2020
-
[3]
I.; Nenkova, A
Agarwal, O.; Durupınar, F.; Badler, N. I.; Nenkova, A. Word Embeddings (Also) Encode Human Personality Stereotypes. Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019). Minneapolis, Minnesota, 2019; pp 205–211
2019
-
[4]
J.; Narayanan, A
Caliskan, A.; Bryson, J. J.; Narayanan, A. Semantics derived automatically from language corpora contain human-like biases. Science 2017, 356, 183–186
2017
-
[5]
Nangia, N.; Vania, C.; Bhalerao, R.; Bowman, S. R. CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Online, 2020; pp 1953–1967
2020
-
[6]
StereoSet: Measuring stereotypical bias in pretrained language models
Nadeem, M.; Bethke, A.; Reddy, S. StereoSet: Measuring stereotypical bias in pretrained language models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Online, 2021; pp 5356–5371
2021
-
[7]
I’m sorry to hear that
Smith, E. M.; Hall, M.; Kambadur, M.; Presani, E.; Williams, A. “I’m sorry to hear that”: Finding New Biases in Language Models with a Holistic Descriptor Dataset. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Abu Dhabi, United Arab Emirates, 2022; pp 9180–9211
2022
-
[8]
BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
Dhamala, J.; Sun, T.; Kumar, V.; Krishna, S.; Pruksachatkun, Y.; Chang, K.-W.; Gupta, R. BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. New York, NY, USA, 2021; p 862–872
2021
Show all 99 references
-
[9]
M.; Curry, A.; Cercas Curry, A.; Abercrombie, G.; Hovy, D
Plaza Del Arco, F. M.; Curry, A.; Cercas Curry, A.; Abercrombie, G.; Hovy, D. Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Lon...
2024
-
[10]
Global Voices, Local Biases: Socio-Cultural Prejudices across Languages
Mukherjee, A.; Raj, C.; Zhu, Z.; Anastasopoulos, A. Global Voices, Local Biases: Socio-Cultural Prejudices across Languages. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023. 2023; pp 15828–15845
2023
-
[11]
Thesis Distillation: Investigating The Impact of Bias in NLP Models on Hate Speech Detection
Elsafoury, F. Thesis Distillation: Investigating The Impact of Bias in NLP Models on Hate Speech Detection. Proceedings of the Big Picture Workshop. Singapore, 2023; pp 53–65
2023
-
[12]
https://www.ohchr.org/en/press-releases/2014/06/marginalized- groups-un-human-rights-expert-calls-end-relegation, 2014
Marginalized groups: UN human rights expert calls for an end to relegation. https://www.ohchr.org/en/press-releases/2014/06/marginalized- groups-un-human-rights-expert-calls-end-relegation, 2014
2014
-
[13]
Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods
Zhao, J.; Wang, T.; Yatskar, M.; Ordonez, V.; Chang, K.-W. Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, V...
2018
-
[14]
Measuring Harmful Sentence Completion in Language Models for LGBTQIA+ Individuals
Nozza, D.; Bianchi, F.; Lauscher, A.; Hovy, D. Measuring Harmful Sentence Completion in Language Models for LGBTQIA+ Individuals. Proceedings of the Second Workshop on Language Technology for Equality, Diversity and Inclusion. Dublin, Ireland, 2022; pp 26–34
2022
-
[15]
R.; Katsigiannis, S.; Ramzan, N
Elsafoury, F.; Wilson, S. R.; Katsigiannis, S.; Ramzan, N. SOS: Systematic Offensive Stereotyping Bias in Word Embeddings. Proceedings of the 29th International Conference on Computational Linguistics. Gyeongju, Republic of Korea, 2022; pp 1263–1274
2022
-
[16]
I’m fully who I am
Ovalle, A.; Goyal, P.; Dhamala, J.; Jaggers, Z.; Chang, K.-W.; Galstyan, A.; Zemel, R.; Gupta, R. “I’m fully who I am”: Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation. Proceedings of the 2023 ACM Conference on Fairness, Accoun...
2023
-
[17]
QueerBench: Quantifying Discrimination in Language Models Toward Queer Identities.arXiv preprint arXiv:2406.12399 2024,
Sosto, M.; Barrón-Cedeño, A. QueerBench: Quantifying Discrimination in Language Models Toward Queer Identities.arXiv preprint arXiv:2406.12399 2024,
2024 arXiv
-
[18]
Detecting and Mitigating LGBTQIA+ Bias in Large Norwegian Language Models
Bergstrand, S.; Gambäck, B. Detecting and Mitigating LGBTQIA+ Bias in Large Norwegian Language Models. Proceedings of the 5th Workshop on Gender Bias in Natural Language Processing (GeBNLP). Bangkok, Thailand, 2024; pp 351–364
2024
-
[19]
Colonial Impulse
Das, D.; Guha, S.; Brubaker, J. R.; Semaan, B. The “Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. New York, NY, USA, ...
2024
-
[20]
R.; Kulkarni, P
Sahoo, N. R.; Kulkarni, P. P.; Ahmad, A.; Goyal, T.; Asad, N.; Garimella, A.; Bhattacharyya, P. IndiBias: A Benchmark Dataset to Measure Social Biases in Language Models for Indian Context. Proceedings of the 2024 Conference of the North American Chapter of the Association for...
2024
-
[21]
Bias Against 93 Stigmatized Groups in Masked Language Models and Downstream Sentiment Classification Tasks
Mei, K.; Fereidooni, S.; Caliskan, A. Bias Against 93 Stigmatized Groups in Masked Language Models and Downstream Sentiment Classification Tasks. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. New York, NY, USA, 2023; p 1699–1710
2023
-
[22]
What is a Refugee? https://www.unrefugees.org/refugee-facts/what-is-a-refugee/#:~:text=A%20refugee%20is%20someone%20who,in%20a% 20particular%20social%20group., 2025
2025
-
[23]
Probing Toxic Content in Large Pre-Trained Language Models
Ousidhoum, N.; Zhao, X.; Fang, T.; Song, Y.; Yeung, D.-Y. Probing Toxic Content in Large Pre-Trained Language Models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Process...
2021
-
[24]
Bias does not equal bias: A socio-technical typology of bias in data-based algorithmic systems
Lopez, P. Bias does not equal bias: A socio-technical typology of bias in data-based algorithmic systems. Internet Policy Review 2021, 10, 1–29
2021
-
[25]
Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services
Hartmann, D.; Oueslati, A.; Staufer, D. Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services. 2024; https://arxiv.org/abs/2406.14154
2024 arXiv
-
[26]
Large Language Model Instruction Following: A Survey of Progresses and Challenges
Lou, R.; Zhang, K.; Yin, W. Large Language Model Instruction Following: A Survey of Progresses and Challenges. Computational Linguistics 2024, 50, 1053–1095
2024
-
[27]
Raiaan, M. A. K.; Mukta, M. S. H.; Fatema, K.; Fahad, N. M.; Sakib, S.; Mim, M. M. J.; Ahmad, J.; Ali, M. E.; Azam, S. A Review on Large Language Models: Architectures, Applications, Taxonomies, Open Issues and Challenges. IEEE Access 2024, 12, 26839–26874
2024
-
[28]
Min, B.; Ross, H.; Sulem, E.; Veyseh, A. P. B.; Nguyen, T. H.; Sainz, O.; Agirre, E.; Heintz, I.; Roth, D. Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey. ACM Comput. Surv. 2023, 56
2023
-
[29]
L.; Lopez, G.; Olteanu, A.; Sim, R.; Wallach, H
Blodgett, S. L.; Lopez, G.; Olteanu, A.; Sim, R.; Wallach, H. Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Confer...
2021
-
[30]
Sap, M.; Swayamdipta, S.; Vianna, L.; Zhou, X.; Choi, Y.; Smith, N. A. Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguisti...
2022
-
[31]
S.; Taylor, S.; Thomas, C.; Weller, J
Webster, C. S.; Taylor, S.; Thomas, C.; Weller, J. M. Social bias, discrimination and inequity in healthcare: mechanisms, implications and recommen- dations. BJA Educ. 2022, 22, 131–137
2022
-
[32]
Brewer, M. B. The Psychology of Prejudice: Ingroup Love and Outgroup Hate? Journal of Social Issues 1999, 55, 429–444
1999
-
[33]
Intergroup bias in third-party punishment stems from both ingroup favoritism and outgroup discrimination
Schiller, B.; Baumgartner, T.; Knoch, D. Intergroup bias in third-party punishment stems from both ingroup favoritism and outgroup discrimination. Evolution and Human Behavior 2014, 35, 169–175
2014
-
[34]
Under the Radar or Into the Spotlight: How Does Social Presence Affect Minorities in Virtual Groups? SIGMIS Database 2024, 55, 98–119
Windeler, J.; Harrison, A.; Sundrup, R. Under the Radar or Into the Spotlight: How Does Social Presence Affect Minorities in Virtual Groups? SIGMIS Database 2024, 55, 98–119
2024
-
[35]
The concept of minority for the study of culture
Laurie, T.; Khan, R. The concept of minority for the study of culture. Continuum 2017, 31, 1–12
2017
-
[36]
https://emergency.unhcr.org/protection/persons-risk/minorities-and-indigenous-peoples, 2024
Minorities and indigenous peoples. https://emergency.unhcr.org/protection/persons-risk/minorities-and-indigenous-peoples, 2024
2024
-
[37]
https://minorityrights.org/ngo-declaration-on-the- framework-convention-for-the-protection-of-national-minorities/, 2008
NGO declaration on the Framework Convention for the Protection of National Minorities. https://minorityrights.org/ngo-declaration-on-the- framework-convention-for-the-protection-of-national-minorities/, 2008
2008
-
[38]
Seyranian, V.; Atuel, H.; Crano, W. D. Dimensions of majority and minority groups. Group Processes and Intergroup Relations 2008, 11, 21–37
2008
-
[39]
https://minorityrights.org/world-map/, 2023
World Directory of Minorities and Indigenous People. https://minorityrights.org/world-map/, 2023
2023
-
[40]
The Alawi capture of power in Syria
Pipes, D. The Alawi capture of power in Syria. Middle Eastern Studies 1989, 25, 429–450
1989
-
[41]
https://www.hrw.org/news/2013/06/27/egypt-lynching-shia-follows-months-hate-speech, 2013
Egypt: Lynching of Shia Follows Months of Hate Speech. https://www.hrw.org/news/2013/06/27/egypt-lynching-shia-follows-months-hate-speech, 2013
2013
-
[42]
HONEST: Measuring Hurtful Sentence Completion in Language Models
Nozza, D.; Bianchi, F.; Hovy, D. HONEST: Measuring Hurtful Sentence Completion in Language Models. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Online, 2021; pp 2398–2406
2021
-
[43]
https://minorityrights.org/country/germany/, 2020
Minority groups in Germany. https://minorityrights.org/country/germany/, 2020
2020
-
[44]
https://reporting.unhcr.org/donors/germany, 2024
Refugees in Germany. https://reporting.unhcr.org/donors/germany, 2024
2024
-
[45]
https://www.antidiskriminierungsstelle.de/SharedDocs/forschungsprojekte/EN/Studie_ DiskrRisiken_fuer_Gefluechtete_en.html, 2016
Risks of discrimination for refugees in Germany. https://www.antidiskriminierungsstelle.de/SharedDocs/forschungsprojekte/EN/Studie_ DiskrRisiken_fuer_Gefluechtete_en.html, 2016
2016
-
[46]
https://www.amnesty.org/en/what-we-do/discrimination/lgbti-rights/, 2022
LGBTI RIGHTS. https://www.amnesty.org/en/what-we-do/discrimination/lgbti-rights/, 2022
2022
-
[47]
David, B. L. C. A. T. B. V. A. Y.-D. How gender norms are perceived across the world. https://cepr.org/voxeu/columns/how-gender-norms-are- perceived-across-world, 2023
2023
-
[48]
https://www.unicef.org/kosovoprogramme/ press-releases/worlds-nearly-240-million-children-living-disabilities-are-being-denied-basic-rights, 2021
The world’s nearly 240 million children living with disabilities are being denied basic rights – UNICEF. https://www.unicef.org/kosovoprogramme/ press-releases/worlds-nearly-240-million-children-living-disabilities-are-being-denied-basic-rights, 2021
2021
-
[49]
Queerinai, O. O. et al. Queer In AI: A Case Study in Community-Led Participatory AI. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. New York, NY, USA, 2023; p 1882–1895
2023
-
[50]
https://ar.wikipedia.org/wiki/ØĺÙĹØğØĺØľ:ÙĚØňØłÙĚØź_ØğÙĎÙĚÙŁÙĚ, 2023
The gateway to the LGBTQ community in Arabic. https://ar.wikipedia.org/wiki/ØĺÙĹØğØĺØľ:ÙĚØňØłÙĚØź_ØğÙĎÙĚÙŁÙĚ, 2023
2023
-
[51]
https://queer-lexikon.net/lexikon/, 2024
Deine Online-Anlaufstelle für sexuelle, romantische und geschlechtliche Vielfalt. https://queer-lexikon.net/lexikon/, 2024. Manuscript submitted to ACM Out of Sight Out of Mind 17
2024
-
[52]
https://www.gov.uk/government/publications/inclusive-communication/ inclusive-language-words-to-use-and-avoid-when-writing-about-disability, 2021
Inclusive language: words to use and avoid when writing about disability. https://www.gov.uk/government/publications/inclusive-communication/ inclusive-language-words-to-use-and-avoid-when-writing-about-disability, 2021
2021
-
[53]
Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification
Borkan, D.; Dixon, L.; Sorensen, J.; Thain, N.; Vasserman, L. Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification. Companion Proceedings of The 2019 World Wide Web Conference. New York, NY, USA, 2019; p 491–500
2019
-
[54]
Hate Personified: Investigating the role of LLMs in content moderation
Masud, S.; Singh, S.; Hangya, V.; Fraser, A.; Chakraborty, T. Hate Personified: Investigating the role of LLMs in content moderation. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Miami, Florida, USA, 2024; pp 15847–15863
2024
-
[55]
Üstün, A. et al. Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model. 2024; https://arxiv.org/abs/2402.07827
2024 arXiv
-
[57]
Chung, H. W. et al. Scaling Instruction-Finetuned Language Models. 2022; https://arxiv.org/abs/2210.11416
2022 arXiv
-
[58]
L.; Bari, M
Muennighoff, N.; Wang, T.; Sutawika, L.; Roberts, A.; Biderman, S.; Scao, T. L.; Bari, M. S.; Shen, S.; Yong, Z.-X.; Schoelkopf, H.; others Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786 2022,
2022 arXiv
-
[59]
CEUR Workshop proceedings
Bassignana, E.; Basile, V.; Patti, V.; others Hurtlex: A multilingual lexicon of words to hurt. CEUR Workshop proceedings. 2018; pp 1–6
2018
-
[60]
Workshop, B. et al. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model. 2022; https://arxiv.org/abs/2211.05100
2022 arXiv
-
[61]
arXiv preprint arXiv:2407.21783 2024,
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; others The llama 3 herd of models. arXiv preprint arXiv:2407.21783 2024,
2024 arXiv
-
[62]
L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; others Gpt-4 technical report
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; others Gpt-4 technical report. arXiv preprint arXiv:2303.08774 2023,
2023 arXiv
-
[63]
Huang, H. et al. AceGPT, Localizing Large Language Models in Arabic. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Mexico City, Mexico, 2024; pp 8139–8163
2024
-
[64]
Sengupta, N. et al. Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models. 2023; https: //arxiv.org/abs/2308.16149
2023 arXiv
-
[65]
LLäMmlein: Compact and Competitive German-Only Language Models from Scratch
Pfister, J.; Wunderle, J.; Hotho, A. LLäMmlein: Compact and Competitive German-Only Language Models from Scratch. 2024; https://arxiv.org/abs/ 2411.11171
2024 arXiv
-
[66]
LEOLM: IGNITING GERMAN-LANGUAGE LLM RESEARCH
Plüster, B. LEOLM: IGNITING GERMAN-LANGUAGE LLM RESEARCH. https://laion.ai/blog/leo-lm/, 2023
2023
-
[67]
AraBART: a Pretrained Arabic Sequence-to-Sequence Model for Abstractive Summarization
Kamal Eddine, M.; Tomeh, N.; Habash, N.; Le Roux, J.; Vazirgiannis, M. AraBART: a Pretrained Arabic Sequence-to-Sequence Model for Abstractive Summarization. Proceedings of the Seventh Arabic Natural Language Processing Workshop (WANLP). Abu Dhabi, United Arab Emirates (Hybrid...
2022
- [68]
-
[69]
AraBERT: Transformer-based Model for Arabic Language Understanding
Antoun, W.; Baly, F.; Hajj, H. AraBERT: Transformer-based Model for Arabic Language Understanding. Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection. Marseille, France, 2020; pp 9–15
2020
-
[70]
Unsupervised Cross-lingual Representation Learning at Scale
Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzmán, F.; Grave, E.; Ott, M.; Zettlemoyer, L.; Stoyanov, V. Unsupervised Cross-lingual Representation Learning at Scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics...
2020
-
[71]
A.; Ardeshna, B.; Bhatt, D
Pandya, H. A.; Ardeshna, B.; Bhatt, D. B. S. Cascading Adaptors to Leverage English Data to Improve Performance of Question Answering for Low-Resource Languages. 2021
2021
-
[72]
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; Zettlemoyer, L. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. Proceedings of the 58th Annual Meeting of the Association f...
2020
-
[73]
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; Soricut, R. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. 2020
2020
-
[74]
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL...
2019
-
[75]
O.; Rossi, R
Gallegos, I. O.; Rossi, R. A.; Barrow, J.; Tanjim, M. M.; Kim, S.; Dernoncourt, F.; Yu, T.; Zhang, R.; Ahmed, N. K. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics 2024, 50, 1097–1179
2024
-
[76]
Evaluation of African American Language Bias in Natural Language Generation
Deas, N.; Grieser, J.; Kleiner, S.; Patton, D.; Turcan, E.; McKeown, K. Evaluation of African American Language Bias in Natural Language Generation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore, 2023; pp 6805–6824
2023
-
[77]
R.; Jurafsky, D.; Goel, S
Koenecke, A.; Nam, A.; Lake, E.; Nudell, J.; Quartey, M.; Mengesha, Z.; Toups, C.; Rickford, J. R.; Jurafsky, D.; Goel, S. Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences 2020, 117, 7684–7689
2020
-
[78]
D.; Howell, N
Meyer, J.; Rauchenstein, L.; Eisenberg, J. D.; Howell, N. Artie Bias Corpus: An Open Dataset for Detecting Demographic Bias in Speech Applications. Proceedings of the Twelfth Language Resources and Evaluation Conference. Marseille, France, 2020; pp 6462–6468
2020
-
[79]
L.; O’Connor, B
Blodgett, S. L.; O’Connor, B. T. Racial Disparity in Natural Language Processing: A Case Study of Social Media African-American English. ArXiv 2017, abs/1707.00061
2017 arXiv
-
[80]
OSIAN: Open Source International Arabic News Corpus - Preparation and Integration into the CLARIN-infrastructure
Zeroual, I.; Goldhahn, D.; Eckart, T.; Lakhouaja, A. OSIAN: Open Source International Arabic News Corpus - Preparation and Integration into the CLARIN-infrastructure. Proceedings of the Fourth Arabic Natural Language Processing Workshop. Florence, Italy, 2019; pp 175–182
2019
-
[81]
El-khair, I. A. 1.5 billion words Arabic Corpus. 2016; https://arxiv.org/abs/1611.04033. Manuscript submitted to ACM 18 Fatma Elsafoury and David Hartmann
2016 arXiv
-
[82]
https://www.internetsociety.org/resources/doc/2020/middle-east-north-africa- internet-infrastructure-report/, 2020
Middle East and North Africa Internet Infrastructure Report. https://www.internetsociety.org/resources/doc/2020/middle-east-north-africa- internet-infrastructure-report/, 2020
2020
-
[83]
https://www.internetsociety.org/resources/doc/2024/connectivity-in-the-middle-east-and- north-africa/, 2024
Connectivity in the Middle East and North Africa. https://www.internetsociety.org/resources/doc/2024/connectivity-in-the-middle-east-and- north-africa/, 2024
2024
-
[84]
The Interplay of Variant, Size, and Task Type in Arabic Pre-trained Language Models
Inoue, G.; Alhafni, B.; Baimukan, N.; Bouamor, H.; Habash, N. The Interplay of Variant, Size, and Task Type in Arabic Pre-trained Language Models. Proceedings of the Sixth Arabic Natural Language Processing Workshop. Kyiv, Ukraine (Virtual), 2021; pp 92–104
2021
-
[85]
The Limitations of Humanity: Differential Refugee Treatment in the EU
Esposito, A. The Limitations of Humanity: Differential Refugee Treatment in the EU. https://hir.harvard.edu/the-limitations-of-humanity- differential-refugee-treatment-in-the-eu/, 2022
2022
-
[86]
Systematic limitations on the integration of Syrian refugees in Egypt and its impact on mental health and well- being
Rayes, D. Systematic limitations on the integration of Syrian refugees in Egypt and its impact on mental health and well- being. https://timep.org/2019/10/18/stuck-in-transit-systematic-limitations-on-the-integration-of-syrian-refugees-in-egypt-and-its-impact-on- the-mental-he...
2019
-
[87]
Sub-Saharan migrants in Egypt subject to increasing abuse and violence
Sanderson, S. Sub-Saharan migrants in Egypt subject to increasing abuse and violence. https://www.infomigrants.net/en/post/21862/subsaharan- migrants-in-egypt-subject-to-increasing-abuse-and-violence, 2020
2020
-
[88]
Said, E. W. Culture and imperialism; Vintage, 1994
1994
-
[89]
The Construction of Arabs as Enemies: Post-September 11 Discourse of George W
Merskin, D. The Construction of Arabs as Enemies: Post-September 11 Discourse of George W. Bush. Mass Communication and Society 2004, 7, 157–175
2004
-
[90]
In Media - Migration - Integration ; Geißler, R., Pöttker, H., Eds.; transcript Verlag: Bielefeld, 2009; pp 181–212
Starck, K. In Media - Migration - Integration ; Geißler, R., Pöttker, H., Eds.; transcript Verlag: Bielefeld, 2009; pp 181–212
2009
-
[92]
J.; Ritter, A.; Xu, W
Naous, T.; Ryan, M. J.; Ritter, A.; Xu, W. Having Beer after Prayer? Measuring Cultural Bias in Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 20...
2024
-
[93]
Deteriorated economy forces Egyptians to endure ‘slavery’, maltreatment in Saudi Arabia and Kuwait
Sakr, T. Deteriorated economy forces Egyptians to endure ‘slavery’, maltreatment in Saudi Arabia and Kuwait. https://www.dailynewsegypt.com/ 2016/08/20/deteriorated-economy-forces-egyptians-endure-slavery-maltreatment-saudi-arabia-kuwait/, 2016
2016
-
[94]
Government restrictions on religion
Majumdar, S. Government restrictions on religion. https://www.pewresearch.org/wp-content/uploads/sites/20/2024/12/PR_2024.12.18_restrictions- on-religion-2022_report.pdf, 2022
2024
-
[95]
In Human Rights in the Middle East: Frameworks, Goals, and Strategies ; Monshipouri, M., Ed.; Palgrave Macmillan US: New York, 2011; pp 153–169
Monshipouri, M.; Whooley, J. In Human Rights in the Middle East: Frameworks, Goals, and Strategies ; Monshipouri, M., Ed.; Palgrave Macmillan US: New York, 2011; pp 153–169
2011
-
[96]
A.; Haider, A
Tair, S. A.; Haider, A. S.; Obeidat, M. M.; Sahari, Y. Challenges in Netflix Arabic subtitling of English nonbinary gender expressions in ‘Degrassi: Next Class’ and ‘One Day at a Time’. Humanities and Social Sciences Communications 2024, 11
2024
-
[97]
https://arabstates.unfpa.org/sites/default/files/pub-pdf/14385_-_disability_in_the_arab_ region_-_final_report_web_version_-_opt.7.pdf, 2021
Disability in the Arab region: A challenged vulnerability. https://arabstates.unfpa.org/sites/default/files/pub-pdf/14385_-_disability_in_the_arab_ region_-_final_report_web_version_-_opt.7.pdf, 2021
2021
-
[98]
Thani, H. A. Disability in the Arab region- Current situation and prospects. Adult Education and Development 2007, 68, 13
2007
-
[99]
Yslas, I. G. A. Queer Reflections: Unveiling the Impact of Media Stereotypes on Adolescent Well-being. Journal of Student Academic Research 2024, 5
2024
-
[100]
R.; Salim, S
Eshelman, L. R.; Salim, S. R.; Bhuptani, P. H.; Saad, M. Sexual Objectification Racial Microaggressions Amplify the Positive Relation Between Sexual Assault and Posttraumatic Stress Among Black Women. Psychology of Women Quarterly 2024, 48, 180–194
2024
-
[101]
P.; Hugenberg, K.; Rule, N
Wilson, J. P.; Hugenberg, K.; Rule, N. O. Racial bias in judgments of physical size and formidability: From size to threat. Journal of Personality and Social Psychology 2017, 113, 59–80. Manuscript submitted to ACM Out of Sight Out of Mind 19 A Appendices A.1 Data Sets SOS Dat...
2017
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.