REVIEW 4 major objections 5 minor 55 references
Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read LLM gender labels beat word-count bias metric by up to 59%
desk verdict New gender-bias dataset and a useful LLM-vs-lexical comparison, but CWEx's fairness claim rests on labels that track genderedness, not bias - and the reported improvements are absolute kappa differences. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Class-wise Weighted Exposure (CWEx) metric, defined as $\mathrm{CWEx} = \alpha \cdot \mathrm{Exposure}_{\text{neutral}} - (1-\alpha) \cdot |\mathrm{Exposure}_{\text{male}} - \mathrm{Exposure}_{\text{female}}|$, where $\mathrm{Exposure}_G$ is the summed position-bias weight $p(i) = 1/\log_2(1+i)$ for documents of class $G$, normalized by the maximum possible exposure, and $\alpha \in [0,1]$ balances promoting neutral documents against reducing male-female disparity. The classification engine is few-shot prompting of LLMs with instructions covering gender-term frequency, balance of information, and lead representation. This machinery converts document-level bias categories into a single interpretable fairness score for a ranked list, and the $\alpha$ parameter makes the trade-off explicit and domain-adjustable.
What would settle it
A user study would settle it: ask people to rate the fairness of many generated rankings and compare their ratings to CWEx. If rankings with high CWEx are not perceived as fairer than low-CWEx ones, or if for a query where no neutral relevant documents exist CWEx marks every ranking as unfair, the document-label-to-ranking link fails.
Extended reading notes
Core claim
The central claim is that gender bias in ranked lists is better measured by semantic document-level labels than by lexical term counts. The authors prompt GPT-4o, Llama-3.1-8B-Instruct, Llama-3.1-8B, Mixtral-8x7B-Instruct, and Qwen2.5-7B-Instruct to classify passages as male-biased, female-biased, or neutral, and show that these labels align with human judgments more closely than the neutrality score of NFaiRR: best Cohen's kappa is 0.858 versus 0.270 on Grep-BiasIR, and 0.572 versus 0.387 on MSMGenderBias. Building on these labels, they define CWEx, a metric that combines normalized neutral exposure with the male-female exposure gap, weighted by a parameter alpha that lets the evaluator trade off neutrality against gender parity. The paper also releases MSMGenderBias, a set of 893 MS MARCO passages labeled by crowdworkers.
Load-bearing premise
The whole measure rests on the assumption that a document's three-way gender label, neutral, male, or female, captures what is unfair about a ranked list, so that a ranking is fair precisely when neutral documents get high exposure and male- and female-labeled documents get equal exposure.
Editorial extensions
If this is right
- Fairness evaluation for ranking can shift from lexical term counting to semantic understanding, catching biases like 'his mother' and 'her son' that current word-based metrics miss.
- CWEx gives practitioners a dial, alpha, to choose between prioritizing neutral language in top ranks and balancing exposure between male- and female-biased documents.
- The same LLM-labeling pipeline extends to other protected attributes such as race, age, and ethnicity, and to non-binary gender groups by replacing the male-female disparity term with the max-min group disparity.
- The released MSMGenderBias dataset provides a public benchmark for comparing future gender-bias detectors and fairness metrics.
- LLMs can produce human-aligned bias labels at scale, reducing reliance on expensive manual annotation for building bias datasets.
Reading between the lines
- The metric's premise that neutral exposure is inherently good may clash with domains where neutral documents are not relevant; a testable extension would weight neutral exposure by relevance before judging fairness.
- Because LLMs themselves carry stereotype biases (the paper finds asymmetry between male- and female-stereotype detection), a single chosen annotator model could inject its own bias into the fairness score; one could ensemble several LLMs and measure disagreement.
- The study's human evaluation covers only top-10 results from 20 queries per query set; scaling to full rankings would check whether CWEx remains stable and meaningful.
- The three-class scheme treats 'mentions a gender' and 'biased toward a gender' as the same signal; a finer label set distinguishing representation from bias could change which rankings count as fair.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes using LLMs as document-level gender-bias detectors with three-class labels (neutral, male-biased, female-biased), introduces a new fairness metric called Class-wise Weighted Exposure (CWEx) that aggregates these labels into a ranked-list fairness score, and releases a new gender-bias annotation subset of MS MARCO called MSMGenderBias. The authors compare LLM labels with human labels on Grep-BiasIR and MSMGenderBias, reporting higher Cohen's kappa agreement than the NFaiRR neutrality score, and they compare ranking models with CWEx and NFaiRR. The core claims are that LLM-based detection is more accurate than lexical baselines and that CWEx provides a more detailed and human-aligned evaluation of fairness in ranked lists.
Significance. If established, the paper would make three useful contributions: a semantic, LLM-based alternative to lexical gender-bias detection for ranking evaluation; a new annotated dataset (MSMGenderBias) for the community; and a class-aware exposure metric that explicitly accounts for neutral, male, and female document categories. The document-level detection experiments are informative and the dataset release is a concrete asset. However, the paper's central ranking-level claim is not yet supported: CWEx is never validated against human judgments of ranking fairness or against rankings with known fairness properties. In addition, the reported kappa improvements are stated as percentages when they are absolute percentage-point differences, and the human annotation design is explicitly aligned with the LLM prompt, which weakens the independence of the human ground truth. These issues are load-bearing for the paper's headline claims, though they appear addressable with additional analysis and new experiments.
major comments (4)
- [Abstract and §5.1, Table 5] The claim of "58.77% improvement" on Grep-BiasIR and "18.51% improvement" on MSMGenderBias is stated as a percentage improvement, but the numbers are absolute percentage-point differences in Cohen's kappa. For Grep-BiasIR, the best LLM kappa is 0.8580 versus NFaiRR's 0.2703, a difference of 0.5877, which is a relative improvement of about 217%, not 58.77%. For MSMGenderBias, 0.5719 versus 0.3868 is a difference of 0.1851, a relative improvement of about 47.9%, not 18.51%. The abstract and Section 5.1 should report the kappa values directly and, if percentages are used, label them as absolute percentage-point differences or relative improvements consistently.
- [§5.3, Tables 7 and 8; Eq. (1)] The central claim that CWEx "effectively distinguishes gender bias in ranking" is not validated at the ranking level. Eq. (1) defines CWEx as a weighted combination of neutral exposure and male-female exposure disparity, but the experiments in Tables 7 and 8 only compare CWEx values across ranking models and against NFaiRR; there is no human judgment of ranking fairness, no simulated ranking with known fairness properties, and no external fairness benchmark. Without such a ranking-level test, the metric's interpretation as a fairness measure, rather than as a measure of gender skew in retrieved content, is unsupported. I recommend adding a controlled experiment with synthetic rankings whose fairness properties are known, or a human study that directly evaluates full ranked lists.
- [Figure 2 and §3.3] The document-level labels appear to conflate genderedness with bias. The annotation guidelines in Figure 2 define Male/Female as documents that "explicitly talk about a person with a specific gender" or that include gender-related terms "more than the other gender," and they label the example containing Alan Turing, Grace Hopper, and other computer scientists as Male despite the presence of a female scientist. This suggests the labels capture the dominant gender of the document's subject matter, not whether the document is biased. Because CWEx (Eq. 1) is built on these labels, a high CWEx score may simply indicate low gender skew in content, which is not the same as absence of bias. The authors should either align the label definition with a bias construct and provide evidence that the labels reflect bias, or explicitly reframe the metric as measuring representation skew.
- [§3.3 and §5.1, Table 1] The human annotation instructions were deliberately designed to align with the LLM prompt (Section 3.3: instructions "designed to align with the instructions provided to the LLM"), and the few-shot examples used in the prompts were selected from Grep-BiasIR, the same dataset used for evaluation (Section 5.1). This creates a risk of circularity: the high kappa agreement between LLM labels and human labels may be partly by construction, because humans were guided to apply the same criteria as the LLM. The paper should report the agreement using independently collected or pre-existing human labels, or at least quantify how much the aligned instructions inflate agreement. This issue directly affects the headline claim that LLM labels are better aligned with human judgment than NFaiRR's neutrality score.
minor comments (5)
- [Title page / author affiliations] There are minor typographical errors in the author affiliations: "Unviersity of Amsterdam" and "The Netherland" should be corrected.
- [§3.2, Eq. (1)] The sentence "The proposed CWEx metric is constrained within the range of alpha to alpha−1" is imprecise; the correct interval is [alpha−1, alpha]. Please rephrase for clarity.
- [Table 5] The paper selects the best prompt setting per model and compares that single best value to NFaiRR without reporting variance or significance tests. Because Cohen's kappa is computed on a finite sample, confidence intervals or a significance test would strengthen the comparison, especially for the smaller MSMGenderBias subset.
- [§3.3] The human evaluation section reports that "the final classification of gender bias... was determined based on the majority vote among the annotators" but also says that "for the documents that each annotator selected one of the classes, we asked an expert annotator to annotate that document." This sentence is unclear: it is not specified when the expert annotation is used versus the majority vote, and the phrase "each annotator selected one of the classes" is ambiguous. Please clarify the adjudication procedure.
- [§4, Dataset] The description of query sampling states that 20 queries were randomly selected from QS1 and QS2, and that BM25, BERT, MiniLM, and TinyBERT retrieved 446 and 447 passages respectively. It would be helpful to report how many unique queries and documents resulted after deduplication across retrieval models, since the statistics in Table 2 suggest some overlap.
Circularity Check
Minor validation circularity: human labels for MSMGenderBias are created with instructions aligned to the LLM prompt; CWEx itself is a definition and no fitted prediction is relabeled as a result.
-
self definitional
[Section 3.3 Human Evaluation; Figure 2 vs Table 1]
"we provided annotators with a comprehensive guideline (Figure 2 in the Appendix) that includes (1) a detailed description of the task, (2) detailed instructions on detecting bias in the given document designed to align with the instructions provided to the LLM (see Table 1) and (3) examples of documents exhibiting different types of bias."
The validation target for the MSMGenderBias half of Table 5 is a set of human labels produced under instructions explicitly 'designed to align' with the LLM prompt in Table 1, and Figure 2 reproduces the same Male/Female/Neutral definition. The LLM is therefore not being checked against an independent operationalization of gender bias; it is being checked against a second application of its own rubric. The reported 18.51% kappa improvement over NFaiRR on MSMGenderBias is partly an artifact of giving the LLM and the human annotators the same custom definition, while NFaiRR's lexical neutrality score was never given that definition. The circularity is partial: agreement is not numerically forced, and the Grep-BiasIR comparison uses external human labels.
full rationale
The paper's main derivation chain is a definition, not an inference from fitted parameters. CWEx (Eq. 1) combines Exposure_G (Eq. 2), maximum exposure (Eq. 3), and position bias p(i) (Eq. 4) with a user-chosen alpha; alpha values are scanned over 0.2/0.5/0.7, not optimized against a target, so no fitted input is renamed as a prediction. The NFaiRR comparison is an external baseline and is not circular. The prompt-design citation to the authors' prior work [1] is corroborated by an independent citation [24] and is not load-bearing; there is no uniqueness theorem or ansatz smuggled in by citation. The one self-referential element is the human evaluation: the MTurk guidelines for the newly released MSMGenderBias are deliberately aligned with the LLM prompt, making the MSMGenderBias human labels a shared-rubric ground truth. This weakens the validity claim for that dataset but does not make the CWEx computation circular. The absence of a ranking-level ground truth for CWEx is a genuine construct-validity gap and a correctness risk, but not a circularity, because no experiment claims to derive ranking fairness from a fitted target. Overall score is 2: minor, partial validation circularity confined to the new dataset, with the central metric and the Grep-BiasIR comparison retaining independent content.
Assumptions & free parameters
free parameters (3)
- alpha (alpha in CWEx) =
0.2, 0.5, 0.7 (user-specified, not fit)
- Top-k cutoff for exposure =
10
- Per-model best prompt setting =
Chosen per model from Table 3 (e.g., CoT for GPT-4o, one-shot for Llama-3.1-8B-Instruct)
assumptions (4)
- domain assumption Document-level gender bias can be captured by a three-class label (neutral, male, female) assigned by human annotators or LLMs.
- domain assumption User attention follows the logarithmic position bias p(i)=1/log2(1+i).
- domain assumption Human annotations created with instructions aligned to the LLM prompt are an unbiased gold standard.
- ad hoc to paper The few-shot examples used in prompts do not constitute test leakage.
Cite this review
Pith. "Pith review of Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement." pith.science (2026). https://pith.science/paper/PQTURJN2
@misc{pith2026250622372,
author = {Pith},
title = {Pith review of: Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement},
year = {2026},
howpublished = {\url{https://pith.science/paper/PQTURJN2}},
note = {Machine review of arXiv:2506.22372}
}
read the original abstract
The presence of social biases in Natural Language Processing (NLP) and Information Retrieval (IR) systems is an ongoing challenge, which underlines the importance of developing robust approaches to identifying and evaluating such biases. In this paper, we aim to address this issue by leveraging Large Language Models (LLMs) to detect and measure gender bias in passage ranking. Existing gender fairness metrics rely on lexical- and frequency-based measures, leading to various limitations, e.g., missing subtle gender disparities. Building on our LLM-based gender bias detection method, we introduce a novel gender fairness metric, named Class-wise Weighted Exposure (CWEx), aiming to address existing limitations. To measure the effectiveness of our proposed metric and study LLMs' effectiveness in detecting gender bias, we annotate a subset of the MS MARCO Passage Ranking collection and release our new gender bias collection, called MSMGenderBias, to foster future research in this area. Our extensive experimental results on various ranking models show that our proposed metric offers a more detailed evaluation of fairness compared to previous metrics, with improved alignment to human labels (58.77% for Grep-BiasIR, and 18.51% for MSMGenderBias, measured using Cohen's Kappa agreement), effectively distinguishing gender bias in ranking. By integrating LLM-driven bias detection, an improved fairness metric, and gender bias annotations for an established dataset, this work provides a more robust framework for analyzing and mitigating bias in IR systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, and Mohammad Aliannejadi
-
[2]
Amin Abolghasemi, Leif Azzopardi, Arian Askari, Maarten de Rijke, and Suzan Verberne. 2024. Measuring Bias in a Ranked List using Term-based Representa- tions. doi:10.48550/arXiv.2403.05975 arXiv:2403.05975 [cs]
work page Pith review arXiv doi:10.48550/arxiv.2403.05975 2024
-
[3]
Jaimeen Ahn and Alice Oh. 2021. Mitigating Language-Dependent Ethnic Bias in BERT. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 533–549. doi:10...
-
[4]
Amin Bigdeli, Negar Arabzadeh, Shirin Seyedsalehi, Bhaskar Mitra, Morteza Zihayat, and Ebrahim Bagheri. 2023. De-biasing Relevance Judgements for Fair Ranking. In Advances in Information Retrieval (Lecture Notes in Computer Science) , Jaap Kamps, Lorraine Goeuriot, Fabio Crestani, Maria Maistro, Hideo Joho, Brian Davis, Cathal Gurrin, Udo Kruschwitz, and ...
-
[5]
Rishi Bommasani, Percy Liang, and Tony Lee. 2023. Holistic evaluation of lan- guage models. Annals of the New York Academy of Sciences1525, 1 (2023), 140–146
work page 2023
-
[6]
Shikha Bordia and Samuel R. Bowman. 2019. Identifying and Reducing Gender Bias in Word-Level Language Models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop, Sudipta Kar, Farah Nadeem, Laura Burdick, Greg Durrett, and Na-Rae Han (Eds.). Association for Computa...
-
[7]
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science 356, 6334 (2017), 183–186
2017
-
[8]
Le Chen, Ruijun Ma, Anikó Hannák, and Christo Wilson. 2018. Investigating the Impact of Gender on Rank in Resume Search Engines. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–14. doi:10.1145/3173574.3174225
arXiv 2018
Show all 55 references
-
[9]
Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2020. On measuring and mitigating biased inferences of word embeddings. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 7659–7666
2020
-
[10]
Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2021. OSCaR: Orthog- onal Subspace Correction and Rectification of Biases in Word Embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huan...
2021 doi
-
[11]
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruk- sachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Tr...
2021
-
[12]
Tommaso Dolci, Fabio Azzalini, and Mara Tanelli. 2023. Improving gender-related fairness in sentence encoders: A semantics-based approach. Data Science and Engineering 8, 2 (2023), 177–195
2023
-
[13]
Michael D Ekstrand, Graham McDonald, Amifa Raj, and Isaac Johnson. 2023. Overview of the TREC 2022 fair ranking track. arXiv preprint arXiv:2302.05558 (2023)
2023 arXiv
-
[14]
Alessandro Fabris, Alberto Purpura, Gianmaria Silvello, and Gian Antonio Susto
-
[15]
Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023. Perspectives on Large Language Models for Relevance Judgment. In Proceedings o...
2023
-
[16]
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2024. Bias and fairness in large language models: A survey.Computational Linguistics (2024), 1–79
2024
-
[17]
Sahin Cem Geyik, Stuart Ambler, and Krishnaram Kenthapadi. 2019. Fairness- Aware Ranking in Search & Recommendation Systems with Application to LinkedIn Talent Search. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Anchorag...
2019
-
[18]
Anthony G Greenwald, Debbie E McGhee, and Jordan LK Schwartz. 1998. Mea- suring individual differences in implicit cognition: the implicit association test. Journal of personality and social psychology 74, 6 (1998), 1464
1998
-
[19]
Wei Guo and Aylin Caliskan. 2021. Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (Virtual Event, USA) (AIES ’21). Association for Co...
2021
-
[20]
Maria Heuss, Daniel Cohen, Masoud Mansoury, Maarten de Rijke, and Carsten Eickhoff. 2023. Predictive Uncertainty-based Bias Mitigation in Ranking. In Pro- ceedings of the 32nd ACM International Conference on Information and Knowledge Management (CIKM ’23) . Association for Com...
2023
-
[21]
Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli. 2020. Reducing Senti- ment Bias in Language Models via Counterfactual Evaluation. In Findings of the Association for Computational Linguistics: EMN...
2020 doi
-
[22]
Masahiro Kaneko and Danushka Bollegala. 2022. Unmasking the mask– evaluating social biases in masked language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 11954–11962
2022
-
[23]
Klara Krieg, Emilia Parada-Cabaleiro, Gertraud Medicus, Oleg Lesota, Markus Schedl, and Navid Rekabsaz. 2023. Grep-biasir: A dataset for investigating gender representation bias in information retrieval results. In Proceedings of the 2023 Conference on Human Information Intera...
2023
-
[24]
Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Beckage, Hsuan Su, Hung-yi Lee, and Lama Nachman
Shachi H. Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Beckage, Hsuan Su, Hung-yi Lee, and Lama Nachman
-
[25]
Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019. Measuring Bias in Contextualized Word Representations. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing , Marta R. Costa- jussà, Christian Hardmeier, Will Radford,...
2019 doi
-
[26]
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, hai zhao, and Pengfei Liu. 2024. Generative Judge for Evaluating Alignment. In The Twelfth Inter- national Conference on Learning Representations . https://openreview.net/forum? id=gtkFw6sZGS
2024
- [27]
-
[28]
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov
-
[29]
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A Python toolkit for reproducible infor- mation retrieval research with sparse and dense representations. In Proceedings of the 44th International ACM SIGIR Conferenc...
2021
-
[30]
Tao Li, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar
-
[31]
In Find- ings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.)
UNQOVERing Stereotyping Biases via Underspecified Questions. In Find- ings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Online, 3475–3489. doi:10.18653/v1/2020.findings-emn...
2020 doi
-
[32]
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM com- puting surveys (CSUR) 54, 6 (2021), 1–35
2021
-
[33]
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereo- typical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Interna- tional Joint Conference on Natural Language...
2021 doi
-
[34]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Lan- guage Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber,...
2020
-
[35]
Bowman, and Rachel Rudinger
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. On Measuring Social Biases in Sentence Encoders. In Pro- ceedings of the 2019 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Human Language Techno...
2019 doi
-
[36]
Michal Měchura. 2022. A Taxonomy of Bias-Causing Ambiguities in Machine Translation. InProceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP), Christian Hardmeier, Christine Basta, Marta R. Costa-jussà, Gabriel Stanovsky, and Hila Gonen (Eds.). ...
2022 doi
-
[37]
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022. BBQ: A hand- built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: ACL 2022 , Smaranda Mure...
2022 doi
-
[38]
Hossein A Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, and Daniel Campos. 2024. Synthetic test collections for retrieval evaluation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2647–2651
2024
-
[39]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Confer- ence on Natural Language Processing (EMNLP-IJ...
2019 doi
-
[40]
Navid Rekabsaz, Simone Kopeinik, and Markus Schedl. 2021. Societal Biases in Retrieved Contents: Measurement Framework and Adversarial Mitigation of BERT Rankers. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval...
2021 doi
-
[41]
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. In Proceedings of the Workshop on Cogni- tive Computation: Integrating neural and symbolic approaches 2016 ...
2016
-
[42]
Debora Nozza, Federico Bianchi, and Dirk Hovy. 2021. HONEST: Measuring Hurtful Sentence Completion in Language Models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Kristina To...
2021
-
[43]
Robertson, Alan Mislove, and Christo Wilson
Piotr Sapiezynski, Wesley Zeng, Ronald E. Robertson, Alan Mislove, and Christo Wilson. 2019. Quantifying the Impact of User Attention on Fair Group Represen- tation in Ranked Lists. http://arxiv.org/abs/1901.10437 arXiv:1901.10437
2019 arXiv
-
[44]
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The Woman Worked as a Babysitter: On Biases in Language Generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on N...
2019 doi
-
[45]
Anthony Sicilia and Malihe Alikhani. 2023. Learning to Generate Equitable Text in Dialogue from Biased Training Data. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoak...
2023 doi
-
[46]
Sijie Tao, Tetsuya Sakai, Junjie Wang, Hanpei Fang, Yuxiang Zhang, Haitao Li, Yiteng Tu, Nuo Chen, Maria Maistro, et al . 2025. Overview of the NTCIR-18 FairWeb-2 Task. In Proceedings of the 18th NTCIR Conference on Evaluation of Information Access Technologies
2025
-
[47]
Navid Rekabsaz and Markus Schedl. 2020. Do Neural Ranking Models Intensify Gender Bias?. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’20) . Association for Computing Machinery, New York, NY, USA, 206...
2020 doi
-
[48]
Nguyen, and Katrin Kirchhoff
Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020. Masked Language Model Scoring. In Proceedings of the 58th Annual Meeting of the Associ- ation for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Associa...
2020 doi
-
[49]
George Zerveas, Navid Rekabsaz, Daniel Cohen, and Carsten Eickhoff. 2022. Mitigating Bias in Search Results Through Contextual Document Reranking and Neutrality Regularization. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Informa...
2022
-
[53]
Alex Wang and Kyunghyun Cho. 2019. BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model. InProceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation , Antoine Bosselut, Asli Celikyilmaz, Marjan Ghazvininejad, S...
2019 doi
-
[54]
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020. Measuring and reducing gendered correlations in pre-trained models. arXiv preprint arXiv:2010.06032 (2020)
2020 arXiv
-
[1967]
doi:10.18653/v1/2020.emnlp-main.154
2020 doi
-
[2020]
Information Processing & Management 57, 6 (2020), 102377
Gender stereotype reinforcement: Measuring the gender bias conveyed by ranking algorithms. Information Processing & Management 57, 6 (2020), 102377
2020
-
[2021]
In International Conference on Machine Learning
Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning . PMLR, 6565–6576
-
[2024]
Can We Use Large Language Models to Fill Relevance Judgment Holes? arXiv preprint arXiv:2405.05600 (2024)
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.