REVIEW 3 major objections 5 minor 134 references
Understanding Gender Bias in AI-Generated Product Descriptions
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that AI-generated product descriptions systematically exhibit gender bias in identifiable categories—body-size assumptions, target-group exclusion and assumptions, stereotyped feature emphasis, product–activity…
desk verdict A genuinely new taxonomy of gender bias in AI-generated product descriptions, with solid expert-driven discovery and mostly careful measurement—but the headline persuasion disparity is confounded by product mix and should be re-analyzed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the six-category taxonomy itself, built by a four-stage pipeline: start from five general bias themes in existing frameworks, flag potentially biased descriptions in a 10,000-example sample of real generations via human annotation and GPT-4o, solicit open-ended reviews from four expert reviewers on 120 flagged descriptions, and synthesize reviews into minimally overlapping categories. The quantitative analyses then use two instruments: vocabulary-based phrase detection for body size, gendered terms, and call-to-action phrases, and counterfactual input pairs of 50 products whose only differing attribute is the gender label, with 500 generated descriptions per pair per model. A simple bigram classifier on the counterfactual outputs, with gendered terms masked, predicts the gender label with over 90% accuracy, which is what lets the paper attribute word-level differences such as 'adventure' versus 'flattering' to the gendered framing itself rather than to product mix.
What would settle it
A counterfactual test on the persuasion category—generating descriptions for the same product with only the male/female label changed and comparing call-to-action rates—would settle whether the gap is model bias or product mix; if matched pairs show no gap, the persuasion-disparity claim collapses.
Extended reading notes
Core claim
On its own terms, the paper's central claim is that e-commerce product description generation exhibits a distinct profile of gender bias, different from the occupation stereotypes and pronoun errors usually studied in LLMs. The paper names six categories and anchors each in expert-reviewed examples: descriptions assume 'regular' or 'all' body sizes, repeat gendered targeting phrases such as 'designed exclusively for men,' attach gendered groups to gender-neutral products like baby bottles, emphasize appearance and 'flattering' language for women's clothing while emphasizing durability for men's, associate women's products with errands and lounging while associating men's with outdoor activities, and produce calls to action more often for men's products. The quantitative results include a 5.5 percentage-point gap in call-to-action frequency for GPT-3.5, exclusionary body-size phrases in 14.3% of one model's clothing descriptions, and a counterfactual experiment in which a simple bigram classifier identifies the labelled gender of the product from the description with over 90% accuracy. The paper presents these as forms of exclusionary norms, stereotyping, and disparate performance, and argues they are detectable and worth mitigating in e-commerce.
Load-bearing premise
The persuasion-disparity claim assumes that the different call-to-action rates for men's and women's products measure the model's gendered treatment, rather than the different categories of products that happen to be marketed to men and women.
Editorial extensions
If this is right
- Automated quality checks that score fluency, fidelity, or attractiveness will not catch these harms, because all six categories can appear in fluent, faithful, attractive text.
- E-commerce platforms that deploy LLMs for listing generation need task-specific evaluation suites built around the taxonomy, not just general-purpose toxicity or stereotyping detectors.
- The 5.5 percentage-point persuasion gap, if it generalises, means sellers of men's and women's items do not get equally persuasive promotional language from the same model.
- The counterfactual method gives a minimal audit design: change only the gender label in the input and compare outputs, so platforms can test their own models for stereotyping before launch.
- The data-driven taxonomy process transfers to other text-generation tasks, giving a template for finding task-specific bias categories instead of reusing general ones.
Reading between the lines
- The persuasion-disparity estimate is computed on all men's and women's products without matching across product categories, so part or all of the 5.5 percentage-point gap could reflect the different mix of products marketed to each group; a counterfactual version of the call-to-action analysis would settle this.
- The same vocabulary and counterfactual machinery could be pointed at other demographic dimensions the expert reviews surfaced, such as body size, skin tone, religion, and culture, producing analogous taxonomies.
- Because human advertising copy already shows similar stereotype patterns, part of what the models do may be inheritance from training data rather than a model-specific invention; comparing LLM outputs to human-written descriptions for the same products would separate the two.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a data-driven taxonomy of gender bias in AI-generated product descriptions, grounding it in existing general-purpose harm taxonomies. The authors use human annotation, GPT-4o filtering, and expert reviews of flagged examples to identify six categories: body size assumptions, target group exclusion, target group assumptions, bias in advertised features, product-activity associations, and persuasion disparities. They quantify each category on two large real-world datasets (50,000 generated descriptions per model for GPT-3.5 and an internal e-commerce LLM), using a counterfactual-pair design for the stereotype categories and phrase-list detection for prevalence estimates. Headline findings include body-size exclusionary language in roughly 10-14% of clothing descriptions, explicit target-group phrases in 8-11%, and a 5.5-percentage-point higher call-to-action rate for men's than women's descriptions generated by GPT-3.5.
Significance. If the quantitative claims are properly supported, the paper makes a valuable contribution: it identifies e-commerce-specific manifestations of gender bias (body size, target group, persuasion) that are absent from general analyses, and it demonstrates a reproducible process for data-driven taxonomy development. The counterfactual-pair design in Sections 4.4-4.5 is a notable strength, as is the inclusion of confidence intervals for the body-size estimates and the public release of phrase lists in the appendix. The main reservation is that the persuasion disparity result, which is also a headline statistic, rests on an aggregate comparison that does not control for product category distribution; this weakens one of the six taxonomy categories and needs to be addressed before the full set of claims is accepted.
major comments (3)
- [Section 4.6 and Appendix B] The persuasion disparity analysis compares call-to-action frequencies across all men's products versus all women's products without any adjustment for product category. The category distribution in Appendix B shows substantial differences across item types (e.g., Clothing, Shoes & Accessories is 22.01% of the 10,000-item sample while Sports Mem, Cards & Fan Shop is 15.52%), and call-to-action language may have different base rates in different categories. The 5.5-percentage-point gap for GPT-3.5 (27.0% vs. 21.5%, Z=4.632) and the 2.9-point gap for the internal model (24.2% vs. 21.3%, Z=2.250) may therefore reflect product mix rather than gender-based disparate performance. The authors should stratify by product category, include category fixed effects, or apply the same counterfactual-pair design used in Sections 4.4-4.5. Until this is done, the 'persuasion disparities' category and the abstract's headline statistic are not established.
- [Section 4.2 and Appendix E.2] The quantitative evidence for target group exclusion relies on two metrics that are not fully convincing. First, the average number of gendered terms per description (1.90 for GPT-3.5, 1.99 for the internal model) is not benchmarked against neutral descriptions and likely counts department labels or title echoes, so it is a weak measure of 'excessive emphasis.' Second, the 'explicitly exclusive phrases' list includes items such as 'designed for women' and 'made for men' that are not necessarily exclusionary in context; the expert-review examples focus on stronger language like 'designed exclusively for men.' The phrase list should be validated by human review of flagged descriptions or restricted to phrases with clear exclusivity markers. Without this, the reported 8.6% and 11.4% prevalence figures may substantially overstate the target-group-exclusion category.
- [Section 3.3.3 and Appendix K] The claim that the GPT-4o flagging step has 'no false negatives' is based on a manual review of only 200 descriptions that were not flagged by GPT-4o. Given that human annotators flagged 7,527 of 10,000 descriptions and GPT-4o reduced the flagged set to 120, the number of unflagged descriptions is large, and a 200-example check cannot support a no-false-negatives claim with meaningful confidence. The authors should either report a confidence interval for the false-negative rate, perform a larger validation sample, or soften the claim to 'no false negatives were found in a small validation sample.' This matters because the taxonomy's completeness depends on the flagging process, and the current language in Section 3.4 and Appendix K is stronger than the evidence supports.
minor comments (5)
- [Section 4.4] The counterfactual sample size is described ambiguously: 'we generated 500 descriptions for each pair of inputs (for a total of 25,000 descriptions per model)' should clarify whether this means 500 descriptions per input condition (which would give 50,000 per model) or 250 per input condition.
- [Appendix G and Appendix H] Appendix G duplicates the annotator information in Appendix C, and Appendix H duplicates the expert reviewer information in Appendix D; these should be consolidated to avoid redundancy.
- [Section 4.6 and Abstract] Since only call-to-action phrase frequency is measured, the paper should consistently describe this finding as a 'call-to-action frequency disparity' rather than a broad 'persuasion disparity' in section headings and the abstract, or explicitly justify the proxy.
- [Section 4.1] The abstract's statement that body-size exclusionary language appears in 'over 14%' of clothing descriptions from the internal model should be reported with the combined confidence interval for the women's (14.3%) and men's (14.2%) estimates, rather than presenting the point estimate without uncertainty.
- [Section 4.5] The product-activity association results report predictive words from the bigram classifier but do not report effect sizes, confidence intervals, or the accuracy of the classifier separately for the two models; adding these statistics would strengthen the claim.
Circularity Check
No significant circularity: the paper's quantitative analyses are independent corpus measurements and counterfactual experiments, not derivations that reduce to the taxonomy they illustrate.
full rationale
The paper's central claims are empirical findings about the frequency of expert-identified language patterns, not predictions derived from fitted parameters or from a theorem. Section 3 develops the taxonomy from human annotation, GPT-4o flagging, and open-ended expert review of 120 examples; Section 4 then operationalizes each category with explicit phrase lists (Appendix E) or with gender-counterfactual generation. Measuring how often phrases such as 'all bodies' or 'order today' occur in 50,000-description datasets is a straightforward corpus count; the 14.3% body-size prevalence and the 5.5-point call-to-action gap are not forced by the construction of the category, since the counts could in principle have been near zero or reversed. The counterfactual bigram classifier in Section 4.4 is a genuine distinguishability experiment holding product inputs fixed. The persuasion-disparity comparison (Section 4.6) is unadjusted for product-category mix, which is a validity threat, but that is a confounding concern, not a circularity: no parameter is fitted to the claimed disparity. The paper's self-citations ([18], [104], [105]) are contextual references to prior frameworks and literature reviews, and they are not load-bearing for the taxonomy or the measurements. No uniqueness theorem, ansatz-citation, or definitional reduction appears. Accordingly, there are no circular steps to report.
Assumptions & free parameters
assumptions (5)
- domain assumption Aggregate comparisons of calls to action across men's and women's products are not confounded by differences in product category distributions.
- domain assumption The handcrafted phrase lists comprehensively capture the target bias categories.
- domain assumption GPT-4o flagging introduces no systematic blind spots.
- domain assumption The 120 expert-reviewed descriptions are representative of biased descriptions in the broader dataset.
- domain assumption Expert reviewers' judgments constitute valid ground truth for bias categories.
Cite this review
Pith. "Pith review of Understanding Gender Bias in AI-Generated Product Descriptions." pith.science (2026). https://pith.science/paper/LSDY74P3
@misc{pith2026250605390,
author = {Pith},
title = {Pith review of: Understanding Gender Bias in AI-Generated Product Descriptions},
year = {2026},
howpublished = {\url{https://pith.science/paper/LSDY74P3}},
note = {Machine review of arXiv:2506.05390}
}
read the original abstract
While gender bias in large language models (LLMs) has been extensively studied in many domains, uses of LLMs in e-commerce remain largely unexamined and may reveal novel forms of algorithmic bias and harm. Our work investigates this space, developing data-driven taxonomic categories of gender bias in the context of product description generation, which we situate with respect to existing general purpose harms taxonomies. We illustrate how AI-generated product descriptions can uniquely surface gender biases in ways that require specialized detection and mitigation approaches. Further, we quantitatively analyze issues corresponding to our taxonomic categories in two models used for this task -- GPT-3.5 and an e-commerce-specific LLM -- demonstrating that these forms of bias commonly occur in practice. Our results illuminate unique, under-explored dimensions of gender bias, such as assumptions about clothing size, stereotypical bias in which features of a product are advertised, and differences in the use of persuasive language. These insights contribute to our understanding of three types of AI harms identified by current frameworks: exclusionary norms, stereotyping, and performance disparities, particularly for the context of e-commerce.
Figures
Reference graph
Works this paper leans on
-
[1]
[n. d.]. 51 eCommerce Statistics In 2025. https://www.sellerscommerce.com/ blog/ecommerce-statistics/. Accessed: 2025-01-16
2025
-
[2]
[n. d.]. Online Shopping Statistics. https://capitaloneshopping.com/research/ online-shopping-statistics/. Accessed: 2025-01-16
2025
-
[3]
[n. d.]. Retail e-commerce sales worldwide from 2014 to 2027. https://www. statista.com/statistics/379046/worldwide-retail-e-commerce-sales/. Accessed: 2025-01-21
2014
-
[4]
Martin Adam, Michael Wessel, and Alexander Benlian. 2021. AI-based chatbots in customer service and their effects on user compliance. Electronic Markets 31, 2 (2021), 427–445
2021
-
[5]
Jaimeen Ahn and Alice Oh. 2021. Mitigating language-dependent ethnic bias in BERT. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 533–549
2021
-
[6]
Jack J Amend, Albatool Wazzan, and Richard Souvenir. 2021. Evaluating gender- neutral training data for automated image captioning. In 2021 IEEE Interna- tional Conference on Big Data (Big Data) . 1226–1235. https://doi.org/10.1109/ BigData52589.2021.9671774
arXiv 2021
-
[7]
Luis Arango, Stephen Pragasam Singaraju, and Outi Niininen. 2023. Consumer responses to AI-generated charitable giving ads. Journal of Advertising 52, 4 (2023), 486–503
2023
-
[8]
Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. 2023. Easily accessible text-to-image generation amplifies demo- graphic stereotypes at large scale. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 1493–1504
2023
Show all 134 references
-
[9]
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Lan- guage (technology) is power: A critical survey of" bias" in NLP. arXiv preprint arXiv:2005.14050 (2020)
2020 arXiv
-
[10]
Julia M Bristor, Renee Gravois Lee, and Michelle R Hunt. 1995. Race and ideology: African-American images in television advertising. Journal of Public Policy & Marketing 14, 1 (1995), 48–59
1995
-
[11]
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science 356, 6334 (2017), 183–186
2017
-
[12]
Colin Campbell, Kirk Plangger, Sean Sands, and Jan Kietzmann. 2022. Preparing for an era of deepfakes and AI-generated ads: A framework for understanding responses to manipulated advertising. Journal of Advertising 51, 1 (2022), 22–38
2022
-
[13]
Christina M Capodilupo, Kevin L Nadal, Lindsay Corman, Sahran Hamit, Oliver B Lyons, and Alexa Weinberg. 2010. The manifestation of gender mi- croaggressions. (2010)
2010
-
[14]
Zhangming Chan, Xiuying Chen, Yongliang Wang, Juntao Li, Zhiqiang Zhang, Kun Gai, Dongyan Zhao, and Rui Yan. 2019. Stick to the facts: Learning towards a fidelity-oriented e-commerce product description generation. In Proceedings of the 2019 Conference on Empirical Methods in ...
2019
-
[15]
Aadi Chauhan, Taran Anand, Tanisha Jauhari, Arjav Shah, Rudransh Singh, Arjun Rajaram, and Rithvik Vanga. 2024. Identifying race and gender bias in stable diffusion AI image generation. In 2024 IEEE 3rd International Conference on AI in Cybersecurity (ICAIC) . 1–6. https://doi...
2024
-
[16]
Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023. Marked personas: Us- ing natural language prompts to measure stereotypes in language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 1504–1532
2023
-
[17]
Zhibo Chu, Zichong Wang, and Wenbin Zhang. 2024. Fairness in large language models: A taxonomic survey. ACM SIGKDD explorations newsletter 26, 1 (2024), 34–48
2024
-
[18]
Marios Constantinides, Edyta Bogucka, Daniele Quercia, Susanna Kallio, and Mohammad Tahaei. 2024. RAI Guidelines: Method for Generating Responsible AI Guidelines Grounded in Regulations and Usable by (Non-)Technical Roles. Proc. ACM Hum.-Comput. Interact. 8, CSCW2, Article 388...
2024 doi
-
[19]
Anthony J Cortese. 2015. Provocateur: Images of women and minorities in adver- tising. Rowman & Littlefield
2015
-
[20]
WTL Cox and PG Devine. 2015. Stereotypes possess heterogeneous directional- ity: A theoretical and empirical exploration of stereotype structure and content. PLoS ONE 10, 3 (2015), e0122292
2015
-
[21]
Lei Cui, Shaohan Huang, Furu Wei, Chuanqi Tan, Chaoqun Duan, and Ming Zhou. 2017. Superagent: A customer service chatbot for e-commerce websites. In Proceedings of ACL 2017, system demonstrations . 97–102
2017
-
[22]
Judy Foster Davis. 2018. Selling whiteness?–A critical review of the literature on marketing and racism. Journal of Marketing Management 34, 1-2 (2018), 134–177
2018
-
[23]
Leon Derczynski, Hannah Rose Kirk, Vidhisha Balachandran, Sachin Kumar, Yu- lia Tsvetkov, and Saif Mohammad. 2023. Assessing language model deployment with risk cards. arXiv preprint arXiv:2303.18190 (2023)
2023 arXiv
-
[24]
Sunipa Dev, Akshita Jha, Jaya Goyal, Dinesh Tewari, Shachi Dave, and Vin- odkumar Prabhakaran. 2023. Building stereotype repositories with LLMs and community engagement for scale and depth. Cross-Cultural Considerations in NLP@ EACL 84 (2023)
2023
-
[25]
Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2020. On measuring and mitigating biased inferences of word embeddings. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 7659–7666
2020
-
[26]
Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff M Phillips, and Kai-Wei Chang. 2021. Harms of gender exclusivity and chal- lenges in non-binary representation in language technologies. arXiv preprint arXiv:2108.12084 (2021)
2021 arXiv
-
[27]
Mark Díaz, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle
-
[28]
Jad Doughman, Wael Khreich, Maya El Gharib, Maha Wiss, and Zahraa Berjawi
-
[29]
Duo Du, Yanling Zhang, and Jiao Ge. 2023. Effect of AI generated content advertising on consumer engagement. In International Conference on Human- Computer Interaction. Springer, 121–129
2023
-
[31]
Seliem El-Sayed, Canfer Akbulut, Amanda McCroskery, Geoff Keeling, Zachary Kenton, Zaria Jalan, Nahema Marchal, Arianna Manzini, Toby Shevlane, Shan- non Vallor, et al. 2024. A mechanism-based approach to mitigating harms from persuasive generative AI. arXiv preprint arXiv:240...
2024 arXiv
-
[32]
CV Evans, ES Johnson, and JS Lin. [n. d.]. Assessing algorithmic bias and fairness in clinical prediction models for preventive services. A health equity methods project for the US Preventive Services Task Force. 2023
2023
-
[33]
Janice L Farlow, Marianne Abouyared, Eleni M Rettig, Alexandra Kejner, Rusha Patel, and Heather A Edwards. 2024. Gender bias in artificial intelligence-written letters of reference. Otolaryngology–Head and Neck Surgery (2024)
2024
-
[34]
Fabio Fasoli, Federica Durante, Silvia Mari, Cristina Zogmaister, and Chiara Volpato. 2018. Shades of sexualization: When sexualization becomes sexual objectification. Sex Roles 78 (2018), 338–351
2018
-
[35]
Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, and Hanna Wallach. 2023. Fair- Prism: Evaluating fairness-related harms in text generation. In Proceedings of the 61st Annual Meeting of the Association for Comp...
2023 doi
-
[36]
I wouldn’t say offensive but
Vinitha Gadiraju, Shaun Kane, Sunipa Dev, Alex Taylor, Ding Wang, Emily Denton, and Robin Brewer. 2023. "I wouldn’t say offensive but... ": Disability- centered perspectives on large language models. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Tr...
2023
-
[37]
Gallegos, Ryan A
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. Bias and fairness in large language models: A survey. Computational Linguistics (07 2024), 1–83. https://doi.org/10.1162/coli_a_00524
2024 doi
-
[38]
Noa Garcia, Yusuke Hirota, Yankun Wu, and Yuta Nakashima. 2023. Uncurated image-text datasets: Shedding light on demographic bias. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE Computer Society, 6957–6966
2023
-
[39]
F Gasparini, I Erba, E Fersini, S Corchs, et al. 2018. Multimodal classification of sexist advertisements. In ICETE 2018-Proceedings of the 15th International Joint Conference on e-Business and Telecommunications-Volume 2, Vol. 1. SciTePress, 399–406
2018
-
[40]
Mary C Gilly. 1988. Sex roles in advertising: A comparison of television adver- tisements in Australia, Mexico, and the United States. Journal of marketing 52, 2 (1988), 75–85
1988
-
[41]
Joelle Sano Gilmore and Amy Jordan. 2012. Burgers and basketball: Race and stereotypes in food and beverage advertising aimed at children in the US.Journal of Children and Media 6, 3 (2012), 317–332
2012
-
[42]
Evelyn Nakano Glenn. 2008. Yearning for lightness: Transnational cir- cuits in the marketing and consumption of skin lighteners. Gender & Society 22, 3 (2008), 281–302. https://doi.org/10.1177/0891243208316089 arXiv:https://doi.org/10.1177/0891243208316089 FAccT ’25, June 23–2...
2008 doi
-
[43]
Nicole Gross. 2023. What ChatGPT tells us about gender: A cautionary tale about performativity and gender biases in AI. Social Sciences 12, 8 (2023). https://doi.org/10.3390/socsci12080435
2023 doi
-
[44]
Darrell Y Hamamoto. 1994. Monitored peril: Asian Americans and the politics of TV representation. U of Minnesota Press
1994
-
[45]
Camille Harris, Matan Halevy, Ayanna Howard, Amy Bruckman, and Diyi Yang
-
[46]
Lucy Havens, Melissa Terras, Benjamin Bach, and Beatrice Alex. 2022. Un- certainty and inclusivity in gender bias annotation: An annotation taxon- omy and annotated datasets of British English text. In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processi...
2022 doi
-
[47]
Yusuke Hirota, Yuta Nakashima, and Noa Garcia. 2022. Quantifying societal bias amplification in image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13450–13459
2022
-
[48]
Yasmeen Hitti, Eunbee Jang, Ines Moreno, and Carolyne Pelletier. 2019. Proposed taxonomy for gender bias in text; a filtering methodology for the gender gener- alization subtype. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing. Association fo...
2019 doi
-
[49]
Thong, and Kar Yan Tam
Weiyin Hong, James Y.L. Thong, and Kar Yan Tam. 2004. Designing product listing pages on e-commerce websites: an examination of presentation mode and information format. International Journal of Human-Computer Studies 61, 4 (2004), 481–503. https://doi.org/10.1016/j.ijhcs.2004.01.006
2004 doi
-
[50]
Tamanna Hossain, Sunipa Dev, and Sameer Singh. 2023. MISGENDERED: Limits of large language models in understanding pronouns. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and...
2023 doi
-
[51]
Flora Huang. 2024. Understanding Asian stereotyping and bias in LLMs
2024
-
[52]
Guanxiong Huang and Sai Wang. 2023. Is artificial intelli- gence more persuasive than humans? A meta-analysis. Jour- nal of Communication 73, 6 (08 2023), 552–562. https://doi. org/10.1093/joc/jqad024 arXiv:https://academic.oup.com/joc/article- pdf/73/6/552/54463456/jqad024_su...
2023 doi
-
[53]
Jennifer Jacobs Henderson and Gerald J Baldasty. 2003. Race, advertising, and prime-time television. Howard Journal of Communications 14, 2 (2003), 97–112
2003
-
[54]
Liqiang Jing, Xuemeng Song, Xuming Lin, Zhongzhou Zhao, Wei Zhou, and Liqiang Nie. 2023. Stylized data-to-text generation: A case study in the e- commerce domain. ACM Transactions on Information Systems 42, 1 (2023), 1–24
2023
-
[55]
Annamma Joy and Alladi Venkatesh. 1994. Postmodernism, feminism, and the body: The visible and the invisible in consumer research. International Journal of research in Marketing 11, 4 (1994), 333–357
1994
-
[56]
Mahammed Kamruzzaman, Hieu Minh Nguyen, and Gene Louis Kim
-
[57]
Mahammed Kamruzzaman, Md Shovon, and Gene Kim. 2024. Investigating subtler biases in LLMs: Ageism, beauty, institutional, and nationality bias in generative models. In Findings of the Association for Computational Linguistics ACL 2024. 8940–8965
2024
-
[58]
Deanna M Kaplan, Roman Palitsky, Santiago J Arconada Alvarez, Nicole S Pozzo, Morgan N Greenleaf, Ciara A Atkinson, and Wilbur A Lam. 2024. What’s in a name? Experimental evidence of gender bias in recommendation letters generated by ChatGPT. J Med Internet Res 26 (5 Mar 2024)...
2024 doi
-
[59]
Aneel Karnani. 2007. Doing well by doing good—case study:‘Fair & Lovely’whitening cream. Strategic management journal 28, 13 (2007), 1351– 1357
2007
-
[60]
Jan Kietzmann, Jeannette Paschen, and Emily Treen. 2018. Ar- tificial Intelligence in Advertising. Journal of Advertising Re- search 58, 3 (2018), 263–267. https://doi.org/10.2501/JAR-2018-035 arXiv:https://www.journalofadvertisingresearch.com/content/58/3/263.full.pdf
2018 doi
-
[61]
Eugenia Kim, De’Aira Bryant, Deepak Srikanth, and Ayanna Howard. 2021. Age bias in emotion detection: An analysis of facial emotion recognition performance on young, middle-aged, and older adults. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society . 638–644
2021
-
[62]
Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Fred- eric Dreyer, Aleksandar Shtedritski, and Yuki Asano. 2021. Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular gener- ative language models. In Advances in ...
2021
-
[63]
Haein Kong, Yongsu Ahn, Sangyub Lee, and Yunho Maeng. 2024. Gender bias in LLM-generated interview responses. In Workshop on Socially Responsible Language Modelling Research
2024
-
[64]
Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in large language models. In Proceedings of The ACM Collective Intelligence Conference (Delft, Netherlands) (CI ’23). Association for Computing Machinery, New York, NY, USA, 12–24. https://doi.org/10....
2023
-
[65]
Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2022. Can pretrained language models generate persuasive, faithful, and informative ad text for product descrip- tions?. In Proceedings of the Fifth Workshop on e-Commerce and NLP (ECNLP 5) , Shervin Malmasi, Oleg Rokhlenko, Nicola...
2022 doi
-
[66]
Claire Kramsch. 2014. Language and culture. AILA review 27, 1 (2014), 30–55
2014
-
[67]
Tonny Krijnen. 2017. Feminist theory and the media. The International Encyclo- pedia of Media Effects (2017), 1–12
2017
-
[68]
Tao Li, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar
-
[69]
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michi- hiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2023. Holistic evaluation of language models. Transactions on Machine Learning Research (2023)
2023
-
[70]
Sue Lim, Hee Jung Cho, Moonsun Jeon, Xiaoran Cui, and Ralf Schmaelzle. 2024. Using VR and eye-tracking to study attention to and retention of AI-generated ads in outdoor advertising environments. bioRxiv (2024), 2024–08
2024
-
[71]
Eric Justin Liu, Wonyoung So, Peko Hosoi, and Catherine D’Ignazio. 2024. Racial steering by large language models: A prospective audit of GPT-4 on housing recommendations. In Proceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization...
2024
-
[72]
Floridi Luciano. 2024. Hypersuasion–On AI’s persuasive power and how to deal with it. Philosophy & Technology 37, 2 (2024), 1–10
2024
-
[73]
They only care to show us the wheelchair
Kelly Avery Mack, Rida Qadri, Remi Denton, Shaun K. Kane, and Cynthia L. Bennett. 2024. “They only care to show us the wheelchair”: disability repre- sentation in text-to-image AI models. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu...
2024
-
[74]
Liam Magee, Lida Ghahremanlou, Karen Soldatic, and Shanthi Robertson. 2021. Intersectional bias in causal language models. arXiv preprint arXiv:2107.07691 (2021)
2021 arXiv
-
[75]
Masato Mita, Soichiro Murakami, Akihiko Kato, and Peinan Zhang. 2024. Strik- ing gold in advertising: Standardization and exploration of ad text generation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 955–972
2024
-
[76]
Talé A Mitchell. 2020. Critical Race Theory (CRT) and colourism: A manifes- tation of whitewashing in marketing communications? Journal of Marketing Management 36, 13-14 (2020), 1366–1389
2020
-
[77]
Saif Mohammad. 2022. Ethics sheets for AI tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 8368–8379
2022
-
[78]
Kasey Lynn Morris, Jamie Goldenberg, and Patrick Boyd. 2018. Women as animals, women as objects: Evidence for two forms of objectification.Personality and Social Psychology Bulletin 44, 9 (2018), 1302–1314
2018
-
[79]
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Pro...
2021
-
[80]
Roberto Navigli, Simone Conia, and Björn Ross. 2023. Biases in large language models: Origins, inventory, and discussion. J. Data and Information Quality 15, 2, Article 10 (jun 2023), 21 pages. https://doi.org/10.1145/3597307
2023 doi
-
[81]
Punam Ohri-Vachaspati, Zeynep Isgor, Leah Rimkus, Lisa M Powell, Dianne C Barker, and Frank J Chaloupka. 2015. Child-directed marketing inside and on the exterior of fast food restaurants. American journal of preventive medicine 48, 1 (2015), 22–30
2015
-
[82]
Maciej Osowski, Aleksandra Krasnodebska, Paweł Drozda, and Rafał Scherer
-
[83]
I’m fully who I am
Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023. “I’m fully who I am”: Towards centering transgender and non-binary voices to measure biases in open language generation. In Proceedings of the 2023...
2023
-
[84]
not-so-silent partner
Hye Jin Paek and Hemant Shah. 2003. Racial ideology, model minorities, and the "not-so-silent partner": Stereotyping of Asian Americans in US magazine advertising. Howard Journal of Communications 14, 4 (2003), 225–243
2003
-
[85]
Chester Palen-Michel, Ruixiang Wang, Yipeng Zhang, David Yu, Canran Xu, and Zhe Wu. 2024. Investigating LLM applications in e-commerce. arXiv preprint arXiv:2408.12779 (2024)
2024 arXiv
-
[86]
Shramay Palta and Rachel Rudinger. 2023. FORK: A bite-sized test set for probing culinary cultural biases in commonsense reasoning models. In Findings of the Association for Computational Linguistics: ACL 2023 . 9952–9962
2023
-
[87]
Jessica K Paulus and David M Kent. 2020. Predictably unequal: understanding and addressing concerns that algorithmic clinical prediction may increase health disparities. NPJ digital medicine 3, 1 (2020), 99
2020
-
[88]
Dana Pessach and Barbara Poblete. 2024. Gender representation across on- line retail products. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (Rio de Janeiro, Brazil) (FAccT ’24). Asso- ciation for Computing Machinery, New York, NY, USA...
2024
-
[89]
In Asian Conference on Intelligent Information and Database Systems
Professionally diverse: AI-generated faces for targeted advertising. In Asian Conference on Intelligent Information and Database Systems . Springer, 171–183
-
[90]
Rida Qadri, Renee Shelby, Cynthia L Bennett, and Emily Denton. 2023. AI’s regimes of representation: A community-centered study of text-to-image models in South Asia. In Proceedings of the 2023 ACM Conference on Fairness, Account- ability, and Transparency. 506–517
2023
-
[91]
Anandi Ramamurthy and Kalpana Wilson. 2013. Racism, appropriation and resistance in advertising. Colonial Advertising & Commodity Racism 69, 4 (2013), 2
2013
-
[92]
Rabia Rauf, Sohail Kamran, and Najeeb Ullah. 2019. Marketing Of skin fairness creams And consumer vulnerability. CITY UNIVERSITY RESEARCH JOURNAL 9, 3 (Oct. 2019). https://www.cusitjournals.com/index.php/CURJ/article/view/262
2019
-
[93]
Maribeth Rauh, John Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, Iason Gabriel, et al. 2022. Characteristics of harmful text: Towards rigorous bench- marking of language models. Advances in Neural In...
2022
-
[94]
Jessica Ringrose and Kaitlyn Regehr. 2020. Feminist counterpublics and public feminisms: Advancing a critique of racialized sexualization in London’s public advertising. Signs: Journal of Women in Culture and Society 46, 1 (2020), 229–257
2020
-
[95]
Alexander Rogiers, Sander Noels, Maarten Buyl, and Tijl De Bie. 2024. Persuasion with Large Language Models: a Survey. arXiv preprint arXiv:2411.06837 (2024)
2024 arXiv
-
[96]
Scott Plous and Dominique Neptune. 1997. Racial and gender biases in magazine advertising: A content-analytic study.Psychology of women quarterly 21, 4 (1997), 627–644
1997
-
[97]
Victoria L. Rubin. 2022. Manipulation in Marketing, Advertising, Propaganda, and Public Relations. Springer International Publishing, Cham, 157–205. https: //doi.org/10.1007/978-3-030-95656-1_6
2022 doi
-
[98]
Fabrizio Santoniccolo, Tommaso Trombetta, Maria Noemi Paradiso, and Luca Rollè. 2023. Gender and media representations: A review of the literature on gender stereotypes, objectification and sexualization. International Journal of Environmental Research and Public Health 20, 10...
2023
-
[99]
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Ros- tamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk. 2023. Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction. In Proceed...
2023
-
[100]
Zara Siddique, Liam Turner, and Luis Espinosa Anke. 2024. Who is better at math, Jenny or Jingzhen? Uncovering stereotypes in large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 18601–18619
2024
-
[101]
Alexandra A. Siegel. 2020. Online Hate Speech . Cambridge University Press, 56–88
2020
-
[102]
Irene Solaiman, Zeerak Talat, William Agnew, Lama Ahmad, Dylan Baker, Su Lin Blodgett, Canyu Chen, Hal Daumé III, Jesse Dodge, Isabella Duan, et al. 2023. Evaluating the social impact of generative AI systems in systems and society. arXiv preprint arXiv:2306.05949 (2023)
2023 arXiv
-
[103]
Konstantinos I Roumeliotis, Nikolaos D Tselikas, and Dimitrios K Nasiopoulos
-
[104]
Natural Language Processing Journal 6 (2024), 100056
LLMs in e-commerce: a comparative analysis of GPT and LLaMA models in product review evaluation. Natural Language Processing Journal 6 (2024), 100056
2024
-
[105]
Mohammad Tahaei, Daricia Wilkinson, Alisa Frik, Michael Muller, Ruba Abu- Salma, and Lauren Wilcox. 2024. Surveys Considered Harmful? Reflecting on the Use of Surveys in AI Research, Development, and Governance. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society...
2024 doi
-
[106]
CWJ van Miltenburg. 2016. Stereotyping and bias in the flickr30k dataset. In 11th workshop on multimodal corpora: computer vision and language processing
2016
-
[107]
Daniel Van Niekerk, María Peréz-Ortiz, John Shawe-Taylor, Davor Orlic, Jackie Kay, Noah Siegel, Katherine Evans, Nyalleng Moorosi, Tina Eliassi-Rad, Leonie Maria Tanczer, et al. 2024. Challenging systematic prejudices: An inves- tigation into bias against women and girls. (2024)
2024
-
[108]
Akshaj Kumar Veldanda, Fabian Grob, Shailja Thakur, Hammond Pearce, Ben- jamin Tan, Ramesh Karri, and Siddharth Garg. 2023. Are Emily and Greg still more employable than Lakisha and Jamal? Investigating algorithmic hiring bias in the era of ChatGPT. arXiv:2310.05135 [cs.CL] ht...
2023 arXiv
-
[109]
Kelly is a warm person, Joseph is a role model
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. “Kelly is a warm person, Joseph is a role model”: Gender biases in LLM-generated reference letters. InFindings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, ...
2023 doi
-
[110]
Jinpeng Wang, Yutai Hou, Jing Liu, Yunbo Cao, and Chin-Yew Lin. 2017. A statistical framework for product description generation. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers), Greg Kondrak and Taro Watanabe...
2017
-
[111]
Luhang Sun, Mian Wei, Yibing Sun, Yoo Ji Suh, Liwei Shen, and Sijia Yang
-
[112]
Journal of Computer-Mediated Communication 29, 1 (02 2024), zmad045
Smiling women pitching down: Auditing representational and presen- tational gender biases in image-generative AI. Journal of Computer-Mediated Communication 29, 1 (02 2024), zmad045. https://doi.org/10.1093/jcmc/zmad045
2024 doi
-
[113]
Mohammad Tahaei, Marios Constantinides, Daniele Quercia, and Michael Muller
-
[114]
Luming Yang, Min Xu, and Lin Xing. 2022. Exploring the core factors of online purchase decisions by building an e-commerce network evolution model.Journal of Retailing and Consumer Services 64 (2022), 102784. https://doi.org/10.1016/j. jretconser.2021.102784
2022
-
[115]
Tao Zhang, Jin Zhang, Chengfu Huo, and Weijun Ren. 2019. Automatic generation of pattern-controlled product description in e-commerce. In The World Wide Web Conference (San Francisco, CA, USA) (WWW ’19) . Associ- ation for Computing Machinery, New York, NY, USA, 2355–2365. htt...
2019
-
[116]
Dora Zhao, Angelina Wang, and Olga Russakovsky. 2021. Understanding and evaluating racial biases in image captioning. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 14830–14840
2021
-
[117]
Li Zhou, Taelin Karidi, Wanlong Liu, Nicolas Garneau, Yong Cao, Wenyu Chen, Haizhou Li, and Daniel Hershcovich. 2024. Does mapo tofu contain coffee? prob- ing LLMs for food-related cultural knowledge. arXiv preprint arXiv:2404.06833 (2024)
2024 arXiv
-
[118]
If it wasn’t secure, they would not use it in the movies
Verena Zimmermann and Nina Gerber. 2017. “If it wasn’t secure, they would not use it in the movies” – Security perceptions and user acceptance of authen- tication technologies. In Human Aspects of Information Security, Privacy and Trust, Theo Tryfonas (Ed.). Springer Internati...
2024
-
[121]
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William...
2022
-
[122]
Robert Wolfe, Yiwei Yang, Bill Howe, and Aylin Caliskan. 2023. Contrastive language-vision AI models pretrained on web-scraped multimodal data exhibit sexual objectification bias. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 1174–1185
2023
-
[123]
Liang Wu, Diane Hu, Liangjie Hong, and Huan Liu. 2018. Turning clicks into purchases: Revenue optimization for product search in e-commerce. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval (Ann Arbor, MI, USA) (SIGIR ’18). Asso...
2018
-
[129]
yes. ” Note that we are especially interested in examples where the group is implicit (see the first example below, where the term “LGBTQ+
Does the input information suggest, in any way, that the item is associated with or describes a certain social or demographic group? (See the samples below for examples.) If any group (see Definitions) is directly associated or implied as associated with the item, respond “yes...
2025
-
[130]
hot and sexy sluts
Does the description contain toxic language or hate speech? Toxic language and hate speech: hostile and malicious language that attacks, threatens, or incites hate against a certain group, as well as slurs, insults, and other derogatory words or phrases that demean or belittle...
2025
-
[131]
flattering fit,
Does the description contain stereotyping or objectification? Stereotyping and objectification: generalizations about particular groups of people, which include implicit or explicit associations between a group and a behavior, trait, occupation, role, item, or other idea, incl...
2025
-
[132]
mom, dad, and the kids
Does the description contain exclusionary norms? Exclusionary norms: expressions of what is normal or typical that implicitly exclude certain groups. In other words, these are statements that could lead to people feeling left out or excluded. Input Generated Description Exclus...
1984
-
[133]
Groups” or “groups of people
Does the description contain any other potential bias? Use this category to flag anything else that seems potentially biased and/or harmful to certain groups. Note that this does not include general issues (e.g., low-quality text or inaccurate statements) that do not pertain t...
2023
-
[134]
INPUT: information about the item that was input to an AI system, such as the item’s title, category, and specific details
-
[135]
yes” if the generated description contains[term] (based on the definition provided below). Respond “no
GENERATED DESCRIPTION: the corresponding item description, generated by the AI system. For the provided record, please respond “yes” if the generated description contains[term] (based on the definition provided below). Respond “no” if it does not. If the answer is nuanced or a...
2007
-
[2018]
In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems
Addressing age-related bias in sentiment analysis. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . 1–14
2018
-
[2020]
In Find- ings of the Association for Computational Linguistics: EMNLP 2020
UNQOVERing stereotyping biases via underspecified questions. In Find- ings of the Association for Computational Linguistics: EMNLP 2020 . 3475–3489
2020
-
[2021]
In Proceedings of the 3rd Workshop on Gender Bias in Natural Language Processing , Marta Costa- jussa, Hila Gonen, Christian Hardmeier, and Kellie Webster (Eds.)
Gender bias in text: Origin, taxonomy, and implications. In Proceedings of the 3rd Workshop on Gender Bias in Natural Language Processing , Marta Costa- jussa, Hila Gonen, Christian Hardmeier, and Kellie Webster (Eds.). Association for Computational Linguistics, Online, 34–44....
-
[2022]
In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT ’22)
Exploring the role of grammar and word choice in bias toward African American English (AAE) in hate speech classification. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT ’22). Association for Computing ...
2022
-
[2023]
arXiv:2302.05284 [cs.HC] https://arxiv.org/abs/2302.05284
A Systematic Literature Review of Human-Centered, Ethical, and Respon- sible AI. arXiv:2302.05284 [cs.HC] https://arxiv.org/abs/2302.05284
-
[2024]
Global is good, local is bad?
"Global is good, local is bad?": Understanding brand bias in LLMs. arXiv:2406.13997 [cs.CL] https://arxiv.org/abs/2406.13997
-
[4968]
https://doi.org/10.18653/v1/D19-1501
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.