REVIEW 3 major objections 4 minor 48 references
A Close Reading Approach to Gender Narrative Biases in AI-Generated Stories
T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Gender bias in AI-generated stories persists at the narrative level even when character counts look balanced.
desk verdict Useful narratological audit of LLM stories, but 15 stories and post hoc coding make the strong quantitative claims illustrative rather than evidential. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a standardized generation-and-reading protocol: a zero-shot prompt that asks each model to write a roughly 500-word story containing five Propp-derived characters—main character (MC), villain (V), helper (H), desired character (DC), dispatcher (D)—and to follow Freytag's five-phase arc (exposition, rise, climax, return/fall, catastrophe). Characters and phases must not be named in the story. This yields comparable outputs across models. The analysis side uses close reading with a reading form that records prompt adherence, gender distribution of each role, physical and psychological descriptors, action types, and plot-level relationships, allowing composite or implic
What would settle it
Generate, say, fifty stories per model with the same prompt, have independent coders who are blind to the hypothesis label each character's gender and code who rescues, guides, or acts decisively at the climax; if villains are not overwhelmingly male or female main characters are not rescued by men more often than role-gender chance would predict, the paper's central pattern fails. An even sharper test: rerun the prompt with an explicit instruction that the villain is female and the desired character is male; if the plot-level rescue structure inverts or disappears, the observed bias is partly
Extended reading notes
Core claim
On the paper's own terms, the discovery is that a structured, five-role prompt (main character, villain, helper, desired character, dispatcher) plus a five-phase Freytag arc reliably produces stories in which narrative agency and moral polarity are gendered. Across all fifteen generated stories, villains are 100% male; helpers are 67% male; dispatchers 62% male; main characters are 73% female, with Gemini and Claude choosing a female main character every time and ChatGPT in only one of five. The close reading shows that female main characters act with endurance, exploration, and moral resolve but are frequently saved, guided, or rescued by male characters at plot level, and female desired ch
Load-bearing premise
The load-bearing assumption is that five stories per model, read interpretively by the authors without a second independent rater, are enough to reveal each model's consistent narrative tendencies, and that the authors' unstated close-reading procedure reliably separates bias in the text from bias in the reader's expectations.
Editorial extensions
If this is right
- Gender-count audits of LLM storytelling are insufficient on their own; a bias score based on pronoun or character counts would have missed the plot-level rescue pattern and the male-villain connotation found here.
- Making the main character female does not neutralize narrative bias: Gemini and Claude always produced female main characters and still placed male guides or saviors at key turning points, suggesting role gender alone is not the lever.
- Debiasing efforts need to target relational structure—who rescues, who guides, who acts, who is passive—not just lexical choices or character gender ratios.
- If these patterns are representative, users of AI story generators in classrooms and creative writing inherit plots that consistently code authority, villainy, and rescue as male and beauty, fragility, and resilience as female.
- Claude's near-balance at the distribution level (52% male, 48% female) shows that balanced counts can coexist with strongly stereotyped storytelling, supporting the paper's call for multi-level assessment.
Reading between the lines
- A natural extension would be to reverse the role-gender assignment in the prompt—for example, explicitly ask for a female villain or a male desired character—to disentangle model-internal bias from the Proppian scaffold itself, which historically genders the hero male and the prize female.
- The protocol could be scaled cheaply: more models, more stories per model, and two or more independent coders with a pre-registered coding scheme would let the close-reading findings be tested quantitatively while preserving the interpretive sensitivity that surfaced the implicit patterns.
- If these results hold, benchmark suites for LLM debiasing should include narrative-relational probes—for example, measuring the gender of the character who performs the climactic rescue—because single-sentence or embedding-level tests cannot detect the bias this paper identifies.
- The study's finding that female strength is consistently framed as internal (resilience, empathy, endurance) while male agency is external (combat, repair, destruction) suggests a further testable claim: in AI-generated stories, the same trait will be described through different lexical and action frames depending on character gender.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a close-reading methodology for detecting gender narrative bias in LLM-generated stories. It prompts ChatGPT, Gemini, and Claude with a zero-shot prompt specifying five Propp-inspired character roles (MC, V, H, DC, D) and a Freytag five-phase plot structure, collecting five stories per model. The analysis covers prompt adherence, gender distribution, character descriptions, actions, and plot. The main empirical claims are: (i) gender distribution is imbalanced, with villains 100% male and main characters 73% female; (ii) descriptions follow gendered patterns (female beauty/internal strength, male strength/deformity); (iii) implicit plot-level bias persists, e.g., male saviors and female damsels even when explicit gender counts are balanced; and (iv) models differ, with ChatGPT most explicitly biased, Claude least gender-imbalanced but formulaic, and Gemini showing implicit bias despite female main characters.
Significance. If the qualitative findings are reliable, the paper makes a useful contribution as a human-centered complement to large-scale computational bias audits: it shows that gender distribution alone can underestimate bias and that narrative functions and plot structure are worth analyzing. The theoretical grounding in Propp and Freytag is appropriate, and the zero-shot neutral prompt is a methodological strength. The paper is also transparent in acknowledging sample-size and inference limitations in the Conclusion. However, the load-bearing quantitative and qualitative results rest on a small, single-coder interpretive process without a released corpus or coding protocol, which limits reproducibility. The specific observations are falsifiable in principle, but the current evidentiary basis is not yet strong enough for the abstract's generalizing claims.
major comments (3)
- [Section IV-B, Table III] The headline 'V shows the highest value, with 100% males' and all role-gender percentages depend on the authors' post hoc role assignment. The text states: 'we assigned each character the role that best matched them... making adjustments where there were obvious misinterpretations.' No decision rules, codebook, or raw character-role-gender annotations are provided. For Claude, the paper itself reports ambiguity among H, D, and DC. Thus Table III is not independently checkable and could shift under alternative plausible codings. This is load-bearing for the abstract's 'persistence of biases' claim. Please provide the story corpus with annotated role assignments and an explicit coding protocol, or reframe the percentages as exploratory single-coder frequencies rather than stable estimates.
- [Sections III and IV-E] The implicit-bias conclusions, such as the 'damsel in distress' readings and 'male characters still play guide or savior,' are produced by close reading by the authors alone. No inter-rater reliability, second coder, or audit trail is reported, and the story corpus is not published. The Conclusion acknowledges 'the representativeness of the sample and the inference of characters' functions,' but the Discussion nevertheless generalizes ('the models fail to update the portrayal') beyond what the evidence supports. To make the qualitative findings load-bearing, release the 15 stories and the reading form mentioned in Section III, and ideally have an independent coder code a subset. Otherwise, these should be framed as hypotheses or illustrations, not confirmatory results.
- [Section IV-B, Tables II and III] With n=15 (five per model), the percentages in Tables II and III are unstable: reclassifying a single character changes an entry by roughly 7 percentage points, and the '100% male V' figure is based on 15 cases. No confidence intervals, significance tests, or per-model raw counts are reported. This sample size is acceptable for a qualitative pilot, but the wording 'the results reveal' and 'V shows the highest value' overstates precision. Please use cautious language and provide per-model raw counts so readers can judge the stability of the estimates.
minor comments (4)
- [Section IV-B] There is a typo: 'female DH' should likely be 'female DC.' Also, the sentence 'female MCs have 55% of female H' lacks the underlying counts, making it hard to interpret.
- [Section II / Figure 1] Figure 1 (Freytag's pyramid) is never referenced in the text; please add an in-text citation or remove the figure.
- [Table IV] The caption 'RELATION BETWEEN AI MODELS AND BIAS EXPOSURES' is unclear. Consider renaming to 'Summary of bias levels and models affected' or similar.
- [Section III] The prompt is deliberately based on Propp's fairy-tale framework, which is historically gendered. Please discuss as a boundary condition whether the observed role-gender patterns might be partly activated by the prompt's genre/schema rather than reflecting model bias in unconstrained generation.
Circularity Check
No significant circularity: the gender-bias findings are empirical observations of LLM-generated stories, not consequences of the prompt or the authors' prior framework.
full rationale
The paper's derivation chain is empirical rather than formal. A deliberately gender-neutral prompt asks for five Propp-derived character roles and Freytag's five plot phases; stories are generated by ChatGPT, Gemini, and Claude and then analyzed via close reading. The central results—e.g., 'V shows the highest value, with 100% males' (Section IV-B) and the persistence of implicit bias in plots (Section IV-E)—are observations of the generated texts, not quantities constructed from the prompt or from the authors' own prior work. No fitted parameter is later relabeled as a prediction; no load-bearing self-citation appears; no uniqueness theorem is imported from the authors. The acknowledged limitations in the Conclusion ('the representativeness of the sample and the inference of characters' functions on their connotation') are methodological constraints on generalizability and coding reliability, not circular reductions. The post-hoc role assignment for Claude and the interpretive nature of implicit-bias identification could affect reproducibility, but they do not make the conclusions equivalent to the inputs by definition. The analysis is self-contained as an interpretive empirical study, so the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Propp's character functions and Freytag's pyramid are valid structures for analyzing LLM-generated short stories.
- domain assumption The gender of each character can be reliably inferred by human readers from names, pronouns, and descriptions.
- domain assumption Close reading by the authors, without inter-rater reliability, can detect implicit narrative bias.
- domain assumption Five stories per model are representative of each model's typical story-generation behavior under this prompt.
Cite this review
Pith. "Pith review of A Close Reading Approach to Gender Narrative Biases in AI-Generated Stories." pith.science (2026). https://pith.science/paper/2672AN3R
@misc{pith2026250809651,
author = {Pith},
title = {Pith review of: A Close Reading Approach to Gender Narrative Biases in AI-Generated Stories},
year = {2026},
howpublished = {\url{https://pith.science/paper/2672AN3R}},
note = {Machine review of arXiv:2508.09651}
}
read the original abstract
The paper explores the study of gender-based narrative biases in stories generated by ChatGPT, Gemini, and Claude. The prompt design draws on Propp's character classifications and Freytag's narrative structure. The stories are analyzed through a close reading approach, with particular attention to adherence to the prompt, gender distribution of characters, physical and psychological descriptions, actions, and finally, plot development and character relationships. The results reveal the persistence of biases - especially implicit ones - in the generated stories and highlight the importance of assessing biases at multiple levels using an interpretative approach.
Figures
Reference graph
Works this paper leans on
-
[1]
Linguistic bias, which arises from the use of certain lan- guage characteristics, such as the correlation of extended masculine or gender-coded words. This type of bias can Daniel Raffini, Agnese Macori, Tiziana Catarci are with Sapienza Uni- versity of Rome, Italy: E-mail: {raffini, macori, catarci }@diag.uniroma1.it. Marco Angelini is with Link Universi...
-
[2]
[3]. The extensive use of LLMs for content creation and text generation makes this issue increasingly urgent. Regarding gender bias, studies have explored different aspects, such as the correlation between gender and occupation [4] [5], personas [6] [7], or the use of adjectives [8]. Many of these studies also compared LLMs’ correlations with official soc...
-
[3]
Interpretative bias, when bias affects the understanding of a text and influences its interpretation. This applies to tasks such as summarizing, text analysis, information ex- traction, classification, and answering statements-based questions
-
[4]
A Close Reading Approach to Gender Narrative Biases in AI-Generated Stories
Narrative bias , when stereotypes emerge not from a single linguistic element, but from a narrative that involves multiple passages, descriptions, and actions. This bias typically arises through the free generation of stories in response to a specific prompt. In our study, we focus on narrative bias. Among existing methodologies, open-ended generation is ...
work page Pith review arXiv 2025
-
[5]
Lys stayed with him, quiet and kind, but her eyes often searched the woods
“Lys stayed with him, quiet and kind, but her eyes often searched the woods”. Here, there is a striking contrast between the woman choosing to stay with the hero who saved her and her gaze that seems to search for something elsewhere
-
[6]
He seemed the embodiment of the artistic success she so desperately desired
“Lira spoke less each day, her light dimming. Daryn stayed close, watching the lanterns burn low. One night, she whispered, ‘You didn’t save me. You only broke the cage.’ And Daryn, watching the soot fall again, finally understood. The city would not heal. Some cages, even when shattered, remain”’. The DC’s words lead the MC to a revelation that challenge...
-
[7]
I. O. Gallegos, R. A. Rossi, J. Barrow, M. Tanjim, S. Kim, F. Der- noncourt, T Yu, R. Zhang, N. K. Ahmed, ”Bias and fairness in Large Language Models: a survey”. 2023. doi: 10.48550/arXiv.2309.00770
-
[8]
N.Torres, C. Ulloa, I. Araya, et al., ”A comprehensive analysis of gender, racial, and prompt-induced biases in large language models” Int J Data Sci Anal, 2024. doi: 10.1007/s41060-024-00696-6
Show all 48 references
-
[9]
H. Zhou, D. Inkpen, B. Kantarci. 2024. Evaluating and mitigating gender bias in generative large language models, International Journal of Computers Communications & Control , V ol. 19, No. 6, 2024. doi: 10.15837/ijccc.2024.6.6853
2024 doi
-
[10]
Kotek, R
H. Kotek, R. Dockum, D. Sun, ”Gender bias and stereotypes in Large Language Models”. In Proceedings of The ACM Collective Intelligence Conference , Association for Computing Machinery, New York, 2023, pp 12–24. doi: 10.1145/3582269.3615599
2023
-
[11]
D ¨oll, M
N. D ¨oll, M. D ¨ohring, A. M ¨uller, ”Evaluating Gender Bias in Large Language Models”, 2024. arXiv:2411.09826
2024 arXiv
-
[12]
R, Ranjan, R. S. Gupta, S. N. Singh. ”Gender Biases in LLMs: Higher intelligence in LLM does not necessarily solve gender bias and stereotyping”, 2024. ArXiv:2409.19959
2024 arXiv
-
[13]
Durmus, D
M Cheng, E. Durmus, D. Jurafsky, ”Marked personas: Using natural language prompts to measure stereotypes in language models”. In Pro- ceedings of the 61st annual meeting of the association for computational linguistics (Long papers) , Association for Computational Linguistics, 2023
2023
-
[14]
Soundararajan, S
S. Soundararajan, S. J. Delany. 2024. ”Investigating Gender Bias in Large Language Models Through Text Generation”. In Proceedings of the 7th International Conference on Natural Language and Speech Processing, Association for Computational Linguistics, 2024, pp. 410- 424
2024
-
[15]
Bas, ”Assessing Gender Bias in LLMs: Comparing LLM Outputs with Human Perceptions and Official Statistics”, 2024
T. Bas, ”Assessing Gender Bias in LLMs: Comparing LLM Outputs with Human Perceptions and Official Statistics”, 2024. arXiv:2411.13738
2024 arXiv
-
[16]
Zayed, G, Mordido, S
A. Zayed, G, Mordido, S. Shabanian, I. Baldini, S. Chandar. ”Fairness- aware structured pruning in transformers”. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 38, 2024, pages 22484–22492
2024
-
[17]
Y . Li, M. Du, R. Song, X. Wang, Y . Wang. ”A survey on fairness in large language models”. 2023. arXiv:2308.10149
2023 arXiv
-
[18]
Z. Chu, Z. Wang, W. Zhang, ”Fairness in Large Language Models: A Taxonomic Survey”, SIGKDD Explor. Newsl . 26, 1, 2024, 34–48. doi: 10.1145/3682112.3682117
2024
-
[19]
Challenging systematic prejudices: an Investigation into Gender Bias in Large Language Models
UNESCO, IRCAI, “Challenging systematic prejudices: an Investigation into Gender Bias in Large Language Models”, 2024
2024
-
[20]
Navigli, S
R. Navigli, S. Conia, B. Ross. 2023. ”Biases in Large Language Models: Origins, Inventory, and Discussion”.J. Data and Information Quality 15, 2, 2023. doi: 10.1145/3597307
2023 doi
-
[21]
Ranjan, S
R. Ranjan, S. Gupta, S. N. Singh. ”A Comprehensive Survey of Bias in LLMs: Current Landscape and Future Directions”, 2024. arXiv:2409.16430
2024 arXiv
-
[22]
S. H. Kumar, S. Sahay, S. Mazumder, E. Okur, R. Manuvinakurike, N. Beckage, H. Su, H. Lee, L. Nachman, ”Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models”, 2024. arXiv:2408.03907
2024 arXiv
-
[23]
Lucy, D, Bamman
L. Lucy, D, Bamman. ”Gender and Representation Bias in GPT-3 Generated Stories”. In Proceedings of the Third Workshop on Narrative Understanding, pages 48–55. Association for Computational Linguistics. 2021
2021
-
[24]
B. J. Smith, ”What was ’close reading’? A century of method in literary studies”, The Minnesota Review 2016.87 (2016): 57-75
2016
-
[25]
Brummett, Techniques of close reading , Sage Publications, 2018
B. Brummett, Techniques of close reading , Sage Publications, 2018
2018
-
[26]
Huang, F
T. Huang, F. Brahnam, V . Shwartz, and S. Chaturvedi, ”Uncovering Implicit Gender Bias in Narratives through Commonsense Inference”, In Findings of the Association for Computational Linguistics: EMNLP 2021, Association for Computational Linguistics, 2021, pp. 3866-3873
2021
-
[27]
Caliskan, J
A. Caliskan, J. J. Bryson, A. Narayanan, ”Semantics derived auto- matically from language corpora contain human-like biases”. Science, 356(6334), 2017, pp. 183–186. doi: 10.1126/science.aal4230
2017 doi
-
[28]
Baker-Sperry and L
L. Baker-Sperry and L. Grauerholz, ”The Pervasiveness and Persistence of the Feminine Beauty Ideal in Children’s Fairy Tales,” Gender & Soci- ety, vol. 17, no. 5, pp. 711-726, 2003. doi: 10.1177/0891243203255605
2003 doi
-
[29]
Gottschall, ”The Heroine with a Thousand Faces: Universal Trends in the Characterization of Female Folk Tale Protagonists,” Evolutionary Psychology, vol
J. Gottschall, ”The Heroine with a Thousand Faces: Universal Trends in the Characterization of Female Folk Tale Protagonists,” Evolutionary Psychology, vol. 3, no. 1, 2005. doi: 10.1177/147470490500300108
2005 doi
-
[30]
Ragan, ”What Happened to the Heroines In Folktales?: An Analysis by Gender of a Multicultural Sample of Published Folktales Collected from Storytellers,” Marvels & Tales, vol
K. Ragan, ”What Happened to the Heroines In Folktales?: An Analysis by Gender of a Multicultural Sample of Published Folktales Collected from Storytellers,” Marvels & Tales, vol. 23, no. 2, pp. 227-247, 2009. doi: 10.1353/mat.2009.a369130
2009 doi
-
[31]
Gottschall et al
J. Gottschall et al. , ”The ’Beauty Myth’ Is No Myth,” Hum Nat , vol. 19, pp. 174–188, 2008. doi: 10.1007/s12110-008-9035-3
2008 doi
-
[32]
C. J. Francemone, M. Grizzard, K. Fitzgerald, J. Huang, and C. Ahn, ”Character Gender and Disposition Formation in Narratives: The Role of Competing Schema,” Media Psychology, vol. 25, no. 4, pp. 547-564,
-
[33]
Todorov, Po´etique de la prose
T. Todorov, Po´etique de la prose . Paris, France: `Edition du Seuil, 1971
1971
-
[34]
E. R. Nusbaumer, ”The Cinderella Stereotype: A Comparative Study of Love in a Fallen City and Cinderella,” Comparative Liter- ature: East & West , vol. 8, no. 2, pp. 221-233, 2024. doi: 10.1080/25723618.2024.2441001
2024
-
[35]
Chang, ”Love in a Fallen City,” in Love in a Fallen City and Other Stories, trans
E. Chang, ”Love in a Fallen City,” in Love in a Fallen City and Other Stories, trans. K. S. Kingsbury and E. Chang, London, U.K.: Penguin, 2007
2007
-
[36]
Propp, Morphology of the Folk Tale
V . Propp, Morphology of the Folk Tale. Austin, TX, USA: University of Texas Press, 1968
1968
-
[37]
Dinan, A
E. Dinan, A. Fan, A. Williams, J. Urbanek, D. Kiela, and J. Weston, ”Queens are Powerful too: Mitigating Gender Bias in Dialogue Genera- tion,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Ling...
2020
-
[38]
D. Gala, M. O. Khursheed, H. Lerner, B. O’Connor, and M. Iyyer, ”Analyzing Gender Bias within Narrative Tropes,” in Proceedings of the Fourth Workshop on Natural Language Processing and Computational Social Science , Association for Computational Linguistics, 2020, pp. 212-217
2020
-
[39]
A. J. Greimas, S´emantique structurale : recherche de m ´ethode. Paris, France: Larousse, 1966
1966
-
[40]
Daulay, ”Heroine and Princess: Women Image Portrayed in selected Disney’s Stories”, J
R. Daulay, ”Heroine and Princess: Women Image Portrayed in selected Disney’s Stories”, J. Basis V ol. 8 No. 1, 2021
2021
-
[41]
Cavalloro, Leggere storie
V . Cavalloro, Leggere storie. Introduzione all’analisi del testo narrativo. Roma, Italy: Carocci, 2014
2014
-
[42]
Freytag, Technique of the Drama: An Exposition of Dramatic Composition and Art
G. Freytag, Technique of the Drama: An Exposition of Dramatic Composition and Art . Chicago, IL, USA: Scott, Foresman & Company, 1900
1900
-
[43]
Donini, Ed
Aristotele, Poetica, P. Donini, Ed. Torino, Italy: Einaudi, 2008
2008
-
[44]
E. Chen, R. Zhan, Y . Lin, H. Chen, ”From Structured Prompts to Open Narratives: Measuring Gender Bias in LLMs Through Open-Ended Storytelling”, 2025. arXiv:2503.15904
2025
-
[45]
G ´omez-Rodr´ıguez, P
C. G ´omez-Rodr´ıguez, P. Williams. 2023. ”A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing”. In Findings of the Association for Computational Linguistics , pages 14504–14528, Association for Computational Linguistics, 2023
2023
-
[46]
C. Stover, ”Damsels and Heroines: The Conundrum of the Post-Feminist Disney Princess”, LUX: A Journal of Transdisciplinary Writing and Research from Claremont Graduate University , V ol. 2m Iss. 1, 2013
2013
-
[48]
Aupitak, ”Feminist Quest Heroine: Deconstruction of Male Heroism in the 21st Century Fairy Tale Narrative”, In N
T. Aupitak, ”Feminist Quest Heroine: Deconstruction of Male Heroism in the 21st Century Fairy Tale Narrative”, In N. Le Clue (Ed.), Gender and the Male Character in 21st Century Fairy Tale Narratives , Emerald Publishing Limited, Leeds, 2024, pp. 39-50. doi:10.1108/978-1-83753...
2024 doi
-
[2021]
doi: 10.1080/15213269.2021.2006718
2021
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.