{"id":"795e47ac-d8ad-4706-aa15-9e2ad20ac5f6","arxiv_id":"2508.09651","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A close reading of 15 AI-generated stories finds that even when character counts are balanced, narrative roles, descriptions, and plot dynamics remain gender-stereotyped (e.g., every villain is male).","lead":"This paper applies close reading, a literary-analysis method, to 15 stories generated by ChatGPT, Gemini, and Claude under a structured Propp-and-Freytag prompt. It reports persistent gendered patterns, including all-male villains and female main characters who are beautiful but often still rescued by men.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post hoc role assignment and unverified coding are load-bearing for the 100% male-villain and implicit-bias conclusions.","rationale":"The reader's weakest_assumption correctly points to the small sample and lack of inter-rater reliability. I agree that these are the most vulnerable aspects, but I want to sharpen the concern to a concrete technical step: the post hoc role assignment in Section IV-B. The paper explicitly says that for Claude they 'assigned each character the role that best matched them' and made 'adjustments where there were obvious misinterpretations.' This is a subjective coding decision with no transparency. Since the headline numeric claims (Table II and III) are all role-conditioned, any instability in role assignment directly undermines the quantitative part of the central claim. Moreover, the qualitative implicit-bias claims rely on the same kind of interpretive judgment without a second coder. I do not think the paper is fatally flawed—it is a qualitative study with acknowledged limitations and a clear prompt, and the authors provide substantial quotes from the stories. But the current evidence is not strong enough for ACCEPT. The reader's CONDITIONAL verdict is appropriate, provided the authors release the data and coding materials. My concrete test would settle whether the concern actually lands: if independent coding reproduces the tables, the concern is mitigated; if not, the 100% male villain and other claims are artifacts of subjective coding. I chose 'partial' agreement because the reader framed the issue as sample size and coding reliability generically, while I identify the post hoc role reassignment as the specific load-bearing element. The paper's own limitation statement supports this, so I am not introducing a novel objection outside the authors' awareness. However, the authors' acknowledgment in the conclusion does not fully temper the strong generalizations in the Discussion, which is why the test is needed before the claims are accepted as robust.","tokens_in":12109,"tokens_out":3116,"duration_ms":34944,"concrete_test":"Release the 15 generated stories and the reading form. Have two independent annotators, blind to the paper's hypotheses, code each character for gender and Proppian role (MC, V, H, DC, D) using a written codebook that defines each role. Compute Cohen's kappa for role and gender assignments. Then recompute Table II and Table III for each annotator's coding. If the 100% male V or the MC gender distribution changes across annotators, or if kappa is below 0.6, the quantitative claims are not robust. Additionally, ask the annotators to identify instances of implicit bias (e.g., male rescue of female MC) and compare their findings with the paper's narrative interpretations. This directly tests whether the role-assignment and close-reading steps are reproducible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that LLMs reproduce gendered narrative structures, with villains always male and implicit bias persisting even when gender counts are balanced—rests on two unsecured steps. First, Section IV-B states that for Claude, where roles are ambiguous, the authors 'assigned each character the role that best matched them... making adjustments where there were obvious misinterpretations.' No codebook or decision rules for this assignment are provided, so the quantitative gender-by-role tables (e.g., V 100% male, MC 73% female) depend on subjective judgments that could change the contingency if done differently. Second, the close-reading evidence for implicit bias (Section IV-E) is produced by the authors without inter-rater reliability; the paper's own conclusion acknowledges 'the representativeness of the sample and the inference of characters' functions' as limitations, but the generality of the discussion (e.g., 'the models fail to update the portrayal') goes beyond what five stories per model can support. The sample size (n=15) is adequate for a qualitative pilot, but the presentation of percentages as if they were stable estimates, combined with an unreported coding process, makes the strongest claims vulnerable to observer expectation. If another reader coded the same stories, the 100% male-V result and the 'damsel in distress' plot readings might not replicate.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a close-reading methodology for detecting gender narrative bias in LLM-generated stories. It prompts ChatGPT, Gemini, and Claude with a zero-shot prompt specifying five Propp-inspired character roles (MC, V, H, DC, D) and a Freytag five-phase plot structure, collecting five stories per model. The analysis covers prompt adherence, gender distribution, character descriptions, actions, and plot. The main empirical claims are: (i) gender distribution is imbalanced, with villains 100% male and main characters 73% female; (ii) descriptions follow gendered patterns (female beauty/internal strength, male strength/deformity); (iii) implicit plot-level bias persists, e.g., male saviors and female damsels even when explicit gender counts are balanced; and (iv) models differ, with ChatGPT most explicitly biased, Claude least gender-imbalanced but formulaic, and Gemini showing implicit bias despite female main characters.","tokens_in":12334,"tokens_out":4610,"duration_ms":52939,"significance":"If the qualitative findings are reliable, the paper makes a useful contribution as a human-centered complement to large-scale computational bias audits: it shows that gender distribution alone can underestimate bias and that narrative functions and plot structure are worth analyzing. The theoretical grounding in Propp and Freytag is appropriate, and the zero-shot neutral prompt is a methodological strength. The paper is also transparent in acknowledging sample-size and inference limitations in the Conclusion. However, the load-bearing quantitative and qualitative results rest on a small, single-coder interpretive process without a released corpus or coding protocol, which limits reproducibility. The specific observations are falsifiable in principle, but the current evidentiary basis is not yet strong enough for the abstract's generalizing claims.","major_comments":[{"comment":"The headline 'V shows the highest value, with 100% males' and all role-gender percentages depend on the authors' post hoc role assignment. The text states: 'we assigned each character the role that best matched them... making adjustments where there were obvious misinterpretations.' No decision rules, codebook, or raw character-role-gender annotations are provided. For Claude, the paper itself reports ambiguity among H, D, and DC. Thus Table III is not independently checkable and could shift under alternative plausible codings. This is load-bearing for the abstract's 'persistence of biases' claim. Please provide the story corpus with annotated role assignments and an explicit coding protocol, or reframe the percentages as exploratory single-coder frequencies rather than stable estimates.","section":"Section IV-B, Table III"},{"comment":"The implicit-bias conclusions, such as the 'damsel in distress' readings and 'male characters still play guide or savior,' are produced by close reading by the authors alone. No inter-rater reliability, second coder, or audit trail is reported, and the story corpus is not published. The Conclusion acknowledges 'the representativeness of the sample and the inference of characters' functions,' but the Discussion nevertheless generalizes ('the models fail to update the portrayal') beyond what the evidence supports. To make the qualitative findings load-bearing, release the 15 stories and the reading form mentioned in Section III, and ideally have an independent coder code a subset. Otherwise, these should be framed as hypotheses or illustrations, not confirmatory results.","section":"Sections III and IV-E"},{"comment":"With n=15 (five per model), the percentages in Tables II and III are unstable: reclassifying a single character changes an entry by roughly 7 percentage points, and the '100% male V' figure is based on 15 cases. No confidence intervals, significance tests, or per-model raw counts are reported. This sample size is acceptable for a qualitative pilot, but the wording 'the results reveal' and 'V shows the highest value' overstates precision. Please use cautious language and provide per-model raw counts so readers can judge the stability of the estimates.","section":"Section IV-B, Tables II and III"}],"minor_comments":[{"comment":"There is a typo: 'female DH' should likely be 'female DC.' Also, the sentence 'female MCs have 55% of female H' lacks the underlying counts, making it hard to interpret.","section":"Section IV-B"},{"comment":"Figure 1 (Freytag's pyramid) is never referenced in the text; please add an in-text citation or remove the figure.","section":"Section II / Figure 1"},{"comment":"The caption 'RELATION BETWEEN AI MODELS AND BIAS EXPOSURES' is unclear. Consider renaming to 'Summary of bias levels and models affected' or similar.","section":"Table IV"},{"comment":"The prompt is deliberately based on Propp's fairy-tale framework, which is historically gendered. Please discuss as a boundary condition whether the observed role-gender patterns might be partly activated by the prompt's genre/schema rather than reflecting model bias in unconstrained generation.","section":"Section III"}],"recommendation":"major_revision","confidential_remarks":"The paper fits a human-centered AI venue but the evidentiary standard for quantitative claims is below what the abstract implies. I would support acceptance after major revision that includes releasing the story corpus and coding protocol, adding inter-rater reliability or framing the analysis as exploratory, and softening the generalizing language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look, but treat the numbers as illustrations, not measurements. The paper does something genuinely useful: it brings narratology (Propp's roles, Freytag's pyramid) into a structured prompt, generates stories from ChatGPT, Gemini, and Claude, and then reads them closely for gender bias across character roles, descriptions, actions, and plot. The central observation—that even when explicit gender counts are fairly balanced, narrative roles still skew male (villains, helpers, dispatchers) and female characters still get described as beautiful, resilient, and sometimes rescued—is important and probably right.\n\nThe soft spots are the expected ones. Fifteen stories is a pilot. The percentages in Tables II and III come with no uncertainty, and '100% male villains' is 15 out of 15, which sounds stronger than it is. The role assignment for Claude required ad hoc adjustments (Section IV-B), and the close reading of plots has no second coder or codebook, so the implicit-bias readings are hard to check. The story corpus isn't released. The stress-test note is fair on this.\n\nThe authors aren't overselling: the conclusion explicitly acknowledges the representativeness of the sample and the inference of characters' functions. But the discussion does drift into general claims like 'the models fail to update the portrayal,' which go beyond what 15 stories support. A revision should either temper that language or add more data and coding validation.\n\nAs a qualitative pilot, it's a solid contribution. It demonstrates a method and produces a plausible, falsifiable hypothesis for larger studies. I'd send it to peer review, expecting reviewers to ask for the corpus, a second coder, and a sharper separation of measurement from interpretation. I'd also cite it as an example of narrative-level bias auditing, though not as evidence for the specific percentages.","headline":"Useful narratological audit of LLM stories, but 15 stories and post hoc coding make the strong quantitative claims illustrative rather than evidential.","tokens_in":12867,"tokens_out":3837,"would_cite":true,"duration_ms":39669,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Gender bias in AI-generated stories persists at the narrative level even when character counts look balanced.","keywords":["gender bias","large language models","AI-generated stories","close reading","Proppian character functions","Freytag's pyramid","narrative bias","generative AI"],"falsifier":"Generate, say, fifty stories per model with the same prompt, have independent coders who are blind to the hypothesis label each character's gender and code who rescues, guides, or acts decisively at the climax; if villains are not overwhelmingly male or female main characters are not rescued by men more often than role-gender chance would predict, the paper's central pattern fails. An even sharper test: rerun the prompt with an explicit instruction that the villain is female and the desired character is male; if the plot-level rescue structure inverts or disappears, the observed bias is partly","tokens_in":11971,"feed_emoji":"📖","tokens_out":6041,"duration_ms":56073,"temperature":0.7,"pith_summary":"This paper argues that gender bias in AI-generated stories cannot be measured by counting how many characters are male or female. Analyzing fifteen zero-shot stories produced by ChatGPT, Gemini, and Claude with a prompt based on Propp's character roles and Freytag's narrative arc, the authors read each text closely and find that the villain is male in every story, while helpers and dispatchers are mostly male and main characters are mostly female. They find that female main characters are repeatedly paired with male rescuers or guides, female desired characters are largely passive, and descriptions consistently code women as beautiful and emotionally resilient and men as strong, damaged, or dominant. The core claim is that implicit narrative bias persists even where explicit gender distribution appears balanced, and that only an interpretive, multi-level reading can expose it. If true, quantitative audits of LLM output miss the most insidious form of stereotype propagation.","feed_headline":"Male villains, rescued women persist in AI stories","feed_subtitle":"Beyond character counts: plot and description analysis finds stereotypes in ChatGPT, Gemini, and Claude.","key_machinery":"The central mechanism is a standardized generation-and-reading protocol: a zero-shot prompt that asks each model to write a roughly 500-word story containing five Propp-derived characters—main character (MC), villain (V), helper (H), desired character (DC), dispatcher (D)—and to follow Freytag's five-phase arc (exposition, rise, climax, return/fall, catastrophe). Characters and phases must not be named in the story. This yields comparable outputs across models. The analysis side uses close reading with a reading form that records prompt adherence, gender distribution of each role, physical and psychological descriptors, action types, and plot-level relationships, allowing composite or implic","core_discovery":"On the paper's own terms, the discovery is that a structured, five-role prompt (main character, villain, helper, desired character, dispatcher) plus a five-phase Freytag arc reliably produces stories in which narrative agency and moral polarity are gendered. Across all fifteen generated stories, villains are 100% male; helpers are 67% male; dispatchers 62% male; main characters are 73% female, with Gemini and Claude choosing a female main character every time and ChatGPT in only one of five. The close reading shows that female main characters act with endurance, exploration, and moral resolve but are frequently saved, guided, or rescued by male characters at plot level, and female desired ch","pith_inferences":["A natural extension would be to reverse the role-gender assignment in the prompt—for example, explicitly ask for a female villain or a male desired character—to disentangle model-internal bias from the Proppian scaffold itself, which historically genders the hero male and the prize female.","The protocol could be scaled cheaply: more models, more stories per model, and two or more independent coders with a pre-registered coding scheme would let the close-reading findings be tested quantitatively while preserving the interpretive sensitivity that surfaced the implicit patterns.","If these results hold, benchmark suites for LLM debiasing should include narrative-relational probes—for example, measuring the gender of the character who performs the climactic rescue—because single-sentence or embedding-level tests cannot detect the bias this paper identifies.","The study's finding that female strength is consistently framed as internal (resilience, empathy, endurance) while male agency is external (combat, repair, destruction) suggests a further testable claim: in AI-generated stories, the same trait will be described through different lexical and action frames depending on character gender."],"forward_implications":["Gender-count audits of LLM storytelling are insufficient on their own; a bias score based on pronoun or character counts would have missed the plot-level rescue pattern and the male-villain connotation found here.","Making the main character female does not neutralize narrative bias: Gemini and Claude always produced female main characters and still placed male guides or saviors at key turning points, suggesting role gender alone is not the lever.","Debiasing efforts need to target relational structure—who rescues, who guides, who acts, who is passive—not just lexical choices or character gender ratios.","If these patterns are representative, users of AI story generators in classrooms and creative writing inherit plots that consistently code authority, villainy, and rescue as male and beauty, fragility, and resilience as female.","Claude's near-balance at the distribution level (52% male, 48% female) shows that balanced counts can coexist with strongly stereotyped storytelling, supporting the paper's call for multi-level assessment."],"supporting_citations":[{"why":"Supplies Propp's character functions, adapted in the prompt into the five roles the generated stories must contain.","marker":"[29]"},{"why":"Supplies Freytag's pyramid, the five-phase narrative structure the prompt requires the stories to follow.","marker":"[35]"},{"why":"Provides the prior UNESCO large-scale story-generation study of gender bias that this close reading approach contrasts with and extends.","marker":"[13]"},{"why":"Documents the persistent feminine-beauty ideal in fairy tales, the historical frame against which the paper interprets female character descriptions.","marker":"[22]"},{"why":"Earlier analysis of gender representation in GPT-3 generated stories that focused only on the main character, which this study extends to multiple character roles.","marker":"[17]"},{"why":"Demonstrates detection of implicit gender bias in narratives through inference, supporting the paper's claim that implicit bias requires interpretive methods.","marker":"[20]"},{"why":"Supports the open-ended storytelling methodology for measuring gender bias in LLMs, which the paper adopts for narrative-level analysis.","marker":"[37]"},{"why":"Establishes that implicit associations in language can influence beliefs, motivating the paper's concern with narrative bias in AI-generated stories.","marker":"[21]"}],"fun_headline_variants":["AI stories: all villains male, women often rescued","Male villains dominate AI tales, women get rescued","AI narratives: male villains, female leads rescued","Study finds AI stories default to male villains, female rescues","In AI fiction, villains are male and women need saving"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing assumption is that five stories per model, read interpretively by the authors without a second independent rater, are enough to reveal each model's consistent narrative tendencies, and that the authors' unstated close-reading procedure reliably separates bias in the text from bias in the reader's expectations.","fun_headline_variants_meta":{"raw":{"variants":["AI stories: all villains male, women often rescued","Male villains dominate AI tales, women get rescued","AI narratives: male villains, female leads rescued","Study finds AI stories default to male villains, female rescues","In AI fiction, villains are male and women need saving"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1036,"prompt_tokens":607,"completion_tokens":429,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":351,"completion_tokens_details":{"reasoning_tokens":352}},"tokens_in":351,"tokens_out":429,"duration_ms":4679,"temperature":1.0,"reasoning_tokens":352,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:54:18.471461+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate, say, fifty stories per model with the same prompt, have independent coders who are blind to the hypothesis label each character's gender and code who rescues, guides, or acts decisively at the climax; if villains are not overwhelmingly male or female main characters are not rescued by men more often than role-gender chance would predict, the paper's central pattern fails. An even sharper test: rerun the prompt with an explicit instruction that the villain is female and the desired character is male; if the plot-level rescue structure inverts or disappears, the observed bias is partly","supporting_citations":[{"cited_title":"Gottschall, ”The Heroine with a Thousand Faces: Universal Trends in the Characterization of Female Folk Tale Protagonists,” Evolutionary Psychology, vol","cited_arxiv_id":null,"evidence_quote":"Supplies Propp's character functions, adapted in the prompt into the five roles the generated stories must contain."},{"cited_title":"Chang, ”Love in a Fallen City,” in Love in a Fallen City and Other Stories, trans","cited_arxiv_id":null,"evidence_quote":"Supplies Freytag's pyramid, the five-phase narrative structure the prompt requires the stories to follow."},{"cited_title":"Durmus, D","cited_arxiv_id":null,"evidence_quote":"Provides the prior UNESCO large-scale story-generation study of gender bias that this close reading approach contrasts with and extends."},{"cited_title":"Dinan, A","cited_arxiv_id":null,"evidence_quote":"Supports the open-ended storytelling methodology for measuring gender bias in LLMs, which the paper adopts for narrative-level analysis."}],"review_version":1}