{"id":"7b7c5d8e-684b-49f1-8a59-41990bae1b06","arxiv_id":"2507.19156","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In Italian, both ChatGPT (gpt-4o-mini) and Gemini (gemini-1.5-flash) associate higher-status professional roles with male pronouns and subordinate roles with female pronouns, e.g., 97-100% of 'she' responses pointed to the assistant.","lead":"Researchers asked two AI chatbots, ChatGPT and Gemini, to answer Italian workplace questions that left ambiguous the gender of a professional. Both models overwhelmingly matched female pronouns with subordinate jobs like assistant and male pronouns with senior jobs like manager.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Paper's summary claim of 'constantly associating leadership roles with males' is contradicted by its own JP2 results; the data support a weaker, pair-dependent bias.","rationale":"The reader's weakest_assumption was that the job titles might not be lexically gender-neutral, which could contaminate the observed pronoun-role associations with grammatical gender effects. That concern is real but secondary: for JP1 and JP2 the role nouns are effectively invariable in Italian, and for JP3 both 'chef' and 'sous chef' are masculine loanwords, so morphology alone cannot explain the strong female-pronoun preference for 'sous chef'. The more immediate problem is internal to the paper's own data: the abstract and Section 4 claim a 'constant' male-leadership association, but JP2 clearly shows the opposite for male pronouns. Since the paper's headline finding is precisely this summary, the overstatement is load-bearing. The data do support a clear and consistent association of female pronouns with subordinate roles, so the paper is not invalidated; it requires a careful rewrite of the central claim. The reader's verdict of CONDITIONAL already captures the need for revision, so the stress-test does not change it.","tokens_in":11437,"tokens_out":11898,"duration_ms":120482,"concrete_test":"Recompute from Tables 3-8, for each model and job pair, the four conditionals P(leader|He), P(follower|He), P(leader|She), P(follower|She). If in any pair P(follower|He) > P(leader|He) (as occurs for JP2 in both models), the summary phrase 'constantly associating leadership roles with males' is false for that pair. The manuscript should be revised to report the pair-specific directional pattern, and the abstract should not use 'constantly'. This requires no new data collection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4's summary answer to RQ1 states that the models 'constantly' associate leadership roles with males and subordinate ones with women. The 'constantly' is not supported by the paper's own Tables 3-8. For JP2 (Principal-Professor), both models with the male pronoun 'he' choose the subordinate role 'Professor' more often than 'Principal': Gemini 0.64 vs 0.36 (Table 4), ChatGPT 0.68 vs 0.32 (Table 7). For JP3 (Chef-Sous Chef), male pronouns yield an almost balanced split (Gemini 0.54/0.46, ChatGPT 0.62/0.38). Only JP1 (Manager-Assistant) shows a clear male-leadership association, and only ChatGPT is near-deterministic (0.94). The only consistent pattern across all pairs and models is the female pronoun -> subordinate role (P >= 0.89 everywhere). The abstract and Section 4 therefore overstate the evidence: the male-leadership claim is not 'constant' and in JP2 is reversed. A further terminological issue: the prompts are called 'ungendered' but include an explicit gendered pronoun (lui/lei), so the reported metrics are anaphora-resolution probabilities P(role|pronoun), not open-ended generation from gender-neutral prompts. The headline conclusion should be restated accordingly.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a black-box audit of two proprietary LLMs (OpenAI ChatGPT gpt-4o-mini and Google Gemini gemini-1.5-flash) for gender-stereotype associations in Italian. Using prompts built from three hierarchical job-title pairs (Manager-Assistant, Principal-Professor, Chef-Sous Chef), five base scenarios, two orders, and two pronouns (lui/lei), the authors collected 3,600 API responses and computed conditional probabilities P(role|pronoun) and P(pronoun|role). They report strong asymmetric associations, e.g. Gemini associating 'she' with Assistant rather than Manager 100% of the time, and conclude that both models reflect traditional gender norms. The paper also provides a public repository with code, prompts, and detailed results.","tokens_in":11671,"tokens_out":3245,"duration_ms":35399,"significance":"If the conclusions were stated accurately, this would be a useful, reproducible contribution to the small but growing literature on LLM bias in gendered non-English languages. The experimental design is transparent: model versions are named, the prompt inventory is fully listed, counts are reported, and the anonymous GitHub link makes replication straightforward. The paper also honestly acknowledges several limitations in Section 6. However, the central quantitative finding as stated is broader than the data support: the male-leadership association is not constant across job pairs, and the prompts are not ungendered because each embeds an explicit pronoun. With a corrected interpretation and uncertainty quantification, the study can still stand as a modest empirical data point, but its current headline claims overreach.","major_comments":[{"comment":"The prompts are repeatedly described as 'ungendered', but every prompt contains an explicit gendered pronoun (lui/lei), e.g. Table 2, P1-A: 'perché lui era in ritardo'. The measured quantities are therefore anaphora-resolution probabilities P(role|pronoun): given that the model is told the late person is 'he' or 'she', which role does it pick? This is a substantially different construct from open-ended generation from a gender-neutral prompt, and it affects the interpretation of RQ1 and the abstract's claim that the study examines 'responses to ungendered prompts'. The terminology should be corrected throughout, and the discussion should be framed as pronoun-to-role association rather than unconstrained generation.","section":"Section 3.2, Prompt Design; Section 4, Summary of RQ1; Abstract"},{"comment":"The claim that 'both Gemini and ChatGPT reflected traditional gender norms by constantly associating leadership roles with males and subordinate ones with women' is contradicted by the paper's own tables. For JP2 (Principal-Professor), the male pronoun selects the subordinate role (Professor) more often than the leadership role: Gemini P(Professor|He)=0.64 versus P(Principal|He)=0.36 (Table 4), and ChatGPT P(Professor|He)=0.68 versus P(Principal|He)=0.32 (Table 7). For JP3 (Chef-Sous Chef), the male-pronoun split is near balanced for both models (Gemini 0.54/0.46, ChatGPT 0.62/0.38). The only pattern that is consistent across all three pairs and both models is the female-pronoun-to-subordinate-role association (P>=0.89 everywhere). The abstract and Section 4 should be revised to state this pair-dependent pattern rather than claiming a constant male-leadership association.","section":"Section 4, Summary of the Answer to RQ1; Abstract; Section 5"},{"comment":"No confidence intervals or significance tests accompany any of the reported conditional probabilities, even though sample sizes are on the order of 300 per condition and exact counts are already in the tables. For example, Table 3 reports P(Manager|She)=0.00 from 296 observations, but without an interval the reader cannot assess the precision of this estimate or how it compares with, say, P(Manager|He)=0.69. The 2% anomaly exclusion is also reported only globally; the tables show that denominators vary (e.g. Table 4 has 293 in the 'she' column, Table 5 has 287), so the per-condition exclusion counts should be disclosed. Adding binomial confidence intervals or a simple test of association would materially strengthen the paper's claims of systematic bias.","section":"Section 4, Tables 3-8; Section 3.2, Bias Quantification Metrics"},{"comment":"The load-bearing premise that the three job-title pairs are 'as neutral as possible' in Italian is asserted without linguistic evidence or a control. If any title carries a masculine lexical default (for instance, 'chef' and 'sous chef' are grammatically masculine nouns in Italian), then part of the Chef-He association in JP3 could reflect grammatical gender rather than social stereotyping. This is especially relevant because the observed male-pronoun associations are the weakest and most inconsistent findings. The authors should either provide evidence for lexical neutrality of the Italian titles or explicitly acknowledge that the design cannot separate grammatical gender from stereotypical association, and temper the conclusions accordingly.","section":"Section 3.2, Job Pair Selection"}],"minor_comments":[{"comment":"The text says 'in a way similar to JP2' but the preceding discussion concerns JP1; this should read 'in a way similar to JP1'.","section":"Section 4.1, Google Gemini, JP2 bullet"},{"comment":"The column headers are dense and partly redundant (e.g. 'Y = he(lui)' followed by three sub-columns each labeled 'he(lui)'). The tables would be much easier to read if the top header distinguished 'counts' from 'P(Y|B)' and 'P(B|Y)' explicitly, and if the total row were separated from the data rows.","section":"Tables 3-8"},{"comment":"The abstract says 'a range of 3600 responses' but does not report how many responses were excluded as anomalies or how many entered the probability calculations; since the tables show varying totals, the final usable count should be stated.","section":"Abstract"},{"comment":"The sentence 'To adhere to rate constraints and preserve reproducibility, brief delays (sleep) were added in between API calls' is good practice, but the paper should also state the API access dates and the exact model snapshot identifiers returned by the APIs, since proprietary models are updated over time.","section":"Section 3.2, Experimental Setup"},{"comment":"The citation to Kotek et al. [5] says LLMs are '3-6 times more likely to assign stereotypical professional roles when forced to answer in a gendered manner', but the precise conditions under which this ratio was measured are not summarized; a one-sentence context would help the reader judge comparability with this study.","section":"Section 2, Related Work"}],"recommendation":"major_revision","confidential_remarks":"This is a straightforward empirical audit with fully reported counts, so the path to a publishable paper is clear. The main risk is that the authors may prefer to retain their current framing; as a referee I would insist that the conclusions be brought in line with the pair-dependent data and that the 'ungendered prompt' terminology be abandoned. The paper is within scope for a fairness/auditing venue but is too narrowly framed for a general CS venue without a stronger methodological contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a small, honest measurement paper on gender stereotypes in Italian LLM outputs. It is not a new phenomenon, and the paper cites the prior Italian work (Ruzzetti et al., 2023) that already established the general direction. Its value is concrete: reproducible, model-specific probabilities for gpt-4o-mini and gemini-1.5-flash in a WinoBias-style coreference setup with three hierarchical job pairs. The female-pronoun asymmetry is robust across all pairs: P(subordinate | lei) is between 0.89 and 1.00. That part of the conclusion holds.\n\nWhat is actually new: the specific measurements for these two model versions, the Italian job-pair set, and a public repo with prompts, responses, and code. The conditional probability calculations are straightforward and the tables support the main empirical claim. This is a clean black-box audit with no fitted parameters and no reliance on the authors' own prior results.\n\nWhere I part ways with the paper and with the stress-test note. The stress-test says the 'constantly associating leadership roles with males' claim is contradicted by the JP2 results. I think that is too strong. Read as P(male | leader), the data do support the association: P(He|Manager)=1.00, P(He|Principal)=1.00, P(He|Chef)=0.89 for Gemini, and 0.97, 0.81, 0.84 for ChatGPT. What the JP2 data actually show is pair dependence: given a male pronoun, the model does not reliably pick the leader. The robust effect is female pronoun -> subordinate. The abstract and Section 4 overstate the male-leadership result, but it is not contradicted.\n\nReal soft spots: (1) the prompts are called 'ungendered' even though each one contains an explicit lui/lei; only the job titles are intended to be neutral. That wording should be fixed. (2) No confidence intervals, significance tests, or decoding parameters are reported, so the probabilities look more exact than they are. (3) The ~2% of ambiguous responses are excluded with no sensitivity analysis; the paper asserts this does not affect validity but does not demonstrate it. (4) Only three job pairs limits generalizability, and the authors acknowledge this.\n\nWho this is for: people working on non-English bias auditing, EU AI Act risk assessment, or Italian NLP. It is not a breakthrough, but it is a usable data point and the repo makes it easy to build on. I would send it to referees with an expectation of light-to-moderate revision on framing, statistical reporting, and the 'ungendered' terminology. I would engage with it as a referee.","headline":"A modest, reproducible Italian-language bias audit whose female-pronoun result is robust; the authors overstate the male-leadership finding and mislabel their prompts as ungendered.","tokens_in":12188,"tokens_out":3529,"would_cite":true,"duration_ms":35794,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Given ungendered Italian prompts, ChatGPT and Gemini both attach leadership roles to male pronouns and subordinate roles to female pronouns at near-deterministic rates.","keywords":["gender bias","large language models","Italian","stereotypes","professional roles","conditional probability","black-box auditing","prompt design"],"falsifier":"Rerun the exact experiment with the leadership titles spelled with explicit feminine agreement in the prompt (e.g., 'la preside', 'la manager', 'la chef') while keeping the subordinate titles unchanged; if the models then attach 'lei' to the leadership roles at rates comparable to the original 'lui' rates, the leader-male association is substantially a product of Italian grammatical gender defaults rather than social stereotyping.","tokens_in":11266,"feed_emoji":"🤖","tokens_out":8875,"duration_ms":83319,"temperature":0.7,"pith_summary":"The paper tries to establish that two widely used commercial chatbots, OpenAI ChatGPT (gpt-4o-mini) and Google Gemini (gemini-1.5-flash), reproduce traditional gender stereotypes when answering ungendered prompts in Italian about professional hierarchies. Using zero-shot, no-context prompts that pair a leadership title ('manager', 'preside', 'chef') with a subordinate title ('assistente', 'insegnante', 'sous chef') and a gendered pronoun 'lui' or 'lei', the authors measured how often each pronoun is attached to each role across 3,600 API responses. They report that both models associate leadership roles with masculine pronouns and subordinate roles with feminine pronouns at near-deterministic rates: for instance, Gemini attached 100% (ChatGPT 97%) of 'she' answers to the assistant rather than the manager. The finding matters because these models are entering hiring, education, and public administration, where such associations could reinforce inequality, and because Italian's rich grammatical gender makes the test a sharp, non-English probe.","feed_headline":"Chatbots tie Italian leadership roles to men nearly every time","feed_subtitle":"In gender-neutral Italian prompts, both models give 'she' to assistants and 'he' to managers in nearly all answers.","key_machinery":"The load-bearing mechanism is a controlled prompt corpus combined with two conditional probability metrics, $P(Y\\mid B)$ (probability that a job title $Y$ is output given pronoun $B$) and $P(B\\mid Y)$ (probability that a pronoun $B$ is used given output title $Y$). The prompts are zero-shot, no-context workplace scenarios: five templates, each instantiated with a hierarchical job pair and a gendered pronoun 'lui' or 'lei', in both title orders, producing 60 distinct prompts and 3,600 collected responses. The three job pairs — manager/assistant, preside/insegnante, chef/sous chef — were chosen as 'as neutral as possible' in Italian while preserving a clear hierarchy, so that an unbiased model would have no lexical reason to attach one pronoun to one title. The metrics convert raw co-occurrence counts into the probabilities the paper interprets as measurable gender bias.","core_discovery":"The paper's central claim is that, given prompts with no gender marking on the job titles, both ChatGPT and Gemini behave as if professional seniority and gender were coupled: the higher-ranking role is treated as male and the lower-ranking one as female. The evidence is a set of conditional probabilities computed over $3{,}600$ responses. Across all three job pairs, the probability of the leadership title given a female pronoun is very low — for Gemini, $P(\\text{manager}\\mid\\text{she})=0$, $P(\\text{preside}\\mid\\text{she})=0$, $P(\\text{chef}\\mid\\text{she})=0.07$ — while the complementary probabilities for subordinate titles with 'she' and leadership titles with 'he' are correspondingly high; ChatGPT shows the same direction with values such as $P(\\text{manager}\\mid\\text{she})=0.03$ and $P(\\text{assistant}\\mid\\text{she})=0.97$. The paper reads these numbers as models 'reflecting traditional gender norms' and notes the pattern is stable across both chatbots, with ChatGPT slightly more unbalanced than Gemini.","pith_inferences":["The authors do not separate grammatical from social causes: titles such as 'preside' and 'chef' may carry masculine lexical defaults in Italian training data, so part of the observed association could be morphological; rerunning the protocol with feminized agreement (e.g., 'la preside', 'la chef') would test this directly.","Extending the same battery to a language with little grammatical gender (e.g., English) or to Italian with gender-neutral endings (schwa) would reveal whether the near-deterministic pattern is driven by Italian morphology or by occupational stereotypes encoded in the training corpus.","Because the paper reports that the position of the job titles inside the prompt changes responses, the 'leadership equals male' effect may include a syntactic priming component; analysing the data separately by permutation, which the paper's repository makes possible, could estimate how much of the bias is order-driven."],"forward_implications":["In Italian-language uses of these chatbots (hiring drafts, career advice, reference letters), prompts that do not specify gender will tend to yield text placing men in leadership and women in support roles.","The bias is not an artifact of one vendor: two independently built proprietary systems show the same directional pattern, pointing to a shared root in training data or task formulation rather than a single company's fine-tuning.","Because the pattern appears under zero-shot, no-context prompting, users would need explicit counter-stereotypical context or mitigation to avoid it; adding neutral context alone may not suffice.","The conditional-probability protocol is a reproducible black-box audit that can be re-run on other models, languages, and job pairs without internal model access."],"supporting_citations":[{"why":"Supplies the WinoBias coreference benchmark and the co-reference prompt design that this experiment adapts to Italian.","marker":"[22]"},{"why":"Demonstrates that ungendered prompts still produce stereotypical role assignment in LLMs and gives the methodological framework for probing this bias.","marker":"[5]"},{"why":"Provides the Italian-language baseline on gender bias in LLM outputs, showing that gendered job titles yield asymmetrical responses and motivating the Italian case.","marker":"[18]"},{"why":"Establishes the foundational finding that word embeddings encode gendered occupational associations, which this paper extends to generative models.","marker":"[1]"},{"why":"Shows GPT-4 exhibits strong gender-occupation associations in a cover-letter generation task, one of the real-world contexts the paper's implications address.","marker":"[12]"},{"why":"Documents gender bias in LLM-generated reference letters, another downstream hiring-related use the paper points to.","marker":"[21]"}],"fun_headline_variants":["Italian AI tests: managers get 'he', assistants get 'she' nearly always","Both chatbots assign Italian leadership roles male pronouns in neutral prompts","Gender-neutral Italian prompts lead LLMs to stereotype job roles","In Italian, AI ties 'she' to assistants: 97% ChatGPT, 100% Gemini","LLMs default to male bosses in Italian even when prompts are gender-neutral"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The three job titles are genuinely gender-neutral in Italian as the prompts use them, so a bias-free model would have no statistical reason to attach 'lui' to the leader and 'lei' to the subordinate.","fun_headline_variants_meta":{"raw":{"variants":["Italian AI tests: managers get 'he', assistants get 'she' nearly always","Both chatbots assign Italian leadership roles male pronouns in neutral prompts","Gender-neutral Italian prompts lead LLMs to stereotype job roles","In Italian, AI ties 'she' to assistants: 97% ChatGPT, 100% Gemini","LLMs default to male bosses in Italian even when prompts are gender-neutral"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000748,"raw_usage":{"total_tokens":3388,"prompt_tokens":1055,"completion_tokens":2333,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":2235}},"tokens_in":671,"tokens_out":2333,"duration_ms":17766,"temperature":1.0,"reasoning_tokens":2235,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:59:19.847294+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the exact experiment with the leadership titles spelled with explicit feminine agreement in the prompt (e.g., 'la preside', 'la manager', 'la chef') while keeping the subordinate titles unchanged; if the models then attach 'lei' to the leadership roles at rates comparable to the original 'lui' rates, the leader-male association is substantially a product of Italian grammatical gender defaults rather than social stereotyping.","supporting_citations":[{"cited_title":"In: CLiC-it 2023: 9th Italian Conference on Computational Linguistics (2023)","cited_arxiv_id":null,"evidence_quote":"Provides the Italian-language baseline on gender bias in LLM outputs, showing that gendered job titles yield asymmetrical responses and motivating the Italian case."},{"cited_title":"In:AdvancesinNeuralInformationProcessingSystems.vol.29.CurranAssociates, Inc","cited_arxiv_id":null,"evidence_quote":"Establishes the foundational finding that word embeddings encode gendered occupational associations, which this paper extends to generative models."},{"cited_title":"In: ICML 2024 Next Generation of AI Safety Workshop (Jul 2024),https://openreview","cited_arxiv_id":null,"evidence_quote":"Shows GPT-4 exhibits strong gender-occupation associations in a cover-letter generation task, one of the real-world contexts the paper's implications address."}],"review_version":2}