{"id":"2db780cd-a184-424f-b0f7-8300249e568a","arxiv_id":"2506.23101","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Dual-character narrative prompts reveal gender biases in six multimodal LLMs that are largely invisible in single-character evaluations, and GENRES provides a structured benchmark to measure them.","lead":"This paper introduces GENRES, a benchmark that asks multimodal AI models to write profiles and stories about a man and a woman in different social relationships, then measures whether the model treats the two genders differently. The authors find biases such as warmth, emotional expression, and status that appear only when two characters interact, not when characters are described alone.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central dual-vs-single claim rests on a two-model pilot with no statistical comparison; without a full controlled single-character baseline across all six models, the headline finding is not established.","rationale":"The reader's weakest assumption focuses on possible gender bias in the generated images, which is a legitimate confound for attributing measured bias to the MLLMs. However, the more load-bearing gap is the lack of a systematic single-vs-dual comparison in the main experiments. The paper's motivating claim is that dual-character interactions surface gender bias that is 'not evident in single-character settings,' yet the only direct evidence is a small pilot with two models and no inferential statistics. If that comparison fails under a controlled test, the benchmark's central novelty is undermined regardless of image quality. The image audit would be a valuable additional check, but the controlled single-character baseline directly targets the headline claim. Since the paper is otherwise well-structured and the concern is addressable with additional experiments, the reader's CONDITIONAL verdict remains appropriate; no verdict change is needed.","tokens_in":24921,"tokens_out":4377,"duration_ms":54056,"concrete_test":"Create single-character variants of all 1,440 GENRES NEPs, one per gender, preserving the same age, scenario, and relationship description where applicable, and generate matched single-person images with the same SDXL pipeline and filtering criteria. Run all six MLLMs under identical decoding settings, compute the same M1-M8 metrics for the single-character outputs, and test the dual-vs-single difference using paired bootstrap or a mixed-effects model with scenario and model as random effects. If the dual-condition bias is not significantly larger than the single-condition bias for a majority of metrics (e.g., the 95% confidence interval of the difference excludes zero in the predicted direction), the paper's central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that MLLMs exhibit stronger gender bias in dual-character narratives than in single-character settings, so interpersonal contexts surface bias that existing benchmarks miss. The only evidence for this comparison is the preliminary study in Figure 1, which uses just two models (Qwen-3B and Janus-7B), reports no sample sizes, confidence intervals, or significance tests, and provides no description of how the single-character prompts and images were constructed to be comparable to the dual-character ones. The main GENRES evaluation (Section 4) includes only dual-character NEPs, so it cannot by itself demonstrate that interaction contexts reveal bias 'not evident in single-character settings.' This matters because the benchmark's novelty and the title's 'from individuals to interactions' framing depend on showing that the dual setting adds signal beyond what single-character evaluation would find. The image-bias concern raised by the reader is real but secondary: even if the images were perfectly gender-neutral, the dual-vs-single comparison remains unsupported without a controlled experiment. Moreover, the single-character pilot may differ in prompt length, task demands, and image content, so any measured difference could reflect task complexity rather than social relationship dynamics.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GENRES, a benchmark that measures gender bias in MLLM-generated narratives about male-female dyads. The dataset contains 1,440 narrative elicitation pairs built from Fiske's four relational models; each pair includes a text prompt and an SDXL-generated image. Models are asked to write character profiles and a roughly 500-word story, and the outputs are scored with eight metrics (M1-M8) covering profile traits, agency/role, emotional expression, and narrative framing, using both NLP tools and three LLM evaluators. The main experiments report bias patterns across six MLLMs, and the paper claims that dual-character interaction contexts surface gender biases that are not evident in single-character settings.","tokens_in":25178,"tokens_out":5341,"duration_ms":54750,"significance":"If the comparative claim were adequately supported, GENRES would make a useful contribution by shifting gender-bias evaluation from isolated entities to relationships. The benchmark is theory-grounded, the metric definitions are transparent, the pipeline is documented in detail, code and data are released, and the use of three independent LLM evaluators is a robustness strength. At present, however, the headline result rests on a small uncontrolled pilot, and the main evaluation cannot by itself demonstrate an advantage over single-character benchmarks; the significance is therefore conditional on the additional controlled evidence recommended below.","major_comments":[{"comment":"The central claim that gender bias is stronger in dual-character narratives and 'not evident in single-character settings' is supported only by the preliminary study in Figure 1, which uses two models (Qwen-3B and Janus-7B) and reports no sample size, error bars, confidence intervals, or significance tests, and no description of how the single-character prompts and images were constructed to match the dual-character ones. The main GENRES evaluation in Section 4 contains only dual-character NEPs. Please add a controlled single-character baseline across all six models, with matched text prompts and images and a statistical comparison (e.g., paired tests with effect sizes), or explicitly reframe the paper's contribution to avoid the comparative claim.","section":"§1 and Figure 1; Abstract; §5"},{"comment":"The image generation and filtering pipeline is described as avoiding gender bias through prompt design, CLIP filtering at threshold 0.25, and manual review, but the paper provides no audit of the final image set. Because the task is multimodal and all metrics are computed from MLLM responses to image-plus-text inputs, systematic gendered cues in appearance, clothing, or setting could produce the observed differences without reflecting model-level relational bias. Please provide an audit of the final images (e.g., human annotations of gender-role cues, automated attribute analysis, or a counterfactual image-swapping experiment) to support the assumption that the images do not introduce gendered artifacts.","section":"§3.2, Appendix B.2"},{"comment":"The M4 analysis applies a post-hoc two-thirds evaluator-agreement filter, but the paper does not report how many samples were excluded per model and relationship, nor how sensitive the conclusions are to the agreement threshold. In addition, Table 2 reports M1-M6 and TBS as point estimates without error bars or confidence intervals even though results are averaged over four independent evaluations. Without variance information or a sensitivity analysis, the cross-model ranking and the claim that TBS rankings are 'highly consistent' across evaluators are not fully supported.","section":"§4.2, Eq. for M4; Table 2; Figure 6"}],"minor_comments":[{"comment":"'Natrual Language Processing' should be 'Natural Language Processing'.","section":"§3.3"},{"comment":"The female name list contains 'Amelia' twice, so the curated list of 15 names actually has only 14 unique female names; please fix and report whether the duplicate affected name sampling.","section":"Table 9"},{"comment":"The left panel's y-axis label 'Warmth-related Word (%)' is unclear as to whether values are percentages or proportions; please state units and add sample sizes.","section":"Figure 1"},{"comment":"The term 'stereoscopic' is used for characters who have both positive and negative traits; please define it at first use in the main text rather than only in Appendix C.2.","section":"§3.3"},{"comment":"p-values are reported for M7 and M8 but not for M1-M6; consider reporting confidence intervals or at least significance flags for all metrics.","section":"Table 2"},{"comment":"Appendix B.1 states that names are assigned in varying order to minimize positional bias, but this detail is absent from the main-text description of text generation in §3.2; adding one sentence would help readers assess the counterbalancing.","section":"§3.2 and Appendix B.1"}],"recommendation":"major_revision","confidential_remarks":"The resource has value and the metric suite is thoughtfully designed, but the paper currently overclaims the dual-versus-single advantage. The requested controlled baseline and image audit are feasible within the scope of a revision; if the authors cannot provide them, the comparative claim should be removed from the abstract and conclusions. I would also encourage the authors to make the four evaluation runs per model available as separate results so variance can be assessed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: GENRES is a genuinely new benchmark artifact — dual-character multimodal narrative prompts grounded in Fiske's relational models, with a multi-metric evaluation suite. If you work on MLLM fairness, it's worth knowing about. But the paper's headline claim, that interpersonal contexts surface bias 'not evident in single-character settings,' is not supported by the evidence actually presented. The stress-test note is right: the dual-vs-single comparison comes from a two-model pilot (Qwen-3B, Janus-7B) with no sample sizes, no error bars, and no significance tests. The pilot also doesn't describe how the single-character prompts and images were matched to the dual ones. So the 'from individuals to interactions' framing is plausible but not established. That's the main soft spot.\n\nWhat the paper does well: the benchmark design is thoughtful. The relationship taxonomy (CS/EM/MP/AR) is applied in a way that avoids the obvious gendered role pairs (doctor/nurse), and the prompt template randomizes name order and avoids explicit role-to-gender assignment. The metrics are mostly transparent difference scores, computed from external dictionaries and LLM judges rather than fitted to the results, so there's no circularity. Using three different LLM evaluators is a reasonable robustness check, and the release of code and data helps. The relationship-specific analyses (e.g., warmth words concentrated in Communal Sharing) are the most interesting part, and the authors honestly acknowledge the binary-gender and scale limitations.\n\nSoft spots beyond the main claim: Table 2 has no error bars on M1-M6 or TBS, just an average over four runs. M4 is analyzed only after a post-hoc two-thirds agreement filter across evaluators; that's defensible as a reliability filter, but it should be reported as a sensitivity analysis, not just the main result. The images are generated with deliberate attempts to omit gender cues, but there's no formal audit of the final image set, so image-related confounds can't be fully ruled out. Those are real but secondary; the controlled single-vs-dual baseline is the load-bearing missing piece.\n\nVerdict: this deserves a serious referee — the benchmark is a useful contribution and the evaluation is mostly careful — but the paper needs a full single-character baseline across all six models, error bars, and a clearer treatment of the M4 filtering before the headline claim can stand. For a fairness-auditing reader, the relationship-specific findings are still worth citing with caution.","headline":"A solid new benchmark for interaction-level gender bias, but the 'not evident in single-character settings' claim needs a real controlled baseline.","tokens_in":25641,"tokens_out":2757,"would_cite":true,"duration_ms":29846,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Gender bias in multimodal large language models surfaces more strongly in dual-character narratives, where a man and a woman interact under a defined social relationship, than in single-character settings; the GENRES benchmark measures…","keywords":["gender bias","multimodal large language models","benchmark","social relationships","relational models theory","narrative generation","stereotype content model","bias evaluation metrics"],"falsifier":"Regenerate or re-pair the images so that gender cues are removed or swapped, for example by neutralizing the visual prompts used for image generation or by swapping which image accompanies which gendered name, and rerun the evaluation; if the warmth-word gap (metric M1) and emotion-word gap (metric M5) shrink or vanish, the bias lives in the benchmark images, while if they persist, the interaction-driven model bias is real.","tokens_in":24727,"feed_emoji":"👥","tokens_out":8640,"duration_ms":81227,"temperature":0.7,"pith_summary":"This paper argues that gender bias in multimodal large language models is best detected not by examining how a model describes a man or a woman in isolation, but by putting both characters in a scene together and letting their social relationship frame the story. To test this, the authors built GENRES, a benchmark of 1,440 image-text pairs in which the model must write individual profiles and a narrative about a mixed-gender pair bound by one of four relationship types drawn from relational models theory. Across six open- and closed-source MLLMs, the paper finds persistent, context-sensitive biases that are weaker or absent in single-character settings: warmth-related traits and emotional language gravitate to female characters, while male characters receive greater psychological depth. If the argument is right, interaction-based evaluation captures a form of gender bias that single-character benchmarks systematically miss.","feed_headline":"Pairing two characters exposes gender bias single portraits hide","feed_subtitle":"A 1,440-prompt benchmark shows vision-language models cast women as warm, emotive leads and men as deeper characters.","key_machinery":"The carrying mechanism is a dual-character profile-and-narrative elicitation task organized by relational models theory, which supplies four fundamental relationship types (Communal Sharing, Authority Ranking, Equality Matching, Market Pricing) that structure 288 scenario entries; with five images per entry this yields 1,440 narrative elicitation pairs. Each pair presents the model with a text template using randomly assigned character names and a Stable Diffusion-generated image, and asks for brief individual profiles plus a roughly 500-word story. The evaluation suite converts the outputs into eight male-minus-female metrics spanning four dimensions: profile assignment (warmth words, positive score), agency and role (subject counts, high-status allocation), emotional expression (emotion lexicon, sentence sentiment), and narrative framing (character stereoscopicity, main-character assignment), which are z-scored and summed into a single Total Bias Score for ranking models.","core_discovery":"The central discovery is that interpersonal context reliably surfaces latent gender bias: when a model narrates a story about a male and a female character engaged in a defined social relationship, its gendered choices diverge more sharply than when it describes either character alone, as shown by a preliminary comparison of warmth-related word usage and gender-stereoscopic portrayal correlations. On the GENRES benchmark, the paper reports that all evaluated models except Gemini assign more warmth-related personality traits and more emotional language to female characters; that female characters are more often cast as the main protagonist and given higher social status while male characters are portrayed as more psychologically stereoscopic; and that larger, more capable models are not necessarily fairer, with GPT-4o and Gemini showing the highest total bias scores and the smaller Qwen models the lowest. The paper also finds the bias is relationship-sensitive: warmth stereotyping is most extreme in communal-sharing contexts, and models introduce social hierarchies even in relationship types that should imply equal standing.","pith_inferences":["Beyond the paper: the dual-character effect should be directly testable within a single model by comparing matched prompts that differ only in whether the same image is framed as one character or two, a controlled contrast the preliminary study gestures toward but does not report systematically.","Beyond the paper: the SDXL-generated image set is an uncontrolled confound; an audit for systematic gendered differences in clothing, pose, or setting, or an image-swap condition, would separate model bias from benchmark artifact.","Beyond the paper: applying the same relational-models frame to same-gender and non-binary pairings (which the authors list as future work) would show whether the warmth/competence split is specific to mixed-gender pairs or tied to the relationship type itself."],"forward_implications":["Single-character benchmarks understate gender bias: a model that looks fair on isolated portraits can still stereotype in interpersonal narratives, so evaluations should include dual-character tasks to diagnose bias.","Bias is context-sensitive: the same model can favor one gender in communal settings and a different one in workplace or market settings, so aggregate scores alone can mislead.","Model capability and fairness do not move together: the largest evaluated models carry the highest total bias while the smaller ones carry the lowest, implicating training data and alignment choices rather than scale.","The recurring pattern of warm, emotionally expressive female leads alongside psychologically deeper male characters defines a specific failure mode (a surface-level elevation of female characters) that debiasing should target.","Models impose hierarchies the prompt never asked for, assigning unequal social status in relationship types designed to be egalitarian, which is itself a measurable bias.","pith_inferences"],"supporting_citations":[{"why":"Grounds the four relational models that structure every scenario and relationship type in the benchmark.","marker":"[9]"},{"why":"Provides the warmth/competence stereotype dimensions that define the first profile-assignment metric.","marker":"[10]"},{"why":"Supplies the stereotype-content dictionary used to count warmth words and compute positive scores in profiles.","marker":"[28]"},{"why":"Provides the NRC Emotion Lexicon used to measure emotion-related words attached to each character.","marker":"[25]"},{"why":"Generates the images that are paired with text prompts in every narrative elicitation pair.","marker":"[30]"},{"why":"Used to produce the visual scene descriptions for image generation and is itself one of the evaluated closed-source models.","marker":"[2]"},{"why":"Filters image-text alignment during dataset construction, alongside manual review.","marker":"[32]"}],"fun_headline_variants":["Dual-character stories unmask gender stereotypes in AI","When AI pairs characters, gender bias emerges","Social context reveals bias single portraits hide","Interactions expose gender bias in multimodal models","Benchmark shows AI genders characters in relationships"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the benchmark images themselves are free of gender bias: the pictures were generated by Stable Diffusion from GPT-4o descriptions and filtered by a CLIP similarity threshold plus manual review, but the paper never audits the final image set for systematic gendered cues, so if the photos encode gendered clothing, settings, or poses, the measured bias could be coming from the benchmark rather than from the models.","fun_headline_variants_meta":{"raw":{"variants":["Dual-character stories unmask gender stereotypes in AI","When AI pairs characters, gender bias emerges","Social context reveals bias single portraits hide","Interactions expose gender bias in multimodal models","Benchmark shows AI genders characters in relationships"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000646,"raw_usage":{"total_tokens":2957,"prompt_tokens":926,"completion_tokens":2031,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":1964}},"tokens_in":542,"tokens_out":2031,"duration_ms":15171,"temperature":1.0,"reasoning_tokens":1964,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:49:31.053412+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Regenerate or re-pair the images so that gender cues are removed or swapped, for example by neutralizing the visual prompts used for image generation or by swapping which image accompanies which gendered name, and rerun the evaluation; if the warmth-word gap (metric M1) and emotion-word gap (metric M5) shrink or vanish, the bias lives in the benchmark images, while if they persist, the interaction-driven model bias is real.","supporting_citations":[{"cited_title":"The four elementary forms of sociality: framework for a unified theory of social relations","cited_arxiv_id":null,"evidence_quote":"Grounds the four relational models that structure every scenario and relationship type in the benchmark."},{"cited_title":"A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition","cited_arxiv_id":null,"evidence_quote":"Provides the warmth/competence stereotype dimensions that define the first profile-assignment metric."},{"cited_title":"Comprehensive stereotype content dictionaries using a semi-automated method","cited_arxiv_id":null,"evidence_quote":"Supplies the stereotype-content dictionary used to count warmth words and compute positive scores in profiles."},{"cited_title":"Nrc emotion lexicon","cited_arxiv_id":null,"evidence_quote":"Provides the NRC Emotion Lexicon used to measure emotion-related words attached to each character."}],"review_version":1}