Pith. sign in

REVIEW 4 major objections 5 minor 43 references

The paper claims that text-to-image models resolve ambiguous two-person scenarios in stereotyped ways more often when asked to produce a storyboard or comic than a single photo, and that narrative formats expose bias modes—event sequencing,

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Across six text-to-image models, stereotyped outputs rise from 25.9% of single photos to about 36% of storyboards and 44% of four-panel comics, with bias expressed through plot, character placement, and dialogue.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection Narrative formats surface T2I bias that photos leave ambiguous, but the paper's 'comics amplify stereotypes' claim doesn't survive a conditional read of its own Table 1. the 4 major comments →

arxiv 2608.01780 v1 pith:6ZNKMKAF submitted 2026-08-03 cs.CV cs.AI

Investigating Social Bias in Narrative Image Generation

classification cs.CV cs.AI
keywords text-to-image generationsocial biasstereotypesnarrative image generationstoryboard generationcomic generationmultilingual bias evaluationBBG
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper compares how six text-to-image models express social bias in three output formats: a single photorealistic image, a storyboard, and a four-panel comic. It finds that for four proprietary models, 25.9% of photo outputs were judged biased on average, while storyboards added 9.6 percentage points and comics added 18.2. The paper argues that photos encode stereotypes through subtle visual cues, whereas storyboards and comics make the same stereotypes more explicit through event order, character placement, story endings, and text inside panels. It also finds Korean prompts draw more biased outputs than English prompts, and that two open models look neutral mainly because they fail to follow the prompt closely enough to depict the specified people or context. If correct, the result means single-image bias evaluations are missing a large share of the bias users would encounter in real narrative-generation settings.

Core claim

Using 140 ambiguous two-person contexts drawn from the BBG benchmark, the authors generate photos, storyboards, and four-panel comics with six text-to-image models and have annotators label each output as biased, counter-biased, or neutral. Across four proprietary models, the biased-output rate rises from 25.9% in photo generation to 35.5% in storyboards and 44.1% in comics. Qualitatively, biases that are hard to read from appearance alone—such as mental illness or religious identity—become explicit in narratives through speech bubbles, narration, and panels that show who struggles and who succeeds. The paper also documents layered bias: beyond assigning a role, models give different explana

What carries the argument

The central machinery is BBG, a social-bias evaluation framework originally built for text generation, adapted here to image generation. BBG supplies 140 deliberately ambiguous contexts, each describing two people with different attributes and leaving an outcome unspecified; for each context a question asks which person the outcome applies to, with a stereotypical and a counter-stereotypical answer option. When a text-to-image model is prompted with such a context, the missing attribute must be filled by the model, so the assignment of the role exposes its stereotypic priors. The authors add three format-specific prompts—photo, storyboard, and comic—each with a note making the mention order

Load-bearing premise

The load-bearing premise is that a 'biased' label means the same thing in a photo, storyboard, and comic; the paper's own Table 1 shows that 'neutral/cannot determine' labels collapse from roughly 55-70% of photo outputs to 12-30% of comic outputs, so part of the rise in biased labels may reflect annotators simply having more panels and text to base a judgment on rather than stronger stereotyping by the model.

What would settle it

Compute the conditional bias rate, bias/(bias + counter-biased), only on outputs where annotators chose a definite answer, for photo versus storyboard versus comic. If this conditional rate stays flat while the raw bias rate rises, the paper's central claim—that narrative formats amplify stereotypes—is not supported; the increase would instead be an artifact of fewer 'cannot determine' labels. The relevant numbers are already in Table 1 and can be checked without new experiments.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Bias evaluation for text-to-image systems should include sequential, text-bearing formats such as storyboards and comics, not just single photos.
  • Narrative formats reveal bias modes—event sequencing, character positioning, narrative resolution, and embedded text—that photo-only tests cannot detect.
  • Stereotypical visual associations such as clothing, symbols, and facial expressions persist across formats, so mitigating them requires more than changing the output format.
  • Models sometimes add moralizing endings or verbal denials of stereotypes, but these surface-level avoidance strategies leave underlying stereotypical visual associations intact.
  • Bias patterns in text generation do not directly predict bias in image generation, so multimodal narrative evaluation adds information that neither text-only nor photo-only evaluation provides.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • My inference: if this holds, deployed text-to-image services—which increasingly offer storyboard and comic modes—should be audited in those modes before being certified as fair, not only in photo mode.
  • My inference: the Korean-language findings point to a separable cultural-capability bias: a model can fail by rendering Western settings or mixed scripts even when it does not stereotype the characters, suggesting fairness metrics should track cultural grounding separately from role assignment.
  • My inference: the 18.2 percentage-point comic increase may partly be an information-density effect—annotators have more panels and text to base a judgment on—so a controlled test that blanks out speech bubbles in comic outputs could separate the contribution of embedded text from that of visual narrative structure.
  • My inference: automatic bias judges for narrative images will need explicit reasoning about sequence and text; the paper's own automatic judge reached only 69.5% accuracy on proprietary models, so scaling this evaluation will require new narrative-aware classifiers.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper extends the BBG text-generation bias benchmark to text-to-image generation, comparing photo, storyboard, and four-panel comic outputs from six T2I models using 140 English/Korean prompts across seven bias categories. The headline claim is that proprietary models produce a higher proportion of stereotypically biased outputs in narrative formats than in photo generation, with comics showing the largest increase, and that qualitative analysis reveals format-specific bias modes such as event sequencing, character positioning, narrative resolution, and embedded text. The paper also reports a video-generation case study and automatic-evaluation results. The central quantitative claim rests on the marginal 'biased output' ratio in Table 1, which the authors interpret as comparable across formats.

Significance. If the central claim were established, this would be a useful contribution: it would show that single-image T2I bias evaluation misses bias that emerges in sequential, text-bearing visual narratives, and it provides a concrete multilingual, multi-model measurement. The paper's strengths include the relatively large prompt set, manual annotation with a reported high-agreement pilot, qualitative thematic coding that surfaces concrete bias mechanisms, and an exploratory video extension. However, the main quantitative conclusion is currently not supported by the data as analyzed, because the marginal bias ratio is not comparable across formats: the neutral/'cannot be determined' category collapses dramatically from photos to comics, and the conditional bias ratio among answerable outputs falls in every proprietary model-language cell. The paper also contains internal numerical inconsistencies between the abstract and Table 1. These issues are load-bearing for the paper's contribution and require substantive revision rather than small edits.

major comments (4)
  1. [§4.1, Table 1] The headline claim that 'comics amplify stereotypical interpretations' depends on treating the marginal bias ratio as comparable across formats. Table 1 shows this is not safe. In every proprietary model-language cell, the neutral/'cannot determine' category collapses sharply from photo to comic (e.g., GPT-Image-2 KO: 0.7031 → 0.1471; Nano Banana 2 EN: 0.7246 → 0.1884). At the same time, the conditional ratio bias/(bias+counter-biased) decreases from photo to comic in all 8 cells: for example, Nano Banana Pro EN 0.739 → 0.509, GPT-Image-2 KO 0.947 → 0.569, Nano Banana 2 KO 0.800 → 0.579. This is the signature of a measurement-artifact alternative: comics force the model to resolve the ambiguous BBG context, so annotators can now assign a label, but among resolvable outputs the model is not more likely to choose the stereotypical option. Please report the conditional ratios, paired prompt
  2. [Abstract and §1 vs. Table 1] The abstract states that proprietary models generate 25.9% biased outputs in photo generation, with increases of 9.6pp in storyboard and 18.2pp in comic generation. Averaging the eight proprietary values in Table 1 gives 26.4% for photo, 36.9% for storyboard, and 44.3% for comic, i.e., increases of +10.5pp and +17.9pp. The introduction's '35.2% biased outputs on average across settings' also does not match the table-derived value of 35.9%. Please correct the reported numbers or explain the exact aggregation rule; as written, the central numerical claims are internally inconsistent.
  3. [§3.2, Appendix A] The three generation settings differ in more than visual format: the storyboard prompt explicitly instructs 'No text', while the comic prompt naturally permits speech bubbles and narration. This is a confound for the comparisons in Table 1 and for the qualitative claim that comics expose bias 'through textual elements.' The higher answerability of comic outputs could be driven by the presence of text rather than by sequential narrative structure per se. Please address this directly---e.g., by ablating the text constraint in storyboards, by analyzing text-free comic panels separately, or by explicitly modeling format/instruction as a factor. At minimum, the interpretation in §4.1 and §4.2.1 must be tempered to acknowledge this confound.
  4. [§3.3, Table 1] No statistical inference is reported. Table 1 gives only fractions, without denominators, confidence intervals, or paired tests across the 140 prompts. The claim that 'narrative image generation produces more biased outputs than photo generation across all proprietary models' is based on point estimates and may not survive prompt-level paired comparison, especially given the conditional-ratio pattern above. Please provide per-prompt paired analyses (e.g., McNemar-type or mixed-effects model) with effect sizes and CIs, and state how the 'cannot be determined' labels are treated in the denominator.
minor comments (5)
  1. [Appendix A.5] Typo: 'Fo the first' should be 'For the first'. Also 'Both a optometrist' should be 'Both an optometrist.'
  2. [§3.3] The single-annotator design is justified by a pilot κ=0.9804 on 16% of the data, but the main annotations are still single-annotator. Given that the central measure is an annotation judgment, please report per-category agreement or at least state clearly that the pilot covered all seven bias categories and both languages.
  3. [Table 1] The table reports only 'bias' and 'neutral' ratios; 'counter-biased' is implicit. Showing all three categories would make the conditional analysis transparent and would help readers see the cannot-determine collapse directly.
  4. [§4.1 and Appendix D] The open-model results are described as 'near-zero bias score' due to weak instruction following. This conflation of model failure with low bias should be explicitly framed as an inability-to-measure result, not evidence of fairness; the current text mostly does this, but the phrase 'near-zero bias score' in §4.1 could be misread.
  5. [General] The paper does not state whether prompts, generated images, and annotation labels will be released. For a measurement study of this kind, releasing the dataset and annotations would substantially strengthen reproducibility.

Circularity Check

0 steps flagged

No significant circularity: the cross-format bias ratios are newly measured from external model outputs and human annotations; the BBG self-citation supplies prompts/definitions but does not entail the conclusion.

full rationale

This paper is a measurement study rather than a derivation. The central claim—that storyboard and comic generation yield higher biased-output ratios than photo generation—comes from human annotation of images produced by six external T2I APIs; no equation is fitted and no parameter is calibrated to produce the reported numbers. The only author-overlapping citation is the BBG framework (Jin et al., 2025; Jiho Jin is a co-author of both works). BBG is used as an external measurement instrument: it provides ambiguous-context prompts and the stereotypical/counter-stereotypical QA definitions. It does not assert, prove, or imply the observed photo-to-storyboard-to-comic ordering, so this self-citation is not load-bearing in a circular way. No uniqueness theorem is invoked, no ansatz is smuggled in through citation, and the paper does not rename a known result as a new derivation. The skeptic's concern that the marginal bias increase may be driven by a collapse of the neutral/cannot-determine category (visible in Table 1) is a question of metric comparability and statistical validity, not circularity: the measured bias ratios are not equivalent to their inputs by construction. Lack of significance testing is a correctness/evidence limitation, not a circular-derivation step. Therefore no circular step is identified.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 0 invented entities

This is a measurement study, so no parameters are fitted to make a derivation work. The entries above capture the analyst choices that function like parameters: how indeterminate annotations are pooled, and how the storyboard prompt was engineered. The design also assumes BBG's stereotype assignments, single-annotator reliability, and format-comparability of the bias metric, the last of which is contradicted by the paper's own conditional ratios in Table 1.

free parameters (2)
  • Pooling of 'cannot be determined' into neutral
    Annotator option C ('cannot be determined') is pooled with neutral, so any format that makes outputs more answerable mechanically raises the biased fraction. This pooling choice drives a large part of the photo-to-comic gap. Section 3.3, Table 1.
  • Per-format prompt differences (storyboard forbids text)
    Storyboards are prompted with 'No text' while comics allow dialogue and narration; these engineered differences control how explicit bias can become and are not treated as covariates in the comparison. Appendix A.1-A.3.
axioms (3)
  • domain assumption BBG's stereotypical versus counter-stereotypical interpretations remain valid when the same prompts are rendered as images.
    The paper imports BBG stereotype assignments (Jin et al., 2025) without re-validating them for visual output or for 2026-era models. Section 3.
  • domain assumption A single non-blind author annotation per image reliably measures bias in multi-panel images.
    Main-set annotations are single-annotator, done by authors who know the hypothesis; the reported kappa of 0.9804 comes from a 16% pilot, and the qualitative thematic coding is likewise non-blind. Section 3.3.
  • ad hoc to paper The 'biased output ratio' is comparable across formats that differ in answerability.
    The central comparison assumes a photo's 'biased' label is commensurable with a comic's. Table 1 shows neutral judgments collapse in comics while conditional stereotype ratios fall, so this premise is the fragile one the headline rests on. Section 4.1.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Investigating Social Bias in Narrative Image Generation." pith.science (2026). https://pith.science/paper/6ZNKMKAF

@misc{pith2026260801780,
  author       = {Pith},
  title        = {Pith review of: Investigating Social Bias in Narrative Image Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6ZNKMKAF}},
  note         = {Machine review of arXiv:2608.01780}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Text-to-image (T2I) generation models are increasingly embedded in applications such as media content creation and education, raising concerns about how their outputs may reproduce social biases. Prior work has shown that T2I models exhibit social biases, yet existing evaluations largely focus on a photo generation task. As a result, it remains unclear whether and how such biases manifest in more narrative visual formats, such as storyboards and comics, where characters and events are presented across multiple panels. In this work, we compare bias expression across photo, storyboard, and comic generation in six T2I models by adapting BBG, a text-based bias evaluation framework, to image generation. Our results show that proprietary models generate 25.9% biased outputs in photo generation on average, with biased outputs increasing by 9.6pp in storyboard generation and 18.2pp in comic generation. We also find that photos mainly encode biases through subtle visual cues, while storyboards and comics reveal them more explicitly through event sequencing, character positioning, narrative resolution, and textual elements. These findings show that biases that remain less visible in photo generation may surface in narrative visual formats, highlighting the importance of evaluating T2I systems with diverse visual formats beyond photo generation.

Figures

Figures reproduced from arXiv: 2608.01780 by Euna Jang, Gahyeon Bae, Hwajung Hong, Hyunseung Lim, Jiho Jin, Junyeong Park, Soobin Kim, Sowon Min.

Figure 1
Figure 1. Figure 1: Bias evaluation in image generation. Adapting the BBG bias evaluation frame￾work (Jin et al., 2025) to image generation, we compare photo, storyboard, and comic generation to examine how models represent different individuals. focus leaves open whether conclusions drawn from photorealistic outputs generalize to narrative visual formats, which differ in how they construct meaning through character continuit… view at source ↗
Figure 2
Figure 2. Figure 2: Mental illness is expressed more explicitly in narrative generation settings. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Creative person visualization across image generation settings. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Reliance of a blind person vs. female nurse [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Incompetence of an optometrist vs. truck driver [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Attempts to avoid biased image generation [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Examples of linguistic misalignment (mixed English, Korean, and Japanese) and [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Comparison of bias generation (left) and neutral generation (right) ratio between [PITH_FULL_IMAGE:figures/full_fig_p008_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Image generation samples. (a) Flux-1-Dev: Images of Asian girls, food (b) SDXL: Japanese/Asian imagery [PITH_FULL_IMAGE:figures/full_fig_p020_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Image generation failure samples (in Korean context). [PITH_FULL_IMAGE:figures/full_fig_p020_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

43 extracted references · 26 canonical work pages · 1 internal anchor

  1. [5]

    Accessed: 2026-06-25

    URL https://huggingface.co/black-forest-labs/ FLUX.1-dev. Accessed: 2026-06-25. Stephen Brade, Bryan Wang, Mauricio Sousa, Sageev Oore, and Tovi Grossman. Promptify: Text-to-image generation through interactive prompt exploration with large language models. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST ’23,...

  2. [6]

    ISBN 9798400701320

    Association for Computing Machinery. ISBN 9798400701320. doi: 10.1145/3586183.3606725. URL https://doi.org/10.1145/ 3586183.3606725. Virginia Braun and Victoria Clarke. Using thematic analysis in psychology.Qualitative Research in Psychology, 3(2):77–101,

  3. [8]

    Collaborative comic generation: Integrating visual narrative theories with ai models for enhanced creativity.arXiv preprint arXiv:2409.17263,

    Yi-Chun Chen and Arnav Jhala. Collaborative comic generation: Integrating visual narrative theories with ai models for enhanced creativity.arXiv preprint arXiv:2409.17263,

  4. [9]

    Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models

    4https://chatgpt.com/; https://cursor.com/; https://asta.allen.ai/ 10 Published at the GenAI4World workshop at COLM 2026 Jaemin Cho, Abhay Zala, and Mohit Bansal. Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. InProceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV), pp. 3043–3054, October

  5. [10]

    Neil Cohn

    doi: 10.1109/ICCV51070.2023.00283. Neil Cohn. The architecture of visual narrative comprehension: The interaction of narrative structure and page layout in understanding comics.Frontiers in Psychology, 5:680,

  6. [12]

    Edwards, Brandon Man, and Faez Ahmed

    Kristen M. Edwards, Brandon Man, and Faez Ahmed. Sketch2prototype: rapid conceptual design exploration and prototyping with generative ai.Proceedings of the Design Society, 4: 1989–1998,

  7. [13]

    Will Eisner.Comics and sequential art: Principles and practices from the legendary cartoonist

    doi: 10.1017/pds.2024.201. Will Eisner.Comics and sequential art: Principles and practices from the legendary cartoonist. WW Norton & Company,

  8. [14]

    ISBN 979-8-89176-251-0

    Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long.966. URLhttps://aclanthology.org/2025.acl-long.966/. Silin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler, Mengjie Zhao, Hiromi Wakaki, Yuki Mitsufuji, Syrielle Montariol, and Antoine Bosselut. Vinabench: Benchmark for faith- ful and consistent visual narratives...

  9. [15]

    Sanjana Gautam, Pranav Narayanan Venkit, and Sourojit Ghosh

    doi: 10.1109/ CVPR52729.2023.00672. Sanjana Gautam, Pranav Narayanan Venkit, and Sourojit Ghosh. From melting pots to misrepresentations: Exploring harms in generative ai,

  10. [16]

    Google Deepmind

    URL https://arxiv.org/ abs/2403.10776. Google Deepmind. Gemini 3 pro image model card,

  11. [17]

    googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Image-Model-Card.pdf

    URL https://storage. googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Image-Model-Card.pdf . Accessed: 2026-06-25. Google Deepmind. Gemini 3.5 Flash Model Card, 2026a. URLhttps://storage.googleapis. com/deepmind-media/Model-Cards/Gemini-3-5-Flash-Model-Card.pdf . Accessed: 2026- 06-25. Google Deepmind. Veo 3.1 Lite, 2026b. URL https://storage.google...

  12. [18]

    ISBN 9781450394215

    Association for Computing Machinery. ISBN 9781450394215. doi: 10.1145/3544548.3580744. URL https://doi.org/ 10.1145/3544548.3580744. 11 Published at the GenAI4World workshop at COLM 2026 John Hart.The Art of the Storyboard: A filmmaker’s introduction. Routledge,

  13. [19]

    doi: 10.18653/v1/2024.acl-long.667

    Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.667. URL https://aclanthology.org/2024.acl-long.667/. Jiho Jin, Woosung Kang, Junho Myung, and Alice Oh. Social bias benchmark for generation: A comparison of generation and QA-based evaluations. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.),F...

  14. [20]

    Generating coherent comic with rich story using ChatGPT and Stable Diffusion

    Association for Computational Linguistics. ISBN 979-8-89176-256-5. doi: 10.18653/v1/ 2025.findings-acl.585. URLhttps://aclanthology.org/2025.findings-acl.585/. Ze Jin and Zorina Song. Generating coherent comic with rich story using chatgpt and stable diffusion.arXiv preprint arXiv:2305.11067,

  15. [21]

    ISBN 9798400701061

    Association for Computing Machinery. ISBN 9798400701061. doi: 10.1145/3581641.3584078. URLhttps://doi.org/10.1145/3581641.3584078. Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Teufel, Marco Bellagente, Min- guk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Fei-Fei Li, Jiajun W...

  16. [22]

    Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, and Jianfeng Gao

    URL https://proceedings.neurips.cc/paper files/paper/2023/ file/dd83eada2c3c74db3c7fe1c087513756-Paper-Datasets and Benchmarks.pdf. Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, and Jianfeng Gao. Storygan: A sequential conditional gan for story visualization. InProceedings of the IEEE/CVF conference on ...

  17. [23]

    URLhttps://doi.org/10.3390/info16050341

    doi: 10.3390/info16050341. URLhttps://doi.org/10.3390/info16050341. Adyasha Maharana, Darryl Hannan, and Mohit Bansal. Storydall-e: Adapting pretrained text-to-image transformers for story continuation. InEuropean conference on computer vision, pp. 70–87. Springer,

  18. [25]

    ISBN 9798400702310

    Association for Computing Machinery. ISBN 9798400702310. doi: 10.1145/3600211.3604711. URL https://doi.org/10.1145/3600211. 3604711. OpenAI. The new chatgpt images is here,

  19. [26]

    Accessed: 2026-06-25

    URL https://openai.com/index/ new-chatgpt-images-is-here/. Accessed: 2026-06-25. OpenAI. Sora 2 System Card,

  20. [27]

    Accessed: 2026-06-25

    URL https://deploymentsafety.openai.com/sora-2. Accessed: 2026-06-25. 12 Published at the GenAI4World workshop at COLM 2026 OpenAI. Introducing chatgpt images 2.0,

  21. [29]

    Accessed: 2026-06-25

    URL https://deploymentsafety.openai.com/ gpt-5-5/gpt-5-5.pdf. Accessed: 2026-06-25. Ville Paananen, Jonas Oppenlaender, and Aku Visuri. Using text-to-image generation for architectural design ideation.International Journal of Architectural Computing, 22 (3):458–474,

  22. [30]

    URL https://doi.org/10.1177/ 14780771231222783

    doi: 10.1177/14780771231222783. URL https://doi.org/10.1177/ 14780771231222783. Xichen Pan, Pengda Qin, Yuhong Li, Hui Xue, and Wenhu Chen. Synthesizing coherent story with auto-regressive latent diffusion models. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 2920–2930,

  23. [31]

    Sdxl: Improving latent diffusion models for high- resolution image synthesis

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high- resolution image synthesis. In B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun (eds.),International Conference on Learning Representations, volume 2024, pp. 1862–1874,

  24. [32]

    Tanzila Rahman, Hsin-Ying Lee, Jian Ren, Sergey Tulyakov, Shweta Mahajan, and Leonid Si- gal

    URL https://proceedings.iclr.cc/paper files/paper/2024/file/ 081b08068e4733ae3e7ad019fe8d172f-Paper-Conference.pdf. Tanzila Rahman, Hsin-Ying Lee, Jian Ren, Sergey Tulyakov, Shweta Mahajan, and Leonid Si- gal. Make-a-story: Visual memory conditioned consistent story generation. InProceedings of the IEEE/CVF conference on computer vision and pattern recogn...

  25. [33]

    Exposing blindspots: Cultural bias evaluation in generative image models.arXiv preprint arXiv:2510.20042,

    Huichan Seo, Sieun Choi, Minki Hong, Yi Zhou, Junseo Kim, Lukman Ismaila, Naome Etori, Mehul Agarwal, Zhixuan Liu, Jihie Kim, et al. Exposing blindspots: Cultural bias evaluation in generative image models.arXiv preprint arXiv:2510.20042,

  26. [34]

    The bias amplification paradox in text-to- image generation

    Preethi Seshadri, Sameer Singh, and Yanai Elazar. The bias amplification paradox in text-to- image generation. In Kevin Duh, Helena Gomez, and Steven Bethard (eds.),Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 6367–6384, Mexico Ci...

  27. [35]

    doi: 10.18653/v1/ 2024.naacl-long.353

    Association for Computational Linguistics. doi: 10.18653/v1/ 2024.naacl-long.353. URLhttps://aclanthology.org/2024.naacl-long.353/. Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk. Sociotechnical harms of algorithmic systems: Scopi...

  28. [36]

    ISBN 9798400702310

    Association for Comput- ing Machinery. ISBN 9798400702310. doi: 10.1145/3600211.3604673. URL https: //doi.org/10.1145/3600211.3604673. Xudong Shen, Chao Du, Tianyu Pang, Min Lin, Yongkang Wong, and Mohan Kankanhalli. Finetuning text-to-image diffusion models for fairness. In B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun (eds.),Internat...

  29. [37]

    13 Published at the GenAI4World workshop at COLM 2026 Mattea Sim, Natalie Grace Brigham, Tadayoshi Kohno, Tessa E

    URL https://proceedings.iclr.cc/paper files/paper/2024/file/6d0bf1265ea9635fb4f9d56f16d7efb2-Paper-Conference.pdf. 13 Published at the GenAI4World workshop at COLM 2026 Mattea Sim, Natalie Grace Brigham, Tadayoshi Kohno, Tessa E. S. Charlesworth, and Aylin Caliskan. Biased ai outputs can impact humans’ implicit bias: A case study of the impact of gender-b...

  30. [38]

    URL https://ojs.aaai.org/index.php/AIES/article/view/36723

    doi: 10.1609/aies.v8i3.36723. URL https://ojs.aaai.org/index.php/AIES/article/view/36723. Yixin Wan, Arjun Subramonian, Anaelia Ovalle, Zongyu Lin, Ashima Suvarna, Christina Chance, Hritik Bansal, Rebecca Pattichis, and Kai-Wei Chang. Survey of bias in text-to- image generation: Definition, evaluation, and mitigation,

  31. [39]

    Jialu Wang, Xinyue Liu, Zonglin Di, Yang Liu, and Xin Wang

    URL https://arxiv.org/ abs/2404.01030. Jialu Wang, Xinyue Liu, Zonglin Di, Yang Liu, and Xin Wang. T2IAT: Measuring valence and stereotypical biases in text-to-image generation. In Anna Rogers, Jordan Boyd- Graber, and Naoaki Okazaki (eds.),Findings of the Association for Computational Linguistics: ACL 2023, pp. 2560–2574, Toronto, Canada, July

  32. [40]

    doi: 10.18653/v1/2023.findings-acl.160

    Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.160. URL https://aclanthology.org/ 2023.findings-acl.160/. Yiping Wang, Xuehai He, Kuan Wang, Luyao Ma, Jianwei Yang, Shuohang Wang, Si- mon Shaolei Du, and Yelong Shen. Is your world simulator a good story presenter? a consecutive events-based benchmark for future long video ge...

  33. [41]

    Junchen Zhu, Huan Yang, Huiguo He, Wenjing Wang, Zixi Tuo, Wen-Huang Cheng, Lianli Gao, Jingkuan Song, and Jianlong Fu

    URL https://proceedings.neurips.cc/paper files/paper/2023/ file/91f18a1287b398d378ef22505bf41832-Paper-Datasets and Benchmarks.pdf. Junchen Zhu, Huan Yang, Huiguo He, Wenjing Wang, Zixi Tuo, Wen-Huang Cheng, Lianli Gao, Jingkuan Song, and Jianlong Fu. Moviefactory: Automatic movie creation from text using large generative models for language and images. I...

  34. [42]

    ISBN 9798400701085

    Association for Computing Machinery. ISBN 9798400701085. doi: 10.1145/3581783. 3612707. URLhttps://doi.org/10.1145/3581783.3612707. Cailin Zhuang, Ailin Huang, Yaoqi Hu, Jingwei Wu, Wei Cheng, Jiaqi Liao, Hongyuan Wang, Xinyao Liao, Weiwei Cai, Hengyuan Xu, et al. Vistorybench: Comprehensive benchmark suite for story visualization. InProceedings of the IE...

  35. [43]

    In each experiment, we include an explicit note within the prompt to mitigate the potential positional bias

    14 Published at the GenAI4World workshop at COLM 2026 A Experiment Prompts A.1 Photo Generation The following prompts are used for realistic photo generation tasks. In each experiment, we include an explicit note within the prompt to mitigate the potential positional bias. Image Generation Prompt (EN) Generate a realistic photo of these people. Context:{C...

  36. [44]

    Cannot be determined

    with Gemini-3.5-Flash (Google Deepmind, 2026a). Compared against human annotations, auto- matic bias annotation achieves an average accuracy of 69.5% on proprietary models and 17 Published at the GenAI4World workshop at COLM 2026 Table 4: Qualitative bias evaluation for video generation models. Context PromptVeo-3.1 Sora-Pro-2 Creativity Prompt 1 Biased B...

  37. [1993]

    Social biases through the text-to-image generation lens

    Ranjita Naik and Besmira Nushi. Social biases through the text-to-image generation lens. InProceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’23, pp. 786–808, New York, NY, USA,

  38. [2006]

    URL https: //doi.org/10.1191/1478088706qp063oa

    doi: 10.1191/1478088706qp063oa. URL https: //doi.org/10.1191/1478088706qp063oa. Emanuele Bugliarello, H Hernan Moraldo, Ruben Villegas, Mohammad Babaeizadeh, Mo- hammad Taghi Saffar, Han Zhang, Dumitru Erhan, Vittorio Ferrari, Pieter-Jan Kinder- mans, and Paul Voigtlaender. Storybench: A multifaceted benchmark for continuous story visualization.Advances i...

  39. [2017]

    Story2board: a training-free approach for expressive storyboard generation.arXiv preprint arXiv:2508.09983,

    David Dinkevich, Matan Levy, Omri Avrahami, Dvir Samuel, and Dani Lischinski. Story2board: a training-free approach for expressive storyboard generation.arXiv preprint arXiv:2508.09983,

  40. [2023]

    ISBN 9798400701924

    Association for Computing Machinery. ISBN 9798400701924. doi: 10.1145/3593013.3594095. URL https://doi.org/10.1145/ 3593013.3594095. Black Forest Labs. FLUX.1 [dev],

  41. [2024]

    URL https://ojs.aaai.org/ index.php/AAAI/article/view/30373

    doi: 10.1609/aaai.v38i21.30373. URL https://ojs.aaai.org/ index.php/AAAI/article/view/30373. Ayan Banerjee, Josep Llad´os, Umapada Pal, and Anjan Dutta. Talediffusion: Multi-character story generation with dialogue rendering.arXiv preprint arXiv:2509.04123,

  42. [2025]

    Hritik Bansal, Da Yin, Masoud Monajatipoor, and Kai-Wei Chang. How well can text-to- image generative models understand ethical natural language interventions? In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (eds.),Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 1358–1370, Abu Dhabi, United Arab Emirates, December

  43. [2026]

    Accessed: 2026-06-25

    URL https://openai.com/index/ introducing-chatgpt-images-2-0/. Accessed: 2026-06-25. OpenAI. GPT-5.5 System Card,

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.