REVIEW 4 major objections 5 minor 43 references
The paper claims that text-to-image models resolve ambiguous two-person scenarios in stereotyped ways more often when asked to produce a storyboard or comic than a single photo, and that narrative formats expose bias modes—event sequencing,
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Across six text-to-image models, stereotyped outputs rise from 25.9% of single photos to about 36% of storyboards and 44% of four-panel comics, with bias expressed through plot, character placement, and dialogue.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection Narrative formats surface T2I bias that photos leave ambiguous, but the paper's 'comics amplify stereotypes' claim doesn't survive a conditional read of its own Table 1. the 4 major comments →
Investigating Social Bias in Narrative Image Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Using 140 ambiguous two-person contexts drawn from the BBG benchmark, the authors generate photos, storyboards, and four-panel comics with six text-to-image models and have annotators label each output as biased, counter-biased, or neutral. Across four proprietary models, the biased-output rate rises from 25.9% in photo generation to 35.5% in storyboards and 44.1% in comics. Qualitatively, biases that are hard to read from appearance alone—such as mental illness or religious identity—become explicit in narratives through speech bubbles, narration, and panels that show who struggles and who succeeds. The paper also documents layered bias: beyond assigning a role, models give different explana
What carries the argument
The central machinery is BBG, a social-bias evaluation framework originally built for text generation, adapted here to image generation. BBG supplies 140 deliberately ambiguous contexts, each describing two people with different attributes and leaving an outcome unspecified; for each context a question asks which person the outcome applies to, with a stereotypical and a counter-stereotypical answer option. When a text-to-image model is prompted with such a context, the missing attribute must be filled by the model, so the assignment of the role exposes its stereotypic priors. The authors add three format-specific prompts—photo, storyboard, and comic—each with a note making the mention order
Load-bearing premise
The load-bearing premise is that a 'biased' label means the same thing in a photo, storyboard, and comic; the paper's own Table 1 shows that 'neutral/cannot determine' labels collapse from roughly 55-70% of photo outputs to 12-30% of comic outputs, so part of the rise in biased labels may reflect annotators simply having more panels and text to base a judgment on rather than stronger stereotyping by the model.
What would settle it
Compute the conditional bias rate, bias/(bias + counter-biased), only on outputs where annotators chose a definite answer, for photo versus storyboard versus comic. If this conditional rate stays flat while the raw bias rate rises, the paper's central claim—that narrative formats amplify stereotypes—is not supported; the increase would instead be an artifact of fewer 'cannot determine' labels. The relevant numbers are already in Table 1 and can be checked without new experiments.
If this is right
- Bias evaluation for text-to-image systems should include sequential, text-bearing formats such as storyboards and comics, not just single photos.
- Narrative formats reveal bias modes—event sequencing, character positioning, narrative resolution, and embedded text—that photo-only tests cannot detect.
- Stereotypical visual associations such as clothing, symbols, and facial expressions persist across formats, so mitigating them requires more than changing the output format.
- Models sometimes add moralizing endings or verbal denials of stereotypes, but these surface-level avoidance strategies leave underlying stereotypical visual associations intact.
- Bias patterns in text generation do not directly predict bias in image generation, so multimodal narrative evaluation adds information that neither text-only nor photo-only evaluation provides.
Where Pith is reading between the lines
- My inference: if this holds, deployed text-to-image services—which increasingly offer storyboard and comic modes—should be audited in those modes before being certified as fair, not only in photo mode.
- My inference: the Korean-language findings point to a separable cultural-capability bias: a model can fail by rendering Western settings or mixed scripts even when it does not stereotype the characters, suggesting fairness metrics should track cultural grounding separately from role assignment.
- My inference: the 18.2 percentage-point comic increase may partly be an information-density effect—annotators have more panels and text to base a judgment on—so a controlled test that blanks out speech bubbles in comic outputs could separate the contribution of embedded text from that of visual narrative structure.
- My inference: automatic bias judges for narrative images will need explicit reasoning about sequence and text; the paper's own automatic judge reached only 69.5% accuracy on proprietary models, so scaling this evaluation will require new narrative-aware classifiers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends the BBG text-generation bias benchmark to text-to-image generation, comparing photo, storyboard, and four-panel comic outputs from six T2I models using 140 English/Korean prompts across seven bias categories. The headline claim is that proprietary models produce a higher proportion of stereotypically biased outputs in narrative formats than in photo generation, with comics showing the largest increase, and that qualitative analysis reveals format-specific bias modes such as event sequencing, character positioning, narrative resolution, and embedded text. The paper also reports a video-generation case study and automatic-evaluation results. The central quantitative claim rests on the marginal 'biased output' ratio in Table 1, which the authors interpret as comparable across formats.
Significance. If the central claim were established, this would be a useful contribution: it would show that single-image T2I bias evaluation misses bias that emerges in sequential, text-bearing visual narratives, and it provides a concrete multilingual, multi-model measurement. The paper's strengths include the relatively large prompt set, manual annotation with a reported high-agreement pilot, qualitative thematic coding that surfaces concrete bias mechanisms, and an exploratory video extension. However, the main quantitative conclusion is currently not supported by the data as analyzed, because the marginal bias ratio is not comparable across formats: the neutral/'cannot be determined' category collapses dramatically from photos to comics, and the conditional bias ratio among answerable outputs falls in every proprietary model-language cell. The paper also contains internal numerical inconsistencies between the abstract and Table 1. These issues are load-bearing for the paper's contribution and require substantive revision rather than small edits.
major comments (4)
- [§4.1, Table 1] The headline claim that 'comics amplify stereotypical interpretations' depends on treating the marginal bias ratio as comparable across formats. Table 1 shows this is not safe. In every proprietary model-language cell, the neutral/'cannot determine' category collapses sharply from photo to comic (e.g., GPT-Image-2 KO: 0.7031 → 0.1471; Nano Banana 2 EN: 0.7246 → 0.1884). At the same time, the conditional ratio bias/(bias+counter-biased) decreases from photo to comic in all 8 cells: for example, Nano Banana Pro EN 0.739 → 0.509, GPT-Image-2 KO 0.947 → 0.569, Nano Banana 2 KO 0.800 → 0.579. This is the signature of a measurement-artifact alternative: comics force the model to resolve the ambiguous BBG context, so annotators can now assign a label, but among resolvable outputs the model is not more likely to choose the stereotypical option. Please report the conditional ratios, paired prompt
- [Abstract and §1 vs. Table 1] The abstract states that proprietary models generate 25.9% biased outputs in photo generation, with increases of 9.6pp in storyboard and 18.2pp in comic generation. Averaging the eight proprietary values in Table 1 gives 26.4% for photo, 36.9% for storyboard, and 44.3% for comic, i.e., increases of +10.5pp and +17.9pp. The introduction's '35.2% biased outputs on average across settings' also does not match the table-derived value of 35.9%. Please correct the reported numbers or explain the exact aggregation rule; as written, the central numerical claims are internally inconsistent.
- [§3.2, Appendix A] The three generation settings differ in more than visual format: the storyboard prompt explicitly instructs 'No text', while the comic prompt naturally permits speech bubbles and narration. This is a confound for the comparisons in Table 1 and for the qualitative claim that comics expose bias 'through textual elements.' The higher answerability of comic outputs could be driven by the presence of text rather than by sequential narrative structure per se. Please address this directly---e.g., by ablating the text constraint in storyboards, by analyzing text-free comic panels separately, or by explicitly modeling format/instruction as a factor. At minimum, the interpretation in §4.1 and §4.2.1 must be tempered to acknowledge this confound.
- [§3.3, Table 1] No statistical inference is reported. Table 1 gives only fractions, without denominators, confidence intervals, or paired tests across the 140 prompts. The claim that 'narrative image generation produces more biased outputs than photo generation across all proprietary models' is based on point estimates and may not survive prompt-level paired comparison, especially given the conditional-ratio pattern above. Please provide per-prompt paired analyses (e.g., McNemar-type or mixed-effects model) with effect sizes and CIs, and state how the 'cannot be determined' labels are treated in the denominator.
minor comments (5)
- [Appendix A.5] Typo: 'Fo the first' should be 'For the first'. Also 'Both a optometrist' should be 'Both an optometrist.'
- [§3.3] The single-annotator design is justified by a pilot κ=0.9804 on 16% of the data, but the main annotations are still single-annotator. Given that the central measure is an annotation judgment, please report per-category agreement or at least state clearly that the pilot covered all seven bias categories and both languages.
- [Table 1] The table reports only 'bias' and 'neutral' ratios; 'counter-biased' is implicit. Showing all three categories would make the conditional analysis transparent and would help readers see the cannot-determine collapse directly.
- [§4.1 and Appendix D] The open-model results are described as 'near-zero bias score' due to weak instruction following. This conflation of model failure with low bias should be explicitly framed as an inability-to-measure result, not evidence of fairness; the current text mostly does this, but the phrase 'near-zero bias score' in §4.1 could be misread.
- [General] The paper does not state whether prompts, generated images, and annotation labels will be released. For a measurement study of this kind, releasing the dataset and annotations would substantially strengthen reproducibility.
Circularity Check
No significant circularity: the cross-format bias ratios are newly measured from external model outputs and human annotations; the BBG self-citation supplies prompts/definitions but does not entail the conclusion.
full rationale
This paper is a measurement study rather than a derivation. The central claim—that storyboard and comic generation yield higher biased-output ratios than photo generation—comes from human annotation of images produced by six external T2I APIs; no equation is fitted and no parameter is calibrated to produce the reported numbers. The only author-overlapping citation is the BBG framework (Jin et al., 2025; Jiho Jin is a co-author of both works). BBG is used as an external measurement instrument: it provides ambiguous-context prompts and the stereotypical/counter-stereotypical QA definitions. It does not assert, prove, or imply the observed photo-to-storyboard-to-comic ordering, so this self-citation is not load-bearing in a circular way. No uniqueness theorem is invoked, no ansatz is smuggled in through citation, and the paper does not rename a known result as a new derivation. The skeptic's concern that the marginal bias increase may be driven by a collapse of the neutral/cannot-determine category (visible in Table 1) is a question of metric comparability and statistical validity, not circularity: the measured bias ratios are not equivalent to their inputs by construction. Lack of significance testing is a correctness/evidence limitation, not a circular-derivation step. Therefore no circular step is identified.
Axiom & Free-Parameter Ledger
free parameters (2)
- Pooling of 'cannot be determined' into neutral
- Per-format prompt differences (storyboard forbids text)
axioms (3)
- domain assumption BBG's stereotypical versus counter-stereotypical interpretations remain valid when the same prompts are rendered as images.
- domain assumption A single non-blind author annotation per image reliably measures bias in multi-panel images.
- ad hoc to paper The 'biased output ratio' is comparable across formats that differ in answerability.
Cite this review
Pith. "Pith review of Investigating Social Bias in Narrative Image Generation." pith.science (2026). https://pith.science/paper/6ZNKMKAF
@misc{pith2026260801780,
author = {Pith},
title = {Pith review of: Investigating Social Bias in Narrative Image Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/6ZNKMKAF}},
note = {Machine review of arXiv:2608.01780}
}
read the original abstract
Text-to-image (T2I) generation models are increasingly embedded in applications such as media content creation and education, raising concerns about how their outputs may reproduce social biases. Prior work has shown that T2I models exhibit social biases, yet existing evaluations largely focus on a photo generation task. As a result, it remains unclear whether and how such biases manifest in more narrative visual formats, such as storyboards and comics, where characters and events are presented across multiple panels. In this work, we compare bias expression across photo, storyboard, and comic generation in six T2I models by adapting BBG, a text-based bias evaluation framework, to image generation. Our results show that proprietary models generate 25.9% biased outputs in photo generation on average, with biased outputs increasing by 9.6pp in storyboard generation and 18.2pp in comic generation. We also find that photos mainly encode biases through subtle visual cues, while storyboards and comics reveal them more explicitly through event sequencing, character positioning, narrative resolution, and textual elements. These findings show that biases that remain less visible in photo generation may surface in narrative visual formats, highlighting the importance of evaluating T2I systems with diverse visual formats beyond photo generation.
Figures
Reference graph
Works this paper leans on
-
[5]
URL https://huggingface.co/black-forest-labs/ FLUX.1-dev. Accessed: 2026-06-25. Stephen Brade, Bryan Wang, Mauricio Sousa, Sageev Oore, and Tovi Grossman. Promptify: Text-to-image generation through interactive prompt exploration with large language models. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST ’23,...
work page 2026
-
[6]
Association for Computing Machinery. ISBN 9798400701320. doi: 10.1145/3586183.3606725. URL https://doi.org/10.1145/ 3586183.3606725. Virginia Braun and Victoria Clarke. Using thematic analysis in psychology.Qualitative Research in Psychology, 3(2):77–101,
-
[8]
Yi-Chun Chen and Arnav Jhala. Collaborative comic generation: Integrating visual narrative theories with ai models for enhanced creativity.arXiv preprint arXiv:2409.17263,
-
[9]
Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models
4https://chatgpt.com/; https://cursor.com/; https://asta.allen.ai/ 10 Published at the GenAI4World workshop at COLM 2026 Jaemin Cho, Abhay Zala, and Mohit Bansal. Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. InProceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV), pp. 3043–3054, October
work page 2026
- [10]
-
[12]
Edwards, Brandon Man, and Faez Ahmed
Kristen M. Edwards, Brandon Man, and Faez Ahmed. Sketch2prototype: rapid conceptual design exploration and prototyping with generative ai.Proceedings of the Design Society, 4: 1989–1998,
work page 1989
-
[13]
Will Eisner.Comics and sequential art: Principles and practices from the legendary cartoonist
doi: 10.1017/pds.2024.201. Will Eisner.Comics and sequential art: Principles and practices from the legendary cartoonist. WW Norton & Company,
-
[14]
Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long.966. URLhttps://aclanthology.org/2025.acl-long.966/. Silin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler, Mengjie Zhao, Hiromi Wakaki, Yuki Mitsufuji, Syrielle Montariol, and Antoine Bosselut. Vinabench: Benchmark for faith- ful and consistent visual narratives...
-
[15]
Sanjana Gautam, Pranav Narayanan Venkit, and Sourojit Ghosh
doi: 10.1109/ CVPR52729.2023.00672. Sanjana Gautam, Pranav Narayanan Venkit, and Sourojit Ghosh. From melting pots to misrepresentations: Exploring harms in generative ai,
arXiv 2023
-
[16]
URL https://arxiv.org/ abs/2403.10776. Google Deepmind. Gemini 3 pro image model card,
-
[17]
googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Image-Model-Card.pdf
URL https://storage. googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Image-Model-Card.pdf . Accessed: 2026-06-25. Google Deepmind. Gemini 3.5 Flash Model Card, 2026a. URLhttps://storage.googleapis. com/deepmind-media/Model-Cards/Gemini-3-5-Flash-Model-Card.pdf . Accessed: 2026- 06-25. Google Deepmind. Veo 3.1 Lite, 2026b. URL https://storage.google...
work page 2026
-
[18]
Association for Computing Machinery. ISBN 9781450394215. doi: 10.1145/3544548.3580744. URL https://doi.org/ 10.1145/3544548.3580744. 11 Published at the GenAI4World workshop at COLM 2026 John Hart.The Art of the Storyboard: A filmmaker’s introduction. Routledge,
arXiv 2026
-
[19]
doi: 10.18653/v1/2024.acl-long.667
Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.667. URL https://aclanthology.org/2024.acl-long.667/. Jiho Jin, Woosung Kang, Junho Myung, and Alice Oh. Social bias benchmark for generation: A comparison of generation and QA-based evaluations. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.),F...
-
[20]
Generating coherent comic with rich story using ChatGPT and Stable Diffusion
Association for Computational Linguistics. ISBN 979-8-89176-256-5. doi: 10.18653/v1/ 2025.findings-acl.585. URLhttps://aclanthology.org/2025.findings-acl.585/. Ze Jin and Zorina Song. Generating coherent comic with rich story using chatgpt and stable diffusion.arXiv preprint arXiv:2305.11067,
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[21]
Association for Computing Machinery. ISBN 9798400701061. doi: 10.1145/3581641.3584078. URLhttps://doi.org/10.1145/3581641.3584078. Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Teufel, Marco Bellagente, Min- guk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Fei-Fei Li, Jiajun W...
-
[22]
URL https://proceedings.neurips.cc/paper files/paper/2023/ file/dd83eada2c3c74db3c7fe1c087513756-Paper-Datasets and Benchmarks.pdf. Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, and Jianfeng Gao. Storygan: A sequential conditional gan for story visualization. InProceedings of the IEEE/CVF conference on ...
work page 2023
-
[23]
URLhttps://doi.org/10.3390/info16050341
doi: 10.3390/info16050341. URLhttps://doi.org/10.3390/info16050341. Adyasha Maharana, Darryl Hannan, and Mohit Bansal. Storydall-e: Adapting pretrained text-to-image transformers for story continuation. InEuropean conference on computer vision, pp. 70–87. Springer,
-
[25]
Association for Computing Machinery. ISBN 9798400702310. doi: 10.1145/3600211.3604711. URL https://doi.org/10.1145/3600211. 3604711. OpenAI. The new chatgpt images is here,
-
[26]
URL https://openai.com/index/ new-chatgpt-images-is-here/. Accessed: 2026-06-25. OpenAI. Sora 2 System Card,
work page 2026
-
[27]
URL https://deploymentsafety.openai.com/sora-2. Accessed: 2026-06-25. 12 Published at the GenAI4World workshop at COLM 2026 OpenAI. Introducing chatgpt images 2.0,
work page 2026
-
[29]
URL https://deploymentsafety.openai.com/ gpt-5-5/gpt-5-5.pdf. Accessed: 2026-06-25. Ville Paananen, Jonas Oppenlaender, and Aku Visuri. Using text-to-image generation for architectural design ideation.International Journal of Architectural Computing, 22 (3):458–474,
work page 2026
-
[30]
URL https://doi.org/10.1177/ 14780771231222783
doi: 10.1177/14780771231222783. URL https://doi.org/10.1177/ 14780771231222783. Xichen Pan, Pengda Qin, Yuhong Li, Hui Xue, and Wenhu Chen. Synthesizing coherent story with auto-regressive latent diffusion models. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 2920–2930,
-
[31]
Sdxl: Improving latent diffusion models for high- resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high- resolution image synthesis. In B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun (eds.),International Conference on Learning Representations, volume 2024, pp. 1862–1874,
work page 2024
-
[32]
Tanzila Rahman, Hsin-Ying Lee, Jian Ren, Sergey Tulyakov, Shweta Mahajan, and Leonid Si- gal
URL https://proceedings.iclr.cc/paper files/paper/2024/file/ 081b08068e4733ae3e7ad019fe8d172f-Paper-Conference.pdf. Tanzila Rahman, Hsin-Ying Lee, Jian Ren, Sergey Tulyakov, Shweta Mahajan, and Leonid Si- gal. Make-a-story: Visual memory conditioned consistent story generation. InProceedings of the IEEE/CVF conference on computer vision and pattern recogn...
work page 2024
-
[33]
Huichan Seo, Sieun Choi, Minki Hong, Yi Zhou, Junseo Kim, Lukman Ismaila, Naome Etori, Mehul Agarwal, Zhixuan Liu, Jihie Kim, et al. Exposing blindspots: Cultural bias evaluation in generative image models.arXiv preprint arXiv:2510.20042,
-
[34]
The bias amplification paradox in text-to- image generation
Preethi Seshadri, Sameer Singh, and Yanai Elazar. The bias amplification paradox in text-to- image generation. In Kevin Duh, Helena Gomez, and Steven Bethard (eds.),Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 6367–6384, Mexico Ci...
work page 2024
-
[35]
doi: 10.18653/v1/ 2024.naacl-long.353
Association for Computational Linguistics. doi: 10.18653/v1/ 2024.naacl-long.353. URLhttps://aclanthology.org/2024.naacl-long.353/. Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk. Sociotechnical harms of algorithmic systems: Scopi...
doi:10.18653/v1/ 2024
-
[36]
Association for Comput- ing Machinery. ISBN 9798400702310. doi: 10.1145/3600211.3604673. URL https: //doi.org/10.1145/3600211.3604673. Xudong Shen, Chao Du, Tianyu Pang, Min Lin, Yongkang Wong, and Mohan Kankanhalli. Finetuning text-to-image diffusion models for fairness. In B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun (eds.),Internat...
arXiv 2024
-
[37]
URL https://proceedings.iclr.cc/paper files/paper/2024/file/6d0bf1265ea9635fb4f9d56f16d7efb2-Paper-Conference.pdf. 13 Published at the GenAI4World workshop at COLM 2026 Mattea Sim, Natalie Grace Brigham, Tadayoshi Kohno, Tessa E. S. Charlesworth, and Aylin Caliskan. Biased ai outputs can impact humans’ implicit bias: A case study of the impact of gender-b...
work page 2024
-
[38]
URL https://ojs.aaai.org/index.php/AIES/article/view/36723
doi: 10.1609/aies.v8i3.36723. URL https://ojs.aaai.org/index.php/AIES/article/view/36723. Yixin Wan, Arjun Subramonian, Anaelia Ovalle, Zongyu Lin, Ashima Suvarna, Christina Chance, Hritik Bansal, Rebecca Pattichis, and Kai-Wei Chang. Survey of bias in text-to- image generation: Definition, evaluation, and mitigation,
-
[39]
Jialu Wang, Xinyue Liu, Zonglin Di, Yang Liu, and Xin Wang
URL https://arxiv.org/ abs/2404.01030. Jialu Wang, Xinyue Liu, Zonglin Di, Yang Liu, and Xin Wang. T2IAT: Measuring valence and stereotypical biases in text-to-image generation. In Anna Rogers, Jordan Boyd- Graber, and Naoaki Okazaki (eds.),Findings of the Association for Computational Linguistics: ACL 2023, pp. 2560–2574, Toronto, Canada, July
Pith/arXiv arXiv 2023
-
[40]
doi: 10.18653/v1/2023.findings-acl.160
Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.160. URL https://aclanthology.org/ 2023.findings-acl.160/. Yiping Wang, Xuehai He, Kuan Wang, Luyao Ma, Jianwei Yang, Shuohang Wang, Si- mon Shaolei Du, and Yelong Shen. Is your world simulator a good story presenter? a consecutive events-based benchmark for future long video ge...
-
[41]
URL https://proceedings.neurips.cc/paper files/paper/2023/ file/91f18a1287b398d378ef22505bf41832-Paper-Datasets and Benchmarks.pdf. Junchen Zhu, Huan Yang, Huiguo He, Wenjing Wang, Zixi Tuo, Wen-Huang Cheng, Lianli Gao, Jingkuan Song, and Jianlong Fu. Moviefactory: Automatic movie creation from text using large generative models for language and images. I...
work page 2023
-
[42]
Association for Computing Machinery. ISBN 9798400701085. doi: 10.1145/3581783. 3612707. URLhttps://doi.org/10.1145/3581783.3612707. Cailin Zhuang, Ailin Huang, Yaoqi Hu, Jingwei Wu, Wei Cheng, Jiaqi Liao, Hongyuan Wang, Xinyao Liao, Weiwei Cai, Hengyuan Xu, et al. Vistorybench: Comprehensive benchmark suite for story visualization. InProceedings of the IE...
-
[43]
14 Published at the GenAI4World workshop at COLM 2026 A Experiment Prompts A.1 Photo Generation The following prompts are used for realistic photo generation tasks. In each experiment, we include an explicit note within the prompt to mitigate the potential positional bias. Image Generation Prompt (EN) Generate a realistic photo of these people. Context:{C...
work page 2026
-
[44]
with Gemini-3.5-Flash (Google Deepmind, 2026a). Compared against human annotations, auto- matic bias annotation achieves an average accuracy of 69.5% on proprietary models and 17 Published at the GenAI4World workshop at COLM 2026 Table 4: Qualitative bias evaluation for video generation models. Context PromptVeo-3.1 Sora-Pro-2 Creativity Prompt 1 Biased B...
work page 2026
-
[1993]
Social biases through the text-to-image generation lens
Ranjita Naik and Besmira Nushi. Social biases through the text-to-image generation lens. InProceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’23, pp. 786–808, New York, NY, USA,
work page 2023
-
[2006]
URL https: //doi.org/10.1191/1478088706qp063oa
doi: 10.1191/1478088706qp063oa. URL https: //doi.org/10.1191/1478088706qp063oa. Emanuele Bugliarello, H Hernan Moraldo, Ruben Villegas, Mohammad Babaeizadeh, Mo- hammad Taghi Saffar, Han Zhang, Dumitru Erhan, Vittorio Ferrari, Pieter-Jan Kinder- mans, and Paul Voigtlaender. Storybench: A multifaceted benchmark for continuous story visualization.Advances i...
-
[2017]
David Dinkevich, Matan Levy, Omri Avrahami, Dvir Samuel, and Dani Lischinski. Story2board: a training-free approach for expressive storyboard generation.arXiv preprint arXiv:2508.09983,
-
[2023]
Association for Computing Machinery. ISBN 9798400701924. doi: 10.1145/3593013.3594095. URL https://doi.org/10.1145/ 3593013.3594095. Black Forest Labs. FLUX.1 [dev],
-
[2024]
URL https://ojs.aaai.org/ index.php/AAAI/article/view/30373
doi: 10.1609/aaai.v38i21.30373. URL https://ojs.aaai.org/ index.php/AAAI/article/view/30373. Ayan Banerjee, Josep Llad´os, Umapada Pal, and Anjan Dutta. Talediffusion: Multi-character story generation with dialogue rendering.arXiv preprint arXiv:2509.04123,
-
[2025]
Hritik Bansal, Da Yin, Masoud Monajatipoor, and Kai-Wei Chang. How well can text-to- image generative models understand ethical natural language interventions? In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (eds.),Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 1358–1370, Abu Dhabi, United Arab Emirates, December
work page 2022
-
[2026]
URL https://openai.com/index/ introducing-chatgpt-images-2-0/. Accessed: 2026-06-25. OpenAI. GPT-5.5 System Card,
work page 2026
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.