REVIEW 2 major objections 2 minor 14 references
Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities
T0 review · 2 major / 2 minor · reviewed 2026-05-22 · grok-4.3
Pith's one-line read Text-to-image models generate traditional attire for many nationalities even during everyday activities
desk verdict The 28.4% traditional attire rate and region associations are new in scale but rest on human labels with no reported reliability checks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Region- and income-level statistical associations between manually or model-labeled traditional attire and the nationality named in the prompt, together with alignment scoring on image-prompt pairs.
What would settle it
A re-labeling study that uses multiple independent raters with a written definition of traditional attire and impracticality and finds no significant regional difference would falsify the reported associations.
Extended reading notes
Core claim
When aggregating across activities and models, 28.4 percent of the images depicted individuals wearing traditional attire, including attire that is impractical for the specified activities in several cases. This pattern was statistically significantly associated with regions, with the Middle East and North Africa and Sub-Saharan Africa disproportionately affected, and was also associated with World Bank income groups. Images labeled as featuring traditional attire received statistically significantly higher alignment scores when prompts included country names, and one model frequently inserted the word traditional in its revisions.
Load-bearing premise
The labeling of images as containing traditional attire and as impractical for the activity is treated as reliable enough to support region-level statistical associations.
Editorial extensions
If this is right
- Including a country name in the prompt raises the chance that the output will show traditional rather than context-appropriate clothing.
- One model internally revises prompts by adding the word traditional at higher rates for certain nationalities.
- Alignment models themselves assign higher scores to the traditional-attire outputs when country information is present.
- The observed patterns differ by generator model and by the choice of evaluation model.
Reading between the lines
- Users requesting images of people from affected regions may receive systematically less practical or contemporary depictions.
- Removing country names from prompts or adding explicit instructions for modern clothing could reduce the frequency of these outputs.
- Similar region-linked patterns may appear in other generative models that draw on overlapping training data.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates how DALL-E 3 and Gemini 3 Pro Preview represent people from 206 nationalities in five everyday activities by generating 2,060 images. Aggregating results, it finds 28.4% depict traditional attire, with statistically significant associations to regions (disproportionately MENA and Sub-Saharan Africa) and income groups. Similar patterns for impractical attire in athletics. Alignment is scored using CLIP, ALIGN, and GPT-4.1 mini on 9,270 pairs, showing higher scores for traditional images with country names in prompts, and prompt revisions frequently insert 'traditional'.
Significance. Should the labeling methodology be validated with reliability metrics, this work would offer valuable quantitative insights into cultural biases in text-to-image generation. It highlights how model behaviors and prompt elements can lead to stereotypical depictions, contributing to the growing body of research on fairness in generative AI. The scale (206 nationalities) and use of multiple alignment models are strengths that could influence future auditing practices in the field.
major comments (2)
- [§3.2 (Image Labeling)] §3.2 (Image Labeling): The process for determining whether generated images contain 'traditional attire' or attire 'impractical for the activity' is not described in detail, including any guidelines provided to annotators, the number of annotators per image, or measures of inter-rater agreement. This is a load-bearing issue for the reported 28.4% statistic and the region associations, as subjective interpretations could vary systematically by cultural context.
- [§4.1 (Statistical Analysis)] §4.1 (Statistical Analysis): While statistical significance is reported for the region and income associations, the manuscript should specify the exact tests used (e.g., chi-square), degrees of freedom, and any multiple comparison corrections to allow full evaluation of the findings' robustness.
minor comments (2)
- [Abstract] Abstract: The abstract states 'including attire that is impractical for the specified activities in several cases' but does not quantify how many such cases; consider adding this detail for precision.
- [Prompt Revision Analysis] Prompt Revision Analysis: Clarify which model showed the 50.3% insertion rate for 'traditional' and whether this holds across all activities.
Simulated Author's Rebuttal
We thank the referee for their detailed and constructive comments, which highlight important areas for improving the clarity and robustness of our methodology. We address each major comment below and will incorporate revisions as indicated.
read point-by-point responses
-
Referee: [§3.2 (Image Labeling)] The process for determining whether generated images contain 'traditional attire' or attire 'impractical for the activity' is not described in detail, including any guidelines provided to annotators, the number of annotators per image, or measures of inter-rater agreement. This is a load-bearing issue for the reported 28.4% statistic and the region associations, as subjective interpretations could vary systematically by cultural context.
Authors: We agree that the current description in §3.2 lacks sufficient detail on the labeling protocol. In the revised manuscript we will expand this section to specify the exact guidelines provided to annotators (including definitions and examples for 'traditional attire' and 'impractical for the activity'), confirm that each image was independently labeled by three annotators, and report inter-rater agreement using an appropriate metric such as Fleiss' kappa. These additions will directly address concerns about subjectivity and cultural context. revision: yes
-
Referee: [§4.1 (Statistical Analysis)] While statistical significance is reported for the region and income associations, the manuscript should specify the exact tests used (e.g., chi-square), degrees of freedom, and any multiple comparison corrections to allow full evaluation of the findings' robustness.
Authors: We will revise §4.1 to explicitly document the statistical procedures: chi-square tests of independence were used for the region and income-group associations, with degrees of freedom and exact p-values reported; we will also state whether any multiple-comparison correction (e.g., Bonferroni) was applied. These details were present in our analysis code but omitted from the initial text and can be added without altering the reported conclusions. revision: yes
Circularity Check
No circularity; empirical pipeline uses external scorers on newly generated images
full rationale
The paper generates 2,060 new images from fixed prompts specifying nationalities and activities, then applies external models (CLIP, ALIGN, GPT-4.1 mini) to score alignment on 9,270 pairs and performs labeling for traditional attire and impracticality. No parameters are fitted inside the dataset, no predictions are made from fitted subsets, and no self-citations or uniqueness theorems are invoked as load-bearing premises. Statistical associations (region and income-group differences) are computed directly from the labeled outputs without any self-definitional reduction or renaming of known results. The analysis is self-contained against external benchmarks and contains none of the enumerated circularity patterns.
Assumptions & free parameters
assumptions (1)
- standard math Standard assumptions of statistical significance testing (independence, appropriate distribution for p-values)
Cite this review
Pith. "Pith review of Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities." pith.science (2026). https://pith.science/paper/2504.06313
@misc{pith2026250406313,
author = {Pith},
title = {Pith review of: Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities},
year = {2026},
howpublished = {\url{https://pith.science/paper/2504.06313}},
note = {Machine review of arXiv:2504.06313}
}
read the original abstract
This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview, depict people from 206 nationalities when prompted to generate images of individuals engaging in common everyday activities. Five scenarios were developed, and 2,060 images were generated using input prompts that specified nationalities across five activities. When aggregating across activities and models, results showed that 28.4% of the images depicted individuals wearing traditional attire, including attire that is impractical for the specified activities in several cases. This pattern was statistically significantly associated with regions, with the Middle East & North Africa and Sub-Saharan Africa disproportionately affected, and was also associated with World Bank income groups. Similar region- and income-linked patterns were observed for images labeled as depicting impractical attire in two athletics-related activities. To assess image-text alignment, CLIP, ALIGN, and GPT-4.1 mini were used to score 9,270 image-prompt pairs. Images labeled as featuring traditional attire received statistically significantly higher alignment scores when prompts included country names, and this pattern weakened or reversed when country names were removed. Revised prompt analysis showed that one model frequently inserted the word "traditional" (50.3% for traditional-labeled images vs. 16.6% otherwise). These results indicate that these representational patterns can be shaped by several components of the pipeline, including image generator, evaluation models, and prompt revision.
Figures
Reference graph
Works this paper leans on
-
[1]
Table 3 presents summary statistics for the dataset. For each model, 206 images were generated, resulting in 1,030 images per model across the five scenarios (i.e., 1,030 for DALL-E 3 and 1,030 for Gemini 3), and 2,060 images in total. Because each imag e is associated with multiple prompts that were evaluated for alignment scores, a total of 9,270 prompt...
-
[2]
Results 5.1 RQ1: Representations in Generated Images Based on the labeling of images generated across activities and countries, distinct patterns emerged in how the two models represented individuals from different nationalities. When aggregating results across all activities and models, 28.4% of images depicted individuals wearing traditional attire. In ...
-
[3]
a couple cooking a meal in their home,
Discussion 6.1 Implications The findings of this study demonstrate how two T2I models represent people from diverse nationalities when prompted to generate images of people engaging in common everyday activities. The results demonstrate that a substantial proportion of generated images depict individuals in traditional cultural attire and clothing impract...
work page 2024
-
[4]
It is therefore unclear whether changes to the revised prompt generation process between these periods may have affected the observed distributions. Finally, how people from one country wish to be portrayed is not easily agreed upon and disagreements likely exist between members of the same group regarding appropriate representation (Hall, Bell, et al., 2...
work page 2024
-
[5]
Conclusion In this paper, two common approaches to generating images were tested to assess how they represent people from various nationalities performing common activities. Results revealed that generated images frequently depicted people in traditional or impractical attire . Statistical analysis showed significant relationships between these representa...
work page 2024
-
[6]
Compute an **Alignment Score** ∈ [-1, 1] (floating-point, any precision). - **-1.0** → completely unrelated or contradictory - ** 0.0** → neutral / no discernible relation - ** 1.0** → perfectly matches in all salient aspects The score should approximate a cosine similarity between the visual scene and the text
-
[7]
### Reasoning Steps (think but only output final answer)
Write a **Concise Justification** (1 –3 well-formed sentences) that cites concrete visual/textual evidence for *why* you chose that score. ### Reasoning Steps (think but only output final answer)
-
[8]
**Identify** prominent objects, actions, attributes in the image
Show all 14 references
-
[9]
**Parse** key facts in the caption
-
[10]
**Compare** each element, note matches, partials, mismatches
-
[11]
**Aggregate** into a single scalar score
-
[12]
### Output Format (JSON-like list, no additional keys) Figure A.1
**Craft** a brief justification citing the strongest evidence. ### Output Format (JSON-like list, no additional keys) Figure A.1. Prompt used to Instruct GPT-4.1 mini to Generate Alignment Scores 29 Table A.1. Vocabulary for clusters using basic frequency counts Cluster Activi...
-
[13]
W., & Brundage, M
References Agarwal, S., Krueger, G., Clark, J., Radford, A., Kim, J. W., & Brundage, M. (2021). Evaluating clip: Towards characterization of broader capabilities and downstream implications. arXiv Preprint arXiv:2108.02818. Alimardani, M., & Elswah, M. (2021). Digital oriental...
2021 doi
-
[14]
tech nations
https://doi.org/10.1038/s42256-020-00257-z Ghosh, S., Venkit, P. N., Gautam, S., Wilson, S., & Caliskan, A. (2024). Do generative AI models output harm while representing non-western cultures: Evidence from a community-centered approach. Proceedings of the AAAI/ACM Conference ...
2024 doi
Reviewed May 22, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.