Pith. sign in

REVIEW 2 major objections 2 minor 14 references

Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities

T0 review · 2 major / 2 minor · reviewed 2026-05-22 · grok-4.3

Pith's one-line read Text-to-image models generate traditional attire for many nationalities even during everyday activities

desk verdict The 28.4% traditional attire rate and region associations are new in scale but rest on human labels with no reported reliability checks. read the letter →

arxiv 2504.06313 v5 submitted 2025-04-08 cs.CV cs.CY

classification cs.CVcs.CY
keywords text-to-imagemodelsnationalityrepresentationtraditionalattireimagegenerationbiasstereotypesDALL-EGeminipromptengineering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests DALL-E 3 and Gemini 3 Pro Preview on prompts that ask for images of people from 206 nationalities performing five common activities. Across the 2,060 generated images, 28.4 percent show traditional clothing, often unsuitable for the described task. This pattern appears far more often for nationalities from the Middle East and North Africa as well as Sub-Saharan Africa and for lower-income groups. Alignment scores from CLIP, ALIGN, and GPT-4.1 mini rise when country names are in the prompt, and one model tends to insert the word traditional into its internal prompt revisions. The results point to multiple pipeline stages that can shape these representational choices.

What carries the argument

Region- and income-level statistical associations between manually or model-labeled traditional attire and the nationality named in the prompt, together with alignment scoring on image-prompt pairs.

What would settle it

A re-labeling study that uses multiple independent raters with a written definition of traditional attire and impracticality and finds no significant regional difference would falsify the reported associations.

Watch

Extended reading notes

Core claim

When aggregating across activities and models, 28.4 percent of the images depicted individuals wearing traditional attire, including attire that is impractical for the specified activities in several cases. This pattern was statistically significantly associated with regions, with the Middle East and North Africa and Sub-Saharan Africa disproportionately affected, and was also associated with World Bank income groups. Images labeled as featuring traditional attire received statistically significantly higher alignment scores when prompts included country names, and one model frequently inserted the word traditional in its revisions.

Load-bearing premise

The labeling of images as containing traditional attire and as impractical for the activity is treated as reliable enough to support region-level statistical associations.

Editorial extensions

If this is right

  • Including a country name in the prompt raises the chance that the output will show traditional rather than context-appropriate clothing.
  • One model internally revises prompts by adding the word traditional at higher rates for certain nationalities.
  • Alignment models themselves assign higher scores to the traditional-attire outputs when country information is present.
  • The observed patterns differ by generator model and by the choice of evaluation model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Users requesting images of people from affected regions may receive systematically less practical or contemporary depictions.
  • Removing country names from prompts or adding explicit instructions for modern clothing could reduce the frequency of these outputs.
  • Similar region-linked patterns may appear in other generative models that draw on overlapping training data.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper investigates how DALL-E 3 and Gemini 3 Pro Preview represent people from 206 nationalities in five everyday activities by generating 2,060 images. Aggregating results, it finds 28.4% depict traditional attire, with statistically significant associations to regions (disproportionately MENA and Sub-Saharan Africa) and income groups. Similar patterns for impractical attire in athletics. Alignment is scored using CLIP, ALIGN, and GPT-4.1 mini on 9,270 pairs, showing higher scores for traditional images with country names in prompts, and prompt revisions frequently insert 'traditional'.

Significance. Should the labeling methodology be validated with reliability metrics, this work would offer valuable quantitative insights into cultural biases in text-to-image generation. It highlights how model behaviors and prompt elements can lead to stereotypical depictions, contributing to the growing body of research on fairness in generative AI. The scale (206 nationalities) and use of multiple alignment models are strengths that could influence future auditing practices in the field.

major comments (2)
  1. [§3.2 (Image Labeling)] §3.2 (Image Labeling): The process for determining whether generated images contain 'traditional attire' or attire 'impractical for the activity' is not described in detail, including any guidelines provided to annotators, the number of annotators per image, or measures of inter-rater agreement. This is a load-bearing issue for the reported 28.4% statistic and the region associations, as subjective interpretations could vary systematically by cultural context.
  2. [§4.1 (Statistical Analysis)] §4.1 (Statistical Analysis): While statistical significance is reported for the region and income associations, the manuscript should specify the exact tests used (e.g., chi-square), degrees of freedom, and any multiple comparison corrections to allow full evaluation of the findings' robustness.
minor comments (2)
  1. [Abstract] Abstract: The abstract states 'including attire that is impractical for the specified activities in several cases' but does not quantify how many such cases; consider adding this detail for precision.
  2. [Prompt Revision Analysis] Prompt Revision Analysis: Clarify which model showed the 50.3% insertion rate for 'traditional' and whether this holds across all activities.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their detailed and constructive comments, which highlight important areas for improving the clarity and robustness of our methodology. We address each major comment below and will incorporate revisions as indicated.

read point-by-point responses
  1. Referee: [§3.2 (Image Labeling)] The process for determining whether generated images contain 'traditional attire' or attire 'impractical for the activity' is not described in detail, including any guidelines provided to annotators, the number of annotators per image, or measures of inter-rater agreement. This is a load-bearing issue for the reported 28.4% statistic and the region associations, as subjective interpretations could vary systematically by cultural context.

    Authors: We agree that the current description in §3.2 lacks sufficient detail on the labeling protocol. In the revised manuscript we will expand this section to specify the exact guidelines provided to annotators (including definitions and examples for 'traditional attire' and 'impractical for the activity'), confirm that each image was independently labeled by three annotators, and report inter-rater agreement using an appropriate metric such as Fleiss' kappa. These additions will directly address concerns about subjectivity and cultural context. revision: yes

  2. Referee: [§4.1 (Statistical Analysis)] While statistical significance is reported for the region and income associations, the manuscript should specify the exact tests used (e.g., chi-square), degrees of freedom, and any multiple comparison corrections to allow full evaluation of the findings' robustness.

    Authors: We will revise §4.1 to explicitly document the statistical procedures: chi-square tests of independence were used for the region and income-group associations, with degrees of freedom and exact p-values reported; we will also state whether any multiple-comparison correction (e.g., Bonferroni) was applied. These details were present in our analysis code but omitted from the initial text and can be added without altering the reported conclusions. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity; empirical pipeline uses external scorers on newly generated images

full rationale

The paper generates 2,060 new images from fixed prompts specifying nationalities and activities, then applies external models (CLIP, ALIGN, GPT-4.1 mini) to score alignment on 9,270 pairs and performs labeling for traditional attire and impracticality. No parameters are fitted inside the dataset, no predictions are made from fitted subsets, and no self-citations or uniqueness theorems are invoked as load-bearing premises. Statistical associations (region and income-group differences) are computed directly from the labeled outputs without any self-definitional reduction or renaming of known results. The analysis is self-contained against external benchmarks and contains none of the enumerated circularity patterns.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The work is purely empirical and does not introduce new theoretical entities or derivations; it relies on standard statistical testing assumptions and the operational definitions of 'traditional attire' and 'impractical'.

assumptions (1)
  • standard math Standard assumptions of statistical significance testing (independence, appropriate distribution for p-values)
    Invoked when reporting statistically significant associations with regions and income groups.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities." pith.science (2026). https://pith.science/paper/2504.06313

@misc{pith2026250406313,
  author       = {Pith},
  title        = {Pith review of: Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2504.06313}},
  note         = {Machine review of arXiv:2504.06313}
}
read the original abstract

This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview, depict people from 206 nationalities when prompted to generate images of individuals engaging in common everyday activities. Five scenarios were developed, and 2,060 images were generated using input prompts that specified nationalities across five activities. When aggregating across activities and models, results showed that 28.4% of the images depicted individuals wearing traditional attire, including attire that is impractical for the specified activities in several cases. This pattern was statistically significantly associated with regions, with the Middle East & North Africa and Sub-Saharan Africa disproportionately affected, and was also associated with World Bank income groups. Similar region- and income-linked patterns were observed for images labeled as depicting impractical attire in two athletics-related activities. To assess image-text alignment, CLIP, ALIGN, and GPT-4.1 mini were used to score 9,270 image-prompt pairs. Images labeled as featuring traditional attire received statistically significantly higher alignment scores when prompts included country names, and this pattern weakened or reversed when country names were removed. Revised prompt analysis showed that one model frequently inserted the word "traditional" (50.3% for traditional-labeled images vs. 16.6% otherwise). These results indicate that these representational patterns can be shaped by several components of the pipeline, including image generator, evaluation models, and prompt revision.

Figures

Figures reproduced from arXiv: 2504.06313 by the authors.

Figure 3
Figure 3. Classification of regions according to the World Bank. Source of image: World Bank [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 5
Figure 5. Alignments results for CLIP, ALIGN, and GPT-4.1 mini. For better readiability, the y-axis ranges have been changed For the standard prompts used to generate the images (see [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

14 extracted references · 14 canonical work pages

  1. [1]

    traditional

    Table 3 presents summary statistics for the dataset. For each model, 206 images were generated, resulting in 1,030 images per model across the five scenarios (i.e., 1,030 for DALL-E 3 and 1,030 for Gemini 3), and 2,060 images in total. Because each imag e is associated with multiple prompts that were evaluated for alignment scores, a total of 9,270 prompt...

  2. [2]

    traditional

    Results 5.1 RQ1: Representations in Generated Images Based on the labeling of images generated across activities and countries, distinct patterns emerged in how the two models represented individuals from different nationalities. When aggregating results across all activities and models, 28.4% of images depicted individuals wearing traditional attire. In ...

  3. [3]

    a couple cooking a meal in their home,

    Discussion 6.1 Implications The findings of this study demonstrate how two T2I models represent people from diverse nationalities when prompted to generate images of people engaging in common everyday activities. The results demonstrate that a substantial proportion of generated images depict individuals in traditional cultural attire and clothing impract...

  4. [4]

    It is therefore unclear whether changes to the revised prompt generation process between these periods may have affected the observed distributions. Finally, how people from one country wish to be portrayed is not easily agreed upon and disagreements likely exist between members of the same group regarding appropriate representation (Hall, Bell, et al., 2...

  5. [5]

    Results revealed that generated images frequently depicted people in traditional or impractical attire

    Conclusion In this paper, two common approaches to generating images were tested to assess how they represent people from various nationalities performing common activities. Results revealed that generated images frequently depicted people in traditional or impractical attire . Statistical analysis showed significant relationships between these representa...

  6. [6]

    Compute an **Alignment Score** ∈ [-1, 1] (floating-point, any precision). - **-1.0** → completely unrelated or contradictory - ** 0.0** → neutral / no discernible relation - ** 1.0** → perfectly matches in all salient aspects The score should approximate a cosine similarity between the visual scene and the text

  7. [7]

    ### Reasoning Steps (think but only output final answer)

    Write a **Concise Justification** (1 –3 well-formed sentences) that cites concrete visual/textual evidence for *why* you chose that score. ### Reasoning Steps (think but only output final answer)

  8. [8]

    **Identify** prominent objects, actions, attributes in the image

Show all 14 references
  1. [9]

    **Parse** key facts in the caption

  2. [10]

    **Compare** each element, note matches, partials, mismatches

  3. [11]

    **Aggregate** into a single scalar score

  4. [12]

    ### Output Format (JSON-like list, no additional keys) Figure A.1

    **Craft** a brief justification citing the strongest evidence. ### Output Format (JSON-like list, no additional keys) Figure A.1. Prompt used to Instruct GPT-4.1 mini to Generate Alignment Scores 29 Table A.1. Vocabulary for clusters using basic frequency counts Cluster Activi...

  5. [13]

    W., & Brundage, M

    References Agarwal, S., Krueger, G., Clark, J., Radford, A., Kim, J. W., & Brundage, M. (2021). Evaluating clip: Towards characterization of broader capabilities and downstream implications. arXiv Preprint arXiv:2108.02818. Alimardani, M., & Elswah, M. (2021). Digital oriental...

  6. [14]

    tech nations

    https://doi.org/10.1038/s42256-020-00257-z Ghosh, S., Venkit, P. N., Gautam, S., Wilson, S., & Caliskan, A. (2024). Do generative AI models output harm while representing non-western cultures: Evidence from a community-centered approach. Proceedings of the AAAI/ACM Conference ...

Pith tools

Reviewed May 22, 2026 · model on record in the stance chip above.