Pith. sign in

REVIEW 2 cited by

Navigating Cultural Chasms: Exploring and Unlocking the Cultural POV of Text-To-Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01929 v3 pith:UOBLLEFP submitted 2023-10-03 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords culturalmodelscultureevaluationsresearchtext-to-imageacrossagency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-To-Image (TTI) models, such as DALL-E and StableDiffusion, have demonstrated remarkable prompt-based image generation capabilities. Multilingual encoders may have a substantial impact on the cultural agency of these models, as language is a conduit of culture. In this study, we explore the cultural perception embedded in TTI models by characterizing culture across three hierarchical tiers: cultural dimensions, cultural domains, and cultural concepts. Based on this ontology, we derive prompt templates to unlock the cultural knowledge in TTI models, and propose a comprehensive suite of evaluation techniques, including intrinsic evaluations using the CLIP space, extrinsic evaluations with a Visual-Question-Answer (VQA) model and human assessments, to evaluate the cultural content of TTI-generated images. To bolster our research, we introduce the CulText2I dataset, derived from six diverse TTI models and spanning ten languages. Our experiments provide insights regarding Do, What, Which and How research questions about the nature of cultural encoding in TTI models, paving the way for cross-cultural applications of these models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CuRe scores text-to-image systems by how much their output changes as prompts add cultural details, and reports better agreement with human ratings than existing proxies.

  2. RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    RusCode is a new 1,250-prompt Russian/English benchmark for cultural awareness in text-to-image models, with human evaluation showing Russian-trained models outperform general models.

Pith tools