{"id":"0b71e5f4-49a9-4310-a923-6c2a5f9a3752","arxiv_id":"2506.16898","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Diffusion models FLUX 1 and SD 3.5 encode fine-grained US geographic knowledge when prompted with states or capitals, but the generic prompt 'USA' produces a metropolitan stereotype that under-represents rural, frontier, and desert regions.","lead":"This paper asked two image-generation models, FLUX 1 and Stable Diffusion 3.5, to create street-view photos of every US state and state capital. It found that the models can tell US regions apart, but when asked for a generic 'USA' picture, they fall back on a metropolitan stereotype, leaving rural and frontier places out.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The geographic-knowledge claim depends on FID distances computed from singular 384x384 covariance estimates with n=150, and the clusters are never validated against geography; the claim is not yet established.","rationale":"The reader's weakest assumption correctly identifies the n=150 versus d=384 covariance estimation problem in Section 2.2 and Equation (1); that is the most concrete technical flaw. I agree it is load-bearing because the entire geographic-knowledge argument is built on pairwise FID distances. I add a second, equally load-bearing condition: even a perfectly estimated FID matrix would not establish geographic knowledge unless the cluster-geography correspondence is tested quantitatively. The paper specifies no clustering algorithm and no null model, so the named regional groupings are asserted by inspection. The diversity-deficit half of the strongest claim is less threatened: the USA-prompt FID distances to desert, tropical, and frontier states are large and directionally consistent across both models, though the magnitudes may shift with regularization. The proposed check, shrunk covariance plus a Mantel test, directly targets the knowledge claim. If the clusters survive shrinkage and spatial correlation is significant, the paper's qualitative conclusion stands and the reader's conditional verdict is appropriate. If not, the verdict should move toward UNVERDICTED until a validated reanalysis is provided. I keep the verdict UNCHANGED because the current CONDITIONAL verdict already requires exactly this additional validation.","tokens_in":6621,"tokens_out":7348,"duration_ms":80076,"concrete_test":"Recompute the entire pairwise FID matrix using Ledoit-Wolf shrunk covariance estimates (or mean-only Euclidean distances on the same embeddings), re-run the hierarchical clustering, and test the FID-versus-geography correlation with a Mantel permutation test against great-circle distances between state centroids/capitals. If the named regional clusters dissolve or the spatial correlation is not significant (permutation p > 0.05), the geographic-knowledge claim is not supported; if both survive, the concern is settled in the paper's favor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1's central claim, that geographically proximate states and capitals cluster in FID space, requires the pairwise FID matrix of Section 2.3 (Eq. 1) to be a reliable visual-similarity measure, and that is not demonstrated. Each location is summarized by 150 DINO-v2 ViT-S/14 embeddings of dimension 384 (Section 2.2), so each empirical 384x384 covariance matrix is singular and high-variance: n < d means off-diagonal structure is dominated by sampling noise, and the trace term involving the matrix square root of the product of two such covariance estimates is ill-conditioned. No regularization, bootstrap, or confidence interval is reported. The 'clustering' in Figure 2 is then interpreted by inspection: no algorithm, linkage criterion, or quantitative agreement with known U.S. regions is specified, so the Mountain West, Desert Southwest, and New England groupings could reflect estimation noise or subjective cluster reading rather than encoded geography. The diversity-deficit finding (Section 3.3) is somewhat more robust because it compares every location to the same 'USA' reference, but its FID ranking inherits the same covariance-estimation instability. This is not an internal inconsistency, but the headline knowledge claim is currently under-supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether two open text-to-image diffusion models (FLUX 1-schnell and Stable Diffusion 3.5-Large) encode geographic knowledge of the United States and exhibit a national-scale representational bias. For each of the 50 states, their capitals, and a generic \"USA\" prompt, the authors generate 150 street-view images, embed them with DINO-v2 ViT-S/14, and compute pairwise Fréchet Inception Distances (FID) between all location pairs. They report that geographically proximate states and capitals cluster together in FID space, that small capitals with European-sounding names are systematically misgenerated as European cityscapes, and that the generic \"USA\" prompt yields images that are much closer in FID to large metropolitan states than to frontier, desert, tropical, or rural states. The paper concludes that the models possess detailed latent geographic knowledge but deploy a narrow metropolitan stereotype when prompted at the national scale.","tokens_in":6858,"tokens_out":2900,"duration_ms":31027,"significance":"If the central findings are supported, the paper makes a useful empirical contribution to the study of geographic bias in generative models, with implications for urban analytics, fairness, and the evaluation of text-to-image systems. The authors are to be credited for a systematic experimental design: 101 prompts, 150 images per prompt per model, a fixed prompt template, two current open models, and a public embedding model. The headline diversity-deficit result is qualitatively plausible and is partially supported by Table 1, which shows that states such as Hawaii, Alaska, and Arizona have much larger FID to the \"USA\" prompt than states like New Jersey or Illinois. However, the more ambitious claim of implicit geographic knowledge rests on visual inspection of unvalidated clusters computed from ill-conditioned covariance estimates, and the reported FID numbers are not accompanied by uncertainty intervals. The paper is therefore best treated as a valuable preliminary study whose central knowledge claim requires stronger statistical grounding before it can be taken as established.","major_comments":[{"comment":"The FID computation uses each location's 150 DINO-v2 embeddings of dimension 384 to estimate a full 384×384 covariance matrix. With n=150 and d=384, each empirical covariance is singular and high-variance, and the matrix-square-root term in Eq. (1) is ill-conditioned. No regularization, shrinkage, dimensionality reduction, bootstrap, or confidence interval is reported. Because the pairwise FID matrix is the basis for the clustering in Figure 2 and the rankings in Table 1, the apparent geographic structure could partly reflect estimation noise. I request a robustness analysis: for example, regularized covariance estimation, PCA projection to fewer dimensions, or bootstrap confidence intervals around the FID values and cluster assignments.","section":"§2.2 and §2.3, Eq. (1)"},{"comment":"The claim that geographically proximate states and capitals \"cluster together\" is based on visual inspection of dendrograms, but the manuscript does not specify the clustering algorithm, linkage criterion, cutoff, or any quantitative validation against known U.S. regions. Without a test such as comparing within-region versus between-region FID distances, a Mantel-style correlation between FID distance and geographic distance, or at least a reproducible cluster assignment, the groupings described in the text (e.g., the Mountain West cluster, Desert Southwest cluster, New England cluster) cannot be distinguished from arbitrary thresholding or subjective reading of the dendrogram. Please add a defined clustering procedure and a quantitative evaluation of its agreement with geography.","section":"§3.1, Figure 2"},{"comment":"The diversity-deficit ranking is more robust than the clustering claim because it compares every location to the same \"USA\" reference, but it inherits the same covariance-estimation instability. Table 1 reports only point estimates ranked by FID; no bootstrap intervals or significance tests are given, so the reader cannot tell whether the rank ordering of, say, New Jersey versus Hawaii is statistically meaningful. Additionally, the table mixes states and capital cities in a single ranking, which may confound two different prompt types; the interpretation would be cleaner if states and capitals were analyzed separately or explicitly modeled as distinct conditions. Please add uncertainty estimates and clarify whether the combined ranking is appropriate.","section":"§3.3, Table 1"}],"minor_comments":[{"comment":"The sentence \"While it is fundamental to examine the geographic knowledge and biases that models encode is crucial\" contains a grammatical error and should be split into two clauses; there is also a typo in \"spaital\" (spatial).","section":"§1, Introduction"},{"comment":"Reference [13] is identical to reference [12] and both cite Rombach et al. 2022, but Stable Diffusion 3.5-Large is not introduced by that paper; the authors should cite the appropriate Stable Diffusion 3.5 technical report or model card.","section":"References"},{"comment":"The sentence \"For each prompt, we generated images\" does not state the number; the abstract says 150 images per prompt, but the main text should repeat that number explicitly for clarity.","section":"§2.1"},{"comment":"The cluster numbers referenced in the text (e.g., \"Cluster 10\" and \"Cluster 11\") are not easily identifiable in the figure; please annotate the dendrograms with the cluster labels used in the text.","section":"Figure 2"},{"comment":"The claim that misgenerated capitals \"resemble European cities\" is based on visual inspection of clustered images; since this is a secondary finding, it should be framed as an illustrative observation rather than a quantitative result, or supported by a content-based evaluation.","section":"§3.2"},{"comment":"No statement of data availability or code release is provided; sharing the generated-image lists, FID matrices, and analysis scripts would substantially improve reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely topic and the diversity-deficit result is interesting, but the geographic-knowledge claim needs stronger statistical support to be publication-ready. I would encourage the editor to request a revision with a regularized or reduced-dimensional FID analysis, bootstrap uncertainty, and a quantitative cluster-geography validation. The duplicate reference [12]/[13] and the incorrect citation for Stable Diffusion 3.5 should also be corrected before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it this morning. Quick take: the headline finding — both FLUX 1-schnell and SD 3.5-L collapse the generic \"USA\" prompt into a metropolitan stereotype, even though they show some per-state differentiation — is likely right and useful. The weaker link is the paper's other main claim, that the models encode fine-grained geographic knowledge. That claim leans on clustering of FID distances that are computed from singular covariance estimates and then interpreted by eye.\n\nWhat's genuinely new: the per-state and per-capital FID analysis of two current open models, and the use of the same \"USA\" reference to expose which states and capitals are left out of the national stereotype. That's a reasonable, incremental contribution to the geographic-bias evaluation literature. The diversity-deficit table is the most convincing part: the gaps are large (Hawaii at 2594 vs. New Jersey at 331 for FLUX; Arizona at 3352 vs. Illinois at 413 for SD 3.5-L). Effects of that size are unlikely to be pure covariance noise, even if the point estimates for any specific rank are unstable.\n\nThe soft spot is Section 3.1. Each location is summarized by 150 DINO-v2 embeddings of dimension 384, so each empirical covariance is singular and high-variance; the FID trace term is ill-conditioned without regularization. No bootstrap, confidence interval, or alternative distance is reported. The clusters in Figure 2 are described qualitatively — \"Mountain West,\" \"Desert Southwest\" — with no linkage criterion, no algorithm, and no quantitative match against known regions. The small-capital misgeneration section is basically anecdotal (\"Olympia notably characterized by features of ancient Greek architecture\"). No artifacts are released, so others can't re-run or sanity-check. Also, the SD 3.5-L reference points to the 2022 LDM paper, which is the wrong source for that model. That's sloppy in a short paper.\n\nThose are real flaws, but they don't sink the core message. The metropolitan-stereotype finding is robust against the methodological caveats because it's a relative comparison against one reference and the effect sizes dwarf the likely noise. The fine-grained clustering should be treated as suggestive, not established. A serious referee would ask for regularized covariance or bootstrap intervals, a pre-registered clustering analysis with external validation against U.S. Census regions, and at least a sample of generated images released.\n\nWho's this for? People evaluating geographic bias in text-to-image models, and the urban-AI community that actually uses these tools to generate scenario imagery. It's a workshop-grade empirical diagnostic, not a definitive study. I'd send it to peer review rather than desk-reject it; with moderate revisions (robustness checks, reframing the knowledge claim, fixing the citation) it would be a fine short paper. I'd bring it to a reading group if we had one that cared about generative-model evaluation, but I wouldn't treat its clustering results as evidence of structured geographic knowledge yet.","headline":"The diversity-deficit result is probably real and worth publishing as a diagnostic; the finer 'geographic knowledge' claim is under-built because it rests on singular covariance estimates and eyeballed clustering.","tokens_in":7360,"tokens_out":2062,"would_cite":true,"duration_ms":23325,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that FLUX 1-schnell and Stable Diffusion 3.5-L encode implicit geographic knowledge of the United States, yet collapse the generic 'USA' prompt into a metropolitan stereotype that underrepresents rural, frontier, desert…","keywords":["diffusion models","geographic knowledge","geographic bias","urban imagery","text-to-image generation","FID","DINO-v2","representation bias"],"falsifier":"Recompute pairwise FID with a shrinkage-regularized covariance estimator, or replace FID with a nonparametric distance such as maximum mean discrepancy on the same DINO-v2 embeddings; if geographically neighboring states no longer cluster together, the reported geographic knowledge is an artifact of unstable covariance estimation. Alternatively, shuffle location labels across image sets and show that random pairings produce clustering as strong as real geographic neighbors.","tokens_in":6436,"feed_emoji":"🏙️","tokens_out":6594,"duration_ms":59389,"temperature":0.7,"pith_summary":"This paper tests whether two open diffusion image models, FLUX 1-schnell and Stable Diffusion 3.5-L, have internal geographic knowledge of the United States and whether they use it fairly. The authors generate 150 street-view images for each U.S. state, each state capital, and a generic 'USA' prompt, embed them with DINO-v2, and measure pairwise Fréchet Inception Distance (FID). They find that geographically neighboring states and capitals cluster together in FID space, evidence of implicit geographic structure. But when prompted generically with 'USA', both models collapse onto a metropolitan stereotype, leaving frontier, desert, tropical, rural, and small-city environments far from the 'USA' reference. The paper concludes that the models know more geography than they deploy in national-level prompts.","feed_headline":"Image models know US geography but default to metro clichés","feed_subtitle":"FLUX and SD 3.5 cluster nearby states correctly, yet 'USA' yields only big-city streets.","key_machinery":"The argument is carried by the pairwise FID distance matrix built from DINO-v2 ViT-S/14 embeddings of the generated images. FID, a measure of dissimilarity between two sets of image embeddings, compares each location's embedding distribution through its mean and covariance; low FID means visually similar street scenes. Hierarchical clustering of these distances shows whether visual similarity mirrors real geography, while each location's FID to the generic 'USA' images quantifies how much that location is represented in the model's national stereotype.","core_discovery":"The central claim is that FLUX 1-schnell and Stable Diffusion 3.5-L encode fine-grained implicit geographic knowledge of the United States, visible as clustering of geographically proximate states and capitals in FID space, while simultaneously reproducing a narrow metropolitan stereotype for the generic 'USA' prompt. Small state capitals with European-sounding names (Frankfort, Montpelier, Pierre, Dover, Olympia, Bismarck) are systematically misgenerated as European cityscapes, which the authors attribute to toponymic confusion and data sparsity. The result is a demonstrated gap between the geographic diversity the models can produce when asked for specific locations and the diversity they actually produce when asked for a broad region.","pith_inferences":["The same FID pipeline could test whether the 'USA' stereotype is driven by training-data overrepresentation rather than by prompt ambiguity, by comparing with a neutral prompt such as 'a street in a country' across multiple countries.","The clustering patterns might align more with climate and ecoregion boundaries than with state borders; if so, the 'geographic knowledge' may actually be environmental knowledge, and the paper's state-level framing would be one way to see it.","A natural extension is to measure whether this metropolitan collapse also occurs at other scales (continents, global regions) and whether it shrinks when the model is asked for 'rural USA' or 'small town USA' explicitly.","The small-capital misgeneration suggests a testable mitigation: adding the state name to the prompt ('Frankfort, Kentucky') should move those capitals toward correct American clusters; if it does not, the issue is visual training data rather than toponymic priors."],"forward_implications":["If the models indeed encode geographic structure, researchers can use image-generation outputs and embedding distances as a probe of what a model has learned about place, without needing labels or external geographic data.","Any downstream urban-analysis pipeline that queries a model with a country-level prompt will inherit the metropolitan bias; generated scenario images for 'USA' will not represent the diversity of American built environments.","The systematic misgeneration of small capitals implies that place names with strong foreign-language associations need explicit geographic context in prompts to avoid visual cross-country confusion.","The FID-to-USA ranking provides a simple, reusable audit: compute the distance of every region's generated images to the generic country prompt to detect underrepresentation.","The knowledge-diversity gap suggests that dynamic prompting (adding region, biome, or city-size hints) could make models display geographic knowledge they already have."],"supporting_citations":[{"why":"Supplies the Fréchet Inception Distance used to measure visual similarity between location image sets.","marker":"[5]"},{"why":"Defines FLUX 1-schnell, one of the two models under test.","marker":"[8]"},{"why":"Underlies Stable Diffusion 3.5-L, the other model under test.","marker":"[13]"},{"why":"Supplies the DINO-v2 ViT-S/14 embeddings from which all FID statistics are computed.","marker":"[11]"},{"why":"Documents geographic disparities in text-to-image outputs that motivate the bias analysis.","marker":"[4]"},{"why":"Provides analysis of visual stereotypes that supports interpreting the results as representation bias.","marker":"[7]"}],"fun_headline_variants":["AI knows state geography, but 'USA' means big-city streets","Diffusion models cluster states, yet 'USA' is a metro cliché","State-level AI geography, but national prompt is all cities","When AI hears 'USA', it sees only skyscrapers and traffic"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The FID computation assumes the 150 image embeddings for each location form a Gaussian distribution and estimates a full 384x384 covariance matrix from just 150 samples without regularization, so the clusters that reveal 'geographic knowledge' could partly reflect estimation noise rather than true visual structure.","fun_headline_variants_meta":{"raw":{"variants":["AI knows state geography, but 'USA' means big-city streets","Diffusion models cluster states, yet 'USA' is a metro cliché","State-level AI geography, but national prompt is all cities","When AI hears 'USA', it sees only skyscrapers and traffic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000281,"raw_usage":{"total_tokens":1610,"prompt_tokens":835,"completion_tokens":775,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":451,"completion_tokens_details":{"reasoning_tokens":697}},"tokens_in":451,"tokens_out":775,"duration_ms":7999,"temperature":1.0,"reasoning_tokens":697,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:16:09.793969+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute pairwise FID with a shrinkage-regularized covariance estimator, or replace FID with a nonparametric distance such as maximum mean discrepancy on the same DINO-v2 embeddings; if geographically neighboring states no longer cluster together, the reported geographic knowledge is an artifact of unstable covariance estimation. Alternatively, shuffle location labels across image sets and show that random pairings produce clustering as strong as real geographic neighbors.","supporting_citations":[],"review_version":1}