{"id":"c7d69c33-48eb-44ef-965b-3c57675d885c","arxiv_id":"2608.02971","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Ten open-weight language models, rating anonymized 40-indicator city profiles, systematically favor larger, faster-growing, infrastructure-rich, less sparse urban forms.","lead":"Language models rate anonymous profiles of real cities, revealing a consistent default city: larger built area, faster growth, more infrastructure, and less sparse form. The study measures what LLMs consider a typical city, which matters for AI fairness and urban applications.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reliability screen may select the headline portrait; replication reuses the same screen, so it does not test this.","rationale":"The reader's weakest assumption identifies the true load-bearing point: the headline 'Default City' portrait is built only from cells that pass a compound reliability screen, and no analysis compares the resulting directions with those in cells that fail one or more criteria. Because the disjoint replication uses the same eligibility rules, it cannot break the link between the filter and the outcome. The paper's sensitivity checks deliberately vary thresholds around the clean-cell set but never vary the definition of clean itself, leaving the selection concern open. A relaxed-eligibility recomputation is the minimal check that would settle it. Since this concern is already reflected in the reader's CONDITIONAL verdict, no verdict change is needed; the condition should explicitly include the filter-sensitivity analysis.","tokens_in":21619,"tokens_out":5450,"duration_ms":64583,"concrete_test":"Recompute the standardized high-vs-low typicality contrasts and the dependency-balanced Default portrait after relaxing the screen to include every cell with valid response rate ≥ .99 and typicality SD ≥ .15, while lowering or dropping the forward/reverse and reduced/full-form correlation thresholds (e.g., to .60). If the relaxed 40-dimensional sign pattern and the combined percentiles for built extent, population growth, infrastructure, and non-residential capacity reproduce the clean-cell results, the filter-selection concern is settled. If signs weaken or reverse, the portrait is partly an artifact of the reliability filter.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the 'clean' cells defined in Measurement and Aggregation, yet only 162 of 400 cells meet those criteria while 72 are low-dispersion and 166 are order/reduction-sensitive (Figures S6–S7). The paper does not report the high-vs-low typicality directions for cells that fail one or more reliability criteria, so it cannot rule out that the clean-cell filter selects dimensions that support the scale–growth–infrastructure portrait. This is not answered by the replication sample: the 512-city rerun applies the same eligibility rules, so its 285/291 direction agreement and .982 rank correlation measure stability within the filtered set, not absence of filter-induced selection. The sensitivity analyses in Table 3 vary upper-tail thresholds, aggregation choices, leave-one-component comparisons, and language, but never the reliability criteria themselves.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a behavioral audit of how ten open-weight language models rank anonymous profiles of real morphological urban centres on 'typicality of a city.' Using constrained next-token probabilities over a 1–7 scale, 40 audited indicators in seven domains, and a stratified sample of 2,000 centres plus a disjoint 512-centre replication sample, the authors find a recurring scale–growth–infrastructure core: models assign higher typicality to profiles with larger built extent, faster recent population growth, greater mapped infrastructure and non-residential capacity, and less sparse form. The paper also reports geographic affinities (Europe and Northern America above baseline, Sub-Saharan Africa below), which shrink after scale and development adjustment; moderate convergence between marginal dimension ratings and direct joint-profile ratings; and close typicality–desirability alignment in the reliable paired cells. Novel methodological ingredients include prespecified reliability screens, forward/reverse field-order controls, dependency-component-aware aggregation, survey-bootstrap inference, and artifact-bound reproduction.","tokens_in":21731,"tokens_out":11397,"duration_ms":115066,"significance":"If the result holds, this is a significant contribution to the empirical study of geographic and categorical bias in language models: it converts an informal intuition about 'default city' into a population-conditional, reproducible measurement over real urban profiles. The study is unusually careful: reliability screens are prespecified, order effects are controlled, lineages are aggregated through dependency components, a disjoint replication sample is used, and the supplement ships version pins, prompt text, inference settings, and artifact hashes. The finding is falsifiable and directly extends earlier observations of metropolitan-size bias. The main caveat is that the headline portrait is built only on the 162 'clean' cells, so the burden is on the authors to show that the reliability filters do not select the result.","major_comments":[{"comment":"The central portrait is computed exclusively from the 162 clean cells defined in 'Measurement and Aggregation', yet the manuscript never reports the direction of the high-vs-low typicality contrast for the 72 low-dispersion or 166 order/reduction-sensitive cells. Because the replication sample applies the same eligibility rules, its 285/291 direction agreement and .982 rank correlation establish stability within the filtered set, not the absence of filter-induced selection; the sensitivity analyses in Table 3 vary tail thresholds and aggregation but never the reliability criteria. A concrete instance is Table 1's population-growth row, where GLM-4 yields a Default of -0.7%/yr against the combined 2.2%/yr; if that cell is excluded by the screen, the paper should say which criterion excluded it and how many countervailing cells are excluded overall. I request a sign distribution for all 400 cells and a threshold-sweep robustness check (e.g., response ≥ .95, std ≥ .10, forward/reverse ρ ≥ .70) to show the scale-growth-infrastructure portrait is not an artifact of the clean-cell filter.","section":"Measurement and Aggregation; Figures S6–S7; Table 3"},{"comment":"The design assumes that excluding geographic names and source identifiers prevents models from inferring specific places, and this assumption is load-bearing for interpreting the ratings as category centrality rather than memory of named cities. The manuscript does not test for leakage; a profile combining climate, HDI, infrastructure, and morphology may uniquely identify a small number of centres. I recommend a probe that asks the same checkpoints to name the city or region from a subset of anonymous profiles, and a control that appends a non-informative random label to see whether typicality ratings change. If models recover locations, the 'Default City' interpretation needs qualification.","section":"Behavioral Elicitation (Prompt Design)"}],"minor_comments":[{"comment":"The sentence 'we compare every source feature' in the paragraph around Eq. (2) is ambiguous: it is unclear whether Δ is computed for the focal dimension's own source feature or for all 40 features in the profile, and the aggregation across features is not specified.","section":"Methods, Eq. (2) context"},{"comment":"The bootstrap description says 'recompute weights, thresholds, and contrasts'; please specify that 'thresholds' refers to the upper-tail quintiles (q.9/q.1) used in Eq. (2), and clarify that the reliability criteria themselves are fixed rather than re-estimated in each replicate.","section":"Methods, bootstrap description"},{"comment":"Table S9 shows that GLM-4's marginal–joint Spearman ρ is .258 with a 95% CI that includes zero ([-0.013, 0.501]), so the statement that 'all ten checkpoints' show positive agreement overstates the evidence for that checkpoint; the interval crossing zero should be noted.","section":"Results, Marginal–joint section"},{"comment":"The replication section states that 'all 41 displayed directions remain' after scale conditioning, but the origin of the number 41 is not explained; please identify which 41 directions are counted (e.g., a subset of the 40 dimensions plus one derived contrast).","section":"Results, Scale conditioning"},{"comment":"The caption and the figure text use 'T ataouine' with a space; this should be 'Tataouine'.","section":"Figure 3 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for cs.CL and is unusually rigorous for an empirical audit. The clean-cell selectivity issue is the main obstacle to accepting the central claim as stated; the requested sign-distribution and threshold-sweep analyses are feasible within the existing supplement. The anonymity/leakage probe is also important, though secondary. If the authors address both, I would be inclined to accept."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper deserves a serious referee. It is a rare instance of treating LLM judgments as a measurement problem rather than as a source of text to be scored. The 'Default City' framing is genuinely new: the authors elicit typicality ratings for anonymized real-city profiles, avoiding named geographies and evaluative labels, and they build a multidimensional portrait from 40 audited indicators across seven domains.\n\nWhat the paper does well is substantial. The design includes constrained next-token probability ratings, prespecified reliability screens, forward/reverse order controls, a disjoint replication sample, joint-profile validation, and dependency-aware aggregation across ten pinned open-weight checkpoints. The data foundation is real (GHS UCDB plus source-specific panels), the prompting is careful about units and labels, and the appendix is unusually transparent, including a full 400-cell status matrix and artifact receipts. The main analyses are reproducible in principle, though code and data are promised rather than released.\n\nThe soft spot is the one the stress-test note identifies. Only 162 of 400 model–dimension cells are 'clean', and the headline portrait is built on those. The paper does not report typicality contrasts for cells that fail one or more reliability criteria. If low-dispersion or order-sensitive cells tended to show weak or countervailing directions, the clean-cell filter could be selecting dimensions that support the scale–growth–infrastructure story. The replication sample does not address this, because it reuses the same eligibility rules; 285/291 direction agreement and a .982 rank correlation demonstrate stability within the filtered set, not absence of filter-induced selection. This is a real, addressable limitation, not a fatal one. The authors can respond by reporting directions for all cells, or by showing that non-clean cells do not systematically differ. They should also make the code and data public, since several checks depend on the full matrix.\n\nTwo smaller issues: the typicality–desirability coupling rests on only 16 of 60 paired model–dimension cells, so it is suggestive rather than strong; and the 48-city joint-profile validation is small, though acceptable as a coherence check.\n\nOverall, the paper is honest, methodologically serious, and fills a clear gap. The core finding—that LLMs share a scale–growth–infrastructure bias in what they treat as a typical city—is likely real, but the current evidence is conditional on the filter analysis. The paper deserves peer review, not desk rejection, and the reviewers should ask for the supplementary analyses. I would bring it to a reading group and would cite it if I worked on LLM evaluation or geospatial AI.","headline":"A careful, well-designed audit of what LLMs implicitly treat as a typical city; the one load-bearing soft spot is that the clean-cell filter may shape the headline portrait, and the replication does not test that.","tokens_in":22267,"tokens_out":2622,"would_cite":true,"duration_ms":32352,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Language models, shown anonymous profiles of real urban centres, reliably rank larger, faster-growing, better-connected, less sparse cities as more typical of what a city is, and this portrait holds across ten checkpoints and a…","keywords":["language models","urban typicality","Default City","geographic bias","prototype theory","urban indicators","reliability audit","open-weight LLMs"],"falsifier":"Compute the standardized high-versus-low typicality contrasts on the full 400-cell matrix, including all low-dispersion and order/reduction-sensitive cells, and check whether the reported positive contrasts on built extent, infrastructure, and non-residential capacity survive; if those dimensions flip or drop to zero, the Default City portrait is an artifact of the reliability screen.","tokens_in":21382,"feed_emoji":"🏙️","tokens_out":5308,"duration_ms":53113,"temperature":0.7,"pith_summary":"The paper asks what a language model means when it completes an underspecified reference to 'a city.' Instead of letting models describe cities in words, it makes them rate anonymized numeric profiles of 2,000 real urban centres across 40 audited indicators, then checks which profiles receive the highest typicality scores. Ten open-weight checkpoints converge on the same portrait: the default city is larger, faster-growing, more built-up, and better connected than the median real urban centre, with higher human development and more non-residential activity. The authors argue this makes a previously vague notion, the category-central city in a model's learned distribution, empirically traceable and replicable.","feed_headline":"LLMs see the typical city as big, dense, fast-growing","feed_subtitle":"A ten-model audit of anonymous city profiles finds a shared scale-growth-infrastructure portrait of urban typicality.","key_machinery":"The central object is the Default City, defined as a model-specific behavioral ranking over empirical urban profiles conditional on a specified elicitation task and reference population. The machinery that carries the argument is a constrained rating protocol: models assign 1–7 typicality scores via next-token probabilities on anonymized, geography-free profiles, and only cells meeting prespecified reliability screens for response validity, dispersion, forward/reverse order agreement, and reduced-form agreement feed the standardized contrasts and combined percentiles. Shared model lineages are aggregated through dependency components before combining, and three population targets, city-count, resident-weighted, and region-balanced, make the reference distribution explicit.","core_discovery":"The paper establishes a behavioral definition of the Default City as a stable ranking over empirical urban profiles that a given language model, task, and reference population produce. Measured across ten open-weight checkpoints, the ranking reliably places larger developed area, faster recent growth, greater mapped infrastructure, less sparse form, higher human development, and greater non-residential capacity at the typical end. This scale–growth–infrastructure core recurs in a disjoint replication sample, with 285 of 291 city-count directions reproduced, and survives conditioning on city scale and development for many dimensions. Direct whole-profile ratings agree with the indicator-wise construction at a dependency-balanced correlation of 0.50. The finding is presented as a measurement of model behavior, not a claim about real cities.","pith_inferences":["The paper's reliability screens are a strong filter, with only 162 of 400 cells clean; an analysis that forced all low-dispersion and order-sensitive cells through the same contrasts would directly test whether the portrait is an artifact of exclusion.","A natural extension not run in the paper is to probe why marginal indicator ratings and direct whole-profile ratings agree only moderately, since identifying the domains that drive disagreement could sharpen the prototype.","The finding implies a practical risk beyond the paper's stated scope: applications that generate urban descriptions, planning text, or place profiles from language models may silently treat sparse, slow-growing, infrastructure-poor cities as less typical and therefore less legible.","The tight typicality–desirability coupling, though measured on only 16 reliable paired cells, suggests that the category-central city may be confounded with normative ideals, a distinction the paper leaves for future work."],"forward_implications":["If the paper is right, any downstream task that relies on a language model's implicit 'city' default inherits a scale–growth–infrastructure bias, not just a named-place geographic bias.","The Default City portrait is measurable and auditable: the same protocol can be run on new checkpoints, new indicator sets, or new reference populations.","Geographic tilt, with Europe and Northern America ranking high and Sub-Saharan Africa low, shrinks substantially after accounting for city scale and development, indicating that much of the apparent geographic bias is a composition effect.","In reliably measured paired cells, typicality and desirability ratings are tightly coupled, with rank correlations of .904–.997, so the 'typical city' is largely the 'desirable city' for these models.","Replication on a disjoint 512-city sample and moderate agreement with direct whole-profile ratings support the stability of the portrait as a behavioral property."],"supporting_citations":[{"why":"Supplies the GHS Urban Centre Database R2024A V1.2, the globally harmonized morphological urban centres from which all anonymous profiles are drawn.","marker":"[Melchiorri et al., 2024]"},{"why":"Provides the degree-of-urbanisation definition that determines which centres qualify as urban morphological units.","marker":"[Dijkstra et al., 2021]"},{"why":"Documents metropolitan-size bias in a specific labor-market task, the result this paper generalizes to multidimensional typicality judgments.","marker":"[Campanella and van der Goot, 2024]"},{"why":"Supplies the prototype-theory framing that motivates treating typicality as graded category centrality.","marker":"[Rosch, 1975]"},{"why":"Underlies the survey bootstrap used to propagate design weights and thresholds in the reported uncertainty intervals.","marker":"[Rao et al., 1992]"},{"why":"Extends the survey bootstrap to Poisson sampling, supporting the replication and inference machinery of the audit.","marker":"[Beaumont and Patak, 2012]"}],"fun_headline_variants":["LLMs default to big, dense, fast-growing cities","AI's urban assumption: big, fast-growing, dense","Language models favor size, growth, and density in cities","What LLMs think a city is: large, growing, dense","LLM city stereotypes: big, booming, dense"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The clean-cell reliability screen separates measurement noise from genuine model preferences without systematically selecting the dimensions that support the scale–growth–infrastructure portrait.","fun_headline_variants_meta":{"raw":{"variants":["LLMs default to big, dense, fast-growing cities","AI's urban assumption: big, fast-growing, dense","Language models favor size, growth, and density in cities","What LLMs think a city is: large, growing, dense","LLM city stereotypes: big, booming, dense"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000483,"raw_usage":{"total_tokens":2352,"prompt_tokens":877,"completion_tokens":1475,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":1393}},"tokens_in":493,"tokens_out":1475,"duration_ms":12636,"temperature":1.0,"reasoning_tokens":1393,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T04:23:16.172135+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the standardized high-versus-low typicality contrasts on the full 400-cell matrix, including all low-dispersion and order/reduction-sensitive cells, and check whether the reported positive contrasts on built extent, infrastructure, and non-residential capacity survive; if those dimensions flip or drop to zero, the Default City portrait is an artifact of the reliability screen.","supporting_citations":[{"cited_title":"2024 , doi =","cited_arxiv_id":null,"evidence_quote":"Supplies the GHS Urban Centre Database R2024A V1.2, the globally harmonized morphological urban centres from which all anonymous profiles are drawn."},{"cited_title":"and Freire, Sergio and Kemper, Thomas and Melchiorri, Michele and Pesaresi, Martino and Schiavina, Marcello , title =","cited_arxiv_id":null,"evidence_quote":"Provides the degree-of-urbanisation definition that determines which centres qualify as urban morphological units."},{"cited_title":"Proceedings of the First Workshop on Natural Language Processing for Human Resources , pages =","cited_arxiv_id":null,"evidence_quote":"Documents metropolitan-size bias in a specific labor-market task, the result this paper generalizes to multidimensional typicality judgments."},{"cited_title":"Journal of Experimental Psychology: General , volume =","cited_arxiv_id":null,"evidence_quote":"Supplies the prototype-theory framing that motivates treating typicality as graded category centrality."},{"cited_title":"On the Generalized Bootstrap for Sample Surveys with Special Attention to Poisson Sampling , journal =","cited_arxiv_id":null,"evidence_quote":"Extends the survey bootstrap to Poisson sampling, supporting the replication and inference machinery of the audit."}],"review_version":1}