REVIEW 2 major objections 5 minor 39 references
Mapping the City Through the Lens of Language Models
T0 review · 2 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Language models, shown anonymous profiles of real urban centres, reliably rank larger, faster-growing, better-connected, less sparse cities as more typical of what a city is, and this portrait holds across ten checkpoints and a…
desk verdict A careful, well-designed audit of what LLMs implicitly treat as a typical city; the one load-bearing soft spot is that the clean-cell filter may shape the headline portrait, and the replication does not test that. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Default City, defined as a model-specific behavioral ranking over empirical urban profiles conditional on a specified elicitation task and reference population. The machinery that carries the argument is a constrained rating protocol: models assign 1–7 typicality scores via next-token probabilities on anonymized, geography-free profiles, and only cells meeting prespecified reliability screens for response validity, dispersion, forward/reverse order agreement, and reduced-form agreement feed the standardized contrasts and combined percentiles. Shared model lineages are aggregated through dependency components before combining, and three population targets, city-count, resident-weighted, and region-balanced, make the reference distribution explicit.
What would settle it
Compute the standardized high-versus-low typicality contrasts on the full 400-cell matrix, including all low-dispersion and order/reduction-sensitive cells, and check whether the reported positive contrasts on built extent, infrastructure, and non-residential capacity survive; if those dimensions flip or drop to zero, the Default City portrait is an artifact of the reliability screen.
Extended reading notes
Core claim
The paper establishes a behavioral definition of the Default City as a stable ranking over empirical urban profiles that a given language model, task, and reference population produce. Measured across ten open-weight checkpoints, the ranking reliably places larger developed area, faster recent growth, greater mapped infrastructure, less sparse form, higher human development, and greater non-residential capacity at the typical end. This scale–growth–infrastructure core recurs in a disjoint replication sample, with 285 of 291 city-count directions reproduced, and survives conditioning on city scale and development for many dimensions. Direct whole-profile ratings agree with the indicator-wise construction at a dependency-balanced correlation of 0.50. The finding is presented as a measurement of model behavior, not a claim about real cities.
Load-bearing premise
The clean-cell reliability screen separates measurement noise from genuine model preferences without systematically selecting the dimensions that support the scale–growth–infrastructure portrait.
Editorial extensions
If this is right
- If the paper is right, any downstream task that relies on a language model's implicit 'city' default inherits a scale–growth–infrastructure bias, not just a named-place geographic bias.
- The Default City portrait is measurable and auditable: the same protocol can be run on new checkpoints, new indicator sets, or new reference populations.
- Geographic tilt, with Europe and Northern America ranking high and Sub-Saharan Africa low, shrinks substantially after accounting for city scale and development, indicating that much of the apparent geographic bias is a composition effect.
- In reliably measured paired cells, typicality and desirability ratings are tightly coupled, with rank correlations of .904–.997, so the 'typical city' is largely the 'desirable city' for these models.
- Replication on a disjoint 512-city sample and moderate agreement with direct whole-profile ratings support the stability of the portrait as a behavioral property.
Reading between the lines
- The paper's reliability screens are a strong filter, with only 162 of 400 cells clean; an analysis that forced all low-dispersion and order-sensitive cells through the same contrasts would directly test whether the portrait is an artifact of exclusion.
- A natural extension not run in the paper is to probe why marginal indicator ratings and direct whole-profile ratings agree only moderately, since identifying the domains that drive disagreement could sharpen the prototype.
- The finding implies a practical risk beyond the paper's stated scope: applications that generate urban descriptions, planning text, or place profiles from language models may silently treat sparse, slow-growing, infrastructure-poor cities as less typical and therefore less legible.
- The tight typicality–desirability coupling, though measured on only 16 reliable paired cells, suggests that the category-central city may be confounded with normative ideals, a distinction the paper leaves for future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a behavioral audit of how ten open-weight language models rank anonymous profiles of real morphological urban centres on 'typicality of a city.' Using constrained next-token probabilities over a 1–7 scale, 40 audited indicators in seven domains, and a stratified sample of 2,000 centres plus a disjoint 512-centre replication sample, the authors find a recurring scale–growth–infrastructure core: models assign higher typicality to profiles with larger built extent, faster recent population growth, greater mapped infrastructure and non-residential capacity, and less sparse form. The paper also reports geographic affinities (Europe and Northern America above baseline, Sub-Saharan Africa below), which shrink after scale and development adjustment; moderate convergence between marginal dimension ratings and direct joint-profile ratings; and close typicality–desirability alignment in the reliable paired cells. Novel methodological ingredients include prespecified reliability screens, forward/reverse field-order controls, dependency-component-aware aggregation, survey-bootstrap inference, and artifact-bound reproduction.
Significance. If the result holds, this is a significant contribution to the empirical study of geographic and categorical bias in language models: it converts an informal intuition about 'default city' into a population-conditional, reproducible measurement over real urban profiles. The study is unusually careful: reliability screens are prespecified, order effects are controlled, lineages are aggregated through dependency components, a disjoint replication sample is used, and the supplement ships version pins, prompt text, inference settings, and artifact hashes. The finding is falsifiable and directly extends earlier observations of metropolitan-size bias. The main caveat is that the headline portrait is built only on the 162 'clean' cells, so the burden is on the authors to show that the reliability filters do not select the result.
major comments (2)
- [Measurement and Aggregation; Figures S6–S7; Table 3] The central portrait is computed exclusively from the 162 clean cells defined in 'Measurement and Aggregation', yet the manuscript never reports the direction of the high-vs-low typicality contrast for the 72 low-dispersion or 166 order/reduction-sensitive cells. Because the replication sample applies the same eligibility rules, its 285/291 direction agreement and .982 rank correlation establish stability within the filtered set, not the absence of filter-induced selection; the sensitivity analyses in Table 3 vary tail thresholds and aggregation but never the reliability criteria. A concrete instance is Table 1's population-growth row, where GLM-4 yields a Default of -0.7%/yr against the combined 2.2%/yr; if that cell is excluded by the screen, the paper should say which criterion excluded it and how many countervailing cells are excluded overall. I request a sign distribution for all 400 cells and a threshold-sweep robustness check (e.g., response ≥ .95, std ≥ .10, forward/reverse ρ ≥ .70) to show the scale-growth-infrastructure portrait is not an artifact of the clean-cell filter.
- [Behavioral Elicitation (Prompt Design)] The design assumes that excluding geographic names and source identifiers prevents models from inferring specific places, and this assumption is load-bearing for interpreting the ratings as category centrality rather than memory of named cities. The manuscript does not test for leakage; a profile combining climate, HDI, infrastructure, and morphology may uniquely identify a small number of centres. I recommend a probe that asks the same checkpoints to name the city or region from a subset of anonymous profiles, and a control that appends a non-informative random label to see whether typicality ratings change. If models recover locations, the 'Default City' interpretation needs qualification.
minor comments (5)
- [Methods, Eq. (2) context] The sentence 'we compare every source feature' in the paragraph around Eq. (2) is ambiguous: it is unclear whether Δ is computed for the focal dimension's own source feature or for all 40 features in the profile, and the aggregation across features is not specified.
- [Methods, bootstrap description] The bootstrap description says 'recompute weights, thresholds, and contrasts'; please specify that 'thresholds' refers to the upper-tail quintiles (q.9/q.1) used in Eq. (2), and clarify that the reliability criteria themselves are fixed rather than re-estimated in each replicate.
- [Results, Marginal–joint section] Table S9 shows that GLM-4's marginal–joint Spearman ρ is .258 with a 95% CI that includes zero ([-0.013, 0.501]), so the statement that 'all ten checkpoints' show positive agreement overstates the evidence for that checkpoint; the interval crossing zero should be noted.
- [Results, Scale conditioning] The replication section states that 'all 41 displayed directions remain' after scale conditioning, but the origin of the number 41 is not explained; please identify which 41 directions are counted (e.g., a subset of the 40 dimensions plus one derived contrast).
- [Figure 3 caption] The caption and the figure text use 'T ataouine' with a space; this should be 'Tataouine'.
Circularity Check
No significant circularity: the Default City portrait is a direct behavioral measurement of model ratings, not a quantity fitted to or derived from the same ratings by construction.
full rationale
The paper's central claim is that ten checkpoints, when rating anonymous urban profiles, assign higher typicality to profiles with greater built extent, growth, infrastructure, and non-residential capacity. This is a direct elicitation: the ratings themselves define typicality, and the reported contrasts compare those ratings against external morphological indicators. No free parameter is fitted to the outcome and then renamed as a prediction; the reliability screens and aggregation rules in 'Measurement and Aggregation' are prespecified and applied before the contrasts, and the 512-city replication is presented as a stability check rather than as an independent confirmation of a fitted model. The two self-citations (Zhao et al. 2026, Zhang et al. 2025) appear in related-work and interpretive passages and do not supply any load-bearing premise, uniqueness theorem, or ansatz to the measurement. The limitation that only 162 of 400 cells are clean, noted in Figures S6–S7 and the appendix, is a measurement-coverage caveat, not evidence that the clean-cell filter constructs the headline portrait by definition; the paper does not claim that excluded cells support or refute the direction. The checklist in Table S13 describes methodological scope, not a circular step. Thus no derivation step reduces to its own input, and no circularity can be exhibited.
Assumptions & free parameters
free parameters (4)
- Clean-cell reliability thresholds =
valid responses >= .99; sd >= .15; order corr >= .80; reduced-form corr >= .85
- Default upper-tail quintile =
20%
- Dependency component grouping =
7 components (Qwen-derived and Gemma checkpoints grouped)
- Bootstrap seeds and replicate counts =
1,000 and 5,000 replicates; seeds 20260724/20260725
assumptions (4)
- domain assumption The GHS Urban Centre Database R2024A V1.2 provides a valid international set of morphological urban centres.
- domain assumption The 40 audited indicators adequately operationalize the concept of "city" for typicality judgments.
- domain assumption The constrained next-token probability of a 1-7 rating reflects a stable internal category judgment.
- ad hoc to paper Excluding geographic names and source identifiers prevents the models from inferring specific places.
Cite this review
Pith. "Pith review of Mapping the City Through the Lens of Language Models." pith.science (2026). https://pith.science/paper/5YJXWAYK
@misc{pith2026260802971,
author = {Pith},
title = {Pith review of: Mapping the City Through the Lens of Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/5YJXWAYK}},
note = {Machine review of arXiv:2608.02971}
}
read the original abstract
Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and function. We measure those assumptions without naming places. Ten open-weight checkpoints rate anonymized profiles derived from real morphological urban centres across 40 audited indicators and seven domains. The design combines constrained probability-based ratings, prespecified reliability screens, lineage-aware aggregation, multiple population weightings, an independent replication sample, and whole-profile validation. The clearest shared tendency favours urban profiles with larger developed area, faster recent growth, greater mapped infrastructure and non-residential capacity, and less sparse form. Most eligible directions recur in the replication data, and direct ratings of complete profiles show moderate agreement with the indicator-wise construction. Geographic differences shrink after accounting for city scale and development, while reliably measured paired tasks indicate that typicality and desirability are often closely aligned. The framework makes an otherwise vague notion of what models regard as an ordinary city empirically traceable. The resulting evidence delineates a shared yet model-dependent portrait of the city through the lens of language models.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Journal of Experimental Psychology: General , volume =
Rosch, Eleanor , title =. Journal of Experimental Psychology: General , volume =. 1975 , doi =
work page 1975
- [2]
-
[3]
Medin, Douglas L. and Schaffer, Marguerite M. , title =. Psychological Review , volume =. 1978 , doi =
work page 1978
-
[4]
Transactions of the Association for Computational Linguistics , volume =
Malaviya, Chaitanya and Chang, Joseph Chee and Roth, Dan and Iyyer, Mohit and Yatskar, Mark and Lo, Kyle , title =. Transactions of the Association for Computational Linguistics , volume =. 2025 , doi =
work page 2025
-
[5]
Proceedings of the 40th International Conference on Machine Learning , series =
Santurkar, Shibani and Durmus, Esin and Ladhak, Faisal and Lee, Cinoo and Liang, Percy and Hashimoto, Tatsunori , title =. Proceedings of the 40th International Conference on Machine Learning , series =. 2023 , url =
work page 2023
-
[6]
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages =
Ribeiro, Marco Tulio and Wu, Tongshuang and Guestrin, Carlos and Singh, Sameer , title =. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages =. 2020 , doi =
work page 2020
-
[7]
Wang, Peiyi and Li, Lei and Chen, Liang and Cai, Zefan and Zhu, Dawei and Lin, Binghuai and Cao, Yunbo and Kong, Lingpeng and Liu, Qi and Liu, Tianyu and Sui, Zhifang , title =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2024 , doi =
work page 2024
-
[8]
Findings of the Association for Computational Linguistics: NAACL 2024 , pages =
Pezeshkpour, Pouya and Hruschka, Estevam , title =. Findings of the Association for Computational Linguistics: NAACL 2024 , pages =. 2024 , doi =
work page 2024
Show all 39 references
-
[9]
arXiv preprint arXiv:2402.02680 , year =
Manvi, Rohin and Khanna, Samar and Burke, Marshall and Lobell, David and Ermon, Stefano , title =. arXiv preprint arXiv:2402.02680 , year =
-
[10]
Proceedings of the 3rd Workshop on Multi-lingual Representation Learning , pages =
Faisal, Fahim and Anastasopoulos, Antonios , title =. Proceedings of the 3rd Workshop on Multi-lingual Representation Learning , pages =. 2023 , doi =
2023
-
[11]
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation , pages =
Dunn, Jonathan and Adams, Benjamin and Tayyar Madabushi, Harish , title =. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation , pages =. 2024 , url =
2024
-
[12]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages =
Li, Bryan and Haider, Samar and Callison-Burch, Chris , title =. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages =. 2024 , doi =
2024
-
[13]
Transactions on Machine Learning Research , year =
Liang, Percy and Bommasani, Rishi and Lee, Tony and Tsipras, Dimitris and Soylu, Dilara and Yasunaga, Michihiro and Zhang, Yian and Narayanan, Deepak and Wu, Yuhuai and Kumar, Ananya and others , title =. Transactions on Machine Learning Research , year =
- [14]
-
[15]
Proceedings of the Conference on Fairness, Accountability, and Transparency , pages =
Mitchell, Margaret and Wu, Simone and Zaldivar, Andrew and Barnes, Parker and Vasserman, Lucy and Hutchinson, Ben and Spitzer, Elena and Raji, Inioluwa Deborah and Gebru, Timnit , title =. Proceedings of the Conference on Fairness, Accountability, and Transparency , pages =. 2...
2019
-
[16]
Datasheets for Datasets , journal =
Gebru, Timnit and Morgenstern, Jamie and Vecchione, Briana and Vaughan, Jennifer Wortman and Wallach, Hanna and Daum\'. Datasheets for Datasets , journal =. 2021 , doi =
2021
-
[17]
Improving Reproducibility in Machine Learning Research: A Report from the
Pineau, Joelle and Vincent-Lamarre, Philippe and Sinha, Koustuv and Larivi\`ere, Vincent and Beygelzimer, Alina and d'Alch\'. Improving Reproducibility in Machine Learning Research: A Report from the. Journal of Machine Learning Research , volume =. 2021 , url =
2021
-
[18]
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages =
Pattnayak, Priyaranjan and Bhatia, Apoorv , title =. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages =. 2026 , doi =
2026
-
[19]
and Gebru, Timnit and McMillan-Major, Angelina and Shmitchell, Shmargaret , title =
Bender, Emily M. and Gebru, Timnit and McMillan-Major, Angelina and Shmitchell, Shmargaret , title =. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2021 , doi =
2021
-
[20]
Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages =
Weidinger, Laura and Uesato, Jonathan and Rauh, Maribeth and Griffin, Conor and Huang, Po-Sen and Mellor, John and Glaese, Amelia and Cheng, Myra and Balle, Borja and Kasirzadeh, Atoosa and others , title =. Proceedings of the 2022 ACM Conference on Fairness, Accountability, a...
2022
-
[21]
2024 , doi =
Melchiorri, Michele and Mari Rivero, Ines and Florio, Pietro and Schiavina, Marcello and Krasnodebska, Katarzyna and Politis, Panagiotis and Uhl, Johannes and Pesaresi, Martino and others , title =. 2024 , doi =
2024
-
[22]
and Freire, Sergio and Kemper, Thomas and Melchiorri, Michele and Pesaresi, Martino and Schiavina, Marcello , title =
Dijkstra, Lewis and Florczyk, Aneta J. and Freire, Sergio and Kemper, Thomas and Melchiorri, Michele and Pesaresi, Martino and Schiavina, Marcello , title =. Journal of Urban Economics , volume =. 2021 , doi =
2021
-
[23]
Rao, J. N. K. and Wu, C. F. J. and Yue, K. , title =. Survey Methodology , volume =
-
[24]
On the Generalized Bootstrap for Sample Surveys with Special Attention to Poisson Sampling , journal =
Beaumont, Jean-Fran. On the Generalized Bootstrap for Sample Surveys with Special Attention to Poisson Sampling , journal =. 2012 , doi =
2012
-
[25]
arXiv preprint arXiv:2406.13945 , year =
Feng, Jie and Zhang, Jun and Liu, Tianhui and Zhang, Xin and Ouyang, Tianjian and Yan, Junbo and Du, Yuwei and Guo, Siqi and Li, Yong , title =. arXiv preprint arXiv:2406.13945 , year =
-
[26]
arXiv preprint arXiv:2504.21027 , year =
Zheng, Yu and Liu, Longyi and Lin, Yuming and Feng, Jie and Zhang, Guozhen and Jin, Depeng and Li, Yong , title =. arXiv preprint arXiv:2504.21027 , year =
-
[27]
arXiv preprint arXiv:2505.13803 , year =
Zhang, Yecheng and Zhao, Rong and Huang, Zimu and Wang, Xinyu and Ma, Yue and Long, Ying , title =. arXiv preprint arXiv:2505.13803 , year =. doi:10.48550/arXiv.2505.13803 , url =
-
[28]
Proceedings of the First Workshop on Natural Language Processing for Human Resources , pages =
Campanella, Charlie and van der Goot, Rob , title =. Proceedings of the First Workshop on Natural Language Processing for Human Resources , pages =. 2024 , doi =
2024
-
[29]
arXiv preprint arXiv:2505.09388 , year =
Yang, An and Li, Anfeng and Yang, Baosong and Zhang, Beichen and Hui, Binyuan and Zheng, Bo and others , title =. arXiv preprint arXiv:2505.09388 , year =
- [30]
-
[31]
2024 , howpublished =
Mistral. 2024 , howpublished =
2024
-
[32]
arXiv preprint arXiv:2408.00118 , year =
Gemma 2: Improving Open Language Models at a Practical Size , author =. arXiv preprint arXiv:2408.00118 , year =
-
[33]
arXiv preprint arXiv:2407.21783 , year =
The Llama 3 Herd of Models , author =. arXiv preprint arXiv:2407.21783 , year =
-
[34]
arXiv preprint arXiv:2601.10387 , year =
Lu, Christina and Gallagher, Jack and Michala, Jonathan and Fish, Kyle and Lindsey, Jack , title =. arXiv preprint arXiv:2601.10387 , year =
-
[35]
ISPRS Journal of Photogrammetry and Remote Sensing , volume =
Hou, Yujun and Quintana, Matias and Khomiakov, Maxim and Yap, Winston and Ouyang, Jiani and Ito, Koichi and Wang, Zeyu and Zhao, Tianhong and Biljecki, Filip , title =. ISPRS Journal of Photogrammetry and Remote Sensing , volume =. 2024 , doi =
2024
-
[36]
arXiv preprint arXiv:2506.04106 , year =
Zhu, Xiao Xiang and Chen, Sining and Zhang, Fahong and Shi, Yilei and Wang, Yuanyuan , title =. arXiv preprint arXiv:2506.04106 , year =
-
[37]
2025 , doi =
Bondarenko, Maksym and Priyatikanto, Rhorom and Tejedor-Garavito, Natalia and Zhang, Wei and McKeen, Taylor and Cunningham, Andrew and Woods, Thomas and Hilton, Jason and Cihan, Doga and Nosatiuk, Bohdan and Brinkhoff, Thomas and Tatem, Andrew and Sorichetta, Alessandro , titl...
2025
-
[38]
Overture Maps Data Documentation: Places, Buildings, and Transportation , howpublished =
-
[39]
General Transit Feed Specification , howpublished =
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.