{"id":"2fb1763a-47ec-438a-8a64-159724b70a09","arxiv_id":"2412.13181","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"U.S. cities from 1900 to 2015 show superlinear temporal scaling for land, buildings, and roads (sprawl), with smaller cities spreading out faster and nearby cities growing alike.","lead":"This paper tracks how the built infrastructure of nearly 650 U.S. metropolitan areas grew relative to their population from 1900 to 2015, and finds that most cities spread out and build larger homes as they grow. It also finds that nearby cities grow in strikingly similar ways, with growth patterns staying correlated over surprisingly long distances.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline superlinear exponents may be an artifact of the house-count population proxy: because population is allocated proportionally to buildings, declining U.S. household size injects a trend that inflates fitted temporal scaling exponents.","rationale":"The reader's weakest assumption is the population proxy, and I agree that it is the most load-bearing concern. The paper's main empirical claims are temporal scaling exponents, and the denominator in every exponent is either directly or indirectly built from housing counts. The specific mechanism I add is quantitative: a declining mean household size creates a systematic gap between housing growth and population growth, and the fitted log-log slope is mathematically inflated by this gap even for a null model in which per-house infrastructure is constant. This directly threatens the headline 'superlinear' finding and the interpretation that cities become less dense and houses get larger. The proposed test is decisive because it replaces the proxy denominator with a directly measured one for the same CBSA-aggregated statistics. If the official census denominator reproduces the reported patterns, the proxy concern is resolved; if not, the central claims require substantial qualification. The reader's CONDITIONAL verdict already captures this uncertainty, so no adjustment is needed. I am not raising concerns about author conduct or novelty; the issue is purely an argument about identifiability of the scaling exponents from the described estimation procedure.","tokens_in":15430,"tokens_out":5415,"duration_ms":58986,"concrete_test":"Recompute the temporal scaling exponents for each CBSA using the official decennial CBSA census population as the denominator (summing all patches within the 2010 CBSA) instead of the house-proportional patch population, for the same building-derived numerators and the same time points. If the median exponents for developed area, indoor area, and footprint area drop to or below 1, or if the negative correlation between exponent and 2015 population weakens materially, the headline patterns are artifacts of the proxy. As a complementary check, simulate a null model with constant per-house developed area and indoor area, official census population, and house counts growing at observed rates; if the proxy-based fitting returns superlinear exponents similar to the reported ones, the demographic-trend explanation is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5 states that population within each urban patch is estimated as the fraction of total houses in a CBSA times the CBSA census population. Thus the temporal population denominator is P(t) = H_patch(t) x [N_census(t)/H_total(t)], where the bracket is the inverse of mean household size (ignoring vacancy). Over 1900–2015 mean U.S. household size fell from about 4.6 to 2.5, a secular downward trend. If log H_patch grows at rate a and the household-size term declines at rate b, then log P grows at rate a-b while a building-derived numerator such as footprint area grows at roughly rate a. The fitted exponent beta in Y ~ P^beta is approximately a/(a-b), which exceeds 1 even when per-house infrastructure is constant. The same building stock contributes to developed area, indoor area, and footprint area, so the superlinear scaling reported in Section 2.1 and the 'houses are getting larger' interpretation may substantially reflect demographic change rather than infrastructure growth. The smaller-city effect (Section 2.2) could also be affected if household-size trends or vacancy differ by city size. The authors acknowledge the proportionality assumption in Section 3.2 but provide no temporal validation; the cross-sectional validation cited from ref 20 does not test the time-varying ratio. The spatial-correlation result is secondary to this concern and is not needed to raise the conditional bar.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses HISDAC-US building records, Microsoft building footprints, and road network data to estimate temporal urban scaling exponents (developed area, indoor area, building footprint area, road length, and road network statistics versus population) for 647 U.S. CBSAs from 1900 to 2015. The authors report three patterns: scaling exponents are often superlinear, implying decreasing density and increasing indoor area per capita as cities grow; larger cities have smaller exponents than smaller cities; and exponents are spatially correlated over distances up to roughly 1000 km, beyond the range of population correlations. Robustness to data-completeness thresholds and time-period splits is reported in the Supplementary Information, and code and data are made available anonymously.","tokens_in":15604,"tokens_out":3688,"duration_ms":39304,"significance":"If the results hold, this would be a valuable long-term, large-sample characterization of temporal urban scaling, extending prior cross-sectional work and providing a new spatial-correlation result that could inform theories of urban growth and sustainability. The paper's strengths include the breadth of the data (over a century, nearly all U.S. metropolitan areas), explicit robustness checks in the SI (S4–S17), high R² values for the power-law fits (SI S18–S19), and the availability of code and data. The central caveat is that population is not measured directly but inferred from building counts, so the temporal scaling exponents may be substantially biased by demographic changes such as declining household size. Because this concern affects the headline superlinearity and the interpretation in terms of per-capita infrastructure growth, the paper in its current form does not yet establish its central claim.","major_comments":[{"comment":"The population denominator is constructed as P_patch(t) = H_patch(t) × [N_census(t)/H_total(t)], where H is the number of houses and N_census is the CBSA census population. Since developed area, indoor area, and building footprint area are derived from the same building-stock data that determines H_patch, the numerator and denominator of the scaling regressions are not independent. Mean U.S. household size declined from roughly 4.6 to 2.5 persons over 1900–2015. If H_patch grows at rate a and the persons-per-house ratio declines at rate b, then a per-house-constant numerator growing at rate a will appear to scale with population with exponent approximately a/(a−b), which exceeds 1 even when per-house infrastructure is constant. The manuscript does not provide any temporal validation of the house-to-population ratio; the dasymetric agreement cited in Section 5 is cross-sectional and does not test the time-varying ratio. Please add an analysis that replaces the building-derived population with a measure not proportional to building counts, or that explicitly corrects for household-size and vacancy trends, and report how the exponents in Figures 2–5 change.","section":"Section 5 (Population estimation) and Section 2.1"},{"comment":"The statement that \"houses are getting larger\" and that \"indoor area per capita increases\" is a direct interpretation of the same non-independent quantities. Because indoor area is roughly per-house indoor area times the number of houses, and population is roughly the number of houses times persons-per-house, the ratio indoor area/population equals (per-house indoor area)/(persons per house). The observed superlinear scaling could therefore be explained entirely by declining household size rather than by larger houses or sprawl. The paper should report per-housing-unit metrics (e.g., indoor area per house or per dwelling) and show whether the superlinear exponents and the size dependence in Figure 5 persist when the population proxy is changed or when household size is controlled for.","section":"Section 2.1 and Abstract"},{"comment":"The \"larger cities have smaller exponents\" pattern and the long-range spatial correlations of exponents could also arise, at least in part, from city-size-dependent or region-dependent trends in household size and vacancy, which propagate directly into the constructed population denominator. The SI robustness checks vary data-completeness thresholds and time periods (S4–S13), but they never vary the population construction itself. I request an additional robustness analysis using CBSA-level census population (rather than patch-level allocation) and/or including household-size and vacancy covariates, to test whether the size dependence and the spatial correlation of exponents survive.","section":"Section 2.2, Figure 5, and Figure 6"}],"minor_comments":[{"comment":"The Methods section states that Pearson correlations are computed for the distance-binned exponent correlations, while the caption of SI Figure S12 refers to Spearman correlations; please reconcile this inconsistency.","section":"Main text (Methods) vs. SI Figure S12 caption"},{"comment":"The caption says red lines are 10 random MSAs and blue dashed lines are 10 random µSAs, but the legend and the main text do not clarify whether the same random cities are used across panels; please state whether the sample is fixed or resampled per panel.","section":"Section 2.1, Figure 2 caption"},{"comment":"The phrase \"Dunns' posthoc test\" should be \"Dunn's post hoc test\" for consistency with the reference.","section":"Section 2.1, text"},{"comment":"The limitation paragraph acknowledges that population is assumed proportional to the number of buildings, but it does not mention the specific risk that a time-varying persons-per-household ratio will bias temporal scaling exponents; adding this explicit caveat would help readers interpret the reported exponents.","section":"Section 3.2, Limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of npj Urban Sustainability and the empirical effort is substantial. The main barrier is the circularity between the building-derived population and the building-derived infrastructure metrics, combined with the lack of a temporal validation of the population proxy. This is fixable with additional analyses, so I do not recommend rejection, but the central claims should not be accepted without a robustness check that addresses the household-size trend."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth reading if you work on urban scaling: it assembles a century of built-environment data for 647 US CBSAs, fits temporal scaling exponents cleanly, and reports two things I have not seen before. First, the scaling exponent decreases with city size, so smaller cities are losing density faster than larger ones. Second, exponents are spatially correlated out to ~1000 km, far beyond population correlations. The robustness checks are thorough — different completeness filters, time splits, MSA/µSA splits — and the power-law fits are genuinely good (SI S18–S19). The data integration is a real achievement, and the code/data link is a plus.\n\nThe soft spots are real but not fatal. The stress-test concern about the population denominator being house-derived is mostly wrong for the CBSA-level analysis: the scaling uses census population summed across patches, not house counts. The circularity is weaker than the reader says. However, the household-size point still lands on the numerator. Indoor area, footprint area, and developed area are all building-derived, so a secular decline in household size means more buildings per person even with fixed building size. That alone would push exponents above one, without any increase in per-house size. The paper's interpretation that \"houses are getting larger\" is therefore conflated with \"more houses per person.\" The authors acknowledge the proportionality assumption in 3.2, but they never test the time-varying ratio, and the cross-sectional validation in ref 20 does not address it. This should be the main revision target.\n\nTwo smaller issues. The spatial-correlation result could still reflect spatially autocorrelated data quality, though the filter tests help. And the novelty demarcation is vague: superlinear temporal scaling is already in Depersin & Barthelemy and Bettencourt et al., and the authors' own 2023 paper covers the same data. What is new — the size-dependence of exponents and the spatial correlation — should be stated more forcefully.\n\nVerdict: this deserves peer review, not desk rejection, but with major revision. I would bring it to a reading group, and I would cite the spatial-correlation result. The household-size confound needs a response before I would trust the quantitative exponents.","headline":"Solid century-scale empirical scaling analysis with two genuinely new patterns, but the superlinear exponents are partly contaminated by the household-size trend and the novelty relative to prior work needs sharper demarcation.","tokens_in":16248,"tokens_out":2996,"would_cite":true,"duration_ms":31893,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Tracking 647 U.S. metropolitan areas over 115 years, this paper finds that per-capita developed land, floor space, and road length grow as cities grow, that smaller cities sprawl fastest, and that nearby cities grow in lockstep.","keywords":["urban scaling","temporal scaling","urban sprawl","U.S. metropolitan areas","historical settlement data","road networks","spatial correlation","urban sustainability"],"falsifier":"Recompute the temporal scaling exponents using an independent historical population estimate per city decade, such as census-tract counts or dasymetric population surfaces rather than house-count allocation, and compare the slopes: if the exponents move to around one or below for developed area and indoor area, the central claim of long-run density decline would be an artifact of the population proxy, whereas if they stay superlinear the claim survives.","tokens_in":15139,"feed_emoji":"🏙️","tokens_out":6444,"duration_ms":57426,"temperature":0.7,"pith_summary":"This paper tries to establish that when individual U.S. metropolitan areas are tracked from 1900 to 2015, their infrastructure does not grow the way cross-sectional city-scaling theory predicts: developed land, indoor floor space, building footprints, and road length typically grow faster than population, so per-capita land and floor space rise as cities grow. It further claims that smaller cities have larger scaling exponents than larger cities, meaning small cities spread out fastest, and that scaling exponents are spatially correlated between cities up to roughly 1,000 km apart, far beyond the short range of population correlations. If these patterns hold, U.S. urban growth is a long-term process of density decline and regional lockstep, with direct consequences for land use, energy demand, and sustainability planning.","feed_headline":"U.S. cities sprawled for 115 years — smaller ones fastest","feed_subtitle":"Per-capita land and floor space rise as U.S. cities grow; nearby metros grow alike.","key_machinery":"The central object is the temporal scaling exponent $\\beta$, obtained by regressing log infrastructure statistic against log population separately for each CBSA over time; $\\beta > 1$ means per-capita infrastructure increases with growth, while $\\beta < 1$ means it decreases. The population side is reconstructed by allocating each CBSA's census population to urban patches in proportion to the number of houses in each patch, an assumption the authors state explicitly. The infrastructure side comes from integrated geospatial datasets: historical settlement layers, building footprints, and road-network data, with road ages inferred from nearby building ages. The spatial-correlation result is carried by a distance-binned Pearson correlation of scaling exponents between city pairs, compared with the same correlation for populations.","core_discovery":"Using historical building, footprint, indoor-area, and road-network reconstructions for 647 U.S. core-based statistical areas with sufficient data, the paper fits a power law $Y \\sim P^{\\beta}$ for each city's infrastructure statistic $Y$ against its estimated population $P$ decade by decade from 1900 to 2015. The central finding is that temporal scaling exponents $\\beta$ are often superlinear ($\\beta > 1$) for developed area, indoor area, building footprint area, and road length, which means each new resident is associated with more developed land, more floor space, and more road per person, so cities become less dense and houses larger as they grow. In contrast, road-intersection and edge counts often scale sublinearly, consistent with sprawl. The paper also finds that larger cities have smaller exponents than smaller cities, so large metropolitan areas show a compounding economy of scale, and that exponents for nearby cities remain positively correlated out to about 1,000 km, a distance at which population correlations have already vanished.","pith_inferences":["Editorial extension: if the house-count population proxy overstates population growth in older cities, for example because household size fell or vacancy rose, the reported superlinear exponents for developed area and indoor area would be too high; re-estimating with vacancy- or household-size-corrected populations is a direct robustness test.","Editorial extension: the spatial correlation length of about 1,000 km suggests that state-level planning regimes, climate zones, or topography could be a shared cause; comparing exponents before and after major federal highway or housing programs would test whether policy regimes shift the correlation.","Editorial extension: applying the same temporal-scaling pipeline to countries with different sprawl histories, such as European or Japanese metropolitan areas, would help separate U.S.-specific car-oriented growth from a universal urban-growth mechanism."],"forward_implications":["If temporal exponents are typically superlinear, U.S. urban growth since 1900 has meant steadily declining population density, so policies aimed at density must contend with a century-long baseline trend, not just recent zoning.","Larger cities' smaller exponents imply a compounding economy of scale: large metros add less new land and floor space per new resident than small ones, so the sustainability burden of sprawl falls disproportionately on small cities.","The roughly 1,000 km spatial correlation of exponents implies that growth patterns are regional, so land-use and transportation policy may be more effective when coordinated across neighboring metropolitan areas rather than city by city.","Sublinear scaling of intersections and edges per capita means road networks become less interconnected per person as cities grow, a signature of car-oriented, low-density expansion.","Regional differences (superlinear in the South and Midwest, sublinear in the Northeast and West) mean a single national growth law does not describe U.S. urbanization; regional context matters."],"supporting_citations":[{"why":"Supplies the historical road-network statistics and the method for dating roads from nearby buildings, used for road length and intersection metrics.","marker":"(19)"},{"why":"Provides the developed area, indoor area, building footprint area, road length, and population estimates used in the temporal scaling fits.","marker":"(20)"},{"why":"The Historical Settlement Data Compilation for the U.S. is the core data source for historical settlement extents.","marker":"(36)"},{"why":"Supplies the building footprint data used for footprint and indoor-area estimates.","marker":"(40)"},{"why":"The National Transportation Dataset provides the current road network from which historical roads are reconstructed.","marker":"(51)"},{"why":"The cross-sectional scaling theory whose predicted exponents the paper's superlinear temporal results are contrasted against.","marker":"(7)"},{"why":"Earlier empirical temporal-scaling analysis that the paper's findings align with and extend.","marker":"(9)"}],"fun_headline_variants":["Sprawl revealed: U.S. cities' footprint outgrows population for 115 years","Big cities grow more efficiently, but all sprawl: 115-year data","Nearby U.S. metros grow alike: 1,000-km similarity in urban scaling","Less dense, not more: U.S. cities' surprising growth pattern","U.S. urban sprawl quantified: each new resident needs more land"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis assumes that a city's population at each decade is proportional to the number of houses in it, so any systematic drift in people per house (smaller households, vacancies, second homes) would change every scaling exponent and could manufacture the reported superlinear sprawl.","fun_headline_variants_meta":{"raw":{"variants":["Sprawl revealed: U.S. cities' footprint outgrows population for 115 years","Big cities grow more efficiently, but all sprawl: 115-year data","Nearby U.S. metros grow alike: 1,000-km similarity in urban scaling","Less dense, not more: U.S. cities' surprising growth pattern","U.S. urban sprawl quantified: each new resident needs more land"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000518,"raw_usage":{"total_tokens":2558,"prompt_tokens":1039,"completion_tokens":1519,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":655,"completion_tokens_details":{"reasoning_tokens":1413}},"tokens_in":655,"tokens_out":1519,"duration_ms":10278,"temperature":1.0,"reasoning_tokens":1413,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:21:08.720527+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the temporal scaling exponents using an independent historical population estimate per city decade, such as census-tract counts or dasymetric population surfaces rather than house-count allocation, and compare the slopes: if the exponents move to around one or below for developed area and indoor area, the central claim of long-run density decline would be an artifact of the population proxy, whereas if they stay superlinear the claim survives.","supporting_citations":[],"review_version":1}