{"id":"684e1b34-4224-4b93-a1d8-b7ddd2ff8f59","arxiv_id":"1908.11549","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A pairwise local exponent and a benchmark-city effective exponent are proposed as fitting-free diagnostics for scaling, yielding new insights on urban datasets.","lead":"This paper introduces a fitting-free method to study scaling laws by comparing pairs of cities directly, defining a local exponent that should converge to the true exponent at large size ratios. The method is applied to urban datasets and reveals threshold effects and conflicting scaling behaviors that standard fits miss.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reliability of β_eff is asserted from in-sample fit; the proposed necessity checks are internal-consistency filters that would also pass under misspecified prefactors, so an out-of-sample test is needed to support the central prediction claim.","rationale":"The reader's verdict is CONDITIONAL, and my analysis supports keeping that verdict rather than moving it. The method is internally coherent: Eqs. (3)-(4) correctly describe the local-exponent construction under the stated assumptions, and the tomography plots are a reasonable exploratory visual tool. The load-bearing gap is the inference from in-sample variance minimization and in-sample f(1/2) to 'reliable' prediction for other cities. This is an unsupported step, not an internal contradiction, and it is closely tied to the reader's weaker assumption about the constant prefactor: both failure modes break the link between the fitted exponent and out-of-sample prediction. The proposed necessary conditions are also too strong because they are phrased as necessary for trusting β_hat, but they only check internal consistency of the data with one scaling exponent. The correct remedy is explicit out-of-sample validation, which would settle whether the method delivers the claimed reliability. Since the paper is already framed as a methodological proposal and the reader already asked for out-of-sample validation, no verdict change is needed; the condition should be made explicit in the acceptance criteria.","tokens_in":15778,"tokens_out":4490,"duration_ms":45383,"concrete_test":"Perform an out-of-sample validation on each dataset: randomly split cities into training (70%) and test (30%); on the training set select the benchmark city by minimizing Eq. (11), compute β_eff and the least-squares β_hat, then compute f(1/2) on the held-out test set for both exponents. Repeat over 100 splits and compare the mean test f(1/2) with the reported in-sample value. Also run one regional split for OECD patents (train on a subset of countries, predict the rest). If held-out f(1/2) is substantially below the in-sample value, or β_eff is not better than β_hat on held-out data, the paper's central reliability assertion is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central practical claim is that the benchmark city and β_eff computed from Eq. (12) allow one to compute 'reliably' properties of other cities, with f(1/2) reported as an accuracy measure. But β_eff is selected by minimizing the variance σ^2(i) on the full dataset and f(1/2) is then evaluated on that same dataset; no train/test split or cross-validation appears. Consequently the reported fractions are in-sample fit quality, not predictive reliability. More sharply, the three 'necessary conditions' in the Discussion are internal consistency checks: β_loc(r)→β_hat, β_eff≈β_hat, and f(1/2)≥50% can all hold while the assumed constant-prefactor form Eq. (1) is false. If the prefactor is not constant but varies systematically with P (e.g., ln a = c ln P), the local-exponent construction and the OLS fit will both converge to the same composite exponent; all criteria pass, yet the exponent is not a stable predictor across regimes or regions, as the paper itself suspects for OECD patents. Conversely, exact scaling with large multiplicative noise can yield f(1/2)<50%, so condition (iii) is not necessary. The proposal is a useful heuristic, but the reliability claim needs stronger evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a set of descriptive tools for studying scaling laws Y ~ P^β without relying on a regression fit. It defines a local exponent β_loc(i,j) = log(Y_j/Y_i)/log(P_j/P_i) for every pair of cities and plots this quantity against the population ratio r, calling the result a 'tomography plot'. It then defines an effective exponent β_eff as the mean local exponent of the benchmark city that minimizes the variance of β_loc, and claims that this exponent can be used to 'reliably' predict the properties of other cities through Eq. (12). The manuscript also proposes three necessary conditions for trusting a fitted exponent β_hat: convergence of β_loc toward β_hat, consistency of β_eff with β_hat, and a fraction f(1/2) of at least 50% of correctly predicted cities. These tools are applied to several urban datasets, showing cases where the method confirms standard fits, cases where it suggests threshold effects, and cases where it indicates that no simple single-exponent scaling holds.","tokens_in":16090,"tokens_out":4289,"duration_ms":41593,"significance":"If the proposed diagnostics are valid, they provide a simple and visually appealing complement to regression-based scaling analyses, with practical value for distinguishing linear, superlinear, and sublinear behavior in noisy, short-range city data. The algebraic derivation of the local exponent and its behavior under multiplicative noise (Eqs. 3-7) is sound and clearly explained. The application to multiple real datasets is a strength, and the paper honestly reports cases where its own diagnostics weaken previously drawn conclusions, such as the European libraries case. However, the central predictive claim is currently supported only by in-sample measures, and the proposed 'necessary conditions' are not logically necessary for the assumed scaling form. These issues do not invalidate the descriptive utility of the tomography plot, but they do require substantial revision of the reliability claims before the method can be recommended for practical prediction.","major_comments":[{"comment":"The effective exponent β_eff is selected by minimizing the variance σ^2(i) over all cities in the dataset, and the quality metric f(1/2) is evaluated on the same cities used for the selection. Consequently, the reported f(1/2) values measure in-sample fit quality, not predictive reliability for other cities. The text's claim that Eq. (12) allows computing 'reliably' properties of other cities needs support from an out-of-sample evaluation, such as a train/test split or leave-one-city-out cross-validation; this is currently absent.","section":"Section 3, Eqs. (10)-(12) and (14)-(15)"},{"comment":"The three proposed necessary conditions are not necessary for the scaling form Eq. (1) to be correct with a constant prefactor. In the multiplicative-noise model used in Eq. (4), exact scaling with large noise can produce f(1/2) < 50%, so condition (iii) is not necessary. Conversely, all three conditions can hold when Eq. (1) is violated by a systematically varying prefactor: if ln a = c ln P, both β_loc and the OLS fit converge to β + c, and β_eff will be consistent with that value, yet the constant-prefactor scaling form is false. The conditions are internal consistency checks among estimators, not tests of the scaling assumption, and the statement that failure of these conditions 'safely' rejects β_hat is not justified.","section":"Discussion, conditions (i)-(iii)"},{"comment":"The threshold interpretations for UK railroads, AIDS cases in Brazil, and US patents are supported mainly by comparing R^2 values of power-law fits and fits with one or two additional parameters, e.g., r2=0.98 versus 0.76 for railroads, and r2=0.93 versus 0.81 for AIDS. Because the threshold models have more parameters, these R^2 differences do not by themselves establish that the scaling form is inadequate. A model-comparison criterion that penalizes complexity (AIC/BIC or a likelihood-ratio test) is needed to substantiate the paper's threshold-related conclusions.","section":"Problematic cases, Figs. 12, 13, 15"}],"minor_comments":[{"comment":"The multiplicative-noise model Y2 = Y1 r^β (1+η) is introduced without stating the support of η; the expansion log(1+η) requires η > -1, and the phrase 'when the noise is not too large' should be made quantitative.","section":"Section 3, Eq. (4)"},{"comment":"The claimed convergence of β_loc toward β_hat is assessed visually from tomography plots, with no formal convergence statistic; a quantitative criterion (e.g., a threshold on the deviation of the binned average from β_hat) would make the first necessary condition testable.","section":"Discussion, condition (i) and Figs. 3, 7"},{"comment":"The caption contains a duplicated phrase: 'We show both the power law ﬁt s the power law ﬁt with exponent...', which should be corrected.","section":"Fig. 13 caption"},{"comment":"The quantity r^2 is used throughout but never defined; stating that it is the coefficient of determination in log-log space would clarify the comparisons.","section":"Table II"},{"comment":"In the definition of SAMIs, the reference quantity Y0 is not defined; please specify what Y0 represents (e.g., the prefactor from the regression or a baseline city value).","section":"Section 3, Eq. (13)"}],"recommendation":"major_revision","confidential_remarks":"The reader's stress-test concern about in-sample evaluation is well founded and goes to the heart of the manuscript's predictive claim. The paper would be much stronger if the authors reframed the effective exponent and f(1/2) as descriptive in-sample summaries, added an out-of-sample or cross-validation analysis, and corrected the 'necessary conditions' language. The descriptive tomography tool itself is a useful addition to the scaling literature, so I do not recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is simple and genuinely useful: instead of fitting a global exponent, look at the local log-log slope between every pair of points and plot it against the size ratio. That tomography plot gives you a cheap visual check of whether a single scaling exponent is plausible, and it correctly flags cases where the standard fit is hiding something. The benchmark-city effective exponent, though, is a softer product.\n\nThe local exponent is elementary — it's just a chord slope in log-log space, and the paper acknowledges the connection to multiscaling. What's new is the packaging: the tomography plot, the benchmark-city minimization of the variance of beta_loc, and the f(1/2) accuracy fraction. Applied to the Leitao et al. datasets, these tools do what they promise: they confirm the clean cases (US GDP, roads), expose problems (UK rail stations, OECD patents), and give a plausible reading of some inconclusive cases. The writing is clear and the algebra in Eqs. 3-7 is correct.\n\nThe soft spots are in the claims around beta_eff. It is selected by minimizing the variance on the full dataset, and f(1/2) is then reported on the same data. That is in-sample fit quality, not predictive reliability. The paper says beta_eff lets you 'reliably' compute other cities' properties, but no out-of-sample test supports that word. The three necessary conditions in the Discussion are internal-consistency filters: if the prefactor varies systematically with P, beta_loc and the OLS fit will both converge to the same composite exponent, all three conditions pass, and the exponent is still not a stable predictor across regimes. The paper itself comes close to saying this for OECD patents. The threshold fits (a + bP) are also exploratory — extra parameters with comparable R^2 — and I read them as suggestive, not conclusive.\n\nNone of this sinks the paper. The descriptive diagnostic is solid and it is honestly presented. What needs fixing is the language of prediction and reliability. A referee should ask for an out-of-sample split or a careful restatement of beta_eff as a descriptive summary. My own verdict: conditional, but worth taking seriously.\n\nWho is this for? Anyone working on urban scaling or other scaling data who wants a quick visual sanity check beyond a fitted line. It is a useful paper for that audience, and it deserves a serious referee. I would probably not cite it in my own work, but I would bring it to reading group.","headline":"A simple, honest diagnostic for scaling data that overstates its own predictive reliability; worth peer review as a methods paper.","tokens_in":16571,"tokens_out":2045,"would_cite":false,"duration_ms":19350,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Pairwise local exponents give a fit-free tomography of scaling, with an effective exponent for reliable prediction.","keywords":["scaling laws","local scaling exponent","effective exponent","tomography plot","city scaling","nonlinearity","power-law fitting","urban scaling"],"falsifier":"Simulate a pure power law with constant prefactor and known $\\beta$; the test is whether the tomography plot converges to that $\\beta$ for large ratios. If it does not, or if a dataset with a known size-dependent prefactor still shows full convergence to one exponent, the central claim fails.","tokens_in":15588,"feed_emoji":"🏙️","tokens_out":8679,"duration_ms":72825,"temperature":0.7,"pith_summary":"Standard regression for scaling laws $Y \\sim P^\\beta$ can be inconclusive when noise is high and the population range is short. This paper replaces the global fit with a pairwise comparison: for any two cities of sizes $P_1$ and $P_2$, define the local exponent $\\beta_{\\mathrm{loc}} = \\log(Y_2/Y_1)/\\log(P_2/P_1)$ and plot it against the size ratio. If a single scaling law with constant prefactor holds, $\\beta_{\\mathrm{loc}}$ converges to the true exponent for widely separated sizes, so the plot works like a tomography scan of the scaling. From the plot one selects a benchmark city whose local exponents fluctuate least, giving an effective exponent that predicts other cities' $Y$ values, with $f(1/2)$ measuring the share of predictions within a factor of two. Applied to city datasets, the method clears up several 'inconclusive' cases and flags threshold effects and multi-regime behaviour that a single fit would miss.","feed_headline":"Two-city ratio reveals true scaling without fitting","feed_subtitle":"A tomography plot of local exponents settles whether city scaling is real or noise.","key_machinery":"The central object is the local scaling exponent $\\beta_{\\mathrm{loc}}(i,j) = \\log(Y_j/Y_i)/\\log(P_j/P_i)$, viewed as a function of the size ratio $r$. It turns the global fitting problem into a family of pairwise comparisons, so the 'tomography plot' of $\\beta_{\\mathrm{loc}}$ versus $r$ shows convergence to a stable exponent when scaling holds, noise-dominated scatter for small $r$, and systematic features such as threshold hyperbolas or drift when the power-law form fails. The benchmark city is selected by minimising the variance of $\\beta_{\\mathrm{loc}}$ around its average, and that average is the effective exponent $\\beta_{\\mathrm{eff}}$ used for predictions via $Y(j) = Y(i_{\\min})(P_j/P_{i_{\\min}})^{\\beta_{\\mathrm{eff}}}$; the fraction $f(1/2)$ of cities predicted within a factor of two quantifies the value of the exponent.","core_discovery":"The central claim is that the reliability of a scaling exponent can be assessed through local, ratio-based exponents instead of trusting a single regression. Defined as $\\beta_{\\mathrm{loc}} = \\log(Y_2/Y_1)/\\log(P_2/P_1)$, the local exponent is the slope of the chord between two cities in log-log space; plotting it against $r = P_2/P_1$ yields the 'tomography plot'. Under the scaling form $Y = a P^\\beta$ with constant prefactor and moderate noise, $\\beta_{\\mathrm{loc}}$ converges to the true $\\beta$ for large $r$ because the noise term is suppressed by $1/\\log r$, and diverges for cities of nearly equal size. The paper defines the benchmark city as the one minimising the variance of $\\beta_{\\mathrm{loc}}$ over all other cities, and its mean local exponent $\\beta_{\\mathrm{eff}}$ as the best single exponent for prediction. Three necessary conditions are proposed for trusting a fitted exponent: convergence of $\\beta_{\\mathrm{loc}}$ to the fit, consistency of $\\beta_{\\mathrm{eff}}$ with the fit, and $f(1/2)$ at least 50%; otherwise the single-power-law form is rejected.","pith_inferences":["A direct extension is to split city pairs into within-region and cross-region sets; if the effective exponent shifts between the two, the prefactor is not constant across regions and the single-exponent picture is misleading.","The decay of the noise envelope in the tomography plot is of order $1/\\log r$, so the persistence of a gap between $\\beta_{\\mathrm{loc}}$ and the fitted exponent for large $r$ could be used as a quantitative test statistic for prefactor variation or wrong functional form.","The benchmark-city criterion (minimum variance of $\\beta_{\\mathrm{loc}}$) is one choice among many; testing sensitivity of $\\beta_{\\mathrm{eff}}$ to alternative criteria (e.g., the city closest to the regression line) would show how much of the conclusion is carried by the selection rule.","Because $\\beta_{\\mathrm{loc}}(r)$ is effectively a chord slope in log-log space, comparing its spectrum across $r$ ranges parallels the multiscaling analysis used in growth kinetics, suggesting a way to look for a continuum of exponents rather than a single one."],"forward_implications":["Scaling claims can be checked without assuming a particular noise model: genuine scaling requires $\\beta_{\\mathrm{loc}}$ to converge to the fitted exponent as the population ratio grows.","The effective exponent provides a direct prediction rule: with the benchmark city's $Y$ value and $\\beta_{\\mathrm{eff}}$, any other city's $Y$ can be estimated, and $f(1/2)$ gives its expected accuracy.","Cases that standard tools call inconclusive can be resolved: external-cause deaths in Brazil and cinema capacity in Europe come out linear, cinema usage in Europe superlinear with large fluctuations.","The method can reject the simple power-law form: UK railroads and Brazilian AIDS cases show threshold-like behaviour, and OECD patents are not described by a single exponent.","If the three necessary conditions fail, the fitted exponent should be abandoned; this gives a practical diagnostic beyond $r^2$ and $p$-values."],"supporting_citations":[{"why":"Supplies the city datasets and the previous statistical classifications of each case as linear, nonlinear, or inconclusive that this paper's tools are tested against.","marker":"[37]"},{"why":"Establishes the scaling form $Y = aP^\\beta$ for cities and the division into $\\beta<1$, $=1$, $>1$, which is the underlying assumption for the local exponent.","marker":"[8]"},{"why":"Provides the SAMI residual concept that the paper contrasts with the benchmark-city approach, and the US patent dataset used in the patent analysis.","marker":"[20]"},{"why":"Shows that UK urban indicators depend on city definition and that many scale linearly, the reference result against which the paper's UK income and patent findings are checked.","marker":"[38]"},{"why":"Raises the objection that nontrivial exponents may be artifacts of using extensive quantities; the paper's GDP-per-capita analysis is designed to answer this.","marker":"[36]"},{"why":"Cited as the generalization of the scaling form with non-constant prefactor, marking the boundary of the constant-prefactor assumption the method relies on.","marker":"[43]"},{"why":"Provides the multiscaling concept that the paper compares $\\beta_{\\mathrm{loc}}(r)$ to, suggesting a possible connection to multiple scales.","marker":"[44]"}],"fun_headline_variants":["Tomography scan of scaling: no regression needed","Find true scaling exponent from two cities alone","Ratio-based exponent outsmarts noisy scaling fits","Local ratio exposes scaling breakdowns without fitting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire method assumes that the scaling law is exactly $Y = a P^\\beta$ with a single constant prefactor $a$; if the prefactor varies systematically with population or region, the local exponent will absorb that variation and the effective exponent will not predict reliably.","fun_headline_variants_meta":{"raw":{"variants":["Tomography scan of scaling: no regression needed","Find true scaling exponent from two cities alone","Ratio-based exponent outsmarts noisy scaling fits","Local ratio exposes scaling breakdowns without fitting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000547,"raw_usage":{"total_tokens":2643,"prompt_tokens":1003,"completion_tokens":1640,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":1583}},"tokens_in":619,"tokens_out":1640,"duration_ms":12598,"temperature":1.0,"reasoning_tokens":1583,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:12:00.771842+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a pure power law with constant prefactor and known $\\beta$; the test is whether the tomography plot converges to that $\\beta$ for large ratios. If it does not, or if a dataset with a known size-dependent prefactor still shows full convergence to one exponent, the central claim fails.","supporting_citations":[{"cited_title":"The plot of this number versus the population of cities is shown in Fig","cited_arxiv_id":null,"evidence_quote":"Supplies the city datasets and the previous statistical classifications of each case as linear, nonlinear, or inconclusive that this paper's tools are tested against."},{"cited_title":"Growth, innovation, scaling, and the pace of life in cities","cited_arxiv_id":null,"evidence_quote":"Establishes the scaling form $Y = aP^\\beta$ for cities and the division into $\\beta<1$, $=1$, $>1$, which is the underlying assumption for the local exponent."},{"cited_title":"Ur- ban scaling and its deviations: Revealing the structure of wealth, innovation and crime across cities","cited_arxiv_id":null,"evidence_quote":"Provides the SAMI residual concept that the paper contrasts with the benchmark-city approach, and the US patent dataset used in the patent analysis."},{"cited_title":"Tomography of scaling","cited_arxiv_id":"1908.11549","evidence_quote":"Shows that UK urban indicators depend on city definition and that many scale linearly, the reference result against which the paper's UK income and patent findings are checked."},{"cited_title":"From global scaling to the dynamics of individual cities","cited_arxiv_id":null,"evidence_quote":"Cited as the generalization of the scaling form with non-constant prefactor, marking the boundary of the constant-prefactor assumption the method relies on."},{"cited_title":"Philosophical writings","cited_arxiv_id":null,"evidence_quote":"Provides the multiscaling concept that the paper compares $\\beta_{\\mathrm{loc}}(r)$ to, suggesting a possible connection to multiple scales."}],"review_version":1}