{"id":"a498d053-603b-4abc-8fe6-c27f6e919f3e","arxiv_id":"2505.09399","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Biographical data from Wikipedia, fed into an elastic net model, produce out-of-sample historical GDP per capita estimates that explain 90% of the variance in known estimates for Europe and North America since 1300.","lead":"This paper trains a machine learning model on the birthplaces, deathplaces, and occupations of over half a million historical figures to estimate GDP per capita for European and North American countries and regions over the past 700 years. If the estimates hold up, they quadruple the available historical income data and give economic historians a new tool for studying long-run growth at regional scale.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The core risk is that the biography-to-GDP mapping learned on richer, better-documented locations may not transfer to the poorer, less-documented unlabeled locations the method targets; the proxy correlations are indirect and confounded by spatial autocorrelation and documentation bias.","rationale":"The paper's central claim requires biographical features to predict GDP per capita in locations that have no GDP data, many of which are poorer and less documented than the training sample. The authors themselves flag this representativeness gap in the Discussion. The out-of-sample R2 is computed on held-out countries that, on average, resemble the labeled set; it does not directly test the extrapolation domain. The external proxy checks are reassuring but incomplete: they are correlations, not calibration tests, and can be inflated by spatial autocorrelation and by the shared documentation gradient. The reader's weakest assumption correctly identifies this. I do not find a more fundamental flaw: the cross-validation design (withholding whole countries) is sound as far as label leakage goes, and the baseline comparison shows a genuine, if modest, incremental signal. The main corrective action is to demonstrate transfer to low-documentation locations directly. Because this is testable with the existing code and data, and the authors have already partially acknowledged the issue, a conditional verdict remains appropriate. If the proposed test fails, the central claim would be unverified or substantially weaker.","tokens_in":18514,"tokens_out":10584,"duration_ms":116009,"concrete_test":"Using the published code and data, stratify the 500 held-out country folds by a documentation proxy (e.g., number of biographies per 100,000 inhabitants in 2000, or 2000 GDP per capita). Report median out-of-sample R2 and MAE separately for the top and bottom terciles. A material drop (e.g., >10 points in R2) for the bottom tercile, or a model trained only on high-documentation locations performing markedly worse on low-documentation held-out countries, would show the mapping does not transfer to the unlabeled target set.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Discussion acknowledges that labeled locations have higher GDP per capita and more famous individuals than unlabeled locations. The central claim requires that the fitted relationship between biographical features and GDP is stable across this representativeness gap. The external validation (Fig. 3E-H) shows correlations with urbanization, height, wellbeing, and church building, and the SI reports similar correlations for labeled and unlabeled observations. But correlations do not establish that the mapping's slope and intercept are correct for poorly documented locations: unlabeled regions in rich, well-documented countries will exhibit high biography counts, high proxy values, and high predicted GDP simultaneously, producing correlation without testing calibration at the low end. Spatial autocorrelation can generate the same pattern even if the biographical signal is weak. Since the model is optimized on labeled locations, there is no direct evidence that the learned weights apply to the target distribution. This is the load-bearing assumption, and it is not settled by the reported tests.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a machine learning method, elastic net regression, that uses features derived from the biographies of 562,962 historical figures (births, deaths, occupations, popularity-weighted counts, SVD factors, economic complexity measures) to estimate historical GDP per capita for countries and regions in Europe and North America between 1300 and 2000. Training data consist of 1,336 known GDP per capita observations from the Maddison project and other regional sources; the model produces 4,364 out-of-sample estimates. The authors report out-of-sample R-squared of 90.1% versus 86.2% for a baseline model that uses lagged GDP and supranational region fixed effects, with mean absolute error improving from 29% to 22.6% of GDP per capita. They validate the estimates by reproducing the Little Divergence between northwestern and southern Europe, linking it to Atlantic trade, and by showing correlations with urbanization, body height, wellbeing, and church building activity. The paper also provides feature importance via Shapley values, robustness checks, and a public dataset and code repository.","tokens_in":18678,"tokens_out":2952,"duration_ms":31135,"significance":"If the estimates are reliable, this work materially expands the availability of historical GDP per capita data, especially at the regional level, and demonstrates a novel use of structured biographical data for economic history. The paper is commendable for shipping the full dataset, confidence intervals, and reproducible code, and for reporting out-of-sample performance against a nontrivial baseline. The external validations against multiple independent proxies and the replication of the Little Divergence are valuable and go beyond simple in-sample fit. However, the central claim depends on the stability of the biography-to-GDP mapping when applied to locations that are systematically poorer and less documented than the labeled training locations; the current evidence for this stability is correlational and indirect. The recursive use of model-generated lagged GDP as a feature also introduces a circularity risk that is not directly addressed. These issues do not invalidate the contribution but they need substantial additional analysis before the estimates can be taken at face value.","major_comments":[{"comment":"The paper acknowledges that labeled locations have higher GDP per capita and more famous individuals than unlabeled locations, but the external validation in Fig. 3E-H and SI 5.3 only shows that correlations with proxies are similar for labeled and unlabeled observations. Correlation similarity does not establish that the learned slope and intercept are correct for poorly documented locations: an unlabeled rich region will simultaneously have high biography counts, high proxy values, and high predicted GDP, producing correlation without testing calibration at the low end. I request a direct calibration test, for example comparing predicted versus actual GDP in held-out low-GDP locations (or in the lower decile of the predicted distribution) and reporting bias and coverage of the confidence intervals there.","section":"Discussion, \"countries and regions for which source data is available are not perfectly representative...\""},{"comment":"The lagged GDP per capita feature is a candidate predictor, and when source data are missing it is filled with the EN model's own estimates from the previous period (and similarly for the baseline model). This makes predictions for unlabeled location-periods recursively dependent on earlier model outputs, so the reported 90.1% R-squared partly reflects internal consistency of the imputation chain rather than independent information from biographies. The paper should report an ablation that excludes the lagged GDP feature entirely, and a variant that uses only source-data lags (without model-imputed lags), to quantify the marginal contribution of biographical features in the absence of recursive imputation.","section":"Materials and Methods, Elastic Net and Model performance"},{"comment":"The full model improves on the baseline by only about 4 percentage points in R-squared (86.2% to 90.1% at the median), and the paper itself reports in SI 5.5.8 that the method does not significantly improve growth-rate prediction. Given that the baseline already captures persistence and regional fixed effects, the incremental signal from biographical features is modest. The paper should clearly state the incremental contribution of the biography-derived features relative to lagged GDP and region dummies alone, and discuss whether the reported 22.6% MAE is economically meaningful for historical inference (e.g., by comparing it to the known uncertainty of Maddison estimates).","section":"Results, Model performance, Fig. 2C-D"}],"minor_comments":[{"comment":"There is a typo: \"the use us of HPI\" should read \"the use of HPI.\"","section":"Robustness of our estimates, sixth item"},{"comment":"The Shapley value equation contains a bracket typo: in the denominator, \"|F]!\" should be \"|F|!\".","section":"Materials and Methods, Shapley values"},{"comment":"The text says \"In a recent publication (37), we tested this proxy by randomly sampling 200 individuals,\" but reference 37 is listed as J. Mokyr, \"Mobility, Creativity, and Technological Development,\" which does not appear to be the authors' own work. Please correct the citation to the authors' relevant paper (or clarify that the validation was done in this manuscript).","section":"Data, Data on historical figures"},{"comment":"The text reports R-squared for the log of GDP per capita but MAE for exponentiated estimates; this distinction should be stated in the main text, not only in the Methods, to avoid confusion when interpreting the 90.1% and 22.6% figures.","section":"Model performance"},{"comment":"The claim \"more than quadrupling the availability of historical economic output data\" is supported by the ratio of 5,700 total observations to 1,336 source observations, but the wording could be made precise by stating that the number of location-year observations increases from 1,336 to 5,700.","section":"Introduction, paragraph 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful empirical contribution with a strong data-release component, but the core methodological claim—that the mapping learned on labeled richer areas transfers to unlabeled poorer areas—needs a direct calibration test and an ablation of the recursive lagged-GDP imputation. The current external validation, while suggestive, does not settle this. The fit to the journal's scope is reasonable, but the novelty relative to the authors' prior Pantheon-related work should be clarified in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Koch, Stojkoski, and Hidalgo have built a genuinely useful thing: a dataset of historical GDP per capita estimates for over 4,300 country/region-year combinations in Europe and North America, using elastic net on features derived from Wikipedia biographies. The feature engineering is thoughtful – HPI weighting, occupation counts, SVD factors, economic complexity – and the authors publish code and data. That alone is worth referee time.\n\nThe out-of-sample validation is real. An R² of 90.1% against a persistence-plus-region fixed-effects baseline at 86.2% is an improvement, but not a dramatic one; the biographies are adding about 4 percentage points of variance explained on top of \"the past predicts the present.\" The growth-rate result is the more telling soft spot: the model does not beat the baseline for growth, which undercuts the method's value for studying change over time.\n\nThe biggest substantive risk is exactly what the authors flag in the Discussion: labeled locations are richer and better documented than unlabeled ones. The external proxy correlations are reassuring but indirect. Correlations with urbanization, height, wellbeing, and church building do not calibrate the slope and intercept of the biography-to-GDP mapping in the poorly documented tail. Spatial autocorrelation and documentation bias could produce similar correlations even if the biographical signal were weak. The SI apparently shows comparable correlations for labeled and unlabeled observations, which helps, but it does not settle the calibration question. This is a load-bearing assumption, and the paper would be strengthened by more direct checks – e.g., leaving out a holdout set of poorer regions and testing calibration there.\n\nThere is also a circularity concern in the final estimates: when lagged GDP is missing, the model imputes it from its own earlier predictions. The out-of-sample test removes entire countries, so the headline R² is not directly contaminated, but the published dataset's recursive structure is worth flagging as a caveat for users.\n\nThe SI is not in the arXiv version, so several robustness claims cannot be checked here. That is minor for now, but it matters for a final verdict.\n\nOverall: a solid, honest paper that will be useful to economic historians and anyone doing long-run growth at regional scale. It deserves a serious referee and, I think, a conditional accept – the generalization and circularity issues should be addressed, but the dataset and method are worth publishing.","headline":"Useful dataset and honest modeling, but the biographical signal adds modestly over persistence and the transfer to poorly documented locations is the open question.","tokens_in":19212,"tokens_out":3289,"would_cite":true,"duration_ms":31713,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Biographical records of famous historical figures contain enough economic signal to estimate GDP per capita for previously unmeasured countries and regions over the past 700 years, with out-of-sample accuracy of 90.1 percent.","keywords":["economic history","GDP per capita estimation","machine learning","elastic net regression","biographical data","Little Divergence","historical economic development"],"falsifier":"Apply the model to a held-out set of regions for which new archival GDP estimates become available after the paper's training cutoff; if the full model's errors on those regions are no better than the persistence-plus-region baseline, or if it systematically overestimates the poorest regions, then the claim that biographical density generalizes to unlabeled locations would be undercut.","tokens_in":18290,"feed_emoji":"📜","tokens_out":10538,"duration_ms":90459,"temperature":0.7,"pith_summary":"This paper tries to establish that the geography of famous historical lives—where people were born, died, and worked—contains enough information about local prosperity that a machine learning model can estimate GDP per capita for European and North American countries and regions over the past 700 years, including times and places with no existing income data. Trained on 1,336 known historical GDP observations and biographical features for more than half a million individuals, the elastic net model explains $R^2=90.1\\%$ of the variance in out-of-sample income levels and misses by $22.6\\%$ of observed GDP per capita on average. The authors argue the estimates are trustworthy because they correlate with four independent proxies of economic development—urbanization, body height, well-being, and church building activity—and because they reproduce the Little Divergence, the faster growth of northwestern Europe relative to southern Europe after 1300. If this holds, the method more than quadruples the available historical income data for the studied regions and turns biographical records into a quantitative source for long-run economic history.","feed_headline":"Machine learning on biographies estimates 700 years of GDP per capita","feed_subtitle":"Historical birth and death records of famous people predict out-of-sample income with 90 percent accuracy.","key_machinery":"The machinery is a supervised feature-construction pipeline followed by elastic net regression. From a database of 562,962 biographies, each country and region receives, for every 50-year period, popularity-weighted counts of famous people born, died, immigrated, and emigrated, split into occupation categories; these counts are log-linearized and supplemented by the first five singular value decomposition factors for each of the four mobility types, economic complexity indices for each type, occupational diversity, and average ubiquity. The elastic net penalty ($\\ell^1$ plus $\\ell^2$ regularization) performs feature selection separately for five historical periods, and the previous period's GDP per capita enters as a persistence feature. The load-bearing mechanism is the correlation between where historically recorded individuals concentrate and local prosperity; the paper is explicit that this channel may be direct or indirect and does not need a causal identification.","core_discovery":"The central claim is that fine-grained biographical data can serve as a legitimate proxy signal for historical income in Europe and North America between 1300 and 2000. The paper builds an elastic net regression for five historical periods, using roughly 250 to 300 candidate features per period: popularity-weighted counts of births, deaths, immigrants, and emigrants broken down by occupation, plus singular value decomposition factors, economic complexity indices, occupational diversity, and the previous period's GDP per capita. In a validation scheme that withholds all observations for a random $20\\%$ of countries, the full model reaches $R^2=90.1\\%$ and a mean absolute error of $22.6\\%$ of observed GDP per capita, improving on a baseline that only uses persistence and supranational-region fixed effects (median $R^2$ rises from $86.2\\%$ to $90.1\\%$). The paper treats the correlation between biographical presence and wealth as sufficient, without requiring a causal direction: wealth may attract talent, talent may create wealth, or wealth may make talent historically visible. It validates the extrapolated estimates by reproducing the Little Divergence, showing that Atlantic-port regions drive much of it, and by finding similar correlations with urbanization, body height, well-being, and church construction for both data-covered and uncovered locations.","pith_inferences":["A natural extension would be to test the same pipeline in world regions beyond Europe and North America once biographical coverage is denser; the paper explicitly avoids this extrapolation, so success there is not established.","The finding that growth rates, unlike levels, could not be predicted better than the baseline suggests that biographical density tracks income levels more than income changes; future work with longer panels or better features might revisit this boundary.","The similar correlations between estimates and proxies for data-covered and uncovered locations offer a template for detecting selection bias in predictive historical reconstruction; an explicit reweighting or selection-correction step could turn that diagnostic into a fix."],"forward_implications":["The released dataset holds 4,364 out-of-sample GDP per capita estimates with confidence intervals, roughly quadrupling the number of location-year observations for Europe and North America between 1300 and 2000.","The estimates support within-country comparisons that were previously unavailable, such as Nuremberg versus other German regions in 1500, Amsterdam versus Rotterdam in 1600, and San Jose and Los Angeles versus Inner London in 1900.","Because the estimates reproduce the Little Divergence and tie much of it to regions with Atlantic ports, they strengthen the evidential base for accounts in which Atlantic trade and associated institutional change drove early modern European growth.","Correlations with body height, well-being, urbanization, and church building activity give independent, non-GDP evidence that the predicted income levels track material living standards rather than merely recording where famous people lived.","The bootstrapped confidence intervals make the new estimates usable in downstream quantitative history even though they are extrapolations."],"supporting_citations":[{"why":"Supplies the gold-standard historical GDP per capita observations used as training labels.","marker":"(27, 28)"},{"why":"Provides the cross-verified database of hundreds of thousands of historical figures with birth and death places and occupations.","marker":"(31)"},{"why":"Defines the Historical Popularity Index used to weight biographical counts by fame.","marker":"(32)"},{"why":"The Atlantic trade and institutional change account that the paper's Little Divergence finding reproduces and extends.","marker":"(49)"},{"why":"Provides the urbanization rate data used as an external validation proxy for the estimates.","marker":"(54)"},{"why":"Provides the 18th-century body height data used as an external validation proxy.","marker":"(55)"},{"why":"Provides the 1850 well-being composite used as an external validation proxy.","marker":"(56)"},{"why":"Provides church building activity data used as an external validation proxy for the 14th and 15th centuries.","marker":"(57)"},{"why":"Introduces the elastic net regression that performs feature selection and out-of-sample prediction.","marker":"(97)"}],"fun_headline_variants":["Biographies reveal 700 years of GDP per capita","Machine learning turns biographies into historical GDP data","Elastic net fills gaps in 700-year GDP history","Historical figures' lives predict past economies","Birth and death records estimate GDP back to 1300"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the link between how many historically recorded famous people are associated with a place and that place's income is the same in places without existing GDP data as in places with it, even though the data-covered places are richer and better documented on average.","fun_headline_variants_meta":{"raw":{"variants":["Biographies reveal 700 years of GDP per capita","Machine learning turns biographies into historical GDP data","Elastic net fills gaps in 700-year GDP history","Historical figures' lives predict past economies","Birth and death records estimate GDP back to 1300"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000277,"raw_usage":{"total_tokens":1701,"prompt_tokens":1046,"completion_tokens":655,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":662,"completion_tokens_details":{"reasoning_tokens":582}},"tokens_in":662,"tokens_out":655,"duration_ms":6508,"temperature":1.0,"reasoning_tokens":582,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:32:18.247613+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the model to a held-out set of regions for which new archival GDP estimates become available after the paper's training cutoff; if the full model's errors on those regions are no better than the persistence-plus-region baseline, or if it systematically overestimates the poorest regions, then the claim that biographical density generalizes to unlabeled locations would be undercut.","supporting_citations":[],"review_version":1}