{"id":"80724f37-7cde-4d48-b34d-27d156bff7e3","arxiv_id":"2608.03716","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A new pipeline generates synthetic firm-level supply networks that match selected empirical network statistics and aggregate exactly to national input-output tables using only public data.","lead":"The authors introduce a fast, open-source method for creating artificial but realistic networks of firms and their supply relationships, using only public statistics and national input-output tables. Such synthetic networks could let economists run large-scale simulations of shocks without access to confidential business data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation only checks targeted moments; known non-targeted properties (reciprocity, clustering, assortativity) deviate by large factors, so the abstract's 'match known properties' overclaims until a holdout country's actual network is used for out-of-sample validation.","rationale":"The reader's weakest assumption identifies the sufficiency and universality of the Ref. [17] target moments. My concern is closely related but more specific: the paper's own Table 1 shows that several known properties not included in the targets deviate substantially, and no validation against any actual firm-level network outside the calibration set is provided. This is not an internal inconsistency in the algorithm — the construction is transparent, reproducible, and the IO aggregation is exact by design. It is a mismatch between the strength of the abstract's claim and the evidence supplied. The appropriate disposition remains conditional acceptance: the method is useful and likely to be adopted, but the headline claim should be tempered to 'matches the selected target moments' and the condition for full acceptance should be an out-of-sample comparison against a real network's full published statistics, using the concrete test above. I therefore keep the reader's CONDITIONAL verdict rather than moving it.","tokens_in":26328,"tokens_out":4587,"duration_ms":44363,"concrete_test":"Using the released open-source code, generate synthetic networks for Belgium and Ecuador (N=100,000, default rounded parameters, 2015 IOTs). Compare all properties reported in Ref. [17] that are NOT in the calibration loss: reciprocity, average and global clustering, degree assortativity, TLS(k_in~k_out), Hill exponent of influence, share of firms with k_out=0, and average path length. If the synthetic values fall outside the empirical ranges in Ref. [17] by the same factors seen in Table 1 (e.g., reciprocity ~0.006 vs 0.03-0.05), then the abstract's claim to match 'known properties' is unsupported and the paper should be revised to claim only the targeted moments. If the untargeted statistics fall inside the reported empirical ranges, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (abstract; Section 2.1) is that synthetic networks 'match both the known properties of firm-level supply networks and the properties of aggregated input-output tables.' The IO-table part is guaranteed by construction (Methods, Section C.2: after rescaling, aggregation is exact). The firm-level part, however, is only demonstrated for the particular moments in the loss function (Eq. 10): degree and strength tail exponents, selected correlations, OLS/TLS slopes, and log-variances. These are calibration targets drawn from Ref. [17], so matching them is expected and does not independently validate realism. What breaks the unconditional version of the claim is Table 1: known properties that were not targeted are far off — reciprocity 0.00671 vs [0.03,0.05], average clustering 0.0987 vs [0.19,0.28], TLS(k_in~k_out) 0.496 vs 0.7, Corr(s_in~k_in) 0.592 vs 0.75, and Hill exponent of influence 1.45 vs [1.2,1.3]. The paper acknowledges in Section 2.2 that the configuration model cannot reproduce these 'important features of real networks.' Thus the strongest claim, as written, is not established: the method matches a selected set of moments, not 'the known properties.' The robustness checks in Section D.2 compare synthetic Belgium and Brunei networks only to input-output tables, not to actual firm-level networks, so there is no out-of-sample evidence that the non-targeted gaps are harmless or that the target set is sufficient. The honest Discussion section proposes a future test using unpublished VAT data; that confirms the need for independent validation rather than supplying it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a method for generating synthetic firm-level supply networks from publicly available data. The pipeline has four stages: sampling in- and out-degree sequences from calibrated Burr XII distributions, reordering them to induce the empirical in-out degree correlation, drawing a binary network from the configuration model, and then assigning weights via a gravity-like formula involving firm fitnesses and degrees, followed by industry assignment and rescaling so that the aggregate matches a target OECD input-output table. The four weight exponents are calibrated by Optuna against selected degree, strength, correlation, OLS/TLS, and variance targets from Ref. [17]. The main demonstration is a 100,000-firm Hungary 2015 network, with robustness checks across countries, network sizes, and parameter perturbations.","tokens_in":26863,"tokens_out":2595,"duration_ms":26773,"significance":"If the claims are appropriately scoped, the paper makes a useful and timely contribution. It is fully reproducible (public code, public data only), computationally scalable (O(N) with sparse tensors), and the industry-level matching is exact by construction. The transparent, modular architecture with separate stages for topology and weights, the careful ablation study of Burr XII discretization/truncation effects, and the provision of default parameter values are all valuable. The paper also frames a useful benchmark for generative graph methods. However, the significance is currently moderated by the gap between the abstract's claim that the networks match 'the known properties of firm-level supply networks' and the actual evidence, which is limited to a selected set of calibration targets.","major_comments":[{"comment":"The abstract's claim that the synthetic networks 'match both the known properties of firm-level supply networks and the properties of aggregated input-output tables' is stronger than what is demonstrated. Table 1 shows that several non-targeted, well-known network properties deviate substantially from empirical benchmarks: reciprocity is 0.00671 versus [0.03, 0.05], average clustering is 0.0987 versus [0.19, 0.28], and the Hill exponent of influence is 1.45 versus [1.2, 1.3]. Section 2.2 acknowledges that the configuration model cannot reproduce these 'important features of real networks.' This is a load-bearing issue because the central claim of the paper is precisely about matching known firm-level properties. The claim should be revised to state that the method matches a selected set of calibration targets, or the method should be extended to incorporate the missing structural properties.","section":"Abstract; Section 2.2, Table 1"},{"comment":"The validation is circular with respect to the targeted statistics. The loss function in Eq. (10) directly minimizes squared discrepancies to the Hill exponents of strengths, the strength-degree correlations, the specified OLS and TLS slopes, and the log-variances of strengths. Matching these statistics in Table 1 is therefore a fitting outcome, not independent evidence of realism. The robustness analysis in Section D.2 compares synthetic Belgium and Brunei networks only against input-output tables, not against actual firm-level networks for those countries. Since the target values themselves come from Ref. [17] and the non-targeted statistics in Table 1 already deviate from known benchmarks, an out-of-sample comparison against a real firm-level network (e.g., a holdout country with VAT data) is needed to establish that the target moment set is sufficient. In the absence of such a test, the Discussion's claim that the method is 'straightforwardly testable' remains a promise rather than a demonstrated property.","section":"Section C.3, Eq. (10); Section D.2"},{"comment":"The paper relies on the universality of the empirical targets: Section 2.1 states 'We rely on the findings in [17] to determine our reconstruction targets,' and Table 2 fixes target values to specific numbers from a small number of countries. If these moments are not universal or not representative of the country being modeled, the synthetic network will confidently reproduce the wrong structure. This is a correctness risk rather than an internal inconsistency, but it should be addressed concretely, for example by reporting sensitivity of the generated networks to plausible variations in the target values, or by validating against a held-out country's actual network. The current robustness checks vary the input-output table but keep the firm-level targets fixed.","section":"Section 2.1; Section A, Table 2"}],"minor_comments":[{"comment":"There is a typo: 'we assign each fitnes variable an exponent' should read 'fitness variable.'","section":"Methods, 'Assignment of weights assuming theta is known'"},{"comment":"Inconsistent typesetting of 'V AT' with a space appears in the Introduction, Section 2.1, and the SI; it should be 'VAT.'","section":"Throughout"},{"comment":"The sentence 'the parameters appear relatively precisely estimated, baring an identification problem' uses 'baring' where 'barring' is intended.","section":"SI Section C.3"},{"comment":"The row 'IOT RMSE (100 MUSD)' reports 8.01e-04 without an explicit unit or normalization; the caption and the main text should clarify whether the RMSE is in units of 100 million USD and how the matrix scale is handled.","section":"Table 1"},{"comment":"The text states that 'as network size increases, average path length rises' and Fig. 12 shows boxplots, but the main text does not quantify the change; adding a sentence with representative values would improve interpretability.","section":"Section D.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest in its Discussion about limitations and even proposes a future test, which makes the overclaim in the abstract fixable by revision. The central methodological contribution is sound and the code is public. My main concern is not the method itself but the evidence for the 'match known properties' claim; a revised version that explicitly scopes the claim, adds a holdout validation against a real firm-level network, or incorporates the missing structural targets would be a solid contribution. I did not find evidence of a load-bearing technical error in the parameter estimation or rescaling procedure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe genuinely new thing in this paper is a complete, open-source pipeline for synthetic firm-level supply networks using only public data: calibrated Burr XII degree sequences, configuration-model topology, fitness-weighted edge attribution, and exact rescaling to OECD input-output tables. That specific combination — public inputs, micro-moment targets, and exact IOT aggregation — is something I have not seen before, and the paper does it cleanly. The SI is unusually thorough: ablations on discretization, truncation, and reconciliation; proofs for the Burr XII parameter mapping; stability checks across country IOTs and network sizes. The code is public and the method is fast enough for practical use.\n\nThe soft spots are real but not fatal. The abstract's claim that the networks \"match the known properties\" is too strong. Table 1 shows that several important non-targeted properties deviate by large factors: reciprocity is about 0.007 versus an empirical range of [0.03, 0.05], average clustering is around 0.1 versus [0.19, 0.28], and the TLS slope for degrees is 0.5 versus 0.7. The authors acknowledge in Section 2.2 that the configuration model cannot reproduce these features, which is candid, but the abstract and introduction still overstate the match. Second, the validation is partly circular: most of the statistics used to judge the network are exactly the terms in the Optuna loss, so the close fit to tail exponents and correlations is expected. The real test would be a holdout country with an actual firm-level network; the robustness section only checks IOT consistency across countries, not firm-level properties. The Discussion explicitly proposes such a test, which is honest but means the evidence is not yet there.\n\nOne smaller issue: the empirical targets are drawn from a review that includes a co-author (Ref. [17]). That alone is not a problem, but the review summarizes statistics from only a handful of countries, so the targets may not be universal. The paper acknowledges some arbitrariness in choosing target values.\n\nOverall, this is a useful, well-engineered methodological contribution that will likely become a baseline for later work. The math and code look solid. I would send it to a serious referee, with the expectation that the authors temper the claims and ideally add an out-of-sample firm-level comparison. If I were refereeing, I would recommend conditional acceptance.","headline":"Useful and well-engineered pipeline, but the abstract overclaims: the match is to a selected set of moments, not all known network properties.","tokens_in":27242,"tokens_out":3149,"would_cite":true,"duration_ms":26295,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A public-data method generates synthetic firm-level supply networks that match both micro statistics and national input-output tables.","keywords":["synthetic populations","supply networks","production networks","input-output tables","configuration model","Burr XII distribution","network reconstruction","agent-based models"],"falsifier":"Take a country whose full firm-to-firm VAT network is known but not yet published, generate a synthetic network using only the public input-output table and the stated targets, then compare the two networks on untargeted statistics—reciprocity, clustering, assortativity, and sectoral concentration; if the observed real values fall far outside the synthetic distributions, the claim that public data suffice to reproduce plausible firm-level structure is refuted.","tokens_in":26091,"feed_emoji":"🏭","tokens_out":10803,"duration_ms":88401,"temperature":0.7,"pith_summary":"The paper proposes a fully public-data method for creating synthetic firm-level supply networks that simultaneously match the statistical fingerprints of real firm-to-firm networks and the aggregate flows of national input-output tables. The authors argue that the bottleneck in high-resolution economic modeling—confidential transaction data—can be bypassed because the distinctive regularities of production networks, such as fat-tailed degree distributions and strength–degree correlations, are enough, when combined with industry-level tables, to pin down plausible micro structure. They demonstrate the pipeline on a 100,000-firm network for Hungary in 2015 and report that the synthetic network reproduces the targeted degree, strength, and weight properties and aggregates exactly to the input-output table. If the method generalizes as claimed, researchers can generate realistic firm networks for any country with public input-output data, initialize agent-based models at one-to-one scale, and run shock-propagation experiments without access to confidential records.","feed_headline":"Public data alone can now generate realistic firm supply networks","feed_subtitle":"A new pipeline reproduces micro-level network statistics and aggregates exactly to national input-output tables.","key_machinery":"The load-bearing objects are the Burr XII distribution—a three-parameter heavy-tailed distribution whose asymptotic tail exponent is the product of two parameters, used to draw in- and out-degree sequences—and the gravity-like weight formula that combines node degrees with lognormal 'fitness' latent variables. Sampling degrees separately and then reordering one sequence with rank-dependent lognormal noise induces the empirical in-out degree correlation of about 0.55 while preserving marginals. The configuration model turns the degree sequences into a binary directed graph; the weight formula then assigns positive values to existing edges; and the rescaling identity $W_{ij} = W^{\\mathrm{init}}_{ij} \\cdot \\mathrm{IOT}_{g_i g_j} / \\sum_{f \\in g_i, h \\in g_j} W^{\\mathrm{init}}_{fh}$ guarantees the industry-level sums equal the target table. The final mechanism is the calibration loop, which searches over the four exponents to match the targeted weighted-network statistics.","core_discovery":"The central claim is that a directed, weighted firm-level production network can be generated in four stages—Burr XII degree sampling, degree-sequence coupling, configuration-model wiring, and weight assignment with input-output rescaling—so that the final object matches a chosen set of firm-level statistics while its industry-level aggregation equals a given input-output table exactly. The weight assignment uses a gravity-like formula on node degrees and latent fitnesses, $W^{\\mathrm{init}}_{ij} = (f^{\\mathrm{out}}_i)^{\\theta_{f,\\mathrm{out}}} (f^{\\mathrm{in}}_j)^{\\theta_{f,\\mathrm{in}}} (k^{\\mathrm{out}}_i)^{\\theta_{k,\\mathrm{out}}} (k^{\\mathrm{in}}_j)^{\\theta_{k,\\mathrm{in}}} A_{ij}$, and the rescaling step multiplies each edge weight by the ratio of the target industry-pair flow to the current aggregated flow, which guarantees exact agreement at the industry level. A calibration loop adjusts the four exponents to minimize squared discrepancies between synthetic and empirical tail exponents, correlations, regression slopes, and log-variances. The paper reports that the calibrated network matches the targeted micro moments, reproduces plausible univariate and joint distributions, and remains stable across network sizes from 10,000 to 500,000 firms and across different countries' input-output tables.","pith_inferences":["The same pipeline could plausibly be extended to worker–firm and household–firm bipartite networks once public moments from administrative or scanner data become available; the paper only gestures at this possibility.","Because the method relies on a small set of cross-country moments, a strongly non-universal country would receive a plausible-looking but wrong network that still passes all stated targets; this is an inference from the paper's reliance on the survey statistics.","Since the configuration-model step yields lower reciprocity and clustering than real networks, downstream applications sensitive to local triadic structure may need a topology-preserving extension; the paper notes the limitation but does not address it.","A direct stress test would compare crisis simulations on synthetic versus observed networks in countries with VAT data, a comparison the paper suggests but leaves to future work."],"forward_implications":["Researchers can generate a plausible firm-level supply network for any country with a published input-output table, using only public data, so confidential administrative records are not needed.","Macroeconomic and agent-based models can be initialized at a one-to-one scale with realistic firm heterogeneity, enabling stress tests and shock propagation studies that previously required private transaction data.","Because the rescaled network aggregates exactly to the input-output table, model outcomes remain consistent with national accounting aggregates by construction.","The same calibrated weight parameters can be reused across countries and network sizes without re-running the optimization, making the pipeline cheap to deploy.","The method scales roughly linearly with the number of firms, up to 500,000 firms, so it remains tractable for large-scale economic models."],"supporting_citations":[{"why":"Supplies the empirical degree, strength, tail, and correlation statistics used as reconstruction targets.","marker":"[17]"},{"why":"Provides the public industry-level input-output tables used as the meso-scale aggregation target.","marker":"[23]"},{"why":"Provides the hyperparameter search procedure used to calibrate the weight exponents against the empirical targets.","marker":"[56]"},{"why":"Supplies the power-law tail estimator used to measure and target tail exponents.","marker":"[57]"},{"why":"Gives the configuration-model properties that explain why reciprocity and clustering are not reproduced.","marker":"[43]"},{"why":"Provides the input-output aggregation algebra used in the rescaling step.","marker":"[58]"}],"fun_headline_variants":["Synthetic supply networks generated from public data","Fast synthetic firm networks from open data","Realistic supply chains from public data alone","Synthetic networks matching input-output tables"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the empirical statistics taken from the survey of firm-level networks are the right and sufficient summary of a real production network, and that input-output tables and VAT-based transaction data are compatible enough that rescaling to industry totals does not distort the micro properties.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic supply networks generated from public data","Fast synthetic firm networks from open data","Realistic supply chains from public data alone","Synthetic networks matching input-output tables"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000212,"raw_usage":{"total_tokens":1391,"prompt_tokens":889,"completion_tokens":502,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":449}},"tokens_in":505,"tokens_out":502,"duration_ms":4633,"temperature":1.0,"reasoning_tokens":449,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:46:57.195935+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a country whose full firm-to-firm VAT network is known but not yet published, generate a synthetic network using only the public input-output table and the stated targets, then compare the two networks on untargeted statistics—reciprocity, clustering, assortativity, and sectoral concentration; if the observed real values fall far outside the synthetic distributions, the claim that public data suffice to reproduce plausible firm-level structure is refuted.","supporting_citations":[{"cited_title":"Firm-level production networks: What do we (really) know?Journal of Economic Dynamics and Control, 187:105313, 2026","cited_arxiv_id":null,"evidence_quote":"Supplies the empirical degree, strength, tail, and correlation statistics used as reconstruction targets."},{"cited_title":"Development of the oecd inter country input-output database 2021.OECD Science, Technology and Industry Working Papers, 2021(14), 2021","cited_arxiv_id":null,"evidence_quote":"Provides the public industry-level input-output tables used as the meso-scale aggregation target."},{"cited_title":"Optuna: A next-generation hy- perparameter optimization framework","cited_arxiv_id":null,"evidence_quote":"Provides the hyperparameter search procedure used to calibrate the weight exponents against the empirical targets."},{"cited_title":"Cambridge university press, 2009","cited_arxiv_id":null,"evidence_quote":"Provides the input-output aggregation algebra used in the rescaling step."}],"review_version":2}