{"id":"1e7756b1-81ef-48c4-95b0-0e9b9f62c48a","arxiv_id":"2412.16122","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A scale-invariant network model, calibrated by link density, reconstructs Dutch firm-to-firm supply networks and preserves density predictions across aggregation levels.","lead":"This paper tests a scale-invariant model for reconstructing firm-to-firm supply networks from partial data, using Dutch bank payment records. If it holds up, it becomes easier to map large supply chains from industry-level statistics rather than exhaustive firm-level data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reconstruction at scale from aggregate data is not demonstrated: the paper's own Discussion admits no unique disaggregation from aggregate IO tables to firm-level stripe fitnesses, yet the reported gains over density baselines require those true stripe fitnesses.","rationale":"The theoretical derivation of the scale-invariant functional form and the conditional-sampling machinery in Appendix A9 are sound, and the empirical comparison in Figures 2 and 6 is honest about mixed performance, so the paper deserves credit for a genuine proof of concept. The soft spot is the operational step: the model's advantage over density-corrected baselines appears to come from stripe fitnesses that encode product-level constraints, but the paper itself does not provide a constructive or validated way to obtain those stripe fitnesses from the aggregate data that the central 'at scale' claim presumes. The reader's weakest_assumption focused on the NACE-as-product proxy and edge independence; my concern is adjacent but more operational: even granting the proxy, the model cannot be applied in the main practical scenario until the aggregate-to-firm stripe disaggregation problem is solved. This does not invalidate the paper, but it means the central claim should be read as conditional on having firm-level sectoral strength information or a future disaggregation method. The reader already issued CONDITIONAL, and my concern reinforces that verdict rather than changing it, so UNCHANGED is the appropriate recommendation. The condition should be stated explicitly as 'requires a workable disaggregation of aggregate stripes', not merely 'better data needed'.","tokens_in":22646,"tokens_out":5388,"duration_ms":54504,"concrete_test":"Use the ABN dataset to construct a 2-digit NACE input-output table, then disaggregate it to firm-level stripe fitnesses using the two best available approximations from the paper (log-normal 'Distribution' sizes and 'Homogeneous' stripes), fit the global delta at the aggregate level, and evaluate firm-level out-degree KS distance and average nearest-neighbour degree against the empirical graph. If these metrics are no better than the dcIN baseline, the 'reconstruct at scale from aggregate data' claim fails in its intended application; if they are close to the true-stripe results, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the invariant multi-scale model 'reliably predicts important topological properties' and is 'a suitable candidate for reconstructing firm-to-firm networks at scale.' The load-bearing condition for that claim is the availability of firm-level stripe fitnesses (in- and out-strengths by NACE sector). Two places in the text make this condition fragile. First, in Appendix A8 ('Estimating the parameters from aggregate data'), the authors show that fitting the stripe model directly to an input-output table is ill-posed: when the observed network coincides with the stripe definitions, the density-matching parameter diverges (delta_alpha -> infinity), so the aggregate fit carries no information about firm-level density. Second, the Discussion explicitly states: 'the additive nature of the parameters does not give us a unique way to obtain the firm-level fitnesses from the estimated aggregate ones. How to solve this in practice is the subject for future work.' The supplementary 'Firm level information effects' section then finds that 'the heterogeneous information given by the true stripes is a clear advantage in terms of reconstruction accuracy.' Taken together, the setting in which the model is shown to outperform the density-corrected baseline is the setting where one already knows the true firm-level sectoral strengths. That is precisely the information a reconstruction method is supposed to provide when only aggregate tables are available. The paper's positive cross-scale results therefore support the mathematical scale-invariance of the functional form, but not the practical claim that the model can reconstruct firm-level networks 'at scale' from the type of public data the introduction motivates.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a multi-scale network model for reconstructing firm-to-firm supply networks, building on the scale-invariant functional form of Garuccio et al. (2023). The model represents each firm by in- and out-strength vectors across NACE sectors, and the link probability is a bilinear exponential form in these vectors. The authors test the model on two Dutch bank payment datasets, comparing density-corrected and stripe-corrected variants against gravity-model baselines in three scenarios: single-scale reconstruction, incorporation of a rest-of-the-world node, and estimation from aggregate input-output data. They claim the model reliably predicts important topological properties at scale and is a suitable candidate for reconstruction. The theoretical derivation of the scale-invariant form is sound, and the empirical evaluation is honest about limitations, but the headline claim is only partially supported by the reported results.","tokens_in":22925,"tokens_out":4583,"duration_ms":40859,"significance":"If the practical claims held, the model would be a significant contribution to economic network reconstruction because its parameters are invariant to node aggregation, which is a useful property when data are available at mixed scales. The paper provides a clear derivation of the functional form, open-source code, and a transparent comparison with existing gravity models. The empirical analysis is careful, and the inclusion of the rest-of-the-world and aggregate-data experiments is valuable. However, the main advantage of the model in reconstructing mesoscopic structure depends on having the true firm-level sectoral strengths, which the paper itself concedes are not uniquely recoverable from aggregate data. The practical scope is therefore narrower than the abstract suggests, and the reported error magnitudes (50-60% density errors in the ROW scenario, KS distances of 0.2-0.6 for degree distributions) do not fully support the term 'reliably predicts'.","major_comments":[{"comment":"The paper's central claim that the model is 'a suitable candidate for reconstructing firm-to-firm networks at scale' is not supported in the concrete scenario of reconstruction from aggregate data. In the aggregate-data scenario (main text, after Fig. 6), the stripe model (scIN) cannot be fitted when the observed network coincides with the stripe definitions, because the density-matching parameter diverges; Appendix A8 states that in this regime the estimated parameter carries no information about firm-level density. The main text therefore fits the density parameter with the dcIN model, which only constrains total link density. The topological properties that give the model its claimed advantage (degree distribution tails, average nearest-neighbour degree, ROC curves in Fig. 2) require the true stripe fitnesses, which are not available from an input-output table. The Discussion explicitly states that the additive nature of the parameters does not give a unique way to obtain firm-level fitnesses from aggregate ones and defers this to future work. The abstract should be qualified to state that the reconstruction advantages are demonstrated only when firm-level sectoral strengths are known.","section":"Working with aggregate data, Appendix A8"},{"comment":"The reported quantitative results do not support the unqualified claim of reliable prediction. In the ROW scenario, the absolute density error of the model is in the range 50-60% (main text, after Fig. 4), and in the aggregation scenario the KS distances between empirical and reconstructed degree distributions are 0.2-0.6 (Fig. 6b). The paper itself describes the ROW performance as 'somewhat disappointing'. These values are not necessarily disqualifying, but they are inconsistent with the abstract's 'reliably predicts important topological properties'. The abstract and conclusions should be revised to reflect the magnitude of the errors and the conditions under which the model is reliable.","section":"Handling the rest of the world"},{"comment":"The evaluation of the model's added value over the density-corrected baseline is circular in a practical sense: the stripe model outperforms the density model only because the true firm-level in- and out-strengths by NACE sector are supplied as inputs. When this information is degraded, as in the Uniform, Distribution, Total, and Homogeneous cases shown in Fig. 8, the model loses its advantage in ROC curves and average nearest-neighbour degree. Since the paper's stated goal is reconstruction from aggregate data, and the paper itself shows that firm-level fitnesses cannot be uniquely recovered from aggregate ones, the comparison with true stripe fitnesses does not establish the model's utility in the reconstruction setting. The paper should either compare against baselines using only information available in the intended application, or explicitly discuss what additional information beyond aggregate tables is needed for the model to deliver its claimed advantages.","section":"Appendix A5, Firm level information effects"}],"minor_comments":[{"comment":"The phrase 'our in and out degree sequences' appears to contain a typo; it should likely read 'the in and out degree sequences' or 'the observed in and out degree sequences'.","section":"Fig. 2 caption"},{"comment":"The term 'stripe' is used frequently but is not defined in the main text before the first use; a brief definition (firm-level in- and out-strengths disaggregated by product/sector layer) would improve readability.","section":"Main text after Eq. (5)"},{"comment":"Equation (A40) defines δ_k as a ratio that appears to depend on the specific pair (il, jl), yet the text refers to a global δ_k; the notation should clarify whether the relation is required to hold for all pairs and how a single δ_k is chosen in that case.","section":"Appendix A8, Eq. (A40)"},{"comment":"The sentence 'This is probably the reason why as the network becomes more dense, our estimation of δ becomes more unreliable as seen in figure 6a' is speculative; it would be clearer to state the connection to the well-posedness issue discussed in Appendix A8 and to refer explicitly to that appendix.","section":"Working with aggregate data, after Fig. 6d"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest in its limitations section, and the theoretical derivation is sound, but the abstract and conclusion oversell the reconstruction capability. The gap between the headline claim and the actual evidence may be a concern for the journal's readership. The model's scale-invariance property is interesting and could support a solid methodological paper, but the practical claims need to be substantially qualified, and the aggregate-data scenario needs to be reframed as a partial result. I recommend major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper is worth engaging with, but the abstract goes further than the results support. The genuinely new contributions are the application of the invariant multi-scale model to two firm-level payment networks, the rest-of-the-world handling, the aggregation scenario, and the conditional fine-graining machinery in Appendix A9. The scale-consistency derivations are correct, the equivalence with the stripe-corrected gravity model at a single scale is a useful check, and the code is released. The authors are also unusually candid about limitations.\n\nThe soft spots are in the gap between what the experiments actually show and what the abstract claims. The ROW scenario has 50-60% density errors; KS distances for degree distributions sit around 0.2-0.6. More fundamentally, the stress-test note is right: the case where the stripe model genuinely outperforms the density baselines is the case where the model is handed the true firm-level sectoral strengths. Appendix A8 shows fitting the stripe model directly to an input-output table is ill-posed, and the Discussion concedes there is no unique way to recover firm-level fitnesses from aggregate ones. The supplementary 'Firm level information effects' section confirms that true stripes are a clear advantage. So the paper demonstrates the functional form's scale-invariance, but not that you can reconstruct firm networks at scale from the public IO data the introduction motivates.\n\nNone of this kills the paper. The math is sound, the empirics are honest, and the conditional fine-graining stuff is a real methodological contribution. The NACE-as-product proxy is a known limitation and they acknowledge it. The main fix is framing: tone down 'reliably predicts' in the abstract and separate the theoretical claim from the open practical problem of disaggregation.\n\nI'd send this to a serious referee. It's a legitimate contribution that needs revision, not a desk reject.","headline":"Sound multi-scale reconstruction paper with honest empirics, but the abstract oversells 'reliably predicts' and the practical claim of reconstructing from aggregate data outruns the evidence.","tokens_in":23482,"tokens_out":3328,"would_cite":true,"duration_ms":28820,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["89.75.Fb","02.50.Tt","89.65.Gh"],"model":"deepseek-v4-flash","headline":"A single scale-invariant link formula reconstructs firm-to-firm supply networks from industry-level data.","keywords":["complex networks","economic systems","financial systems","supply networks","network reconstruction","multi-scale invariance","input-output tables","firm-to-firm networks"],"falsifier":"Use payment or VAT data with true product classification (CPA) for a subset of firms: if two firms with identical NACE-stripe strengths but different product mixes show systematically different connection patterns, the sector-proxy assumption fails. Alternatively, compute the product-layer aggregation relation from disaggregated product data; if the implied layer-aggregated parameter differs substantially from the global $\\delta$ used across scales, product-level scale invariance does not hold.","tokens_in":22445,"feed_emoji":"🔗","tokens_out":6391,"duration_ms":56367,"temperature":0.7,"pith_summary":"The paper argues that a recently proposed scale-invariant probabilistic model can reconstruct large firm-to-firm supply networks from aggregate information, and tests this on payment-based networks of Dutch firms. The model assigns each firm two vectors, sales and purchases split by industrial sector, and links firms with probability $1-\\exp(-\\sum_\\alpha \\delta_\\alpha s^{\\mathrm{out}}_{i,\\alpha}s^{\\mathrm{in}}_{j,\\alpha})$. Because this functional form is unchanged when nodes are merged, the same parameter $\\delta$ can be estimated from coarse industry-level data and then used to predict fine-grained firm-level structure. On the Dutch payment data the model matches degree distributions, average nearest-neighbour degree, and link rankings at least as well as the stripe-corrected gravity benchmark, and including a rest-of-the-world node removes a large systematic bias in density estimates. If correct, this offers a computationally tractable route to supply-chain reconstruction from public input-output tables.","feed_headline":"One scale-invariant formula rebuilds firm supply networks","feed_subtitle":"Fitting one parameter on coarse industry data predicts millions of firm-level links in Dutch payment networks.","key_machinery":"The load-bearing object is the multi-scale invariant functional $p_{ij}=1-e^{-\\theta_i^T B\\theta_j}$, where $\\theta_i$ stacks the firm's sector-level out- and in-strength vectors and $B$ selects the product-layer interaction; with $B$ chosen as a block matrix containing $\\mathrm{diag}(\\delta)$, this reduces to the stripe form. The argument is carried by the bilinearity of $\\log(1-p_{ij})$, which guarantees that if node parameters add under coarse-graining, the probability of an aggregate link is the same whether computed directly or by aggregating micro-links. This identity lets the authors calibrate $\\delta$ on a coarse graph of a few hundred nodes, predict a firm-level graph of roughly $3\\times10^5$ nodes, incorporate rest-of-the-world information as a single aggregate node, and condition fine-grained sampling on an observed macro graph at cost $M+N+1$ instead of $MN$.","core_discovery":"The central claim is that the scale-invariant link probability $p_{ij}=1-e^{-\\sum_\\alpha \\delta_\\alpha s^{\\mathrm{out}}_{i,\\alpha}s^{\\mathrm{in}}_{j,\\alpha}}$, with $s^{\\mathrm{out}}_{i,\\alpha}$ and $s^{\\mathrm{in}}_{i,\\alpha}$ the sector-level sales and purchases of firm $i$ and $\\delta_\\alpha$ a density parameter, is a sufficient and transferable description of directed production links. The model's defining property is that the logarithm of the non-link probability is bilinear in additive node vectors, so coarsening any partition of firms into sectors leaves the functional form intact. The paper shows this invariance is not only formal: fitting the density parameter on graphs aggregated by NACE digits and predicting the firm-level graph gives small density errors, the stripe version reproduces the observed neighbourhood structure better than density-only versions, and treating the unobserved part of the economy as one additional node corrects a roughly 2000% density overestimation that arises when outside trade is ignored.","pith_inferences":["The model's empirical success is entangled with the NACE-code proxy for products; testing with true product categories (CPA) would show how much of the stripe advantage comes from the model rather than from the classification.","The aggregation consistency is formal, and the paper itself finds an outlier when fitting on randomly grouped intermediate levels, suggesting that real partitions with skewed sector mixes may need partition-aware calibration.","The density-calibration choice means $\\delta$ absorbs all network sparsity; applying maximum likelihood to node fitnesses as free parameters is a natural next benchmark that could make the method work when only aggregate totals, not firm sizes, are known.","The conditional fine-graining result offers a scalable sampling algorithm: sample a sparse macro graph, then sample only within existing macro links, which could make full ensemble generation feasible for millions of firms."],"forward_implications":["Input-output tables at 2 to 5 digit NACE resolution are enough to calibrate the model, and the fitted parameter transfers to firm level with small density error across all tested years.","Including a rest-of-the-world aggregate node when only a subgraph is observed removes a systematic density overestimation, so partially observed or multi-country networks can be reconstructed without discarding outside-trade information.","The same functional form supports dyadic factors like geographic distance and conditional fine-grained sampling, extending the approach beyond the density-only version tested here.","Because the ensemble is not optimized for link prediction, its usefulness is measured by reproducing network statistics and functional equivalence rather than by exact link identification.","When sector information is available, the stripe-corrected version captures average nearest-neighbour degree and ranks observed links higher than density-only versions."],"supporting_citations":[{"why":"Supplies the multi-scale invariant model and the functional form adapted here.","marker":"[34]"},{"why":"Provides the stripe-corrected gravity benchmark, the firm-level reconstruction baseline, and the data construction for the Dutch payment networks.","marker":"[30]"},{"why":"Defines the random allocation model that the invariant model generalizes and contains as a special case.","marker":"[38]"},{"why":"Establishes the sparse-network Chung-Lu limit used to explain why density and stripe models coincide in degree reconstruction.","marker":"[41]"},{"why":"Extends the invariance principle underlying the model's consistency under arbitrary node partitions.","marker":"[35]"}],"fun_headline_variants":["Scale-invariant formula rebuilds firm supply networks from coarse data","One scale-free parameter predicts millions of firm links","Scale-invariant model reconstructs large supply networks","Coarse industry data yields firm-level links via scale invariance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument collapses if a firm's buying and selling behaviour is not fully captured by its total purchases and sales split by the seller's NACE industry, since sector codes are used as a proxy for the products actually exchanged; it also assumes links form independently given those strengths.","fun_headline_variants_meta":{"raw":{"variants":["Scale-invariant formula rebuilds firm supply networks from coarse data","One scale-free parameter predicts millions of firm links","Scale-invariant model reconstructs large supply networks","Coarse industry data yields firm-level links via scale invariance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000695,"raw_usage":{"total_tokens":3184,"prompt_tokens":1030,"completion_tokens":2154,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":2090}},"tokens_in":646,"tokens_out":2154,"duration_ms":15783,"temperature":1.0,"reasoning_tokens":2090,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:46:37.698956+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use payment or VAT data with true product classification (CPA) for a subset of firms: if two firms with identical NACE-stripe strengths but different product mixes show systematically different connection patterns, the sector-proxy assumption fails. Alternatively, compute the product-layer aggregation relation from disaggregated product data; if the implied layer-aggregated parameter differs substantially from the global $\\delta$ used across scales, product-level scale invariance does not hold.","supporting_citations":[{"cited_title":"Garuccio, M","cited_arxiv_id":null,"evidence_quote":"Supplies the multi-scale invariant model and the functional form adapted here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the stripe-corrected gravity benchmark, the firm-level reconstruction baseline, and the data construction for the Dutch payment networks."},{"cited_title":"macro” link that we did not sample before there can be no “micro","cited_arxiv_id":null,"evidence_quote":"Defines the random allocation model that the invariant model generalizes and contains as a special case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the sparse-network Chung-Lu limit used to explain why density and stripe models coincide in degree reconstruction."}],"review_version":1}