{"id":"020a4832-a232-4854-87d1-5d88467d70de","arxiv_id":"2607.27505","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":8.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A single observed network suffices for finite-sample valid confidence sets on the strategic-complementarity parameter in network formation models, via realization-wise sandwich inequalities and simulated critical values.","lead":"Many real networks—friendships, trade, contagion—form because links depend on other links, which makes standard statistics invalid. This paper proves that from a single observed network one can still build confidence intervals for the strength of such strategic link interdependence, without simulating equilibria, as long as the unobserved pairwise shocks are independent and the observed traits are exogenous.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 3 (i.i.d. pairwise errors) is the load-bearing premise; if εij carry agent-level heterogeneity, the simulated critical values no longer dominate the test statistic and coverage can fall below 1−α.","rationale":"I read the proof of Theorem 2 carefully and found no internal error. The bounding-by-c inequalities are realization-wise, the control statistic R_n is a function only of ε and Z, and conditional on Z, under Assumption 3, the per-cell empirical CDFs are independent uniform empirical CDFs. The simulation of critical values from cell sizes alone is therefore distributionally correct. The finite-B adjustment and the DKW-based closed-form critical value in Proposition 1 are also valid. The central claim is exactly as strong as Assumptions 1–4, and the paper does not overstate it. The single point where the argument can break in practice is Assumption 3: if pairwise errors are correlated through an unobserved agent component, the simulated control distribution no longer matches, and the finite-sample coverage guarantee is lost. This is a limitation of the maintained assumptions rather than a flaw in the theorem, and the paper explicitly acknowledges the standard nature of the assumption and its own companion work relaxing it only at the cost of dropping node-level covariates. The reader's weakest-assumption analysis identified the same issue. Since the theorem is internally sound and the concern is about the realism of an explicitly stated assumption, the appropriate verdict remains ACCEPT; no revision to the verdict is needed, though the paper could usefully add a sensitivity analysis to node-level unobserved heterogeneity.","tokens_in":34422,"tokens_out":15998,"duration_ms":168169,"concrete_test":"Simulate the Section 4.1 DGP (supermodular common-friend utility, β0 = −1, γ0 = 16, logistic base shocks) but generate εij = a_i + a_j + u_ij with a_i ~ N(0, σ_a²), u_ij ~ Logistic(0,1), for σ_a ∈ {0, 0.25, 0.5, 1}. Compute the semiparametric 95% confidence set of Section 3.1 on K = 200 replications at n = 400 and n = 1600. If empirical coverage drops below 0.95 (e.g., below 0.90 for σ_a = 0.5), Assumption 3 is genuinely load-bearing for the coverage claim; if coverage stays near 0.95, the procedure is more robust than the proof suggests.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 2 (Section 3.1) constructs C_semi_n by comparing the observable statistic Q̂_n(θ0) to a simulated critical value derived from the control statistic R_n. The proof in Appendix A applies the probability integral transform to claim that, conditional on Z, each cell's middle term F̂_n(c; zD) is the empirical CDF of m̂(zD) i.i.d. uniform variables, so R_n is distributionally identical to the simulated R^(b). This step uses Assumption 3 in two inseparable ways: (i) shocks in different dyads are independent, including dyads sharing a node, and (ii) shocks are independent of Z. If the true error process has an agent-level component—e.g., εij = a_i + a_j + u_ij with a_i a random effect—then within a cell the εij are dependent, the empirical CDF is not the empirical CDF of i.i.d. draws from Fε, and the simulated critical value is not the correct quantile of R_n. The sandwich dominance Q̂_n(θ0) ≤ R_n remains valid pointwise because it is algebraic, but the distribution of R_n is no longer the simulated one, so the nominal conditional coverage guarantee fails. The paper states Assumption 3 explicitly and cites the companion paper (Gao, Li, and Xu 2026) that relaxes it only by dropping node-level covariates. This is not an internal inconsistency; it is an untestable, empirically questionable premise on which the headline finite-sample guarantee rests.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a finite-sample valid inference procedure for the structural parameters of a strategic network formation model with pairwise stability, using a single observed network. The key device is a set of realization-wise 'bounding-by-c' inequalities: a formed link with an index below c certifies that the pairwise shock is below c, and an absent link with an index above c certifies the reverse. Averaging these inequalities over exogenous covariate cells sandwiches the empirical CDF of the i.i.d. pairwise errors between two observable envelopes. The paper uses this sandwich to derive a pathwise large-network identified set (Theorem 1) and, more importantly, finite-sample confidence sets whose test statistics are dominated by a control statistic that depends only on exogenous covariates and errors. The conditional distribution of this control statistic is exactly simulable, yielding semiparametric (Theorem 2), parametric (Theorem 3), and generalized nondecreasing-aggregation (Theorem 4) confidence sets with guaranteed coverage under Assumptions 1–4. A closed-form DKW-based critical value is also provided (Proposition 1). The methods are implemented in simulations with networks up to n = 10,000 and in two empirical applications, where positive strategic coefficients are reported.","tokens_in":34779,"tokens_out":17925,"duration_ms":196051,"significance":"If the results hold, this is a substantial contribution: it appears to be the first finite-sample valid inference procedure for a single strategic network without solving, simulating, or enumerating equilibrium network structures, and without restricting density or equilibrium selection. The realization-wise sandwich argument is elegant and the proofs in Appendix A are complete and self-contained. The computational tractability is a genuine strength: the procedure reduces to counting dyads in covariate cells, avoiding the combinatorial explosion of equilibrium-based methods. The paper is also careful to discuss several extensions and improvements (constrained thresholding, studentization, Berk–Jones weights). The main limitation, acknowledged by the authors, is Assumption 3 (i.i.d. pairwise errors), which is load-bearing for the simulated critical values; the paper should make the consequences of relaxing this assumption more prominent.","major_comments":[{"comment":"The abstract and introduction claim that the procedure restricts 'neither its density, nor the dependence structure induced by strategic interaction, nor the equilibrium selection mechanism.' This is accurate for the equilibrium-induced dependence, but it should be qualified: the finite-sample coverage guarantee rests on Assumption 3, i.i.d. pairwise errors. In the proof of Theorem 2, the probability integral transform is used to assert Fhat_n(c;z_D) = G_{z_D}(F_ε(c)), and this equality requires both independence of shocks across dyads and independence of shocks from Z. If the error process has an agent-level component, e.g., ε_ij = a_i + a_j + u_ij with a_i random, then within a cell the shocks are dependent, the simulated critical values are not the correct quantiles of R_n, and nominal coverage can fail. This is not an internal inconsistency, but the manuscript should explicitly flag","section":"Section 3.1 / Theorem 2"}],"minor_comments":[{"comment":"The definition of the semiparametric confidence set does not specify what happens when no cell meets the minimum sample size m, i.e., when \\hat Z_D is empty. A convention such as setting \\hat Q_n = 0 and the critical value to 0 (or requiring n large enough) would make the procedure fully well-defined for all n.","section":"Section 3.1, after Eq. (17)"},{"comment":"The 'pathwise identified set' Θ∞_I is a random object because p∞_L and p∞_U are limits along a realized sequence and may remain random. The paper explains this, but it would help to state explicitly that these limits are not consistently estimable from a single finite network, and that the role of Theorem 1 is to provide identifying restrictions along the path, whereas the finite-sample validity of Section 3 does not depend on this asymptotic object.","section":"Section 2.3 / Theorem 1"},{"comment":"The constrained threshold set C(z_D; θ, Z) is defined using the support of X. In practice, researchers may use a finite grid even when the support is unbounded. It would be useful to state explicitly that using any subset of thresholds preserves validity (because the inequalities hold pointwise) but may affect power; this is implicit but not stated.","section":"Section 3.3"},{"comment":"The table reports 'width' entries that are contaminated by grid truncation at small n; the text acknowledges this, but a table note would help the reader avoid over-interpreting the early rows. The 'sign' rate is the cleaner metric there.","section":"Section 4.3, Table 3"},{"comment":"There are several typographical issues: 'exploit use' in Section 1, 'sophisciated' in the introduction, 'certifies' in the abstract, and minor grammatical slips. A careful proofread is recommended.","section":"General"}],"recommendation":"minor_revision","confidential_remarks":"The paper is technically sound and the main theorems are convincing. The finite-sample inference results are a genuine contribution, and the computational scalability is attractive. My only substantive concern is that Assumption 3 (i.i.d. pairwise errors) is the lynchpin of the simulated critical values; the authors should be explicit that coverage is conditional on this assumption and that the method is not robust to common shocks or agent-level heterogeneity. This is a clarify-and-qualify issue, not a flaw in the proofs. I support publication after minor revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth the read. It gives the first finite-sample valid confidence sets for strategic network formation from a single observed network, with no restrictions on density, equilibrium selection, or dependence. The key device is a bounding-by-c sandwich: the observable statistic is bracketed by a control statistic built only from the i.i.d. pair errors and exogenous covariates, whose exact distribution is simulable. That is a clean, non-obvious move, and the proofs check out. The pathwise identification framing (random identified sets in the limit) is also novel and theoretically interesting.\n\nThe main soft spot, as the paper itself says, is Assumption 3: the pair errors are i.i.d. across all dyads, including dyads that share a node. The finite-sample coverage guarantee rests on this. If the true errors carry an agent-level component, the control statistic's distribution is not the simulated one, and coverage can dip below 1−α. This is not an internal contradiction—the theorem is correct conditional on the assumption—but it is a genuinely untestable premise from a single network. The companion paper relaxing it to fixed effects drops node-level covariates, which is a real cost. So the headline guarantee is exact, but only under a strong exogeneity and homogeneity assumption.\n\nThe paper is also honest about what it is not: no sharpness, conservative sets, no asymptotic distribution. The simulations show the method scales to 10,000 nodes and certifies the sign of the strategic coefficient; the empirical applications are plausible though not a substitute for validation on data with known structure. Missing code is a minor concern.\n\nThe citation pattern is fine. The bounding-by-c technique is credited to Gao and Wang (2026), and the paper extends it to a game-theoretic setting rather than claiming it from scratch. The novelty is real: the single-network finite-sample result is new, and the pathwise identification idea is worth citing.\n\nWho is this for? Anyone doing structural estimation of network formation with a single network. It gives a valid, computation-free inferential procedure under maintained assumptions, and it should be read carefully by the field. I'd send it to peer review, not desk reject. The stress-test point about Assumption 3 is the single most important thing for a referee to probe, but it is a limitation acknowledged in the paper, not a fatal flaw.","headline":"Finite-sample inference for single-network strategic formation is real and new; the i.i.d. pairwise error assumption is the acknowledged price of the guarantee.","tokens_in":35220,"tokens_out":1876,"would_cite":true,"duration_ms":19336,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper develops finite-sample-valid confidence sets for the structural parameters of a strategic network formation model using only one observed network, without restricting network density, dependence, or equilibrium selection.","keywords":["finite-sample inference","strategic network formation","pairwise stability","single network","endogenous covariates","equilibrium multiplicity","partial identification","bounding-by-c"],"falsifier":"Generate pairwise-stable networks from the same model but with a node-level random effect added to every dyad incident to that node, so dyad shocks are dependent; apply the semiparametric procedure to many such networks and measure empirical coverage. A drop below the nominal 1−α would show that the i.i.d. pairwise-error assumption, not the sandwich logic, is the operative ingredient.","tokens_in":34308,"feed_emoji":"🕸️","tokens_out":5046,"duration_ms":58890,"temperature":0.7,"pith_summary":"The paper claims that a single observed network formed by strategically interdependent agents can support exact finite-sample inference on structural parameters, provided pair-level shocks are iid and exogenous. Its key device is a family of \"bounding-by-c\" inequalities: a formed link whose latent index falls below a threshold c certifies that its shock is below c, and an absent link whose index exceeds c certifies that its shock is above c. Averaging these certificates within cells of exogenous covariates sandwiches the unobserved shock CDF between two observable, endogenous frequencies. A test statistic built from this sandwich is dominated by a control statistic that involves only exogenous covariates and shocks, whose exact conditional distribution can be simulated even though the equilibrium is intractable. The payoff is confidence sets with coverage at least 1−α at every sample size, with no restriction on density, dependence, or equilibrium selection.","feed_headline":"One observed network yields exact confidence sets","feed_subtitle":"Bounding-by-c inequalities make uncertainty simulable, so coverage holds at every sample size without solving equilibria.","key_machinery":"The bounding-by-c inequalities are the load-bearing object: for any dyad and any real threshold c, Yij·1{δij ≤ c} ≤ 1{εij ≤ c} ≤ 1 − (1−Yij)·1{δij ≥ c}. This trades the endogenous, equilibrium-determined threshold δij for a deterministic scan threshold c, so all equilibrium complexity is pushed into observable indicator functions while the middle term involves only the exogenous shock. Averaged within cells of exogenous covariates, the inequalities sandwich the empirical CDF of the shocks; conditional on Z, the probability integral transform makes each cell's shock CDF an empirical CDF of iid uniforms, so the control statistic's exact distribution can be simulated from cell sizes alone.","core_discovery":"The central claim is Theorem 2: under i.i.d. agent covariates, i.i.d. dyadic shocks, and exogeneity, the semiparametric confidence set C_semi_n covers the true parameter with conditional probability at least 1−α at every finite n, even when the observed network is one draw from an equilibrium with arbitrary multiplicity and unknown selection. The argument rests on realization-wise \"bounding-by-c\" inequalities that hold for every primitive, every equilibrium, and every selection. Averaging within covariate cells sandwiches the exogenous shock CDF between two observable endogenous frequencies; the middle term is a purely exogenous statistic whose exact conditional distribution is simulable via","pith_inferences":["The i.i.d. pairwise-shock assumption is untestable from a single network; correlated unobserved heterogeneity, such as node-specific effects, would invalidate the simulated critical values, so sensitivity analysis is advisable in empirical work.","Because the method deliberately avoids all dependence assumptions, it is conservative; combining its valid critical values with a central limit theorem under weak dependence could buy sharpness when such conditions are plausible.","The semiparametric critical value shrinks at the dyadic rate 1/n up to a log factor, so the cost of leaving the shock distribution nonparametric is visible in slower concentration relative to the parametric version, giving a practical ordering of assumptions against precision.","Since the identified set may remain random even in the large-network limit, the finite-sample confidence set—rather than a point estimate—may be the more honest inferential target for single large networks."],"forward_implications":["Confidence sets are exact in finite samples, so no weak-dependence, sparsity, or density assumptions are needed; coverage holds even when observable network frequencies fail to converge.","The sign of the strategic interdependence coefficient γ0 can be certified from a single network; simulations show 100% sign certification at n=400 in the parametric version and at larger n in semiparametric versions.","The hypothesis of no strategic interaction, γ0=0, is testable by checking whether zero lies in the confidence set, so the procedure delivers a no-interdependence test as a byproduct.","The same sandwich restrictions extend to many-network settings, nontransferable utility, and general subnetwork configurations, each yielding closed-form counting restrictions.","Computation scales with the number of dyads, O(n^2), rather than the size of the graph space, which is why networks of size 10,000 are routine in the simulations and applications."],"fun_headline_variants":["Single network, exact finite-sample confidence sets","Bounding-by-c gives simulable uncertainty from one network","One observed network is enough for finite-sample inference","Exact coverage with one network, no equilibrium solving"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is Assumption 3: the idiosyncratic pair-specific shocks εij are i.i.d. across all dyads; if shocks sharing a node are correlated, the simulated critical values no longer reproduce the control statistic's conditional distribution and the coverage guarantee fails.","fun_headline_variants_meta":{"raw":{"variants":["Single network, exact finite-sample confidence sets","Bounding-by-c gives simulable uncertainty from one network","One observed network is enough for finite-sample inference","Exact coverage with one network, no equilibrium solving"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000188,"raw_usage":{"total_tokens":1164,"prompt_tokens":736,"completion_tokens":428,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":379}},"tokens_in":480,"tokens_out":428,"duration_ms":4873,"temperature":1.0,"reasoning_tokens":379,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T06:36:49.422498+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate pairwise-stable networks from the same model but with a node-level random effect added to every dyad incident to that node, so dyad shocks are dependent; apply the semiparametric procedure to many such networks and measure empirical coverage. A drop below the nominal 1−α would show that the i.i.d. pairwise-error assumption, not the sandwich logic, is the operative ingredient.","supporting_citations":[],"review_version":1}