{"id":"6d87b427-3868-42c5-8578-a5734ee58ca7","arxiv_id":"2607.26366","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A corrected Gaussian bootstrap makes KS and CvM specification tests for dyadic regression valid in both shared-node and independent-dyad regimes.","lead":"This paper develops specification tests — checks that a linear regression formula fits — for data organized in pairs, such as trade flows or social networks. The key result is a corrected bootstrap that gives valid p-values both when pairs that share a person or country are correlated and when all pairs are independent.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Implemented tests use a fixed grid X_G; Theorem 3 proves consistency only for continuum KS/CvM, so the abstract's 'omnibus' and fixed-alternative consistency claims are not established for the delivered grid-based procedure.","rationale":"I considered the reader's weakest assumption (AHK exchangeability). That is a genuine scope restriction, but it is explicit: the paper defines the problem for undirected, jointly exchangeable and dissociated arrays and states that directed links and general totally degenerate arrays are excluded. It is not an internal gap in the argument. The fixed-grid issue, by contrast, affects the paper's central advertised claim. Theorems 1-3 and all bootstrap theory are developed for the continuum index set X, but the algorithms and the empirical application use a fixed finite grid. The paper notices this in Section 3.2 but does not prove that the grid tests are consistent against all fixed alternatives, nor does the abstract qualify 'omnibus.' This is not a mathematical inconsistency in the proofs, but it is an unsubstantiated claim about the delivered procedure. The concrete test above would demonstrate the gap. The reader's verdict of CONDITIONAL is appropriate; my concern reinforces that condition (qualify the abstract or prove a growing-grid extension). Therefore verdict_should_be is UNCHANGED.","tokens_in":38584,"tokens_out":13944,"duration_ms":117548,"concrete_test":"Construct a DGP with a single regressor W~Unif[0,1], fixed grid X_G={0.1,0.2,...,0.9}, and null-mean model Y=1+W+g(W)+ε, where g(w)=c·sin(5πw) (c≠0). Then g integrates to zero over each interval between consecutive grid points, so Δ(x_g)=E[ε 1{W≤x_g}]=0 for all g, while sup_x |Δ(x)|>0. Run the implemented grid KS and CvM tests (Algorithms 1/2) with n large, e.g., n=500, and 1000 replications. If the abstract's 'omnibus' consistency claim holds, rejection rates should approach 1 for any c≠0. If the fixed-grid concern is valid, rejection rates stay near 5% (size), despite the continuum statistic sup_x |√n bR_n(x)| diverging. Also report results with a finer grid that includes a point inside a period of the sine, where power should appear.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (abstract) is that the resulting KS and CvM tests are consistent against fixed alternatives. The theoretical consistency result, Theorem 3, is stated for the continuum statistics T_KS = sup_{x∈X}|√n bR_n(x)| and T_CvM = n∫_X bR_n^2 dν. The implemented statistics in (27) replace X by a finite deterministic grid X_G. Theorem 3 does not cover these grid statistics: it requires max_g |Δ(x_g)|>0 or ∑ w_g Δ(x_g)^2>0. A fixed grid cannot guarantee this for every fixed alternative. There exist alternatives with Δ(x_g)=0 at every grid point while sup_X |Δ|>0 (e.g., a misspecification whose integral over each grid cell is zero, such as g(w)=c·sin(5πw) on [0,1] with grid {0.1,...,0.9}). For such alternatives, the grid sample statistic converges to 0 and the bootstrap critical values are O_p(1), so the rejection probability tends to at most the nominal level, not to 1. The paper explicitly acknowledges this gap in Section 3.2 ('Establishing consistency of the implemented grid tests additionally requires the grid to capture the departure') but provides no growing-grid or data-dependent-grid theorem, and the abstract retains the unqualified 'omnibus' claim. This is the most load-bearing concern because the headline contribution is an omnibus test, yet the delivered algorithm is not omnibus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops omnibus specification tests for linear conditional-mean models on undirected dyadic data. The dyad array is modeled by an Aldous–Hoover–Kallenberg representation, and the test statistic is a residual-marked empirical process indexed by lower orthants of the regressor distribution. The paper proves a uniform projection theorem that reduces the dyadic process, uniformly over a VC-type class, to its first-order node projections; shows that a raw node-multiplier bootstrap is valid under nondegenerate shared-node dependence but double-counts dyad-specific variation under independent dyads; and proposes a corrected Gaussian bootstrap based on an exact covariance decomposition. The main theorems state uniform weak convergence, asymptotic size of the bootstrap tests, consistency against fixed alternatives for continuum KS and CvM functionals, and local-power results in both variance regimes. The paper also reports simulations and an application to the Lazega law-firm network.","tokens_in":38961,"tokens_out":8656,"duration_ms":80052,"significance":"If the results hold, this is a genuinely useful contribution: it extends residual-based omnibus specification testing to dyadic data while explicitly handling the two variance regimes (node-dominated and independent-dyad). The paper contains substantial technical work: the exact projection identity, the Rademacher/chaos chaining for the uniform remainder, the plug-in stability via a deterministic VC enlargement, and the covariance decomposition (30) are all detailed and internally consistent. The corrected bootstrap is derived from an exact identity rather than curve-fitted, and the factor-of-two inflation of the raw bootstrap under independent dyads is a crisp, falsifiable prediction. The paper also honestly states its scope boundaries (undirected links, AHK structure, exclusion of general totally degenerate arrays). The main shortcoming is that the headline 'omnibus' consistency claim is not established for the implemented fixed-grid procedure, a gap the paper itself acknowledges but does not close.","major_comments":[{"comment":"Theorem 3 proves consistency only for the continuum statistics T_KS = sup_X |√n bR_n(x)| and T_CvM = n∫ bR_n^2 dν. The implemented statistics in (27) replace X by a fixed deterministic grid X_G. For any fixed grid, there exist fixed alternatives with Δ(x_g)=0 at all grid points but sup_X |Δ|>0 (e.g., X=[0,1], Δ(x)=c·sin(5πx), grid {0,0.2,0.4,0.6,0.8,1}). For such alternatives the grid statistic converges to 0 while bootstrap critical values are O_p(1), so the rejection probability does not tend to 1. The paper states in §3.2 that grid consistency would require the grid to capture the departure, but it provides no growing-grid or data-dependent-grid theorem. This is load-bearing because the abstract claims unqualified omnibus consistency for the delivered procedure.","section":"§3.2, §4, Eq. (27)"},{"comment":"The CvM consistency statement is overstated. Theorem 3 gives consistency of the continuum CvM functional only under the extra condition ∫ Δ^2 dν>0. For a fixed finite measure ν, there are fixed alternatives with Δ not identically zero but ∫ Δ^2 dν=0, so the CvM test cannot detect every fixed violation. The abstract's unqualified claim that 'the resulting Kolmogorov-Smirnov and Cramér-von Mises tests are consistent against fixed alternatives' should be qualified by the support condition on ν or by choosing ν with full support.","section":"Abstract, Theorem 3"},{"comment":"The local-power theorem relies on the high-level condition 'Assumptions 1-4 hold uniformly along the local alternatives in (40)' and on the covariance kernel of the centered process converging to that of G_R or G_D. No primitive conditions are given for this uniformity, and the proof in Appendix C simply states that the local perturbation is o(1). This makes the local-power claims conditional on an unverified uniformity assumption. The issue is secondary to the grid gap, but it should be formalized or explicitly deferred.","section":"§4.4, Theorem 6"}],"minor_comments":[{"comment":"The abstract has a typographical artifact: 'Cram\\'er' should be 'Cramér'.","section":"Abstract"},{"comment":"The local-power simulations use γ_n = h/√n, which corresponds to scenario (i). The independent-dyad local-power predictions of Theorem 6(b) are not simulated; adding a small ω=0 power panel would make the boundary behavior of the corrected test more concrete.","section":"§5.1"},{"comment":"The positive-semidefinite projection is described verbally. Please state explicitly that Π_+ replaces negative eigenvalues by zero, and note that the projection is applied to the estimated covariance matrix before the Gaussian draw.","section":"§4.2, Algorithm 2"},{"comment":"The discussion of the naive dyad bootstrap p-value 0.030 in the last row is clear, but consider labeling it explicitly as an invalid benchmark in the table notes to avoid confusion.","section":"§6, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The technical core is sound and the corrected bootstrap is a real contribution. The main problem is that the implemented grid-based procedure is not shown to be consistent against fixed alternatives, despite the abstract's omnibus claim. Since the paper itself identifies the missing grid condition, this is fixable by either adding a growing-grid theorem or carefully qualifying the claims. I would also ask for a more precise treatment of local-power uniformity. The paper is publishable after these points are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has one real contribution that holds up: the covariance-corrected Gaussian bootstrap. The exact finite-sample covariance decomposition in (30) is the right way to see why the raw node multiplier double-counts same-dyad variation when node effects vanish, and the correction that removes one copy of that term is clean and non-fitted. The uniform projection theorem and the associated chaining arguments also look sound; I read the proof of Lemma 2 carefully and the separate control of the Rademacher process and the Rademacher chaos is credible. The simulation design is sensible, and the corrected KS test indeed behaves better than the naive and raw alternatives in the reported scenarios. I would cite this paper for the bootstrap correction and for the careful treatment of the degenerate boundary.\n\nThe soft spot is exactly where the stress-test note points. Theorem 3 proves consistency for continuum KS and CvM statistics, but the implemented procedure uses a fixed grid X_G. The paper admits in Section 3.2 that consistency of the grid test requires the grid to capture the departure, but the abstract still calls the tests 'omnibus' and says they are 'consistent against fixed alternatives' without qualification. That is an overstatement. With a fixed finite grid you can always construct a misspecification that is zero at every grid point but nonzero elsewhere; for such an alternative the test has essentially no power asymptotically. This is not a minor nit: the headline contribution is an omnibus test, and the delivered algorithm is not omnibus on the theory presented. The fix would be a growing-grid result or a data-dependent grid with a proof that the separation condition holds with probability tending to one, or an honest qualification of the claims.\n\nOther concerns are smaller. The grid choice is arbitrary with no guidance beyond 'sufficiently fine,' and no replication code is provided for the simulations or the Lazega application, which makes the empirical results harder to verify. The directed-link exclusion is explicit and is a boundary rather than a hidden flaw. The multiple-testing point in the empirical specification search is real but minor, since the authors report all three specifications.\n\nThis paper deserves a serious referee. The bootstrap theory is valuable and the main gap is fixable; the authors should not be desk-rejected but should be pushed to reconcile the claims with the implemented grid procedure.","headline":"Strong bootstrap theory for dyadic specification tests, but the fixed-grid implementation does not deliver the advertised omnibus consistency.","tokens_in":39399,"tokens_out":1812,"would_cite":true,"duration_ms":20459,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A corrected Gaussian bootstrap makes dyadic specification tests valid in both dependence regimes.","keywords":["dyadic data","specification testing","residual-marked empirical process","exchangeable arrays","multiplier bootstrap","shared-node dependence","Kolmogorov-Smirnov test","Cramér-von Mises test"],"falsifier":"Simulate independent dyads with n=50, compute the empirical rejection rate of the raw node-multiplier KS test at 5% under the null: the paper predicts it will be close to zero (conservative) because its critical values are sqrt(2) too large, while the corrected test should be near 5%; observing the raw test near 5% would contradict Theorem 5(c).","tokens_in":38499,"feed_emoji":"📊","tokens_out":3747,"duration_ms":30690,"temperature":0.7,"pith_summary":"This paper builds omnibus specification tests for linear conditional-mean models on undirected dyadic data, where observations sharing a node are dependent. The authors prove that the dyadic residual-marked process can be reduced uniformly to latent first-order node projections, and that a raw node-multiplier bootstrap works when those projections are nondegenerate. However, that bootstrap double-counts the variation of each dyad when dyads are actually independent. An exact covariance decomposition yields a corrected Gaussian bootstrap that is valid in both regimes. As a result, Kolmogorov-Smirnov and Cramér-von Mises tests based on it are consistent against fixed alternatives and have nontrivial local power, with the corrected KS test showing the steadiest size control in simulations.","feed_headline":"Dyadic regression tests now work with or without shared-node dependence","feed_subtitle":"A corrected Gaussian bootstrap keeps nominal size in both regimes, where the raw node bootstrap fails by a factor of two.","key_machinery":"The Aldous-Hoover-Kallenberg representation Z_ij = tau(U_i, U_j, U_ij) with i.i.d. latent node and dyad variables; the exact projection identity decomposing the dyadic empirical process into a node-average term plus a degenerate second-order remainder; the uniform negligibility of that remainder over VC-type classes (via a Rademacher process and a decoupled Rademacher chaos); and the finite-sample covariance decomposition Var(sqrt(n) P_n a) = 2/(n-1) V0 + 4(n-2)/(n-1) V1 that separates same-dyad from shared-node covariation and motivates the corrected covariance estimator.","core_discovery":"The central claim is that, under the Aldous-Hoover-Kallenberg representation of jointly exchangeable and dissociated undirected dyadic arrays, the feasible residual-marked empirical process is asymptotically driven by first-order node projections, so a bootstrap that multiplies centered incident-dyad averages by node-level multipliers reproduces the null distribution only when that node component is nondegenerate. When dyads are independent, the node projection vanishes and the raw multiplier bootstrap has twice the correct covariance. The paper shows that subtracting exactly one copy of the same-dyad covariance (via the finite-sample decomposition 2/(n-1)V0 + 4(n-2)/(n-1)V1) yields a Gaussi","pith_inferences":["Although the paper restricts to undirected dyads, a similar covariance-correction argument could plausibly be extended to directed dyadic data by working with ordered pairs and separating the two endpoint roles; the factor-of-two counting would become a factor-of-one per direction.","The factor-of-two correction is a general property of any node-level bootstrap that averages each observation through two endpoints; it may transfer to other dyadic inference problems, such as two-way cluster-robust testing, whenever the cluster-level component vanishes.","A testable extension is to allow the node-level and dyad-level variance components to mix in a single asymptotic framework rather than treating them as two separate regimes; the corrected bootstrap might be shown to adapt continuously as the node projection shrinks.","The authors explicitly exclude more general totally degenerate exchangeable arrays where a Gaussian chaos component becomes leading; constructing valid specification tests for those arrays would require a different bootstrap, possibly based on dyad-level multipliers with estimated variance."],"forward_implications":["Practitioners can test the linear conditional-mean specification in dyadic data without knowing whether node-level or dyad-level variation dominates; the corrected KS and CvM tests are asymptotically valid in both regimes.","The raw node-multiplier bootstrap, while valid under nondegenerate shared-node dependence, is asymptotically conservative under independent dyads because it doubles the same-dyad variance; its local power drops accordingly.","The naive dyad-multiplier bootstrap is invalid under shared-node dependence and should not be used when the dependence regime is unknown.","In the Lazega law-firm network, the tests reject additive linear and quadratic models but find no remaining misspecification after adding a shared-office-by-shared-practice interaction, pointing to complementarities among pair characteristics.","The tests are implemented on a fixed finite grid with one OLS estimation and no bandwidth choice, making them straightforward to apply."],"fun_headline_variants":["Corrected bootstrap fixes dyadic specification tests","New tests for dyadic regression work under any dependence","Dyadic tests get a bootstrap that doesn't double-count","One corrected bootstrap, valid dyadic tests in both regimes","Specification tests for dyadic data with a robust bootstrap"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire projection reduction and bootstrap validity rest on Assumption 1: that the undirected dyadic array can be written as a symmetric function tau(U_i, U_j, U_ij) with independent node and dyad latent variables; if links are directed, or if node-level unobservables are correlated with the regressors, the node-projection reduction and the Gaussian bootstrap may fail.","fun_headline_variants_meta":{"raw":{"variants":["Corrected bootstrap fixes dyadic specification tests","New tests for dyadic regression work under any dependence","Dyadic tests get a bootstrap that doesn't double-count","One corrected bootstrap, valid dyadic tests in both regimes","Specification tests for dyadic data with a robust bootstrap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000388,"raw_usage":{"total_tokens":1855,"prompt_tokens":688,"completion_tokens":1167,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":1090}},"tokens_in":432,"tokens_out":1167,"duration_ms":10164,"temperature":1.0,"reasoning_tokens":1090,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T17:19:17.310302+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate independent dyads with n=50, compute the empirical rejection rate of the raw node-multiplier KS test at 5% under the null: the paper predicts it will be close to zero (conservative) because its critical values are sqrt(2) too large, while the corrected test should be near 5%; observing the raw test near 5% would contradict Theorem 5(c).","supporting_citations":[],"review_version":1}