{"id":"0b5f0db0-4590-4e9d-87e3-110703c7cded","arxiv_id":"2509.07343","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A corrected 2SLS estimator for peer effects is valid when network links are randomly misclassified, using new instruments and closed-form misclassification-rate estimates.","lead":"Link misclassification, where survey-reported connections miss real ties or invent fake ones, is shown to break standard two-stage least squares estimation of peer effects. The paper proposes a corrected estimator that recovers the misclassification rates from the data and adjusts the network measure, then applies it to the classic Indian microfinance study.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Adjusted 2SLS consistency rests on A3, which requires link formation to be exogenous to outcome errors; when A3 fails, the misclassification correction cannot restore consistency, and the paper's 'no link-formation model' claim does not address this.","rationale":"The reader identified A3 as the weakest assumption, and this is indeed the most load-bearing concern. Lemma 1 is the foundation of the entire adjusted-2SLS construction; if A3 fails, the exogeneity of X in the adjusted structural form and the validity of the new instruments both collapse, regardless of how well misclassification rates are estimated. A4 and the regularity conditions are also important, but they are secondary: A4 concerns the structure of measurement error, and violations there might be addressed by other instrument choices, whereas a violation of A3 is fatal to the identification strategy. The paper itself flags A3 as 'rules out endogeneity in link formation' in Section 3.1, so this is not an unanticipated oversight, but it is still the point at which the central claim is least secure. The reader's conditional verdict already reflects this concern; my stress-test does not change that verdict.","tokens_in":32718,"tokens_out":5348,"duration_ms":69002,"concrete_test":"Run a Monte Carlo DGP that violates A3 while keeping the misclassification structure of Section 5: draw a group-level or dyadic unobservable U that enters both link formation (e.g., G_ij = 1{γ|X_i-X_j| + δ(U_i+U_j) ≥ c}) and the outcome error (ε_i = ρU_i + e_i), generate misclassified H as in Section 5, and estimate the paper's adjusted 2SLS on many replications. If the average estimate of λ deviates from the true value by more than, say, 10% of |λ|, the central consistency claim fails exactly where A3 fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central consistency claim depends directly on Lemma 1, which asserts E(v|G,X)=0. The proof of Lemma 1 uses E(ε|G,X,H)=0 (A3) to eliminate E(WMε|G,X) and E(GMε|G,X). If A3 fails, i.e., actual links G are correlated with structural errors ε, then E(Gy|G,X) contains E(ε|G,X)-type terms that the misclassification adjustment W cannot remove, because W only corrects for the difference between H and G, not for the endogeneity of G itself. The paper's selling point that it 'does not require structural modeling of link formation' is therefore narrower than it appears: it avoids specifying Pr(G|X), but it does not relax the exogeneity of G. This is not a merely technical condition. In many social network applications, link formation is driven by unobserved shocks (common friends, shared unobserved traits, information spillovers) that also affect outcomes. The simulations in Section 5 always generate G from X plus independent ε, so they impose A3 by construction and cannot reveal this failure. Consequently, the headline result 'consistently estimates (λ,β)' is only valid when A3 holds, and the paper provides no diagnostic or alternative for the empirically common case where it does not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-step adjusted 2SLS estimator for linear social network models y = λGy + Xβ + ε when the observed adjacency matrix H misclassifies the true links at fixed, unknown rates (p0, p1). The central device is the adjusted matrix W(H,p) = (H − p0(ιι′−I))/(1−p0−p1), which satisfies E(W|G,X)=G under assumptions (A1)-(A2). The authors show that the adjusted structural form y = λWy + Xβ + v, with v = ε + λ(G−W)y, has E(v|G,X)=0 under A3 (Lemma 1), and that conventional instruments such as HX are invalid while H′X (single unsymmetrized measure) or the other noisy measure H(2)X (two measures) are valid. They provide closed-form identification of the misclassification rates from moments of H(1), H(2), and X only, propose feasible two-step estimators, state an asymptotic normality result (Proposition 5), report Monte Carlo evidence, and apply the method to estimate peer effects in microfinance participation using Banerjee et al. (2013) data.","tokens_in":33084,"tokens_out":11432,"duration_ms":143734,"significance":"If the results hold, the paper makes a useful and practical contribution: it extends 2SLS estimation of peer effects to the empirically common case of non-vanishing link misclassification without specifying a full parametric model of network formation. The identification of the misclassification rates is based on network-report moments only, not on outcome moments, so the procedure is not circular in the outcome equation. The estimator is closed-form, simple to compute, and the proposed instruments (H′X for a single unsymmetrized measure, and cross-measure instruments for multiple measures) are new and intuitive. The simulation evidence supports the main consistency claims. The principal limitations are that exogeneity of the true network (A3) is maintained, and that the asymptotic distribution is stated under regularity conditions relegated to an Online Appendix that is not included in this preprint; both need attention before the paper is fully acceptable.","major_comments":[{"comment":"Assumption (A3), E(ε|G,X,H)=0, is load-bearing for the central consistency claim. If actual link formation G is correlated with unobserved shocks that also affect outcomes, Lemma 1 fails and the adjusted estimator is inconsistent; W corrects for the difference between H and G, not for endogeneity of G itself. The paper's repeated statement that the method 'does not require structural modeling of link formation' should be qualified: it avoids specifying Pr(G|X), but it does not relax exogeneity of G. The simulations in Section 5 generate G from X plus independent ε, so they impose A3 by construction and cannot reveal the failure. I recommend adding an explicit scope discussion and, if possible, a diagnostic or robustness discussion for the case where A3 is doubtful, especially since many network applications involve link formation driven by unobserved shocks.","section":"Section 3.1, Lemma 1, Eq. (6)"},{"comment":"The asymptotic distribution in Proposition 5 is stated under 'regularity conditions (REG) in the Online Appendix,' but the Online Appendix is not part of this arXiv submission. Since the application in Section 6 reports clustered standard errors based on this proposition, the inference part of the paper is not verifiable from the submitted manuscript. The authors should either include the regularity conditions in the main text or provide the Online Appendix as part of the submission. Similarly, Proposition 4, which underlies the linear-in-means extension in Section 3.6, has its proof only in the Online Appendix; at least a sketch of the induction argument should be in the main text.","section":"Section 4.2, Proposition 5 and Section 3.6, Proposition 4"},{"comment":"Identification of the misclassification rates relies on a binary partition ϕ(X) with π1 ≠ π0. In the application this is same-caste versus different-caste pairings. While this is a weaker assumption than a full link-formation model, it is still a substantive assumption about link formation, and the 'no structural modeling of link formation' claim should be tempered accordingly. Moreover, the application also assumes conditional independence of the misclassification errors in the two symmetrized measures (A4′). Since the two measures come from different survey questions about visits, correlated recall errors across questions could violate this assumption and bias both the misclassification-rate estimates and the IV exogeneity. The paper should discuss the plausibility of this independence in the application and, ideally, provide some robustness evidence.","section":"Section 3.4.1 and Section 6.2"}],"minor_comments":[{"comment":"The reported 'Expected # of peers' values do not match the DGP. For n=25 with π1=0.2, π0=0.1, and P(X_i1=1)=0.5, the expected degree is 24×(0.5×0.2+0.5×0.1)=3.6, not 3.75; for n=50 it is 49×0.15=7.35, not 7.5. Please correct or explain the calculation.","section":"Section 5, Table 1(b)"},{"comment":"The text says the adjusted 2SLS uses W(t′)X as the instrument, while the tables report H(t′)X. Both are valid under Proposition 2, but the notation should be consistent.","section":"Section 4.2/4.3 and Tables 1(b), 1(c)"},{"comment":"The claim that the estimator is 'unlikely to suffer from weak instrument issues' is not backed by first-stage F-statistics or a formal strength analysis. Reporting first-stage F statistics would strengthen the finite-sample section.","section":"Section 5"},{"comment":"The rank conditions in Eq. (9) are called 'primitive,' but they involve M=(I−λG)−1, which depends on the unknown λ and G. They are better described as high-level moment-rank conditions; a more detailed discussion of their content would be helpful.","section":"Section 3.3, Proposition 3"},{"comment":"The caption of Table 5 refers to columns (a)-(e), but it does not define that these correspond to the estimators in Table 4. Please make the cross-reference explicit.","section":"Section 6.3 and Table 5"}],"recommendation":"major_revision","confidential_remarks":"The paper contains a clean and useful correction for link misclassification in linear network models, and the identification of p from network-report moments is a real strength. The main obstacles are that the asymptotic inference relies on an unavailable Online Appendix and the paper's 'no link-formation model' language overstates what is delivered, since exogeneity of G (A3) is still required. If the authors include the missing regularity conditions and proofs, and carefully qualify the scope, the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read of Lewbel-Qu-Tang. The headline: the paper gives applied network people a simple closed-form correction for misclassified links in linear-in-sums peer-effects models, and the misclassification-rate estimator from two noisy measures is genuinely new. It deserves refereeing.\n\nThe core trick—replace H by W=(H-p0(ιι'-I))/(1-p0-p1) and use H'X or the other noisy measure as instrument—is clean. Lemma 1 holds under A1–A3, and the proof that the H'X moments cancel using A4 is right. The closed-form moment equations in Appendix A2 for (p0,p1) are a real contribution: they don't require modeling Pr(G|X), only a binary φ(X) that shifts link propensity. Simulations confirm the bias correction, and the application to microfinance is a nice demonstration, even if the bias there is small.\n\nThe soft spot is A3: E(ε|G,X,H)=0. That's not a minor regularity condition; it says actual link formation is exogenous to the outcome error. The paper's \"does not require structural modeling of link formation\" is true only in a narrow sense—it avoids specifying Pr(G|X), but it does not relax exogeneity of G. If links are endogenous (common shocks that form friendships and affect outcomes), the adjusted estimator is inconsistent and the misclassification correction can't rescue it. The simulations always draw G from X plus independent ε, so they impose A3. Since the whole method is about fixing data problems, a referee worth her salt will ask for a discussion of when A3 is plausible, and possibly a diagnostic or partial-identification statement when A3 fails.\n\nOther concerns are smaller: Proposition 5 relies on REG conditions in an online appendix that isn't in the preprint; the standard errors for the two-measure S2SLS are omitted ('for brevity'); and A4' between the visit-out and visit-in measures in the application is plausible but not tested. None of these are fatal.\n\nBottom line: serious paper, well-executed, with a real new result. The main caveat is that 'no link formation model' is not the same as 'no link endogeneity.' I'd cite it if I worked on measurement error in networks, and I'd send it to a referee with instructions to focus on A3.","headline":"Adjusted 2SLS is a real fix for misclassified links, but it inherits the standard exogeneity of link formation (A3), and the 'no link-formation model' claim doesn't cover that.","tokens_in":33515,"tokens_out":2474,"would_cite":true,"duration_ms":29140,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that an adjusted 2SLS estimator, which rescales the noisy network matrix using estimated misclassification rates and instruments with the transpose of that matrix or a second independent report, consistently estimates peer","keywords":["social networks","peer effects","link misclassification","two-stage least squares","instrumental variables","measurement error","microfinance participation"],"falsifier":"Simulate a many-group design with the same structure as the paper but let dyad link formation depend on the outcome shock, e.g., Pr(G_ij = 1) = Phi(a + b|eps_i + eps_j|), generate two noisy measures with known p0 and p1, and run the adjusted 2SLS. If the peer-effect estimate does not converge to the true lambda as the number of groups grows, the exogeneity assumption E(eps|G,X,H)=0 is the point of failure.","tokens_in":1550,"feed_emoji":"🔗","tokens_out":4164,"duration_ms":139144,"temperature":0.7,"pith_summary":"Survey-based network links are often wrong: people forget or misreport who they connected with, and data entry lapses flip zeros and ones. This paper shows that such two-sided misclassification is not a nuisance that 2SLS can ignore: it makes every covariate endogenous, so conventional network estimators are inconsistent. The proposed fix is an adjusted two-stage least-squares estimator that rescales the noisy adjacency matrix with estimated misclassification rates, then builds instruments from the transpose of the noisy matrix or from a second independent report. The estimator needs no model of how links form, only that link errors are random and independent of outcome shocks, and it recovers misclassification rates in closed form. Applied to household participation in a microfinance program in Indian villages, it finds that an additional linked participant raises a household's participation by about 5.1 percent, while ignoring misclassification biases the peer effect upward.","feed_headline":"Misreported links no longer block peer-effect estimates","feed_subtitle":"A corrected two-stage estimator restores consistent estimates and finds a 5.1% microfinance peer effect in India.","key_machinery":"The adjusted adjacency matrix W(H,p0,p1) and its conditional-expectation property E(W|G,X)=G. This property is what makes Lemma 1 hold — the adjusted structural error is mean-zero given the true network and covariates — and it is what turns the usual invalid instrument HX into the valid instrument H'X, because the problematic conditional covariances vanish in the transpose. For row-normalized linear-in-means models, a row-specific invertible probability transformation fW plays the same role.","core_discovery":"The central claim is that the corrected adjacency matrix W = (H - p0(11' - I))/(1 - p0 - p1), built from the observed noisy matrix H and unknown rates p0 (false links) and p1 (missed links), satisfies E(W|G,X)=G, so the composite error in the adjusted outcome equation is mean-zero conditional on G and X. That restores exogeneity of the covariates X. The remaining endogeneity of the peer outcome Wy is handled by new instruments: with a single unsymmetrized report, H'X is valid even though HX is not; with two conditionally independent reports, one noisy report can instrument the other. The paper also identifies p0 and p1 from covariate-driven variation in link formation and from joint moments","pith_inferences":["Editorial inference: the transpose-instrument mechanism should generalize beyond social networks to any dyadic or interaction regressor measured with independent Bernoulli errors; the practical rule is to use the transpose of the noisy design as an instrument rather than the raw noisy design.","Editorial inference: because the method corrects reporting error but maintains exogeneity of true link formation, a natural next step is to pair W with instruments that are valid for G itself when link formation is endogenous.","Editorial inference: the authors themselves note one uncovered cell — an asymmetric true network with a single unsymmetrized report has valid instruments but no identified misclassification rates without an auxiliary link-formation model; researchers in that setting need a second report or an additional model.","Editorial inference: the covariate phi used to recover misclassification rates must actually shift link formation (pi1 != pi0); checking this first-stage relationship in the observed noisy reports is advisable, since a flat phi would leave the rates weakly identified."],"forward_implications":["Applied researchers can estimate peer effects from misclassified survey links with a closed-form two-step 2SLS, without specifying a link-formation model or a likelihood for the true network.","When two conditionally independent network reports are available, one report can be used as an instrument for the other, regardless of whether the reports are symmetrized.","In the microfinance application, the peer effect is estimated at about 0.051, meaning one additional linked participant raises participation by roughly 5.1 percent; naive 2SLS that ignores misclassification produces upward-biased estimates.","The method extends to linear-in-means social interaction models through a finite invertible transformation of the noisy row, as long as p0 + p1 differs from 1.","Asymptotic normality with group-level clustering is established for the adjusted 2SLS estimators as the number of independent groups grows."],"supporting_citations":[{"why":"Supplies the conventional 2SLS instruments GX and G^2X for peer outcomes that the paper shows become invalid when links are misclassified.","marker":"Bramoullé et al. (2009)"},{"why":"Provides the two symmetrized visit-based network reports and the microfinance participation outcomes used in the empirical application.","marker":"Banerjee et al. (2013)"},{"why":"Establishes the p0 + p1 < 1 positive-correlation condition that the paper relies on so the noisy measure remains informative about the true link.","marker":"Bollinger (1996)"},{"why":"Shows when small link-measurement errors can be ignored; the present paper extends the setting to fixed, non-diminishing misclassification rates.","marker":"Lewbel et al. (2023)"}],"fun_headline_variants":["Adjusted 2SLS corrects bias from wrong links in networks","Noisy survey links fixed for consistent peer effects","Microfinance peer effect confirmed after correcting link errors","Estimating peer effects when network links are misreported","Correcting link errors in network data without structural models"],"cache_read_input_tokens":35328,"weakest_assumption_plain":"The load-bearing premise is that the true network links and the reported noisy links are unrelated to unobserved shocks that also affect outcomes; if people form or report links because of exactly those outcome shocks, the correction does not remove the bias.","fun_headline_variants_meta":{"raw":{"variants":["Adjusted 2SLS corrects bias from wrong links in networks","Noisy survey links fixed for consistent peer effects","Microfinance peer effect confirmed after correcting link errors","Estimating peer effects when network links are misreported","Correcting link errors in network data without structural models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000623,"raw_usage":{"total_tokens":2702,"prompt_tokens":703,"completion_tokens":1999,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":447,"completion_tokens_details":{"reasoning_tokens":1934}},"tokens_in":447,"tokens_out":1999,"duration_ms":19025,"temperature":1.0,"reasoning_tokens":1934,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T22:21:42.892420+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a many-group design with the same structure as the paper but let dyad link formation depend on the outcome shock, e.g., Pr(G_ij = 1) = Phi(a + b|eps_i + eps_j|), generate two noisy measures with known p0 and p1, and run the adjusted 2SLS. If the peer-effect estimate does not converge to the true lambda as the number of groups grows, the exogeneity assumption E(eps|G,X,H)=0 is the point of failure.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the two symmetrized visit-based network reports and the microfinance participation outcomes used in the empirical application."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the p0 + p1 < 1 positive-correlation condition that the paper relies on so the noisy measure remains informative about the true link."},{"cited_title":"Qu, and X","cited_arxiv_id":null,"evidence_quote":"Shows when small link-measurement errors can be ignored; the present paper extends the setting to fixed, non-diminishing misclassification rates."}],"review_version":1}