REVIEW 2 major objections 5 minor 64 references
Estimating Peer Effects Using Partial Network Data
T0 review · 2 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Peer effects can be consistently estimated from partial network data whenever the network-formation parameters are identifiable, and correcting for missing links in Add Health raises the estimated effect by about 50%.
desk verdict A solid, useful paper on estimating peer effects with partial network data, but the main theorem assumes the key global-identification condition rather than proving it; that is the referee's main ask. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the estimated distribution of the true network, constructed from a network formation model estimated on partial data—sampled, censored, or misclassified links. The SGMM machinery is the bias-corrected simulated moment function (4), which uses three independent draws from that distribution: one draw builds the instruments, a second builds the regressors, and a third approximates the bias that would otherwise remain because the simulated network does not converge to the true network within bounded groups. The Bayesian machinery is data augmentation: the network itself is treated as an unknown parameter, updated by a Metropolis-Hastings step whose prior is the estimat
What would settle it
Take a fully observed network, artificially censor or misclassify links at known rates, and compare the SGMM estimate to the estimate on the complete network; if coverage of the true coefficient deteriorates sharply, or if the concentrated objective shows multiple minima on a grid of alpha values, the central claim fails.
Extended reading notes
Core claim
The paper's formal claim is Theorem 1: under exogeneity of the network, bounded independent groups, a parametric link-independent network formation model, and the existence of a √M-consistent estimator of the formation parameters from the partial data, the simulated GMM estimator defined by the moment function (4) is consistent and asymptotically normal for any positive numbers of simulated draws R, S, T. The moment function removes the approximation error caused by substituting simulated networks for the true one: instruments come from one set of draws, the regressors from a second, and a third draw approximates the asymptotic bias term, so that the moment is zero at the true parameters. Th
Load-bearing premise
The result stands on the premise that partial network data identify the network-formation parameters at the usual statistical convergence rate; the additional global identification condition (a unique minimum of the concentrated objective) is supported only by simulation evidence of convexity, not by a proof.
Editorial extensions
If this is right
- A single cross-section of partial network data—sampled pairs, censored nominations, or misclassified links—can be enough to identify the endogenous peer effect, as long as the formation parameters are identified.
- Standard instrumental-variable estimates that treat the observed network as complete are downward-biased; in Add Health, the bias is large enough that the estimated endogenous effect grows by about 50% after reconstruction.
- The SGMM estimator stays centered on the true value in simulations even when half of the links are missing, with precision declining as missingness grows.
- The Bayesian estimator remains valid in finite samples and can accommodate network formation models, such as those with unobserved degree heterogeneity or aggregated relational data, that fall outside the SGMM's asymptotic framework.
- The reconstructed network changes which individuals would be chosen as 'key players' for interventions, not just the size of the social multiplier.
Reading between the lines
- Beyond the paper: if this downward-bias pattern holds in other nomination-based surveys (many ask for a fixed number of friends), published peer-effect coefficients from such data are likely conservative, and re-estimating with reconstructed networks could change policy conclusions.
- Beyond the paper: because the global identification condition is not proved, a practical robustness check is to sweep the concentrated objective over a grid of alpha values in each application and report whether a unique minimum is visible.
- Beyond the paper: the framework suggests a testable extension—using the same reconstruction on data sets where the full network is known but deliberately noised, to quantify how misspecification of the network-formation model propagates into the peer-effect estimate.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies estimation of linear-in-means peer effects when the true adjacency matrix is not fully observed. It proposes an SGMM estimator built from draws of an estimated network formation distribution and a Bayesian data-augmentation estimator. The main theoretical result (Theorem 1) states that, under high-level regularity and a global identification condition, the SGMM is consistent and asymptotically normal for any positive R,S,T. The paper reports simulations for missing and misclassified links and an Add Health application in which correcting for missing nominations raises the estimated endogenous peer effect by about 50%.
Significance. If the theorem were fully established, the paper would make a useful contribution: it unifies sampled, censored, and misclassified networks under one computational framework, avoids numerical integration over networks, and ships an R package. The bias-corrected moment construction is clever, and the treatment of first-stage uncertainty in the variance estimation is careful. The simulations are encouraging. However, the central result is conditional on an unproved global identification assumption, and the asymptotic normality proof has a joint-convergence gap. The empirical 1.5x claim is driven by the Bayesian estimator; the SGMM estimates are not statistically distinguishable. With a strengthened theorem for leading cases, the contribution would be solid.
major comments (2)
- [Theorem 1; Online Appendix C.1] The global identification condition in Theorem 1 is assumed, not derived. Lemma 1 establishes that E[m_m(θ0,ρ0)] = 0, but consistency requires θ0 to be the unique zero of lim E[mbar_M(θ,ρ0)]. Online Appendix C.1 reduces the condition to full rank of Bbar_0(α) plus unique minimization of the concentrated objective Q^c_0(α), and then states that 'simplifying this condition is challenging' and appeals to simulations showing strict convexity. This is not a proof, and nothing in Assumptions 1–5 rules out a second minimum or a flat region for some DGPs (e.g., dense or strongly clustered networks, or α near the boundary). Because the abstract claims that a consistent estimator of the network distribution is sufficient, the paper should either provide primitive identification conditions for the leading logistic-link cases or qualify the claim as conditional on global identification. The empirica
- [Appendix A.2, Eq. (8)] The asymptotic normality proof is incomplete. The text states that, conditional on sqrt(M)(rho_hat - rho_0), a Lyapunov CLT applies to sqrt(M) mbar_M(θ0,ρ0). But rho_hat depends on all M groups, so conditioning on sqrt(M)(rho_hat - rho_0) does not restore independence across the m-th moment terms; the conditioning event is a function of the entire sample. A rigorous joint convergence argument is needed, e.g., an influence-function expansion for rho_hat together with a Cramér-Wold argument for the vector of averages, or adding the first-stage moment conditions to the GMM system. As written, Theorem 1's asymptotic normality and the variance estimator in Online Appendix C.2 rest on an unproven joint distribution.
minor comments (5)
- [Abstract; Section 5] The abstract says network data errors have a large downward bias. In Tables G.2–G.3, the SGMM point estimate moves from 0.455 to 0.683–0.753, but the standard errors are about 0.23–0.25, so the difference is not statistically significant. Footnote 35 admits this only for the Bayesian estimator; the abstract and Section 5 should be qualified.
- [Table G.3] There are typos in the table: '0.03)' should be '0.03' and '0.08)' should be '0.08'. Similar small typos appear in the notes to Figure G.4.
- [Online Appendix C.1] The statement that Q^c_0(α) is 'strictly convex' in numerous simulation exercises would be more credible if a figure or grid evaluation of the concentrated objective were included, at least for one representative DGP.
- [Section 4] The claim that the Bayesian estimator is 'valid in finite samples' is imprecise: the MCMC targets the posterior distribution, but no posterior consistency or finite-sample risk statement is proved. Suggest rewording to 'has a finite-sample Bayesian interpretation'.
- [Section 5 and Table G.2] The text says about 45% of friendship nominations are coded with error, while Table G.2 reports 60% 'proportion of inferred network data'. The denominators differ (named nominations vs. all potential network entries) and should be reconciled explicitly.
Circularity Check
No significant circularity: the core SGMM argument is a standard two-step simulated-moment construction; the assumed global identification condition is an unproven limitation, not a circular reduction.
full rationale
The paper's central claim is that a consistent estimator of the network-formation distribution (Assumption 5) suffices, together with a standard global identification condition, for consistency and asymptotic normality of the SGMM estimator. Lemma 1 shows the moment function has zero mean at (θ0, ρ0) because the simulated draws and the true network share the same distribution at ρ0; this is the usual simulation-based moment validity argument, not a definitional identity. The identification condition (lim E(m̄_M(θ,ρ0)) ≠ 0 for θ≠θ0) is explicitly assumed in Theorem 1, and Online Appendix C.1 admits 'simplifying this condition is challenging' and offers only simulation evidence of convexity. That is a gap between the abstract's 'sufficient' claim and the theorem, but it is a correctness/identification concern, not circularity: the condition is not derived from the object it seeks to identify. The empirical application estimates the network-formation model from the observed partial network data and then uses draws from that estimated distribution to form instruments; this is a two-step simulated GMM, not a fitted parameter being renamed a prediction. The Bayesian estimator's use of the outcome in reconstructing the network is standard data augmentation, and the reported 1.5× ratio is a comparison of estimates, not a construction. The only self-citations (e.g., Houndetoungan and Maoude 2024 for variance estimation; Hsieh, Lee and Boucher 2020; Boucher and Mourifié 2017 for extensions) are technical or peripheral and are not load-bearing for the main theorems. No equation in the paper reduces to its own input by construction.
Assumptions & free parameters
free parameters (3)
- Network formation coefficients ρ =
e.g., same sex 0.31, age diff -0.70 (Add Health)
- Misclassification rates in simulations =
false positive 0-15%, false negative 15-30%
- Bayesian prior hyperparameters =
μ_ᾶ=-1, σ_ᾶ^-2=2, a=b=4, Σ_Λ^-1=I/100
assumptions (6)
- domain assumption Assumption 1: M groups of bounded size, independent across m
- domain assumption Assumption 4: E[ε_m | X_m, A_m, A_bar_m] = 0 (exogenous network)
- domain assumption Assumption 5: there exists a sqrt(M)-consistent estimator of the network formation parameters from partial data
- domain assumption Conditional link independence in network formation: P(A_m|X_m)=Π_ij P(a_ij,m|X_m)
- ad hoc to paper Global identification: for any θ≠θ0, lim_{M→∞} E(¯m_M(θ,ρ0)) ≠ 0
- domain assumption For the Bayesian estimator: ε_m ~ N(0, σ^2 I) and priors on α, β, γ, σ^2
Cite this review
Pith. "Pith review of Estimating Peer Effects Using Partial Network Data." pith.science (2026). https://pith.science/paper/6ERNIRUR
@misc{pith2026250908145,
author = {Pith},
title = {Pith review of: Estimating Peer Effects Using Partial Network Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/6ERNIRUR}},
note = {Machine review of arXiv:2509.08145}
}
read the original abstract
We study the estimation of peer effects through social networks when researchers do not observe the entire network structure. Special cases include sampled networks, censored networks, and misclassified links. We assume that researchers can obtain a consistent estimator of the distribution of the network. We show that this assumption is sufficient for estimating peer effects using a linear-in-means model. We provide an empirical application to the study of peer effects on students' academic achievement using the widely used Add Health database, and show that network data errors have a large downward bias on estimated peer effects.
Figures
Reference graph
Works this paper leans on
-
[1]
Andrews, D. W. (1994): Empirical process methods in econometrics, Handbook of Econometrics, 4, 2247--2294
work page 1994
-
[2]
(2021): Nash equilibria on (un) stable networks, Econometrica, 89, 1179--1206
Badev, A. (2021): Nash equilibria on (un) stable networks, Econometrica, 89, 1179--1206
work page 2021
-
[3]
Calv \'o -Armengol, and Y
Ballester, C., A. Calv \'o -Armengol, and Y. Zenou (2006): Who's who in networks. Wanted: The key player, Econometrica, 74, 1403--1417
2006
-
[4]
Banerjee, A., A. G. Chandrasekhar, E. Duflo, and M. O. Jackson (2013): The diffusion of microfinance, Science, 341, 1236498
2013
-
[5]
Bhamidi, S., G. Bresler, and A. Sly (2008): Mixing time of exponential random graphs, in 2008 49th Annual IEEE Symposium on Foundations of Computer Science, IEEE, 803--812
work page 2008
-
[6]
Boucher, V. and I. Mourifi \'e (2017): My friend far, far away: a random field approach to exponential random graph models, The Econometrics Journal, 20, S14--S46
2017
-
[7]
Bramoull \'e , Y., H. Djebbari, and B. Fortin (2009): Identification of peer effects through social networks, Journal of Econometrics, 150, 41--55
work page 2009
-
[8]
--- -.1pt --- -.1pt --- (2020): Peer effects in networks: A survey, Annual Review of Economics, 12, 603--629
work page 2020
Show all 64 references
-
[9]
(2016): Field experiments, social networks, and development, The Oxford Handbook on the Economics of Networks, 412--439
Breza, E. (2016): Field experiments, social networks, and development, The Oxford Handbook on the Economics of Networks, 412--439
2016
-
[10]
Breza, E., A. G. Chandrasekhar, S. Lubold, T. H. McCormick, and M. Pan (2023): Consistently estimating network statistics using aggregated relational data, Proceedings of the National Academy of Sciences, 120, e2207185120
2023
-
[11]
Breza, E., A. G. Chandrasekhar, T. H. McCormick, and M. Pan (2020): Using aggregated relational data to feasibly identify network structure without network data, American Economic Review, 110, 2454--84
2020
-
[12]
Patacchini, and Y
Calv \'o -Armengol, A., E. Patacchini, and Y. Zenou (2009): Peer effects and social networks in education, The Review of Economic Studies, 76, 1239--1267
2009
-
[13]
Cameron, A. C. and P. K. Trivedi (2005): Microeconometrics: methods and applications, Cambridge University Press
2005
-
[14]
Chandrasekhar, A. and R. Lewis (2011): Econometrics of sampled networks, Unpublished manuscript, MIT.[422]
2011
-
[15]
Diaconis, et al
Chatterjee, S., P. Diaconis, et al. (2013): Estimating and understanding exponential random graph models, The Annals of Statistics, 41, 2428--2461
2013
-
[16]
Chen, and P
Chen, X., Y. Chen, and P. Xiao (2013): The impact of sampling and network topology on the estimation of social intercorrelations, Journal of Marketing Research, 50, 95--110
2013
-
[17]
Conley, T. G. and C. R. Udry (2010): Learning about a new technology: Pineapple in Ghana, American Economic Review, 100, 35--69
2010
-
[18]
(2017): Econometrics of network models, in Advances in Economics and Econometrics: Theory and Applications: Eleventh World Congress (Econometric Society Monographs, ed
De Paula, A. (2017): Econometrics of network models, in Advances in Economics and Econometrics: Theory and Applications: Eleventh World Congress (Econometric Society Monographs, ed. by M. P. B. Honore, A. Pakes and L. Samuelson, Cambridge: Cambridge University Press, 268--323
2017
-
[19]
Rasul, and P
De Paula, A., I. Rasul, and P. C. Souza (2024): Identifying network ties from panel data: Theory and an application to tax competition, Review of Economic Studies, rdae088
2024
-
[20]
Richards-Shubik, and E
De Paula, \'A ., S. Richards-Shubik, and E. Tamer (2018): Identifying preferences in networks with bounded degree, Econometrica, 86, 263--288
2018
-
[21]
Glaeser, E. L., B. I. Sacerdote, and J. A. Scheinkman (2003): The social multiplier, Journal of the European Economic Association, 1, 345--353
2003
-
[22]
Goldsmith-Pinkham, P. and G. W. Imbens (2013): Social networks and the identification of peer effects, Journal of Business & Economic Statistics, 31, 253--264
2013
-
[23]
Gourieroux, A
Gourieroux, M., C. Gourieroux, A. Monfort, D. A. Monfort, et al. (1996): Simulation-based econometric methods, Oxford University Press
1996
-
[24]
Graham, B. S. (2017): An econometric model of network formation with degree heterogeneity, Econometrica, 85, 1033--1063
2017
-
[25]
(2022): Name your friends, but only five? the importance of censoring in peer effects estimates using social network data, Journal of Labor Economics, 40, 779--805
Griffith, A. (2022): Name your friends, but only five? the importance of censoring in peer effects estimates using social network data, Journal of Labor Economics, 40, 779--805
2022
-
[26]
Griffith, A. and J. Kim (2023): The Impact of Missing Links on Linear Reduced-form Network-Based Peer Effects Estimates, Working Paper
2023
-
[27]
Hardy, M., R. M. Heath, W. Lee, and T. H. McCormick (2024): Estimating spillovers using imprecisely measured networks, arXiv preprint arXiv:1904.00136
2024 arXiv
-
[28]
Hausman, J. A., J. Abrevaya, and F. M. Scott-Morton (1998): Misclassification of the dependent variable in a discrete-response setting, Journal of Econometrics, 87, 239--269
1998
-
[29]
Herstad, E. I. (2023): Estimating peer effects and Network formation models with missing network links, Working Paper
2023
-
[30]
Hsieh, C.-S., Y.-C. Hsu, S. I. Ko, J. Kov \'a r \' k, and T. D. Logan (2024): Non-representative sampled networks: Estimation of network structural properties by weighting, Journal of Econometrics, 240, 105689
2024
-
[31]
Hsieh, C.-S., M. D. K \"o nig, and X. Liu (2019): A structural model for the coevolution of networks and behavior, Review of Economics and Statistics, 1--41
2019
-
[32]
Lee, and V
Hsieh, C.-S., L.-F. Lee, and V. Boucher (2020): Specification and estimation of network formation and network interaction models with the exponential probability distribution, Quantitative Economics, 11, 1349--1390
2020
-
[33]
Hsieh, C.-S. and H. Van Kippersluis (2018): Smoking initiation: Peers and personality, Quantitative Economics, 9, 825--863
2018
-
[34]
(2004): Asymptotic distributions of quasi-maximum likelihood estimators for spatial autoregressive models, Econometrica, 72, 1899--1925
Lee, L.-F. (2004): Asymptotic distributions of quasi-maximum likelihood estimators for spatial autoregressive models, Econometrica, 72, 1899--1925
2004
-
[35]
Liu, and X
Lee, L.-f., X. Liu, and X. Lin (2010): Specification and estimation of social interaction models with network structures, The Econometrics Journal, 13, 145--176
2010
-
[36]
Qu, and X
Lewbel, A., X. Qu, and X. Tang (2023): Social networks with unobserved links, Journal of Political Economy, 131, 898--946
2023
-
[37]
--- -.1pt --- -.1pt --- (2024 a ): Estimating Social Network Models with Link Misclassification, Working Paper
2024
-
[38]
--- -.1pt --- -.1pt --- (2024 b ): Ignoring measurement errors in social networks, The Econometrics Journal, 27, 171--187
2024
-
[39]
(2013): Estimation of a local-aggregate network model with sampled networks, Economics Letters, 118, 243--246
Liu, X. (2013): Estimation of a local-aggregate network model with sampled networks, Economics Letters, 118, 243--246
2013
-
[40]
Patacchini, and E
Liu, X., E. Patacchini, and E. Rainone (2017): Peer effects in bedtime decisions among adolescents: a social network model with sampled data, The Econometrics Journal, 20, S103--S125
2017
-
[41]
(2016): Estimating the structure of social interactions using panel data, Working paper
Manresa, E. (2016): Estimating the structure of social interactions using panel data, Working paper
2016
-
[42]
Manski, C. F. (1993): Identification of endogenous social effects: The reflection problem, Review of Economic Studies, 60, 531--542
1993
-
[43]
Manski, C. F. and S. R. Lerman (1977): The estimation of choice probabilities from choice based samples, Econometrica: Journal of the Econometric Society, 1977--1988
1977
-
[44]
(2017): A structural model of Dense Network Formation, Econometrica, 85, 825--850
Mele, A. (2017): A structural model of Dense Network Formation, Econometrica, 85, 825--850
2017
-
[45]
--- -.1pt --- -.1pt --- (2020): Does school desegregation promote diverse interactions? An equilibrium model of segregation within schools, American Economic Journal: Economic Policy, 12, 228--57
2020
-
[46]
Newey, W. K. and D. McFadden (1994): Large sample estimation and hypothesis testing, Handbook of Econometrics, 4, 2111--2245
1994
-
[47]
Reeves, S. W., S. Lubold, A. G. Chandrasekhar, and T. H. McCormick (2024): Model-based inference and experimental design for interference using partial network data, arXiv preprint arXiv:2406.11940
2024 arXiv
-
[48]
Snijders, T. A. (2002): Markov chain Monte Carlo estimation of exponential random graph models, Journal of Social Structure, 3, 1--40
2002
-
[49]
Tanner, M. A. and W. H. Wong (1987): The calculation of posterior distributions by data augmentation, Journal of the American Statistical Association, 82, 528--540
1987
-
[50]
(2019): Identification and estimation of network statistics with missing link data, Working Paper
Thirkettle, M. (2019): Identification and estimation of network statistics with missing link data, Working Paper
2019
-
[51]
Van der Vaart, A. W. (2000): Asymptotic Statistics, vol. 3, Cambridge University Press
2000
-
[52]
and L.-F
Wang, W. and L.-F. Lee (2013): Estimation of spatial autoregressive models with randomly missing data in the dependent variable, The Econometrics Journal, 16, 73--102
2013
-
[53]
(2024): Spillovers of program benefits with mismeasured networks, arXiv preprint arXiv:2009.09614
Zhang, L. (2024): Spillovers of program benefits with mismeasured networks, arXiv preprint arXiv:2009.09614
2024 arXiv
-
[54]
Auerbach, and M
Alidaee, H., E. Auerbach, and M. P. Leung (2020): Recovering network structure from aggregated relational Data using Penalized Regression, arXiv preprint arXiv:2001.06052
2020 arXiv
-
[55]
Gentzkow, and J
Andrews, I., M. Gentzkow, and J. M. Shapiro (2017): Measuring the sensitivity of parameter estimates to estimation moments, The Quarterly Journal of Economics, 132, 1553--1592
2017
-
[56]
Atchad \'e , Y. F. and J. S. Rosenthal (2005): On adaptive Markov chain Monte Carlo algorithms, Bernoulli, 11, 815--828
2005
-
[57]
Brown, and N
Bound, J., C. Brown, and N. Mathiowetz (2001): Measurement error in survey data, in Handbook of Econometrics, Elsevier, vol. 5, 3705--3843
2001
-
[58]
Chib, S. and S. Ramamurthy (2010): Tailored randomized block MCMC methods with application to DSGE models, Journal of Econometrics, 155, 19--38
2010
-
[59]
Hoff, P. D., A. E. Raftery, and M. S. Handcock (2002): Latent space approaches to social network analysis, Journal of the American Statistical Association, 97, 1090--1098
2002
-
[60]
Houndetoungan, A. and A. H. Maoude (2024): Inference for Two-Stage Extremum Estimators, arXiv preprint arXiv:2402.05030
2024 arXiv
-
[61]
Johnson, C. R. and R. A. Horn (1985): Matrix analysis, Cambridge University Press
1985
-
[62]
McCormick, T. H. and T. Zheng (2015): Latent surface models for networks using Aggregated Relational Data, Journal of the American Statistical Association, 110, 1684--1695
2015
-
[63]
Onishi, R. and T. Otsu (2021): Sample sensitivity for two-step and continuous updating GMM estimators, Economics Letters, 198, 109685
2021
-
[64]
Thijssen, B. and L. F. Wessels (2020): Approximating multivariate posterior distribution functions from Monte Carlo samples for sequential Bayesian inference, PloS ONE, 15, e0230101
2020
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.