REVIEW 3 major objections 4 minor 55 references
An Empirical Comparison of Weak-IV-Robust Procedures in Just-Identified Models
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Using AER replication data and simulations, this paper establishes that the classical Anderson-Rubin test typically delivers higher power and shorter confidence intervals than the newer tF t-ratio procedure in just-identified…
desk verdict Useful empirical comparison of AR vs tF, but the CI-length comparison drops exactly the weak-instrument cases where AR intervals are unbounded, so that headline result is not established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The comparison is carried by two inference procedures for a just-identified linear IV model. The Anderson-Rubin test is a weak-identification-robust test of the structural coefficient that remains valid regardless of first-stage strength. The tF procedure, proposed by Lee et al. (2022), preserves the familiar 2SLS t-ratio but calibrates its critical value through a smooth function of the first-stage F-statistic, giving uniform size control. The empirical machinery is the AER replication dataset—the same one used by Lee et al. (2022)—restricted to 151 specifications from 17 studies that are just-identified, public-data, and free of nonlinear structures, alongside 250,000-replication Monte Carlo experiments matching the original tF study's DGP.
What would settle it
A second empirical study using a different large sample of just-identified IV specifications—for example, replication data from other journals or from private-data studies—that found tF rejecting at least as often as AR and yielding shorter intervals in a majority of cases would directly contradict the paper's central empirical claim. Alternatively, a Monte Carlo exercise covering non-normal errors, heteroskedasticity, or asymmetric instrument strength in which tF's power curves dominate AR's would weaken the claimed generality.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is empirical: in a replication sample of just-identified IV specifications from AER articles published 2013–2019, the Anderson-Rubin procedure dominates the tF procedure in both significance testing and interval length. Among specifications insignificant under tF, 47.52% are significant under AR, and there is no specification in which tF is significant at a more stringent level than AR. AR confidence intervals are shorter than tF intervals in 96.85% of specifications at the 5% level and 95.93% at the 1% level; when AR intervals are longer the loss is modest, while when they are shorter the gain is often substantial. Monte Carlo power curves show AR weakly dominating tF at low endogeneity and achieving higher power at low instrument strength.
Load-bearing premise
The 151 specifications from 17 AER studies that survived exclusions (public data, single instrument, linear structure, available CI lengths) are representative of just-identified IV applications generally, so that the observed AR-over-tF ranking is not an artifact of which studies were left out.
Editorial extensions
If this is right
- Applied researchers in just-identified IV settings with weak instruments can expect the AR test to detect more nonzero effects than tF at the same nominal size, since AR rejects 47.52% of specifications that tF does not.
- AR confidence intervals will typically be much shorter than tF intervals, particularly when the first-stage F-statistic is low; the paper estimates AR intervals are shorter in 96.85% of specifications at the 5% level.
- Because tF retains a theoretical expected-length advantage, the two procedures are best viewed as complementary: AR for power-oriented testing and tF for cases where its expected length is preferred.
- The empirical ranking aligns with the recommendations of Keane and Neal (2023, 2024) that applied work should prefer AR-based inference over t-ratio-based inference.
- The same empirical protocol can be applied to refinements such as the VtF procedure or bootstrap versions of tF to see whether the AR advantage persists.
Reading between the lines
- An editorial extension: tF's conservativeness in this sample may result from its smooth critical value function being conservative at low F values, and a bootstrap-calibrated tF could close much of the power gap; testing this would require re-running the rejection counts with bootstrap critical values.
- Another inference: because the 151-specification sample excludes private-data studies, the ranking could change if private-data replications tend to have stronger or weaker instruments; a replication using such datasets would be a direct external-validity check.
- A third inference: the Monte Carlo design uses homoskedastic normal errors, so the power ordering might differ under heteroskedasticity or heavy-tailed errors, where both procedures' asymptotic critical values can be distorted; a simulation with such error structures would be informative.
- Fourth: the CI-length comparison drops cases where AR intervals are unbounded; treating unbounded intervals explicitly (e.g., with a loss function) might reduce AR's apparent dominance.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper compares the Anderson–Rubin (AR) procedure and the tF procedure of Lee et al. (2022) for just-identified instrumental variable models, using the American Economic Review replication dataset (439 specifications, reduced to 151 usable specifications) and Monte Carlo simulations. It reports that AR tends to have higher power and shorter confidence intervals than tF, and concludes that the two procedures are complementary.
Significance. If the empirical findings are robust, they would provide useful guidance for applied researchers choosing among weak-IV-robust procedures. The paper offers a clear empirical comparison of significance outcomes and simulation power curves, which are of interest to the empirical IV literature. However, the main CI-length claim is vulnerable to a sample selection issue that is endogenous to the outcome, which limits the force of the paper's central message.
major comments (3)
- [Section 2.4, Figure 2] The paper states that 'After excluding cases with missing CI length information, 127 specifications remain at the 5% significance level and 123 at the 1% level,' but it never defines what makes CI length information 'missing.' In a just-identified IV model, the AR confidence set is the solution set of a quadratic inequality whose leading coefficient is proportional to the first-stage F-statistic minus the AR critical value; when F is below that critical value (approximately 3.84 at the 5% level), the AR set is unbounded and has infinite length, while the tF interval remains bounded. The reported 96.85% and 95.93% shares of specifications in which AR is shorter are therefore computed only over specifications with bounded AR sets. This is an endogenous selection on the outcome being compared: the dropped specifications are likely those with weak instruments where AR cannot be shorter than tF. The authors should report the number of excluded specifications with unbounded AR sets, their F-statistic distribution, and redo the comparison including these cases (e.g., treating their length as infinite) or provide a sensitivity analysis that does not condition on boundedness.
- [Section 2.4, Figure 3] The heatmap in Figure 3 is based on the same restricted sample of 127 and 123 specifications. The paper claims that AR delivers substantially shorter CIs 'particularly in areas with low first-stage F-statistics,' but this is exactly the region where unbounded AR sets are most likely, so the figure describes only the bounded-AR subset. Consequently, the figure does not support the conclusion that AR is shorter in weak-instrument settings; it may reflect the selection of specifications for which AR happens to be bounded. The authors should either include the unbounded cases in the comparison or clearly temper the interpretation of the heatmap.
- [Section 2.1] The paper does not fully document the sample selection from 439 to 343 to 151 specifications. It states that cases with multiple instruments or nonlinear structures were removed, and that the remaining exclusions were 'primarily due to limitations related to data confidentiality,' but no detailed breakdown by study is provided. Because the 151-specification sample is a non-random subset (e.g., private-data studies are excluded), the representativeness of the significance and CI-length comparisons is unclear. The authors should provide a table listing each excluded study and the reason for exclusion, and discuss how the exclusions might affect the comparison.
minor comments (4)
- [Figure 1] The axis tick labels in Figure 1 appear garbled in the manuscript (e.g., '01.9622.5762∞t2 statistic'); the figure should be redrawn with clear tick labels.
- [Online Appendix A] The simulation section refers to 'the simulations of Lee et al. (2022)' but does not provide the exact DGP equations or the implementation details of the tF procedure; additional details would aid reproducibility.
- [References] The reference to Moreira (2009) contains a typo ('abitrarily' should be 'arbitrarily') and appears to have an unusual citation format ('152:131–140'); please verify the full reference details.
- [Section 3] The conclusion notes that tF has a theoretical advantage in expected CI length, but the paper does not reconcile this with the empirical finding of shorter AR intervals in the sample; a brief discussion of why the theoretical expectation may not hold in observed samples would strengthen the paper.
Circularity Check
No circularity: the comparison of AR and tF is an external empirical and simulation exercise, not a derivation from fitted inputs or self-citations.
full rationale
The paper's central claim is that the Anderson-Rubin (AR) procedure has higher power and yields shorter confidence intervals than the tF procedure in an external AER replication sample and in Monte Carlo simulations. No parameter is fitted and then renamed as a prediction; no equation or result is defined in terms of the conclusion it supports. The tF procedure and the AER dataset are taken from Lee et al. (2022), but the author of the present paper is not an author of that source, so this is a matter of provenance rather than load-bearing self-citation. The power simulations compare rejection rates under independently varied structural and nuisance parameters, and the empirical significance comparison directly computes AR and tF rejections on existing specifications. The only notable concern is that the confidence-interval length comparison excludes cases with 'missing CI length information' (151 to 127 at the 5% level and 123 at the 1% level), and the paper does not state why those lengths are missing; if these are cases where the AR confidence set is unbounded, the reported percentage of shorter AR intervals could be affected. That is a sample-selection and external-validity critique, not a circularity critique, because the reported measure is not constructed so as to force the conclusion. The paper also explicitly acknowledges the complementary theoretical advantage of tF (Lee et al., 2022), which further shows the empirical result is not presupposed. Overall, the derivation chain is self-contained and the analysis is best assessed on statistical grounds rather than circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption Each specification satisfies the IV assumptions: Cov(u,Z)=0 and Cov(Z,X)≠0, as stated in equations (1) and (2).
- domain assumption The tF critical value function from Lee et al. (2022) is correctly implemented and its asymptotic properties hold in the finite samples studied.
Cite this review
Pith. "Pith review of An Empirical Comparison of Weak-IV-Robust Procedures in Just-Identified Models." pith.science (2026). https://pith.science/paper/KPHVZVRW
@misc{pith2026250618001,
author = {Pith},
title = {Pith review of: An Empirical Comparison of Weak-IV-Robust Procedures in Just-Identified Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/KPHVZVRW}},
note = {Machine review of arXiv:2506.18001}
}
read the original abstract
Instrumental variable (IV) regression is recognized as one of the five core methods for causal inference, as identified by Angrist and Pischke (2008). This paper compares two leading approaches to inference under weak identification for just-identified IV models: the classical Anderson-Rubin (AR) procedure and the recently popular tF method proposed by Lee et al. (2022). Using replication data from the American Economic Review (AER) and Monte Carlo simulation experiments, we evaluate the two procedures in terms of statistical significance testing and confidence interval (CI) length. Empirically, we find that the AR procedure typically offers higher power and yields shorter CIs than the tF method. Nonetheless, as noted by Lee et al. (2022), tF has a theoretical advantage in terms of expected CI length. Our findings suggest that the two procedures may be viewed as complementary tools in empirical applications involving potentially weak instruments.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Alesina, A., Stantcheva, S., and Teso, E. (2018). Intergenerational mobility and preferences for redistribution. American Economic Review , 108(2):521--554
work page 2018
-
[2]
Anderson, T. W. and Rubin, H. (1949). Estimation of the parameters of a single equation in a complete system of stochastic equations. The Annals of mathematical statistics , 20(1):46--63
work page 1949
-
[3]
Andrews, I. (2016). Conditional linear combination tests for weakly identified models. Econometrica , 84(6):2155--2182
work page 2016
-
[4]
Andrews, I. and Mikusheva, A. (2016). Andrews-Mikusheva(2016) Conditional inference with a functional nuisance parameter . Econometrica , 84(4):1571--1612
work page 2016
-
[5]
Andrews, I., Stock, J. H., and Sun, L. (2019). Weak instruments in instrumental variables regression: Theory and practice. Annual Review of Economics , 11(1):727--753
work page 2019
-
[6]
Angrist, J. D. and Pischke, J.-S. (2008). Mostly harmless econometrics: An empiricist's companion . Princeton university press
work page 2008
-
[7]
Berman, N., Couttenier, M., Rohner, D., and Thoenig, M. (2017). This mine is mine! how minerals fuel conflicts in africa. American Economic Review , 107(6):1564--1610
work page 2017
-
[8]
Boot, T. and Ligtenberg, J. W. (2023). Identification- and many instrument-robust inference via invariant moment conditions. arXiv:2303.07822
arXiv 2023
Show all 55 references
-
[9]
Brollo, F., Nannicini, T., Perotti, R., and Tabellini, G. (2013). The political resource curse. American Economic Review , 103(5):1759--1796
2013
-
[10]
Campante, F. R. and Do, Q.-A. (2014). Isolated capital cities, accountability, and corruption: Evidence from us states. American Economic Review , 104(8):2456--2481
2014
-
[11]
N., Long, J
Condra, L. N., Long, J. D., Shaver, A. C., and Wright, A. L. (2018). The logic of insurgent electoral violence. American Economic Review , 108(11):3199--3231
2018
-
[12]
Crudu, F., Mellace, G., and S \'a ndor, Z. (2021). Inference in instrumental variable models with heteroskedasticity and many instruments. Econometric Theory , 37(2):281--310
2021
-
[13]
and MacKinnon, J
Davidson, R. and MacKinnon, J. G. (2008). Davidson-Mackinnon(2008) Bootstrap inference in a linear equation estimated by instrumental variables . The Econometrics Journal , 11(3):443--477
2008
-
[14]
and MacKinnon, J
Davidson, R. and MacKinnon, J. G. (2010). Davidson-Mackinnon(2010) Wild bootstrap tests for IV regression . Journal of Business & Economic Statistics , 28(1):128--144
2010
-
[15]
and MacKinnon, J
Davidson, R. and MacKinnon, J. G. (2014). Davidson-Mackinnon(2014b) Bootstrap confidence sets with weak instruments . Econometric Reviews , 33(5-6):651--675
2014
-
[16]
Decarolis, F. (2015). Medicare part d: Are insurers gaming the low income subsidy design? American Economic Review , 105(4):1547--1580
2015
-
[17]
Di Tella, R., Perez-Truglia, R., Babino, A., and Sigman, M. (2015). Conveniently upset: avoiding altruism by distorting beliefs about others' altruism. American Economic Review , 105(11):3416--3442
2015
-
[18]
B., and Mavroeidis, S
Dov \` , M.-S., Kock, A. B., and Mavroeidis, S. (2024). A ridge-regularized jackknifed anderson-rubin test. Journal of Business & Economic Statistics , pages 1--12
2024
-
[19]
Dufour, J.-M. (1997). Some impossibility theorems in econometrics, with applications to structural and dynamic models. Econometrica , 65(6):1365--1389
1997
-
[20]
and Reguant, M
Fabra, N. and Reguant, M. (2014). Pass-through of emissions costs in electricity markets. American Economic Review , 104(9):2872--2899
2014
-
[21]
Field, E., Pande, R., Papp, J., and Rigol, N. (2013). Does the classic microfinance model discourage entrepreneurship among the poor? experimental evidence from india. American Economic Review , 103(6):2196--2226
2013
-
[22]
and Magnusson, L
Finlay, K. and Magnusson, L. M. (2019). Finlay-Magnusson(2019) Two applications of wild bootstrap methods to improve inference in cluster-IV models . Journal of Applied Econometrics , 34(6):911--933
2019
-
[23]
Hornung, E. (2014). Immigration and the diffusion of technology: The huguenot diaspora in prussia. American Economic Review , 104(1):84--122
2014
-
[24]
and Wang, W
Kaffo, M. and Wang, W. (2017). Kaffo-Wang(2017) On bootstrap validity for specification testing with many weak instruments . Economics Letters , 157:107--111
2017
-
[25]
and Neal, T
Keane, M. and Neal, T. (2023). Instrument strength in iv estimation and inference: A guide to theory and practice. Journal of Econometrics , 235(2):1625--1653
2023
-
[26]
Keane, M. P. and Neal, T. (2024). A practical guide to weak instruments. Annual Review of Economics , 16
2024
-
[27]
Kleibergen, F. (2002). Kleibergen(2002) Pivotal statistics for testing structural parameters in instrumental variables regression . Econometrica , 70(5):1781--1803
2002
-
[28]
S., McCrary, J., Moreira, M
Lee, D. S., McCrary, J., Moreira, M. J., and Porter, J. (2022). Valid t-ratio inference for iv. American Economic Review , 112(10):3260--3290
2022
-
[29]
S., McCrary, J., Moreira, M
Lee, D. S., McCrary, J., Moreira, M. J., Porter, J. R., and Yap, L. (2023). What to do when you can't use'1.96'confidence intervals for iv. Technical report, National Bureau of Economic Research
2023
-
[30]
Lim, D., Wang, W., and Zhang, Y. (2024a). A conditional linear combination test with many weak instruments. Journal of Econometrics , 238(2):105602
2024
-
[31]
Lim, D., Wang, W., and Zhang, Y. (2024b). A dimension-agnostic bootstrap anderson-rubin test for instrumental variable regressions. arXiv preprint arXiv:2412.01603
2024
-
[32]
MacKinnon, J. G. (2023). Fast cluster bootstrap methods for linear regression models. Econometrics and Statistics , 26:52--71
2023
-
[33]
and Otsu, T
Matsushita, Y. and Otsu, T. (2024). A jackknife lagrange multiplier test with many weak instruments. Econometric Theory , 40(2):447--470
2024
-
[34]
and Sun, L
Mikusheva, A. and Sun, L. (2022). Inference with many weak instruments. Review of Economic Studies , 89(5):2663--2686
2022
-
[35]
Moreira, M. J. (2003). Moreira(2003) A conditional likelihood ratio test for structural models . Econometrica , 71(4):1027--1048
2003
-
[36]
Moreira, M. J. (2009). Tests with correct size when instruments can be abitrarily weak. 152:131--140
2009
-
[37]
J., Porter, J., and Suarez, G
Moreira, M. J., Porter, J., and Suarez, G. (2009). Moreira-Porter-Suarez(2009) Bootstrap validity for the score test when instruments may be weak . 149(1):52--64
2009
-
[38]
Moser, P., Voena, A., and Waldinger, F. (2014). German jewish \'e migr \'e s and us invention. American Economic Review , 104(10):3222--3255
2014
-
[39]
Navjeevan, M. (2023). An identification and dimensionality robust test for instrumental variables models. arXiv preprint arXiv:2311.14892
2023 arXiv
-
[40]
and Qian, N
Nunn, N. and Qian, N. (2014). Us food aid and civil conflict. American economic review , 104(6):1630--1666
2014
-
[41]
Ottaviano, G. I. P., Peri, G., and Wright, G. C. (2013). Immigration, offshoring, and american jobs. American Economic Review , 103(5):1925--1959
2013
-
[42]
Rao, G. (2019). Familiarity does not breed contempt: Generosity, discrimination, and diversity in delhi schools. American Economic Review , 109(3):774--809
2019
-
[43]
, MacKinnon, J
Roodman, D., Nielsen, M. ., MacKinnon, J. G., and Webb, M. D. (2019). Roodman-Nielsen-MacKinnon-Webb(2019) Fast and wild: Bootstrap inference in Stata using boottest . The Stata Journal , 19(1):4--60
2019
-
[44]
repeat income
Rosenthal, S. S. (2014). Are private markets and filtering a viable source of low-income housing? estimates from a “repeat income” model. American Economic Review , 104(2):687--706
2014
-
[45]
and Stock, J
Staiger, D. and Stock, J. H. (1997). Instrumental variables regression with weak instruments. 65(3):557--586
1997
-
[46]
Steinwender, C. (2018). Real effects of information frictions: When the states and the kingdom became united. American Economic Review , 108(3):657--696
2018
-
[47]
Stock, J. H. and Wright, J. H. (2000). GMM with weak identification. Econometrica , 68(5):1055--1096
2000
-
[48]
Stock, J. H. and Yogo, M. (2002). Testing for weak instruments in linear iv regression
2002
-
[49]
a nder, N. and Voth, H.-J. (2013). How the west “invented
Voigtl \"a nder, N. and Voth, H.-J. (2013). How the west “invented” fertility restriction. American Economic Review , 103(6):2227--2264
2013
-
[50]
and Doko Tchatoka, F
Wang, W. and Doko Tchatoka, F. (2018). On bootstrap inconsistency and bonferroni-based size-correction for the subset anderson--rubin test under conditional homoskedasticity. Journal of Econometrics , 207(1):188--211
2018
-
[51]
and Kaffo, M
Wang, W. and Kaffo, M. (2016). Bootstrap inference for instrumental variable models with many weak instruments. Journal of Econometrics , 192(1):231--268
2016
-
[52]
and Liu, Q
Wang, W. and Liu, Q. (2015). Bootstrap-based selection for instrumental variables model. Economics Bulletin , 35(3):1886--1896
2015
-
[53]
and Zhang, Y
Wang, W. and Zhang, Y. (2024). Wild bootstrap inference for instrumental variables regressions with weak and few clusters. Journal of Econometrics , 241(1):105727
2024
-
[54]
Yap, L. (2023). Valid wald inference with many weak instruments. arXiv preprint arXiv:2311.15932
2023 arXiv
-
[55]
Young, A. (2022). Consistency without inference: Instrumental variables in practical application. European Economic Review , 147:104112
2022
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.