Pith. sign in

REVIEW 3 major objections 4 minor 55 references

An Empirical Comparison of Weak-IV-Robust Procedures in Just-Identified Models

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Using AER replication data and simulations, this paper establishes that the classical Anderson-Rubin test typically delivers higher power and shorter confidence intervals than the newer tF t-ratio procedure in just-identified…

desk verdict Useful empirical comparison of AR vs tF, but the CI-length comparison drops exactly the weak-instrument cases where AR intervals are unbounded, so that headline result is not established. read the letter →

arxiv 2506.18001 v1 pith:KPHVZVRW submitted 2025-06-22 econ.EM

classification econ.EM MSC 62P2062F0362F25
keywords instrumentalvariablesweakidentificationjust-identifiedmodelsAnderson-RubintesttFprocedureconfidenceintervalspowercomparisonAERreplicationdata
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks which of two weak-instrument-robust procedures applied researchers should trust in just-identified instrumental-variable regressions: the classical Anderson-Rubin (AR) test or the newer tF method, which replaces the usual normal critical value with a smooth function of the first-stage F-statistic. Using 151 specifications from 17 AER studies and Monte Carlo simulations, it finds that AR typically rejects the null more often and produces shorter confidence intervals than tF, especially when instruments are weak. The gap is large: among specifications tF fails to reject at the 5% level, AR rejects 47.52% of the time, and AR intervals are shorter in roughly 96% of specifications. The paper nevertheless notes that tF has a theoretical edge in expected interval length and concludes the two procedures are complementary.

What carries the argument

The comparison is carried by two inference procedures for a just-identified linear IV model. The Anderson-Rubin test is a weak-identification-robust test of the structural coefficient that remains valid regardless of first-stage strength. The tF procedure, proposed by Lee et al. (2022), preserves the familiar 2SLS t-ratio but calibrates its critical value through a smooth function of the first-stage F-statistic, giving uniform size control. The empirical machinery is the AER replication dataset—the same one used by Lee et al. (2022)—restricted to 151 specifications from 17 studies that are just-identified, public-data, and free of nonlinear structures, alongside 250,000-replication Monte Carlo experiments matching the original tF study's DGP.

What would settle it

A second empirical study using a different large sample of just-identified IV specifications—for example, replication data from other journals or from private-data studies—that found tF rejecting at least as often as AR and yielding shorter intervals in a majority of cases would directly contradict the paper's central empirical claim. Alternatively, a Monte Carlo exercise covering non-normal errors, heteroskedasticity, or asymmetric instrument strength in which tF's power curves dominate AR's would weaken the claimed generality.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is empirical: in a replication sample of just-identified IV specifications from AER articles published 2013–2019, the Anderson-Rubin procedure dominates the tF procedure in both significance testing and interval length. Among specifications insignificant under tF, 47.52% are significant under AR, and there is no specification in which tF is significant at a more stringent level than AR. AR confidence intervals are shorter than tF intervals in 96.85% of specifications at the 5% level and 95.93% at the 1% level; when AR intervals are longer the loss is modest, while when they are shorter the gain is often substantial. Monte Carlo power curves show AR weakly dominating tF at low endogeneity and achieving higher power at low instrument strength.

Load-bearing premise

The 151 specifications from 17 AER studies that survived exclusions (public data, single instrument, linear structure, available CI lengths) are representative of just-identified IV applications generally, so that the observed AR-over-tF ranking is not an artifact of which studies were left out.

Editorial extensions

If this is right

  • Applied researchers in just-identified IV settings with weak instruments can expect the AR test to detect more nonzero effects than tF at the same nominal size, since AR rejects 47.52% of specifications that tF does not.
  • AR confidence intervals will typically be much shorter than tF intervals, particularly when the first-stage F-statistic is low; the paper estimates AR intervals are shorter in 96.85% of specifications at the 5% level.
  • Because tF retains a theoretical expected-length advantage, the two procedures are best viewed as complementary: AR for power-oriented testing and tF for cases where its expected length is preferred.
  • The empirical ranking aligns with the recommendations of Keane and Neal (2023, 2024) that applied work should prefer AR-based inference over t-ratio-based inference.
  • The same empirical protocol can be applied to refinements such as the VtF procedure or bootstrap versions of tF to see whether the AR advantage persists.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An editorial extension: tF's conservativeness in this sample may result from its smooth critical value function being conservative at low F values, and a bootstrap-calibrated tF could close much of the power gap; testing this would require re-running the rejection counts with bootstrap critical values.
  • Another inference: because the 151-specification sample excludes private-data studies, the ranking could change if private-data replications tend to have stronger or weaker instruments; a replication using such datasets would be a direct external-validity check.
  • A third inference: the Monte Carlo design uses homoskedastic normal errors, so the power ordering might differ under heteroskedasticity or heavy-tailed errors, where both procedures' asymptotic critical values can be distorted; a simulation with such error structures would be informative.
  • Fourth: the CI-length comparison drops cases where AR intervals are unbounded; treating unbounded intervals explicitly (e.g., with a loss function) might reduce AR's apparent dominance.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper compares the Anderson–Rubin (AR) procedure and the tF procedure of Lee et al. (2022) for just-identified instrumental variable models, using the American Economic Review replication dataset (439 specifications, reduced to 151 usable specifications) and Monte Carlo simulations. It reports that AR tends to have higher power and shorter confidence intervals than tF, and concludes that the two procedures are complementary.

Significance. If the empirical findings are robust, they would provide useful guidance for applied researchers choosing among weak-IV-robust procedures. The paper offers a clear empirical comparison of significance outcomes and simulation power curves, which are of interest to the empirical IV literature. However, the main CI-length claim is vulnerable to a sample selection issue that is endogenous to the outcome, which limits the force of the paper's central message.

major comments (3)
  1. [Section 2.4, Figure 2] The paper states that 'After excluding cases with missing CI length information, 127 specifications remain at the 5% significance level and 123 at the 1% level,' but it never defines what makes CI length information 'missing.' In a just-identified IV model, the AR confidence set is the solution set of a quadratic inequality whose leading coefficient is proportional to the first-stage F-statistic minus the AR critical value; when F is below that critical value (approximately 3.84 at the 5% level), the AR set is unbounded and has infinite length, while the tF interval remains bounded. The reported 96.85% and 95.93% shares of specifications in which AR is shorter are therefore computed only over specifications with bounded AR sets. This is an endogenous selection on the outcome being compared: the dropped specifications are likely those with weak instruments where AR cannot be shorter than tF. The authors should report the number of excluded specifications with unbounded AR sets, their F-statistic distribution, and redo the comparison including these cases (e.g., treating their length as infinite) or provide a sensitivity analysis that does not condition on boundedness.
  2. [Section 2.4, Figure 3] The heatmap in Figure 3 is based on the same restricted sample of 127 and 123 specifications. The paper claims that AR delivers substantially shorter CIs 'particularly in areas with low first-stage F-statistics,' but this is exactly the region where unbounded AR sets are most likely, so the figure describes only the bounded-AR subset. Consequently, the figure does not support the conclusion that AR is shorter in weak-instrument settings; it may reflect the selection of specifications for which AR happens to be bounded. The authors should either include the unbounded cases in the comparison or clearly temper the interpretation of the heatmap.
  3. [Section 2.1] The paper does not fully document the sample selection from 439 to 343 to 151 specifications. It states that cases with multiple instruments or nonlinear structures were removed, and that the remaining exclusions were 'primarily due to limitations related to data confidentiality,' but no detailed breakdown by study is provided. Because the 151-specification sample is a non-random subset (e.g., private-data studies are excluded), the representativeness of the significance and CI-length comparisons is unclear. The authors should provide a table listing each excluded study and the reason for exclusion, and discuss how the exclusions might affect the comparison.
minor comments (4)
  1. [Figure 1] The axis tick labels in Figure 1 appear garbled in the manuscript (e.g., '01.9622.5762∞t2 statistic'); the figure should be redrawn with clear tick labels.
  2. [Online Appendix A] The simulation section refers to 'the simulations of Lee et al. (2022)' but does not provide the exact DGP equations or the implementation details of the tF procedure; additional details would aid reproducibility.
  3. [References] The reference to Moreira (2009) contains a typo ('abitrarily' should be 'arbitrarily') and appears to have an unusual citation format ('152:131–140'); please verify the full reference details.
  4. [Section 3] The conclusion notes that tF has a theoretical advantage in expected CI length, but the paper does not reconcile this with the empirical finding of shorter AR intervals in the sample; a brief discussion of why the theoretical expectation may not hold in observed samples would strengthen the paper.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the comparison of AR and tF is an external empirical and simulation exercise, not a derivation from fitted inputs or self-citations.

full rationale

The paper's central claim is that the Anderson-Rubin (AR) procedure has higher power and yields shorter confidence intervals than the tF procedure in an external AER replication sample and in Monte Carlo simulations. No parameter is fitted and then renamed as a prediction; no equation or result is defined in terms of the conclusion it supports. The tF procedure and the AER dataset are taken from Lee et al. (2022), but the author of the present paper is not an author of that source, so this is a matter of provenance rather than load-bearing self-citation. The power simulations compare rejection rates under independently varied structural and nuisance parameters, and the empirical significance comparison directly computes AR and tF rejections on existing specifications. The only notable concern is that the confidence-interval length comparison excludes cases with 'missing CI length information' (151 to 127 at the 5% level and 123 at the 1% level), and the paper does not state why those lengths are missing; if these are cases where the AR confidence set is unbounded, the reported percentage of shorter AR intervals could be affected. That is a sample-selection and external-validity critique, not a circularity critique, because the reported measure is not constructed so as to force the conclusion. The paper also explicitly acknowledges the complementary theoretical advantage of tF (Lee et al., 2022), which further shows the empirical result is not presupposed. Overall, the derivation chain is self-contained and the analysis is best assessed on statistical grounds rather than circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper does not introduce free parameters or invented entities. The analysis relies on standard IV assumptions and on the validity of the AR and tF procedures as previously developed. The simulation parameters (f0, rho) are varied over grids and are not fitted to data.

assumptions (2)
  • domain assumption Each specification satisfies the IV assumptions: Cov(u,Z)=0 and Cov(Z,X)≠0, as stated in equations (1) and (2).
    The AR and tF procedures are valid only under these exclusion and relevance conditions; the paper relies on this for all specifications in the AER sample.
  • domain assumption The tF critical value function from Lee et al. (2022) is correctly implemented and its asymptotic properties hold in the finite samples studied.
    The paper applies the tF procedure without re-deriving it, assuming the published implementation and its size control properties carry over.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Empirical Comparison of Weak-IV-Robust Procedures in Just-Identified Models." pith.science (2026). https://pith.science/paper/KPHVZVRW

@misc{pith2026250618001,
  author       = {Pith},
  title        = {Pith review of: An Empirical Comparison of Weak-IV-Robust Procedures in Just-Identified Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KPHVZVRW}},
  note         = {Machine review of arXiv:2506.18001}
}
read the original abstract

Instrumental variable (IV) regression is recognized as one of the five core methods for causal inference, as identified by Angrist and Pischke (2008). This paper compares two leading approaches to inference under weak identification for just-identified IV models: the classical Anderson-Rubin (AR) procedure and the recently popular tF method proposed by Lee et al. (2022). Using replication data from the American Economic Review (AER) and Monte Carlo simulation experiments, we evaluate the two procedures in terms of statistical significance testing and confidence interval (CI) length. Empirically, we find that the AR procedure typically offers higher power and yields shorter CIs than the tF method. Nonetheless, as noted by Lee et al. (2022), tF has a theoretical advantage in terms of expected CI length. Our findings suggest that the two procedures may be viewed as complementary tools in empirical applications involving potentially weak instruments.

Figures

Figures reproduced from arXiv: 2506.18001 by the authors.

Figure 1
Figure 1. Statistical Significance of t-ratio, AR, and tF Notes: This figure is based on 151 specifications. The vertical axis plots t 2/1.962 1+t 2/1.962 , while the horizontal axis plots F/10 1+F/10 . For the AR test, black circles denote insignificance, blue circles indicate significance at the 5% level only, and red circles represent significance at the 1% level. The solid black line represents the 5% CV for the tF test, … view at source ↗
Figure 2
Figure 2. Distribution of ln  lengthtF lengthAR  Notes: Part (a) is based on 127 specifications at the 5% level, and part (b) on 123 specifications at the 1% level. Each bar represents the frequency of observations within a 0.05-width bin of the log difference in CI length. We find that the AR CI is shorter than that of the tF procedure in 96.85% of specifications at the 5% level and 95.93% at the 1% level. Among cases in w… view at source ↗
Figure 3
Figure 3. presents a heatmap of the log difference in CI lengths, plotted against two key diagnostic measures: the first-stage F-statistic (vertical axis) and the absolute value of the estimated residual correlation, |ρˆ| (horizontal axis). To facilitate visualization, the vertical axis applies the transformation F/10 1+F/10 , which maps the F-statistic onto the unit interval. Representative values of F = 104.67, 10, 2.5762 ,… view at source ↗
Figures from the paper (1 more)
Figure 3
Figure 3. Figure 3: Heatmap of ln  lengthtF lengthAR  (continued) Notes: The vertical axis applies the transformation F/10 1+F/10 to map the first-stage F-statistic onto the unit interval. Under this scaling, F values of 104.67, 10, 2.5762 , and 1.962 correspond to vertical positions of…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

55 extracted references · 51 canonical work pages

  1. [1]

    Alesina, A., Stantcheva, S., and Teso, E. (2018). Intergenerational mobility and preferences for redistribution. American Economic Review , 108(2):521--554

  2. [2]

    Anderson, T. W. and Rubin, H. (1949). Estimation of the parameters of a single equation in a complete system of stochastic equations. The Annals of mathematical statistics , 20(1):46--63

  3. [3]

    Andrews, I. (2016). Conditional linear combination tests for weakly identified models. Econometrica , 84(6):2155--2182

  4. [4]

    and Mikusheva, A

    Andrews, I. and Mikusheva, A. (2016). Andrews-Mikusheva(2016) Conditional inference with a functional nuisance parameter . Econometrica , 84(4):1571--1612

  5. [5]

    H., and Sun, L

    Andrews, I., Stock, J. H., and Sun, L. (2019). Weak instruments in instrumental variables regression: Theory and practice. Annual Review of Economics , 11(1):727--753

  6. [6]

    Angrist, J. D. and Pischke, J.-S. (2008). Mostly harmless econometrics: An empiricist's companion . Princeton university press

  7. [7]

    Berman, N., Couttenier, M., Rohner, D., and Thoenig, M. (2017). This mine is mine! how minerals fuel conflicts in africa. American Economic Review , 107(6):1564--1610

  8. [8]

    and Ligtenberg, J

    Boot, T. and Ligtenberg, J. W. (2023). Identification- and many instrument-robust inference via invariant moment conditions. arXiv:2303.07822

Show all 55 references
  1. [9]

    Brollo, F., Nannicini, T., Perotti, R., and Tabellini, G. (2013). The political resource curse. American Economic Review , 103(5):1759--1796

  2. [10]

    Campante, F. R. and Do, Q.-A. (2014). Isolated capital cities, accountability, and corruption: Evidence from us states. American Economic Review , 104(8):2456--2481

  3. [11]

    N., Long, J

    Condra, L. N., Long, J. D., Shaver, A. C., and Wright, A. L. (2018). The logic of insurgent electoral violence. American Economic Review , 108(11):3199--3231

  4. [12]

    Crudu, F., Mellace, G., and S \'a ndor, Z. (2021). Inference in instrumental variable models with heteroskedasticity and many instruments. Econometric Theory , 37(2):281--310

  5. [13]

    and MacKinnon, J

    Davidson, R. and MacKinnon, J. G. (2008). Davidson-Mackinnon(2008) Bootstrap inference in a linear equation estimated by instrumental variables . The Econometrics Journal , 11(3):443--477

  6. [14]

    and MacKinnon, J

    Davidson, R. and MacKinnon, J. G. (2010). Davidson-Mackinnon(2010) Wild bootstrap tests for IV regression . Journal of Business & Economic Statistics , 28(1):128--144

  7. [15]

    and MacKinnon, J

    Davidson, R. and MacKinnon, J. G. (2014). Davidson-Mackinnon(2014b) Bootstrap confidence sets with weak instruments . Econometric Reviews , 33(5-6):651--675

  8. [16]

    Decarolis, F. (2015). Medicare part d: Are insurers gaming the low income subsidy design? American Economic Review , 105(4):1547--1580

  9. [17]

    Di Tella, R., Perez-Truglia, R., Babino, A., and Sigman, M. (2015). Conveniently upset: avoiding altruism by distorting beliefs about others' altruism. American Economic Review , 105(11):3416--3442

  10. [18]

    B., and Mavroeidis, S

    Dov \` , M.-S., Kock, A. B., and Mavroeidis, S. (2024). A ridge-regularized jackknifed anderson-rubin test. Journal of Business & Economic Statistics , pages 1--12

  11. [19]

    Dufour, J.-M. (1997). Some impossibility theorems in econometrics, with applications to structural and dynamic models. Econometrica , 65(6):1365--1389

  12. [20]

    and Reguant, M

    Fabra, N. and Reguant, M. (2014). Pass-through of emissions costs in electricity markets. American Economic Review , 104(9):2872--2899

  13. [21]

    Field, E., Pande, R., Papp, J., and Rigol, N. (2013). Does the classic microfinance model discourage entrepreneurship among the poor? experimental evidence from india. American Economic Review , 103(6):2196--2226

  14. [22]

    and Magnusson, L

    Finlay, K. and Magnusson, L. M. (2019). Finlay-Magnusson(2019) Two applications of wild bootstrap methods to improve inference in cluster-IV models . Journal of Applied Econometrics , 34(6):911--933

  15. [23]

    Hornung, E. (2014). Immigration and the diffusion of technology: The huguenot diaspora in prussia. American Economic Review , 104(1):84--122

  16. [24]

    and Wang, W

    Kaffo, M. and Wang, W. (2017). Kaffo-Wang(2017) On bootstrap validity for specification testing with many weak instruments . Economics Letters , 157:107--111

  17. [25]

    and Neal, T

    Keane, M. and Neal, T. (2023). Instrument strength in iv estimation and inference: A guide to theory and practice. Journal of Econometrics , 235(2):1625--1653

  18. [26]

    Keane, M. P. and Neal, T. (2024). A practical guide to weak instruments. Annual Review of Economics , 16

  19. [27]

    Kleibergen, F. (2002). Kleibergen(2002) Pivotal statistics for testing structural parameters in instrumental variables regression . Econometrica , 70(5):1781--1803

  20. [28]

    S., McCrary, J., Moreira, M

    Lee, D. S., McCrary, J., Moreira, M. J., and Porter, J. (2022). Valid t-ratio inference for iv. American Economic Review , 112(10):3260--3290

  21. [29]

    S., McCrary, J., Moreira, M

    Lee, D. S., McCrary, J., Moreira, M. J., Porter, J. R., and Yap, L. (2023). What to do when you can't use'1.96'confidence intervals for iv. Technical report, National Bureau of Economic Research

  22. [30]

    Lim, D., Wang, W., and Zhang, Y. (2024a). A conditional linear combination test with many weak instruments. Journal of Econometrics , 238(2):105602

  23. [31]

    Lim, D., Wang, W., and Zhang, Y. (2024b). A dimension-agnostic bootstrap anderson-rubin test for instrumental variable regressions. arXiv preprint arXiv:2412.01603

  24. [32]

    MacKinnon, J. G. (2023). Fast cluster bootstrap methods for linear regression models. Econometrics and Statistics , 26:52--71

  25. [33]

    and Otsu, T

    Matsushita, Y. and Otsu, T. (2024). A jackknife lagrange multiplier test with many weak instruments. Econometric Theory , 40(2):447--470

  26. [34]

    and Sun, L

    Mikusheva, A. and Sun, L. (2022). Inference with many weak instruments. Review of Economic Studies , 89(5):2663--2686

  27. [35]

    Moreira, M. J. (2003). Moreira(2003) A conditional likelihood ratio test for structural models . Econometrica , 71(4):1027--1048

  28. [36]

    Moreira, M. J. (2009). Tests with correct size when instruments can be abitrarily weak. 152:131--140

  29. [37]

    J., Porter, J., and Suarez, G

    Moreira, M. J., Porter, J., and Suarez, G. (2009). Moreira-Porter-Suarez(2009) Bootstrap validity for the score test when instruments may be weak . 149(1):52--64

  30. [38]

    Moser, P., Voena, A., and Waldinger, F. (2014). German jewish \'e migr \'e s and us invention. American Economic Review , 104(10):3222--3255

  31. [39]

    Navjeevan, M. (2023). An identification and dimensionality robust test for instrumental variables models. arXiv preprint arXiv:2311.14892

  32. [40]

    and Qian, N

    Nunn, N. and Qian, N. (2014). Us food aid and civil conflict. American economic review , 104(6):1630--1666

  33. [41]

    Ottaviano, G. I. P., Peri, G., and Wright, G. C. (2013). Immigration, offshoring, and american jobs. American Economic Review , 103(5):1925--1959

  34. [42]

    Rao, G. (2019). Familiarity does not breed contempt: Generosity, discrimination, and diversity in delhi schools. American Economic Review , 109(3):774--809

  35. [43]

    , MacKinnon, J

    Roodman, D., Nielsen, M. ., MacKinnon, J. G., and Webb, M. D. (2019). Roodman-Nielsen-MacKinnon-Webb(2019) Fast and wild: Bootstrap inference in Stata using boottest . The Stata Journal , 19(1):4--60

  36. [44]

    repeat income

    Rosenthal, S. S. (2014). Are private markets and filtering a viable source of low-income housing? estimates from a “repeat income” model. American Economic Review , 104(2):687--706

  37. [45]

    and Stock, J

    Staiger, D. and Stock, J. H. (1997). Instrumental variables regression with weak instruments. 65(3):557--586

  38. [46]

    Steinwender, C. (2018). Real effects of information frictions: When the states and the kingdom became united. American Economic Review , 108(3):657--696

  39. [47]

    Stock, J. H. and Wright, J. H. (2000). GMM with weak identification. Econometrica , 68(5):1055--1096

  40. [48]

    Stock, J. H. and Yogo, M. (2002). Testing for weak instruments in linear iv regression

  41. [49]

    a nder, N. and Voth, H.-J. (2013). How the west “invented

    Voigtl \"a nder, N. and Voth, H.-J. (2013). How the west “invented” fertility restriction. American Economic Review , 103(6):2227--2264

  42. [50]

    and Doko Tchatoka, F

    Wang, W. and Doko Tchatoka, F. (2018). On bootstrap inconsistency and bonferroni-based size-correction for the subset anderson--rubin test under conditional homoskedasticity. Journal of Econometrics , 207(1):188--211

  43. [51]

    and Kaffo, M

    Wang, W. and Kaffo, M. (2016). Bootstrap inference for instrumental variable models with many weak instruments. Journal of Econometrics , 192(1):231--268

  44. [52]

    and Liu, Q

    Wang, W. and Liu, Q. (2015). Bootstrap-based selection for instrumental variables model. Economics Bulletin , 35(3):1886--1896

  45. [53]

    and Zhang, Y

    Wang, W. and Zhang, Y. (2024). Wild bootstrap inference for instrumental variables regressions with weak and few clusters. Journal of Econometrics , 241(1):105727

  46. [54]

    Yap, L. (2023). Valid wald inference with many weak instruments. arXiv preprint arXiv:2311.15932

  47. [55]

    Young, A. (2022). Consistency without inference: Instrumental variables in practical application. European Economic Review , 147:104112

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.