{"id":"2d157191-20a6-443e-85c2-bbd9a89c0ca3","arxiv_id":"2506.18001","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Anderson-Rubin and tF procedures are compared on AER data and simulations, with Anderson-Rubin showing higher power and shorter confidence intervals in most specifications.","lead":"This paper compares two statistical methods for handling weak instruments in causal studies: the classic Anderson-Rubin test and the newer tF method. Using data from the American Economic Review and simulations, it finds the Anderson-Rubin method typically delivers higher power and shorter confidence intervals.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CI-length comparison drops unbounded AR confidence sets; weak-instrument cases where AR length is infinite are excluded, biasing the 'AR shorter' result.","rationale":"The reader's weakest assumption was that the 151-specification subsample is representative. My concern is more specific and more damaging: within that subsample, the CI-length comparison further drops cases with missing length information, and the most plausible mechanism is that these are exactly the weak-instrument cases where AR confidence sets are unbounded. This makes the sample selection endogenous to the outcome being measured, not just a matter of external representativeness. The power simulations in Appendix A and the significance comparisons in Section 2.2 provide independent support for the power component of the paper's claim, and the paper does ultimately acknowledge Lee et al.'s theoretical expected-length result. However, the headline empirical finding on CI length rests on a comparison that likely excludes the very cases where AR performs worst. This strengthens the case for the reader's conditional verdict: the paper needs to report how many AR confidence sets are unbounded, include them in the comparison, and qualify the CI-length conclusion accordingly. The concern does not change the verdict from CONDITIONAL; it sharpens the reason why the paper should not be accepted as is.","tokens_in":11886,"tokens_out":6122,"duration_ms":72071,"concrete_test":"Reconstruct the excluded specifications from the Lee et al. (2022) replication data using the same selection criteria. For each of the 22 (5%) and 24 (1%) dropped specifications, invert the AR statistic to determine whether the AR confidence set is unbounded, empty, or bounded. If most dropped cases are unbounded AR sets with first-stage F below approximately 3.84, code unbounded AR sets as 'not shorter' and recompute the proportion of the full 151-specification sample where AR is shorter at both levels. Report both the full-sample proportion and the bounded-only proportion to quantify the selection bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.4 compares CI lengths after dropping 'missing CI length information' (151 -> 127 at the 5% level, 151 -> 123 at the 1% level), but the paper never states why these lengths are missing. In a just-identified IV model, the AR confidence set is the solution set of a quadratic inequality whose leading coefficient is proportional to the first-stage F-statistic minus the AR critical value. When the first-stage F is below that critical value (roughly 3.84 at the 5% level), the AR confidence set is unbounded and its length is infinite. The tF confidence interval, being a centered t-ratio interval, remains bounded. Thus the excluded specifications are likely concentrated among weak-instrument cases where AR has infinite length and cannot be shorter than tF. Comparing only bounded cases removes exactly the observations where tF is shorter, mechanically inflating the reported 96.85% and 95.93% figures. The same selection affects Figure 3, so the heatmap describes only bounded-AR specifications. This is not merely a representativeness concern: the paper's own measure, ln(length_tF/length_AR), is undefined when AR is unbounded, and the sample restriction is endogenous to the quantity being compared. The central claim that AR yields shorter CIs is therefore not established for the weak-instrument applications that motivate the paper.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares the Anderson–Rubin (AR) procedure and the tF procedure of Lee et al. (2022) for just-identified instrumental variable models, using the American Economic Review replication dataset (439 specifications, reduced to 151 usable specifications) and Monte Carlo simulations. It reports that AR tends to have higher power and shorter confidence intervals than tF, and concludes that the two procedures are complementary.","tokens_in":12097,"tokens_out":4919,"duration_ms":51702,"significance":"If the empirical findings are robust, they would provide useful guidance for applied researchers choosing among weak-IV-robust procedures. The paper offers a clear empirical comparison of significance outcomes and simulation power curves, which are of interest to the empirical IV literature. However, the main CI-length claim is vulnerable to a sample selection issue that is endogenous to the outcome, which limits the force of the paper's central message.","major_comments":[{"comment":"The paper states that 'After excluding cases with missing CI length information, 127 specifications remain at the 5% significance level and 123 at the 1% level,' but it never defines what makes CI length information 'missing.' In a just-identified IV model, the AR confidence set is the solution set of a quadratic inequality whose leading coefficient is proportional to the first-stage F-statistic minus the AR critical value; when F is below that critical value (approximately 3.84 at the 5% level), the AR set is unbounded and has infinite length, while the tF interval remains bounded. The reported 96.85% and 95.93% shares of specifications in which AR is shorter are therefore computed only over specifications with bounded AR sets. This is an endogenous selection on the outcome being compared: the dropped specifications are likely those with weak instruments where AR cannot be shorter than tF. The authors should report the number of excluded specifications with unbounded AR sets, their F-statistic distribution, and redo the comparison including these cases (e.g., treating their length as infinite) or provide a sensitivity analysis that does not condition on boundedness.","section":"Section 2.4, Figure 2"},{"comment":"The heatmap in Figure 3 is based on the same restricted sample of 127 and 123 specifications. The paper claims that AR delivers substantially shorter CIs 'particularly in areas with low first-stage F-statistics,' but this is exactly the region where unbounded AR sets are most likely, so the figure describes only the bounded-AR subset. Consequently, the figure does not support the conclusion that AR is shorter in weak-instrument settings; it may reflect the selection of specifications for which AR happens to be bounded. The authors should either include the unbounded cases in the comparison or clearly temper the interpretation of the heatmap.","section":"Section 2.4, Figure 3"},{"comment":"The paper does not fully document the sample selection from 439 to 343 to 151 specifications. It states that cases with multiple instruments or nonlinear structures were removed, and that the remaining exclusions were 'primarily due to limitations related to data confidentiality,' but no detailed breakdown by study is provided. Because the 151-specification sample is a non-random subset (e.g., private-data studies are excluded), the representativeness of the significance and CI-length comparisons is unclear. The authors should provide a table listing each excluded study and the reason for exclusion, and discuss how the exclusions might affect the comparison.","section":"Section 2.1"}],"minor_comments":[{"comment":"The axis tick labels in Figure 1 appear garbled in the manuscript (e.g., '01.9622.5762∞t2 statistic'); the figure should be redrawn with clear tick labels.","section":"Figure 1"},{"comment":"The simulation section refers to 'the simulations of Lee et al. (2022)' but does not provide the exact DGP equations or the implementation details of the tF procedure; additional details would aid reproducibility.","section":"Online Appendix A"},{"comment":"The reference to Moreira (2009) contains a typo ('abitrarily' should be 'arbitrarily') and appears to have an unusual citation format ('152:131–140'); please verify the full reference details.","section":"References"},{"comment":"The conclusion notes that tF has a theoretical advantage in expected CI length, but the paper does not reconcile this with the empirical finding of shorter AR intervals in the sample; a brief discussion of why the theoretical expectation may not hold in observed samples would strengthen the paper.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The CI-length comparison is the paper's most striking empirical result, and the unexplained exclusion of unbounded AR confidence sets is a serious issue that goes to the validity of the main claim. I would encourage the editor to require the authors to address this concern directly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Wenze Li compares the Anderson-Rubin procedure against Lee et al.'s tF method on a large AER replication sample and simulations. The significance-testing comparison is new and useful: among tF-insignificant specifications, 47.5% are AR-significant, and never the reverse. That part is credible and worth having.\n\nThe power simulations are also fine, though they largely confirm prior work. The paper is honest that tF has a theoretical expected-length advantage.\n\nThe soft spot is the CI-length comparison. The paper drops 'missing CI length information' without saying what that means. In a just-identified model, the AR confidence set is unbounded when the first-stage F is below the AR critical value. Those are precisely the weak-instrument cases that motivate the paper, and their length is infinite, not missing. By excluding them, the 96.85% and 95.93% figures are mechanically inflated: you are comparing only the bounded cases where AR can't lose by infinite length. The heatmap in Figure 3 has the same selection. This is not a minor representativeness caveat; it undermines the central claim that AR yields shorter CIs in weak-instrument applications.\n\nThe paper should either report the frequency of unbounded AR sets, analyze the CI comparison including infinite lengths as a separate category, or restrict the claim to bounded sets. The significance-testing results and power curves are not affected by this problem, so the paper is not without value.\n\nThere is also no replication package, and the sample shrinks from 439 to 151 specifications with only a vague 'data confidentiality' reason. That is fixable but needs addressing.\n\nBottom line: this is a serious empirical exercise that deserves refereeing, but the CI-length section needs a substantial rewrite before the paper can be trusted. I would send it to review with that as the main request.","headline":"Useful empirical comparison of AR vs tF, but the CI-length comparison drops exactly the weak-instrument cases where AR intervals are unbounded, so that headline result is not established.","tokens_in":12658,"tokens_out":1731,"would_cite":false,"duration_ms":19383,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P20","62F03","62F25"],"pacs":[],"model":"deepseek-v4-flash","headline":"Using AER replication data and simulations, this paper establishes that the classical Anderson-Rubin test typically delivers higher power and shorter confidence intervals than the newer tF t-ratio procedure in just-identified…","keywords":["instrumental variables","weak identification","just-identified models","Anderson-Rubin test","tF procedure","confidence intervals","power comparison","AER replication data"],"falsifier":"A second empirical study using a different large sample of just-identified IV specifications—for example, replication data from other journals or from private-data studies—that found tF rejecting at least as often as AR and yielding shorter intervals in a majority of cases would directly contradict the paper's central empirical claim. Alternatively, a Monte Carlo exercise covering non-normal errors, heteroskedasticity, or asymmetric instrument strength in which tF's power curves dominate AR's would weaken the claimed generality.","tokens_in":11645,"feed_emoji":"📊","tokens_out":8627,"duration_ms":75841,"temperature":0.7,"pith_summary":"The paper asks which of two weak-instrument-robust procedures applied researchers should trust in just-identified instrumental-variable regressions: the classical Anderson-Rubin (AR) test or the newer tF method, which replaces the usual normal critical value with a smooth function of the first-stage F-statistic. Using 151 specifications from 17 AER studies and Monte Carlo simulations, it finds that AR typically rejects the null more often and produces shorter confidence intervals than tF, especially when instruments are weak. The gap is large: among specifications tF fails to reject at the 5% level, AR rejects 47.52% of the time, and AR intervals are shorter in roughly 96% of specifications. The paper nevertheless notes that tF has a theoretical edge in expected interval length and concludes the two procedures are complementary.","feed_headline":"Anderson-Rubin test beats tF on power and intervals","feed_subtitle":"In AER replication data, AR rejects 47.5% of specs tF misses and gives shorter intervals in about 96% of cases.","key_machinery":"The comparison is carried by two inference procedures for a just-identified linear IV model. The Anderson-Rubin test is a weak-identification-robust test of the structural coefficient that remains valid regardless of first-stage strength. The tF procedure, proposed by Lee et al. (2022), preserves the familiar 2SLS t-ratio but calibrates its critical value through a smooth function of the first-stage F-statistic, giving uniform size control. The empirical machinery is the AER replication dataset—the same one used by Lee et al. (2022)—restricted to 151 specifications from 17 studies that are just-identified, public-data, and free of nonlinear structures, alongside 250,000-replication Monte Carlo experiments matching the original tF study's DGP.","core_discovery":"On the paper's own terms, the central discovery is empirical: in a replication sample of just-identified IV specifications from AER articles published 2013–2019, the Anderson-Rubin procedure dominates the tF procedure in both significance testing and interval length. Among specifications insignificant under tF, 47.52% are significant under AR, and there is no specification in which tF is significant at a more stringent level than AR. AR confidence intervals are shorter than tF intervals in 96.85% of specifications at the 5% level and 95.93% at the 1% level; when AR intervals are longer the loss is modest, while when they are shorter the gain is often substantial. Monte Carlo power curves show AR weakly dominating tF at low endogeneity and achieving higher power at low instrument strength.","pith_inferences":["An editorial extension: tF's conservativeness in this sample may result from its smooth critical value function being conservative at low F values, and a bootstrap-calibrated tF could close much of the power gap; testing this would require re-running the rejection counts with bootstrap critical values.","Another inference: because the 151-specification sample excludes private-data studies, the ranking could change if private-data replications tend to have stronger or weaker instruments; a replication using such datasets would be a direct external-validity check.","A third inference: the Monte Carlo design uses homoskedastic normal errors, so the power ordering might differ under heteroskedasticity or heavy-tailed errors, where both procedures' asymptotic critical values can be distorted; a simulation with such error structures would be informative.","Fourth: the CI-length comparison drops cases where AR intervals are unbounded; treating unbounded intervals explicitly (e.g., with a loss function) might reduce AR's apparent dominance."],"forward_implications":["Applied researchers in just-identified IV settings with weak instruments can expect the AR test to detect more nonzero effects than tF at the same nominal size, since AR rejects 47.52% of specifications that tF does not.","AR confidence intervals will typically be much shorter than tF intervals, particularly when the first-stage F-statistic is low; the paper estimates AR intervals are shorter in 96.85% of specifications at the 5% level.","Because tF retains a theoretical expected-length advantage, the two procedures are best viewed as complementary: AR for power-oriented testing and tF for cases where its expected length is preferred.","The empirical ranking aligns with the recommendations of Keane and Neal (2023, 2024) that applied work should prefer AR-based inference over t-ratio-based inference.","The same empirical protocol can be applied to refinements such as the VtF procedure or bootstrap versions of tF to see whether the AR advantage persists."],"supporting_citations":[{"why":"Supplies the tF procedure being compared and the AER replication dataset that the empirical analysis reuses.","marker":"Lee et al. (2022)"},{"why":"Defines the Anderson-Rubin test that is the baseline comparison procedure.","marker":"Anderson and Rubin (1949)"},{"why":"Provides the weak-instrument asymptotics under which AR maintains correct size, justifying its role as the reference.","marker":"Staiger and Stock (1997)"},{"why":"Shows AR is the uniformly most powerful unbiased test in just-identified models, giving the paper its theoretical benchmark.","marker":"Moreira (2009)"},{"why":"Recommend AR over t-ratio-based procedures; the paper's findings corroborate this recommendation.","marker":"Keane and Neal (2023, 2024)"},{"why":"Documents the prevalence of weak instruments (F below 10) in AER specifications, motivating the comparison.","marker":"Andrews et al. (2019)"}],"fun_headline_variants":["AR beats tF on power and interval length","Anderson-Rubin outperforms tF in weak-IV specs","AR intervals shorter and more powerful than tF","tF lags AR in significance and interval length","Replication shows AR dominates tF for weak IVs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 151 specifications from 17 AER studies that survived exclusions (public data, single instrument, linear structure, available CI lengths) are representative of just-identified IV applications generally, so that the observed AR-over-tF ranking is not an artifact of which studies were left out.","fun_headline_variants_meta":{"raw":{"variants":["AR beats tF on power and interval length","Anderson-Rubin outperforms tF in weak-IV specs","AR intervals shorter and more powerful than tF","tF lags AR in significance and interval length","Replication shows AR dominates tF for weak IVs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000745,"raw_usage":{"total_tokens":3287,"prompt_tokens":876,"completion_tokens":2411,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":2335}},"tokens_in":492,"tokens_out":2411,"duration_ms":16767,"temperature":1.0,"reasoning_tokens":2335,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:23:15.639713+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A second empirical study using a different large sample of just-identified IV specifications—for example, replication data from other journals or from private-data studies—that found tF rejecting at least as often as AR and yielding shorter intervals in a majority of cases would directly contradict the paper's central empirical claim. Alternatively, a Monte Carlo exercise covering non-normal errors, heteroskedasticity, or asymmetric instrument strength in which tF's power curves dominate AR's would weaken the claimed generality.","supporting_citations":[{"cited_title":"S., McCrary, J., Moreira, M","cited_arxiv_id":null,"evidence_quote":"Supplies the tF procedure being compared and the AER replication dataset that the empirical analysis reuses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Anderson-Rubin test that is the baseline comparison procedure."},{"cited_title":"and Stock, J","cited_arxiv_id":null,"evidence_quote":"Provides the weak-instrument asymptotics under which AR maintains correct size, justifying its role as the reference."},{"cited_title":"and Neal, T","cited_arxiv_id":null,"evidence_quote":"Recommend AR over t-ratio-based procedures; the paper's findings corroborate this recommendation."},{"cited_title":"H., and Sun, L","cited_arxiv_id":null,"evidence_quote":"Documents the prevalence of weak instruments (F below 10) in AER specifications, motivating the comparison."}],"review_version":1}