REVIEW 4 major objections 4 minor 1 cited by
Misspecification-Robust Shrinkage and Selection for VAR Forecasts and IRFs
T0 review · 4 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This paper claims that two information criteria give asymptotically unbiased estimates of prediction and impulse-response risk, so shrinkage, lag length, and estimator type can be jointly selected under misspecification.
desk verdict Solid, genuinely new IRF selection criterion that deserves referee time; the local-drift assumption is the main thing to push on. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a companion-form shrinkage posterior mean that weights the MLE or LFE toward a prior mean with precision $\lambda T$, embedded in a local-drift DGP that makes sampling variance, misspecification bias, and prior-induced bias all $O(T^{-1/2})$. The selection criteria are modified final-prediction-error statistics: an in-sample multi-step squared-error term plus twice an estimated covariance penalty between the candidate estimator and the unshrunk loss-function benchmark. For IRFs, the benchmark is the lag-$q$ local projection $\bar\Psi_T(lfe,0,q)$, and Theorem 4's centering property guarantees that this benchmark has no misspecification bias; the criterion then measures, with the added covariance penalty, how far a candidate VAR or LP estimate is from that correctly centered target.
What would settle it
In the paper's own Monte Carlo design (local DGP with $\alpha=2$, $T=5000$), compute the actual risk differential $R(\hat y_{T+h}(\iota,\lambda))-R(\hat y_{T+h}(\iota',\lambda'))$ and the expected criterion differential $E[PC_T(\iota,\lambda)-PC_T(\iota',\lambda')]$; if their gap does not vanish as $T$ grows, the unbiasedness claim fails. A second check: fix a non-drifting infinite-order VMA DGP and verify that the gap stops vanishing, confirming that the local-drift rate is the boundary of the theory.
Extended reading notes
Core claim
Under the drifting DGP in equations (13)–(15), where the true process is a stationary infinite-order VMA that approaches a finite-order VAR at rate $T^{-1/2}$ and the prior means approach the pseudo-true values at the same rate, the paper proves a limit distribution for the shrinkage estimators $\bar\Psi_T(\iota,\lambda)$ and decomposes the normalized prediction risk into a bias term and a variance term (Theorems 1 and 2). It then constructs $PC_T(\iota,\lambda,p)$ and $IRFC_T(\iota,\lambda,p)$ so that the key convergence results (25) and (51) hold: the expected difference between two criterion values converges to the difference between the corresponding true risks. For impulse responses, Theorem 4 shows that the local-projection estimator is correctly centered if and only if the lag length is at least $p^*+1$, which makes lag augmentation a selection rule rather than an ad hoc choice. The paper's message is that under misspecification, shrinkage, lag length, and estimator choice should be governed by an estimate of the target risk, not by marginal likelihood.
Load-bearing premise
The entire asymptotic argument rests on the load-bearing local-drift assumption: the true DGP must drift toward the finite-order VAR at rate $T^{-1/2}$, the prior means must drift toward pseudo-true values at the same rate, and the prior precision must grow as $\lambda T$; if the misspecification is fixed rather than local, the unbiasedness of $PC_T$ and $IRFC_T$ breaks down.
Editorial extensions
If this is right
- Forecasters using Bayesian VARs can replace marginal-data-density hyperparameter choice with $PC_T$; in the paper's simulations this gives nearly the same performance under correct specification and large risk reductions under misspecification.
- Researchers estimating impulse responses can use $IRFC_T$ to choose between VAR and local projections; the empirical analysis shows that neither estimator dominates, with LP chosen in 60–85% of samples depending on horizon and lag length.
- Lag augmentation becomes operational: because $IRFC_T$ incurs a large bias when $p<p^*$, the criterion will select enough lags to center the LP, while shrinkage is chosen to offset the variance cost.
- The risk-targeting logic also provides a principled way to average across estimators and hyperparameters, an extension the paper explicitly notes at the end.
Reading between the lines
- The same unbiased-risk construction would likely carry over to non-conjugate or data-based priors, provided the posterior mean remains a $T$-consistent weighted average and the prior displacement drifts at $T^{-1/2}$; one could test this by replacing the conjugate prior with a Minnesota-style prior in the same Monte Carlo design.
- Because the criteria estimate risk only up to a constant and remain stochastic in the limit, selection uncertainty is intrinsic; an implied extension is to report the distribution of the argmin over $\lambda$ and to build averaging weights rather than hard selections.
- The same principle could target other plug-in objects—cumulative multipliers, long-run responses, or forecast-error decompositions—by changing the quadratic form and the covariance penalty so that the criterion remains an unbiased estimate of the corresponding risk.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops two information criteria, PC and IRFC, for selecting the shrinkage hyperparameter, the VAR lag length, and the estimator type (MLE/iterated VAR versus LFE/local projection) when the VAR is potentially misspecified. The analysis is conducted in a local-to-VAR drifting DGP framework: the true infinite-order VMA drifts toward a finite-order VAR at rate T^{-1/2}, prior means drift toward pseudo-true values at the same rate, and prior precision scales with T. Under these assumptions, Theorem 1 gives the limit distribution of the shrinkage estimators, Theorem 2 gives the normalized prediction risk, Theorem 3 shows that the in-sample loss differential converges to a random variable whose expectation is the risk differential, and Theorem 4 establishes the lag-augmentation centering property for local projections. The proposed criteria jointly select (ι, λ, p); their finite-sample behavior is studied by Monte Carlo and their use for IRF selection is illustrated on 200 FRED-QD samples.
Significance. The contribution is potentially significant: it extends the local-misspecification selection framework of Schorfheide (2005) from plug-in MLE/LFE predictors to shrinkage estimators and to IRF estimation, and it provides the first criterion that jointly selects among VAR- and LP-based IRF estimates, shrinkage, and lag length. The derivation is internally consistent and the local asymptotics are stated explicitly rather than hidden. The paper also gives a transparent account of the key limitation that PC and IRFC are asymptotically unbiased but not consistent estimators of the relevant risk. The Monte Carlo and empirical sections are useful, and the connection to the recent local-projection literature (MPQW, LPW) is appropriately drawn. The main weakness is that the central claims about 'misspecification robustness' are established only under the local-drift design, and the small-sample evidence at T=100 is less supportive than the abstract suggests.
major comments (4)
- [Section 2.2, equations (13)-(15); Theorems 1-3; equation (25)] The unbiasedness of PC and IRFC is derived under the local-drift assumption that the DGP misspecification and the prior mean deviate from the VAR at rate T^{-1/2} and that prior precision grows as λT. If the misspecification is fixed rather than local, or if the prior center is fixed away from the pseudo-true value, the expansions in Theorem 1 and the risk-normalization underlying (25) and (51) break down. The paper is explicit that it operates in the local framework, so this is not an internal inconsistency; however, the title and abstract claim 'misspecification-robust' selection in a way that suggests broader scope. I recommend adding a concrete sensitivity analysis, for example simulating a fixed (non-drifting) misspecification magnitude and a fixed prior center, and comparing the PC-selected risk with the oracle risk, to show how large the approximation error can be in realistic samples.
- [Section 7.3 and Table A-3] The abstract states that 'once the VAR model is misspecified PC hyperparameter selection works significantly better than an MDD-based selection,' but the simulation evidence at T=100 does not support this unqualified statement. In Table A-3, under α=2 and T=100, MDD selection produces lower risk than PC selection in most configurations; for example, for LFE at h=4, p=1 the PC risk is 19% higher than MDD, and for the PC-selected lag length it is 31% higher. The online appendix acknowledges that 'for small sample sizes, MDD seems to do better than PC,' but this caveat is absent from the abstract and the main-text discussion. The authors should either qualify the claim to the asymptotically relevant sample sizes or provide an explanation of why the small-sample reversal does not undermine the practical recommendation.
- [Section 5, equation (43), Theorem 4, Definition 3] The IRFC criterion requires p* < q, where q is the maximum number of lags used in the benchmark LP. This assumption is needed so that the benchmark LP with q lags has the centering property M' μ(lfe,0,q) M = μ(irf); if p* = q, this equality fails and equation (51) no longer holds. The empirical application in Section 8 selects p=6, which is the maximum lag q, in the majority of samples (Figure 4). Thus the empirically relevant case p* = q is not merely hypothetical. The authors should discuss what happens when p* = q and, if possible, adapt the criterion or state clearly the additional assumptions needed to cover that case.
- [Section 7.1, Figure 3] At T=100, the top row of Figure 3 shows a large wedge between the Monte Carlo risk and both the asymptotic risk and the expected value of PC. The text attributes this to a difference between finite-sample and asymptotic variance of the estimated coefficients, but it does not assess whether this wedge materially degrades the selection properties of PC at small sample sizes. Given that the paper is aimed at empirical researchers who often have samples of this order, the authors should provide a quantitative assessment of the implied selection error, for example by reporting the frequency with which PC selects a λ far from the finite-sample-optimal value at T=100.
minor comments (4)
- [Section 4, after equation (33)] The dimension of the matrix M Υ_q is printed as nq × (n−1)q; it should be nq × n(q−1).
- [Section 1, final paragraph of the introduction] The sentence 'discrediting with widespread belief' is ungrammatical and should read 'discrediting the widespread belief.'
- [Section 7.1, third paragraph] The phrase 'pointmass at λ = ∞' should be written as 'point mass at λ = ∞.'
- [Throughout] The spacing in 'V AR' is inconsistent; sometimes it appears as 'V AR' and sometimes as 'VAR.' A consistent rendering would improve readability.
Circularity Check
No significant circularity: PC and IRFC are derived as asymptotically unbiased risk estimates under explicit local-misspecification assumptions, and the target selection decision is not an input to the risk estimates.
full rationale
The paper derives PC and IRFC from first principles under stated assumptions rather than by fitting the selection outcome. The key unbiasedness results in equations (25) and (51) follow from Theorem 3 and equations (47) and (50), which are proved from the limit distribution in Theorem 1 and the covariance formulas in the Online Appendix. The construction is a URE: the in-sample loss is corrected by a covariance penalty whose expectation converges to the covariance term, so the resulting criterion is unbiased for the normalized risk. The target quantities, namely the selected hyperparameter, lag length, and estimator type, are not used as inputs to define the risk estimates. The local-drift assumptions in equations (13)-(15) are explicit modeling assumptions that make the bias and variance terms of the same order; they are not conclusions derived from the criteria. The paper also relies on the lag-augmentation result in Theorem 4, which is proved directly using the Frisch-Waugh-Lovell theorem rather than assumed. Citations to Schorfheide (2005) and Montiel Olea, Plagborg-Moller, Qian, and Wolf (2024) provide external bases for the local-misspecification framework and centering properties, and the current paper extends those results with its own proofs. The paper honestly notes that PC and IRFC are asymptotically unbiased but not consistent because the misspecification is local, so there is no overclaim of consistency. Overall, the derivation chain is self-contained relative to its assumptions, and no load-bearing step reduces to its own inputs by construction.
Assumptions & free parameters
assumptions (6)
- domain assumption DGP is a stationary infinite-order VMA drifting toward a VAR at rate T^{-1/2}: y_t = F y_{t-1} + ε_t + (α/√T) Σ_{j=1}^∞ A_j ε_{t-j}.
- domain assumption The prior mean drifts to the pseudo-true value at T^{-1/2} and the prior precision scales as λT: Φ_T = F + T^{-1/2}φ, Ψ_T = F^h + T^{-1/2}ψ, λ̃_T = λT.
- domain assumption Assumption 1 regularity conditions: largest eigenvalue of F less than one, Σ j^2 ||A_j|| < ∞, moments and Lipschitz conditions on ε_t.
- domain assumption The asymptotic lag order p* is strictly less than the maximum lag q (p* < q).
- domain assumption The shock identification matrix Ξ is known and the reduced-form covariance matrix is consistently estimated; local misspecification does not affect identification.
- standard math The prediction risk is evaluated with two independent copies of the DGP, one for estimation and one for the forecast target.
Cite this review
Pith. "Pith review of Misspecification-Robust Shrinkage and Selection for VAR Forecasts and IRFs." pith.science (2026). https://pith.science/paper/FDOEJ47T
@misc{pith2026250203693,
author = {Pith},
title = {Pith review of: Misspecification-Robust Shrinkage and Selection for VAR Forecasts and IRFs},
year = {2026},
howpublished = {\url{https://pith.science/paper/FDOEJ47T}},
note = {Machine review of arXiv:2502.03693}
}
read the original abstract
Vector autoregressions (VARs) are vulnerable to dynamic misspecification, forcing researchers to navigate complex bias-variance trade-offs when forecasting or estimating impulse response functions (IRFs). We derive novel, task-specific selection criteria -- PC for forecasting and IRFC for IRF estimation -- based on asymptotically unbiased estimates of frequentist risk under local dynamic misspecification. These criteria give empirical researchers a fully data-driven and misspecification-aware workflow that simultaneously selects among candidate estimators, the degree of Bayesian shrinkage, and the lag length. IRFC is the first criterion allowing researchers to seamlessly choose between iterated-VAR and local-projection IRF estimators while jointly tuning shrinkage and lag length for the task at hand.
Figures
Forward citations
Cited by 1 Pith paper
-
Targeted Local Projections
Targeted local projections shrink LP impulse responses toward SVAR estimates at each horizon with data-driven weights, then use a double bootstrap for inference.
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.