{"id":"a4e650e3-5e2e-4486-87da-2abc02243aad","arxiv_id":"2502.03693","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"New risk-based criteria PC and IRFC select shrinkage, lag length, and estimator type in potentially misspecified VARs, outperforming marginal-data-density selection in simulations.","lead":"This paper derives two information criteria that choose the amount of Bayesian shrinkage, the lag length, and the estimator type (iterated VAR versus direct multi-step or local projection) in vector autoregressions. The criteria target the out-of-sample forecast risk and the impulse response estimation risk directly, which makes them more robust to model misspecification than the standard marginal-likelihood approach.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unbiasedness of PC and IRFC depends on the T^{-1/2} local drift of both the DGP misspecification and the prior mean; under fixed misspecification or a fixed prior center the key expansions (19), (25), and (51) break down.","rationale":"The reader's weakest assumption identifies the T^{-1/2} local drift of both the DGP and the prior mean, and I agree that this is the most load-bearing condition. The paper's formal results are internally consistent under that assumption: Theorem 1 provides the O(T^{-1/2}) expansion, Theorem 3 links the in-sample loss to the risk components, and Definitions 1 and 3 construct the covariance corrections that deliver equations (25) and (51). The concern is not an algebraic error but a question of external validity: the motivating applications involve a single fixed DGP, whereas the unbiasedness theorems describe a triangular array of DGPs that move toward the VAR as T grows. A fixed prior mean is especially problematic because with prior precision growing as lambda*T, the posterior mean is inconsistent unless the prior mean itself drifts toward the pseudo-true value. This is not a hidden flaw, since the paper states the local assumption, but it is the condition on which the central claim rests. I considered whether the gap between pointwise unbiasedness and the properties of the data-driven argmin lambda-hat is more serious, since the paper does not prove argmin consistency; however, the stated central claim is the unbiasedness of PC_T and IRFC_T as risk estimates, and the simulation evidence supports the selection step. The fixed-misspecification check would settle whether the local-drift assumption is essential or merely a technical convenience, so the reader's conditional verdict remains appropriate.","tokens_in":40811,"tokens_out":8860,"duration_ms":91967,"concrete_test":"Run a fixed-misspecification Monte Carlo version of the paper's alpha=2 design: keep the VMA coefficients A_j independent of T (drop the T^{-1/2} scaling) and use a fixed prior mean mu different from F, with prior precision lambda*T. For T in {100, 500, 5000}, compute E[PC_T(iota,lambda) - PC_T(lfe,0,q)] and compare it with the finite-sample risk differential R(iota,lambda) - R(lfe,0,q) over the same lambda grid. If the gap does not shrink with T, or if the argmin lambda shifts systematically as T increases, the unbiasedness claim is specific to the local-drift device and does not carry over to fixed misspecification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is the local-drift triangular array in (13)-(15). The expansions underlying Theorem 1, equation (19), and the unbiasedness results (25) and (51) all require the DGP misspecification term O(T^{-1/2}), prior means of the form F + T^{-1/2}phi and F^h + T^{-1/2}psi, and prior precision scaling as lambda*T. If the misspecification is fixed rather than local, or if the prior mean is a fixed value mu different from F, the sequence T^{1/2}(bar Psi_T - F^h) is no longer Op(1). With fixed prior mean and lambda>0, the prior weight lambda*T makes the posterior mean converge to mu, so T^{1/2}(bar Psi_T - F^h) diverges. With fixed misspecification, the normalized risk T*R diverges, so the risk differences that PC_T and IRFC_T are supposed to estimate are not even defined in the same normalization. The paper is explicit that it uses the local framework, so the internal derivations are not circular; however, the motivating claim of 'misspecification-robust' selection in actual samples relies on the local experiment being a good approximation to a fixed DGP. That is the least secure link between the theorems and the stated purpose of the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops two information criteria, PC and IRFC, for selecting the shrinkage hyperparameter, the VAR lag length, and the estimator type (MLE/iterated VAR versus LFE/local projection) when the VAR is potentially misspecified. The analysis is conducted in a local-to-VAR drifting DGP framework: the true infinite-order VMA drifts toward a finite-order VAR at rate T^{-1/2}, prior means drift toward pseudo-true values at the same rate, and prior precision scales with T. Under these assumptions, Theorem 1 gives the limit distribution of the shrinkage estimators, Theorem 2 gives the normalized prediction risk, Theorem 3 shows that the in-sample loss differential converges to a random variable whose expectation is the risk differential, and Theorem 4 establishes the lag-augmentation centering property for local projections. The proposed criteria jointly select (ι, λ, p); their finite-sample behavior is studied by Monte Carlo and their use for IRF selection is illustrated on 200 FRED-QD samples.","tokens_in":41057,"tokens_out":6304,"duration_ms":63883,"significance":"The contribution is potentially significant: it extends the local-misspecification selection framework of Schorfheide (2005) from plug-in MLE/LFE predictors to shrinkage estimators and to IRF estimation, and it provides the first criterion that jointly selects among VAR- and LP-based IRF estimates, shrinkage, and lag length. The derivation is internally consistent and the local asymptotics are stated explicitly rather than hidden. The paper also gives a transparent account of the key limitation that PC and IRFC are asymptotically unbiased but not consistent estimators of the relevant risk. The Monte Carlo and empirical sections are useful, and the connection to the recent local-projection literature (MPQW, LPW) is appropriately drawn. The main weakness is that the central claims about 'misspecification robustness' are established only under the local-drift design, and the small-sample evidence at T=100 is less supportive than the abstract suggests.","major_comments":[{"comment":"The unbiasedness of PC and IRFC is derived under the local-drift assumption that the DGP misspecification and the prior mean deviate from the VAR at rate T^{-1/2} and that prior precision grows as λT. If the misspecification is fixed rather than local, or if the prior center is fixed away from the pseudo-true value, the expansions in Theorem 1 and the risk-normalization underlying (25) and (51) break down. The paper is explicit that it operates in the local framework, so this is not an internal inconsistency; however, the title and abstract claim 'misspecification-robust' selection in a way that suggests broader scope. I recommend adding a concrete sensitivity analysis, for example simulating a fixed (non-drifting) misspecification magnitude and a fixed prior center, and comparing the PC-selected risk with the oracle risk, to show how large the approximation error can be in realistic samples.","section":"Section 2.2, equations (13)-(15); Theorems 1-3; equation (25)"},{"comment":"The abstract states that 'once the VAR model is misspecified PC hyperparameter selection works significantly better than an MDD-based selection,' but the simulation evidence at T=100 does not support this unqualified statement. In Table A-3, under α=2 and T=100, MDD selection produces lower risk than PC selection in most configurations; for example, for LFE at h=4, p=1 the PC risk is 19% higher than MDD, and for the PC-selected lag length it is 31% higher. The online appendix acknowledges that 'for small sample sizes, MDD seems to do better than PC,' but this caveat is absent from the abstract and the main-text discussion. The authors should either qualify the claim to the asymptotically relevant sample sizes or provide an explanation of why the small-sample reversal does not undermine the practical recommendation.","section":"Section 7.3 and Table A-3"},{"comment":"The IRFC criterion requires p* < q, where q is the maximum number of lags used in the benchmark LP. This assumption is needed so that the benchmark LP with q lags has the centering property M' μ(lfe,0,q) M = μ(irf); if p* = q, this equality fails and equation (51) no longer holds. The empirical application in Section 8 selects p=6, which is the maximum lag q, in the majority of samples (Figure 4). Thus the empirically relevant case p* = q is not merely hypothetical. The authors should discuss what happens when p* = q and, if possible, adapt the criterion or state clearly the additional assumptions needed to cover that case.","section":"Section 5, equation (43), Theorem 4, Definition 3"},{"comment":"At T=100, the top row of Figure 3 shows a large wedge between the Monte Carlo risk and both the asymptotic risk and the expected value of PC. The text attributes this to a difference between finite-sample and asymptotic variance of the estimated coefficients, but it does not assess whether this wedge materially degrades the selection properties of PC at small sample sizes. Given that the paper is aimed at empirical researchers who often have samples of this order, the authors should provide a quantitative assessment of the implied selection error, for example by reporting the frequency with which PC selects a λ far from the finite-sample-optimal value at T=100.","section":"Section 7.1, Figure 3"}],"minor_comments":[{"comment":"The dimension of the matrix M Υ_q is printed as nq × (n−1)q; it should be nq × n(q−1).","section":"Section 4, after equation (33)"},{"comment":"The sentence 'discrediting with widespread belief' is ungrammatical and should read 'discrediting the widespread belief.'","section":"Section 1, final paragraph of the introduction"},{"comment":"The phrase 'pointmass at λ = ∞' should be written as 'point mass at λ = ∞.'","section":"Section 7.1, third paragraph"},{"comment":"The spacing in 'V AR' is inconsistent; sometimes it appears as 'V AR' and sometimes as 'VAR.' A consistent rendering would improve readability.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid contribution to a well-defined literature and the theoretical results appear internally consistent. The main issue is scope: the local-misspecification assumption is explicit and standard in this literature, but the title and abstract overstate the robustness of the procedure relative to what is shown. The T=100 Monte Carlo results are also in tension with the abstract's unqualified claim of superiority over MDD. These are fixable by qualification and additional sensitivity analysis, so I do not recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this paper delivers the first selection criterion that jointly picks the VAR-vs-LP estimator for impulse responses, the shrinkage hyperparameter, and the lag length. The IRF criterion (IRFC) is genuinely new. The prediction criterion (PC) extension to shrinkage is a more direct generalization of Schorfheide (2005), but it is competently done and useful.\n\nWhat's good: the asymptotic risk calculations are internally consistent, proofs are in the online appendix, and the Monte Carlo at T=500 and T=5000 lines up with the theory. The empirical exercise on 200 FRED-QD samples is honest: the conclusion that VAR vs LP is sample-dependent and that LP is not universally preferred is supported by the figures. The paper is careful to distinguish what the criteria can and cannot do.\n\nSoft spots: the whole framework rests on local misspecification and local prior drift (T^{-1/2}). The unbiasedness of PC and IRFC in equations (25) and (51) breaks down if misspecification is fixed rather than local, or if the prior mean is far from the pseudo-true value. That is not a hidden flaw—the paper states it plainly—but it means 'misspecification-robust' is really 'robust to local misspecification'. How well that approximates a fixed DGP in practice is the least secure link. Also, the T=100 simulations show a wedge between finite-sample risk and asymptotic risk, which is worth acknowledging. Minor: no replication code, selection frequencies lack uncertainty quantification, and the claim that PC* and PC perform equally well rests on unreported simulations.\n\nBottom line: this is a method paper with real value for applied macro and for econometricians working on LP/VAR inference. It deserves a serious referee, mainly to push on the local-drift assumption and to ask for code. I'd be happy to see it in a good journal after reasonable revision.","headline":"Solid, genuinely new IRF selection criterion that deserves referee time; the local-drift assumption is the main thing to push on.","tokens_in":41608,"tokens_out":2062,"would_cite":true,"duration_ms":19530,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that two information criteria give asymptotically unbiased estimates of prediction and impulse-response risk, so shrinkage, lag length, and estimator type can be jointly selected under misspecification.","keywords":["vector autoregressions","shrinkage estimation","hyperparameter selection","local projections","impulse response functions","misspecification","prediction risk","information criteria"],"falsifier":"In the paper's own Monte Carlo design (local DGP with $\\alpha=2$, $T=5000$), compute the actual risk differential $R(\\hat y_{T+h}(\\iota,\\lambda))-R(\\hat y_{T+h}(\\iota',\\lambda'))$ and the expected criterion differential $E[PC_T(\\iota,\\lambda)-PC_T(\\iota',\\lambda')]$; if their gap does not vanish as $T$ grows, the unbiasedness claim fails. A second check: fix a non-drifting infinite-order VMA DGP and verify that the gap stops vanishing, confirming that the local-drift rate is the boundary of the theory.","tokens_in":40593,"feed_emoji":"🎯","tokens_out":9618,"duration_ms":87340,"temperature":0.7,"pith_summary":"This paper derives two information criteria, $PC_T$ and $IRFC_T$, that estimate—up to a constant that is irrelevant for ranking—the $h$-step-ahead prediction risk and the impulse-response estimation risk of Bayesian VAR shrinkage estimators. The central formal claim is that the criteria are asymptotically unbiased under local misspecification: for any two configurations, the expected difference of the criteria converges to the difference of the true risks. This matters because the standard Bayesian way of choosing shrinkage hyperparameters, the marginal data density, is not designed for misspecified models, and the paper shows that it can lead to much larger risk than the proposed criteria when the VAR is misspecified. With these criteria, an applied researcher can choose the shrinkage hyperparameter, the lag length, and the estimator type—MLE/LFE for forecasting or VAR/local projection for impulse responses—in a single data-driven step.","feed_headline":"Risk-based criteria set shrinkage, lags, and VAR vs. LP in one step","feed_subtitle":"Unbiased estimates of forecast and impulse-response risk replace marginal-likelihood tuning under misspecification.","key_machinery":"The machinery is a companion-form shrinkage posterior mean that weights the MLE or LFE toward a prior mean with precision $\\lambda T$, embedded in a local-drift DGP that makes sampling variance, misspecification bias, and prior-induced bias all $O(T^{-1/2})$. The selection criteria are modified final-prediction-error statistics: an in-sample multi-step squared-error term plus twice an estimated covariance penalty between the candidate estimator and the unshrunk loss-function benchmark. For IRFs, the benchmark is the lag-$q$ local projection $\\bar\\Psi_T(lfe,0,q)$, and Theorem 4's centering property guarantees that this benchmark has no misspecification bias; the criterion then measures, with the added covariance penalty, how far a candidate VAR or LP estimate is from that correctly centered target.","core_discovery":"Under the drifting DGP in equations (13)–(15), where the true process is a stationary infinite-order VMA that approaches a finite-order VAR at rate $T^{-1/2}$ and the prior means approach the pseudo-true values at the same rate, the paper proves a limit distribution for the shrinkage estimators $\\bar\\Psi_T(\\iota,\\lambda)$ and decomposes the normalized prediction risk into a bias term and a variance term (Theorems 1 and 2). It then constructs $PC_T(\\iota,\\lambda,p)$ and $IRFC_T(\\iota,\\lambda,p)$ so that the key convergence results (25) and (51) hold: the expected difference between two criterion values converges to the difference between the corresponding true risks. For impulse responses, Theorem 4 shows that the local-projection estimator is correctly centered if and only if the lag length is at least $p^*+1$, which makes lag augmentation a selection rule rather than an ad hoc choice. The paper's message is that under misspecification, shrinkage, lag length, and estimator choice should be governed by an estimate of the target risk, not by marginal likelihood.","pith_inferences":["The same unbiased-risk construction would likely carry over to non-conjugate or data-based priors, provided the posterior mean remains a $T$-consistent weighted average and the prior displacement drifts at $T^{-1/2}$; one could test this by replacing the conjugate prior with a Minnesota-style prior in the same Monte Carlo design.","Because the criteria estimate risk only up to a constant and remain stochastic in the limit, selection uncertainty is intrinsic; an implied extension is to report the distribution of the argmin over $\\lambda$ and to build averaging weights rather than hard selections.","The same principle could target other plug-in objects—cumulative multipliers, long-run responses, or forecast-error decompositions—by changing the quadratic form and the covariance penalty so that the criterion remains an unbiased estimate of the corresponding risk."],"forward_implications":["Forecasters using Bayesian VARs can replace marginal-data-density hyperparameter choice with $PC_T$; in the paper's simulations this gives nearly the same performance under correct specification and large risk reductions under misspecification.","Researchers estimating impulse responses can use $IRFC_T$ to choose between VAR and local projections; the empirical analysis shows that neither estimator dominates, with LP chosen in 60–85% of samples depending on horizon and lag length.","Lag augmentation becomes operational: because $IRFC_T$ incurs a large bias when $p<p^*$, the criterion will select enough lags to center the LP, while shrinkage is chosen to offset the variance cost.","The risk-targeting logic also provides a principled way to average across estimators and hyperparameters, an extension the paper explicitly notes at the end."],"supporting_citations":[{"why":"Provides the local-misspecification DGP and the original PC prediction criterion that the paper extends to shrinkage estimators.","marker":"S2005"},{"why":"Final prediction error criterion that motivates the form of the PC risk estimate.","marker":"Shibata (1980)"},{"why":"Defines local projection IRF estimation, the multi-step regression used as the LFE/LP candidate.","marker":"Jordà (2005)"},{"why":"Shows lag-augmented local projections can alleviate serial-correlation inference problems, underlying the centering analysis.","marker":"Montiel Olea and Plagborg-Møller (2021)"},{"why":"Establishes double robustness and correct centering of local projections under the drifting DGP, used for the p≥p*+1 condition.","marker":"MPQW"},{"why":"Large-scale simulation comparison of local projections and VAR IRFs that the paper's empirical selection frequencies are contrasted with.","marker":"LPW"},{"why":"Shows LP and VAR IRFs coincide in population with unrestricted lags, the baseline for comparing estimators.","marker":"Plagborg-Møller and Wolf (2021)"},{"why":"MDD-based hyperparameter selection that PC is designed to replace under misspecification.","marker":"Giannone, Lenza, and Primiceri (2015)"}],"fun_headline_variants":["One criterion picks shrinkage, lags, and estimator under misspecification","Misspecification-proof risk criteria unify VAR and LP selection","Forecast risk beats marginal likelihood for VAR tuning","Shrinkage and lag length chosen by unbiased risk estimates","Risk-based selection replaces marginal likelihood in VARs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire asymptotic argument rests on the load-bearing local-drift assumption: the true DGP must drift toward the finite-order VAR at rate $T^{-1/2}$, the prior means must drift toward pseudo-true values at the same rate, and the prior precision must grow as $\\lambda T$; if the misspecification is fixed rather than local, the unbiasedness of $PC_T$ and $IRFC_T$ breaks down.","fun_headline_variants_meta":{"raw":{"variants":["One criterion picks shrinkage, lags, and estimator under misspecification","Misspecification-proof risk criteria unify VAR and LP selection","Forecast risk beats marginal likelihood for VAR tuning","Shrinkage and lag length chosen by unbiased risk estimates","Risk-based selection replaces marginal likelihood in VARs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000577,"raw_usage":{"total_tokens":2714,"prompt_tokens":927,"completion_tokens":1787,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":1706}},"tokens_in":543,"tokens_out":1787,"duration_ms":10558,"temperature":1.0,"reasoning_tokens":1706,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T04:05:02.106373+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In the paper's own Monte Carlo design (local DGP with $\\alpha=2$, $T=5000$), compute the actual risk differential $R(\\hat y_{T+h}(\\iota,\\lambda))-R(\\hat y_{T+h}(\\iota',\\lambda'))$ and the expected criterion differential $E[PC_T(\\iota,\\lambda)-PC_T(\\iota',\\lambda')]$; if their gap does not vanish as $T$ grows, the unbiasedness claim fails. A second check: fix a non-drifting infinite-order VMA DGP and verify that the gap stops vanishing, confirming that the local-drift rate is the boundary of the theory.","supporting_citations":[{"cited_title":"Local Projection Inference is Easier Than You Think,","cited_arxiv_id":null,"evidence_quote":"Shows lag-augmented local projections can alleviate serial-correlation inference problems, underlying the centering analysis."}],"review_version":1}