{"id":"81cad99c-0b6c-4092-9546-4ddf431d14d6","arxiv_id":"1908.04569","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"New forecast encompassing tests for Expected Shortfall, built on a joint VaR-ES loss function, with misspecification-robust asymptotics and simulation evidence of reasonable size and power.","lead":"The authors introduce three statistical tests to decide whether one expected shortfall (ES) forecast is at least as informative as another, a tool banks and regulators need under Basel III. Two tests need both value-at-risk and ES forecasts; a third works with ES forecasts alone, which matches current regulatory reporting.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Strict ES test's robustness under quantile misspecification is unproven; finite-sample sizes suggest pseudo-true η* drifts from (1,0), so the central robustness claim may fail asymptotically.","rationale":"The paper's central theoretical result, Theorem 2.10, is a conditional statement: under Assumption 2.7, the Wald statistics converge to chi-square under the null hypotheses as defined via the pseudo-true parameter θ*_n. For the joint and auxiliary tests, the quantile link is correctly specified, so the pseudo-true null is the natural encompassing hypothesis. For the strict test, the quantile link is deliberately misspecified in general DGPs, and the leap from 'η*=η0' to 'e1 encompasses e2' is exactly what the paper asserts but does not prove. This is not an external disagreement with consensus; it is an internal gap between the theorem's null and the test's advertised use. The simulation evidence is the only support for the leap, and its most misspecified DGP shows 13.5% rejection at 10% nominal at n=5000, which is precisely the direction the unproven bias would predict. The reader's weakest assumption identifies this same point, so I agree. Because the joint and auxiliary tests, the misspecification-robust asymptotics for the pseudo-true null, and the extensive simulations remain valuable, the verdict should stay conditional rather than unconditional or reject. A revision should either prove the bias is zero or vanishing under realistic conditions, characterize its magnitude, or qualify the strict test's claims.","tokens_in":47675,"tokens_out":7234,"duration_ms":72295,"concrete_test":"Fix the VaR/ES GAS DGP at π=0 and compute the pseudo-true parameter η* = arg min E[ρ(Y, gq(e,β), ge(e,η))] directly, by long simulations or numerical integration over the stationary distribution of the 1F-GAS process, using α=0.025 and the linear links from Section 3.1. If η* is not (1,0) within Monte Carlo error, the strict test is asymptotically a test of a different null and the robustness claim fails. Corroborate by rerunning the size simulation at n=50,000 and n=100,000: rejection rates near 10% support the paper, whereas rates trending to 100% confirm the bias. Also re-estimate the asymptotic covariance without the F_t(gq)≈α approximation and compare size at n=5000 to isolate the misspecification effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strict ES test (Definition 2.6, regression (2.17)) replaces the quantile link gq(ˆq_t,β) with gq(ˆe_t,β). For any DGP outside the pure-scale class, such as the VaR/ES GAS DGP, the quantile equation is misspecified. The paper's robustness claim requires that, whenever e1 genuinely encompasses e2, the pseudo-true parameter η* in (2.15) equals η0=(1,0). Section 2.2 (after Definition 2.6) asserts this bias is 'negligible' but provides no proof; Theorem 2.10 only supplies chi-square limits for the null η*=η0, not for the economic null 'e1 encompasses e2'. If η*≠(1,0), rejection probability goes to 1, so the test is asymptotically invalid for its advertised decision. Table 1 is consistent with this: at 10% nominal size, the strict test rejects 13.5% of the time for the VaR/ES GAS DGP at n=5000, with no evidence of convergence to 10%. The same misspecification also makes the covariance approximation F_t(gq(β*))≈α (Section 2.3) unsupported, but the pseudo-true parameter bias is the load-bearing issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes three forecast encompassing tests for Expected Shortfall (ES) based on the FZ0 joint loss function: a joint VaR-and-ES test, an auxiliary ES test, and a strict ES test that requires only ES forecasts. The authors develop misspecification-robust asymptotic theory for the M-estimator of the combination parameters, derive Wald tests for linear and nonlinear link functions, and investigate finite-sample size and power in simulations calibrated to GARCH, GAS, and CAViaR DGPs. An empirical application to IBM, S&P 500, and DAX returns illustrates the use of the tests for forecast selection and combination.","tokens_in":47915,"tokens_out":9955,"duration_ms":95790,"significance":"If the results hold, the paper makes a useful contribution to risk management by providing the first encompassing tests for ES, and the strict ES test is especially valuable because Basel III requires banks to report only ES forecasts. The misspecification-robust asymptotic theory for nonlinear link functions extends prior work, and the simulations are extensive and based on realistic DGPs. However, the central claim that the strict ES test is robust to quantile misspecification rests on a negligibility argument that is not formally proved, which limits the strength of the headline result.","major_comments":[{"comment":"The strict ES encompassing test in Definition 2.6 tests the restriction η*=η0, where η* is the pseudo-true parameter defined in (2.15). The economic null of interest, that forecast e1 encompasses e2, requires that whenever e1 genuinely encompasses e2, the pseudo-true parameter satisfies η*=(1,0). The paper states after Definition 2.6 that the misspecification bias is 'negligible' but does not provide a proof or a set of sufficient conditions. Theorem 2.10 only gives a chi-squared limit for the test statistic under the hypothesis η*=η0; it does not establish that the economic null implies η*=η0. The simulation results in Table 1 are not decisive: for the VaR/ES GAS DGP at n=5000, the strict test shows rejection rates of 13.5% and 10.6% at the 10% level for the two hypotheses, which could reflect a small asymptotic bias, though the H1 size point corresponds to a correctly specified DGP (the 1F GAS model has colinear VaR and ES). Please provide a formal analysis of the pseudo-true parameters under misspecification, or clearly redefine the null as testing η*=η0 and discuss the economic interpretation accordingly.","section":"Section 2.2 and Theorem 2.10"},{"comment":"The covariance matrix estimator Ω̂ relies on the approximation Ft(gq_t(β*))≈α, which is exact when the quantile link is correctly specified but not in general. For the strict ES test under misspecification, the term Ft(gq_t(β*))-α in Eq. (2.22) is nonzero, and the approximation is justified only by a heuristic 'the degree of misspecification is small' argument. Theorem 2.10 assumes that Ω̂-Ω_n converges to zero, but the consistency of the nonparametric estimators (nid and scl-sp) under the misspecification conditions of Assumption 2.7 is not established. The authors should either prove consistency of Ω̂ or provide a reference that does so in this setting.","section":"Section 2.3, covariance estimation"}],"minor_comments":[{"comment":"The column labels 'Str ES', 'Aux ES', 'VaR ES', and 'VaR' are potentially confusing; the 'VaR ES' column refers to the joint VaR-and-ES test, but the abbreviation is not defined in the table note. I suggest renaming the column to 'Joint VaR/ES' for clarity.","section":"Table 1 and notes"},{"comment":"The displayed expression is split across lines and the bracket opened at the end is closed only in the following display; please reformat to make the expression self-contained.","section":"Equation (2.24)"},{"comment":"The paper uses a fixed forecasting scheme with one-time in-sample estimation, but the asymptotic theory in Section 2 does not explicitly account for estimation error in forecast parameters. The paper cites Giacomini and White (2006), but it should be stated more clearly that the tests are intended for fixed forecast methods rather than models with estimated parameters.","section":"Section 4 and Assumption 2.7"},{"comment":"With 2000 Monte Carlo replications and a nominal size of 10%, the binomial standard error is about 0.9%, so the difference between 13.5% and 10% is just under four standard errors. Reporting standard errors or confidence bands for the size estimates would help the reader judge the significance of the deviations.","section":"Section 3.2, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-written and the core asymptotic theory appears sound for the joint and auxiliary tests. The main concern is the strict ES test's robustness claim, which is not formally proven. This is a fixable issue if the authors provide a rigorous treatment of the pseudo-true parameters or moderate the claims. The simulations and empirical work are thorough and suitable for the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real news here is the first set of forecast encompassing tests for Expected Shortfall, built on the Fissler-Ziegel joint loss. The joint VaR+ES and auxiliary ES variants are straightforward and sound: given correctly specified quantile forecasts, the regression framework is correctly specified, and the Wald tests hit their intended null. The asymptotic theory under misspecification is standard but carefully done, and the paper goes beyond the existing literature by allowing nonlinear link functions and a pseudo-true parameter. The simulation study is extensive across GARCH, GAS, GAS-t, and ES-CAViaR processes, and the empirical application with eight forecasting models is a useful illustration. I would point a colleague to this paper as the natural reference for ES encompassing testing.\n\nThe soft spot is the strict ES test. Its null is defined as the pseudo-true parameter equaling (1,0), not as the economic condition that forecast 1 encompasses forecast 2. In the strict test, the quantile link is fitted using ES forecasts, so any DGP outside the pure-scale class misspecifies the quantile equation. The paper asserts in Section 2.2 that the resulting bias is negligible, but there is no proof. The simulations are consistent with a real problem: at a 10% nominal level, the strict test rejects 13.5% for the VaR/ES GAS DGP at n=5000, down from 29% at n=500 but with no clear evidence of convergence to 10%. If the pseudo-true ES weight drifts from (1,0) under the economic null, the rejection rate goes to one asymptotically, and the test is not testing what it advertises. The joint and auxiliary tests do not suffer from this issue, so the paper's core contribution survives, but the strict test should not be sold as robust without either a proof of negligibility or a substantially larger simulation study showing convergence.\n\nMinor points: the covariance approximation F_t(gq)≈α is heuristic, though likely harmless; and there is no code or data for replication, which should be fixed. These are secondary. The paper deserves serious peer review, but the strict test needs a more honest treatment, and the authors should provide replication artifacts. I would accept it with major revision, not desk reject it.","headline":"The paper gives the first forecast encompassing tests for Expected Shortfall; the joint and auxiliary variants are solid, but the strict ES test's advertised robustness rests on an unproven negligibility claim that the simulations do not fully support.","tokens_in":701,"tokens_out":752,"would_cite":true,"duration_ms":38107,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P05","62F03","62M10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces three forecast encompassing tests for Expected Shortfall, proves the tests stay valid when the risk model is misspecified, and shows one variant needs only the ES forecasts banks already report.","keywords":["forecast encompassing","Expected Shortfall","Value at Risk","joint loss function","elicitability","misspecification robust inference","forecast combination","statistical tests"],"falsifier":"Simulate a data-generating process with pronounced time-varying skewness so the ratio of ES to VaR moves substantially, generate two forecasts where one truly encompasses the other, run the strict ES test at $n=5000$, and check the empirical rejection rate; if it stays far above the 10% nominal level or the estimated weight on the encompassing forecast is not near one, the paper's negligibility argument fails.","tokens_in":47409,"feed_emoji":"📉","tokens_out":15823,"duration_ms":135884,"temperature":0.7,"pith_summary":"Expected Shortfall (ES) is the risk measure the Basel III rules make banks report, but it has no loss function of its own, so forecast evaluation for it has lagged behind. This paper builds on the fact that ES is jointly elicitable with Value at Risk and constructs three encompassing tests—tests that ask whether one forecast already contains the information in any combination of two forecasts: a joint test for the VaR–ES pair, an auxiliary test that tests only the ES parameters, and a strict ES test that uses ES forecasts alone, which is the situation regulators face under current reporting rules. The paper proves that under mild regularity conditions the Wald statistics of all three tests converge to a chi-squared distribution, and it develops the asymptotic theory under model misspecification because the strict test feeds ES forecasts into the quantile link and can therefore be misspecified. An extensive simulation study across GARCH, GAS, and ES-CAViaR data generators finds that the tests are approximately correctly sized in large samples—the empirical sizes at $n=500$ are inflated, but shrink toward the nominal level as $n$ grows—and have good power, with the strict test nearly matching the auxiliary test that has access to more information. Applied to daily returns of IBM stock, the S&P 500, and the DAX 30, the tests show ES forecast combinations frequently beat the stand-alone models for the single stock, while the broad indices benefit less.","feed_headline":"Expected Shortfall forecasts can be tested without VaR forecasts","feed_subtitle":"New tests show which bank risk forecast dominates and when combining forecasts is worth it.","key_machinery":"The central object is the zero-homogeneous joint loss function for the pair (VaR, ES), $\\rho(Y,q,e) = -\\frac{1}{e}\\left(e - q + \\frac{(q-Y)\\mathbf{1}\\{Y\\le q\\}}{\\alpha}\\right) + \\log(-e)$, introduced by Fissler and Ziegel (2016). Because the ES alone is not elicitable, this loss supplies the moment conditions and the semiparametric regressions $Y_{t+1} = g_q(\\hat{q}_t,\\beta) + u^q_{t+1}$ and $Y_{t+1} = g_e(\\hat{e}_t,\\eta) + u^e_{t+1}$, through which the optimal combination weights are estimated by M-estimation. A general link function $g$ maps two competing forecasts and parameters into a linear or nonlinear forecast combination; the null hypothesis that forecast 1 encompasses forecast 2 is $\\theta^* = \\theta_0$, meaning the weight on forecast 2 is zero and the weight on forecast 1 is one. The strict ES test sets the quantile link to $g_q(\\hat{e}_t,\\beta)$, using ES forecasts in place of VaR forecasts, and the misspecification-robust asymptotic theory in Propositions 2.8 and 2.9 plus Theorem 2.10 is what keeps the Wald statistics $\\chi^2$ under the null.","core_discovery":"The paper's central claim is that forecast encompassing for the ES can be tested through M-estimation of a joint VaR–ES regression, and that the resulting Wald tests are valid even when the regression model is misspecified. Theorem 2.10 states that under Assumption 2.7 the test statistics of all three variants converge to a chi-squared distribution with degrees of freedom equal to the number of restricted parameters: four for the joint test, two for the auxiliary and strict tests under linear link functions. The asymptotic theory generalizes the joint quantile–ES M-estimator developed in earlier work to potentially misspecified nonlinear link functions, and uses a misspecification-robust covariance estimator assembled from the nid estimator of the density quantile and the scl-sp estimator of the truncated variance. The simulation evidence belongs to the claim: all three tests show approximately correct size and good power across GARCH, VaR/ES GAS, GAS-t, and ES-CAViaR DGPs, and the strict ES test behaves almost identically to the auxiliary test even though it uses no VaR forecasts, which supports the paper's argument that the misspecification from substituting ES forecasts for quantile forecasts is negligible in realistic financial settings.","pith_inferences":["The strict ES test's clean behavior probably relies on the near proportionality of VaR and ES that holds for scale-type daily return models; in asset classes with strongly time-varying higher moments, the quantile misspecification could be larger than the 13.5% rejection observed here, so a cautious user would validate the test on the target data before trusting it.","The same joint-loss machinery transfers to other jointly elicitable functionals, such as range value at risk or the mean–variance pair, which the paper notes as future work; the misspecification-robust M-estimation theory would need to be reworked for each new functional.","A direct empirical check of the strict test's key assumption is to compare its estimated combination weights with those from the correctly specified joint model on datasets where both VaR and ES forecasts are available; systematic divergence would quantify the information lost by dropping VaR forecasts.","Conditional encompassing tests that use instruments beyond the forecasts themselves would require extending the theory to overidentified GMM with nonsmooth moments, which the paper leaves open; if supplied, such tests could identify which variables drive the misspecification."],"forward_implications":["Regulators who receive only ES forecasts, as under Basel III reporting rules, can run pairwise encompassing tests without needing the underlying VaR forecasts.","When both directional hypotheses are rejected, the estimated combination weights from the joint regression provide a direct way to build a combined ES forecast, and the empirical application shows such combinations frequently beat stand-alone models for the IBM stock.","The tests are implemented for linear, affine, and nonlinear link functions, so practitioners can choose the forecast combination formula that fits their setting.","Because the strict and auxiliary tests behave almost identically in the simulations, the strict test loses little by not using VaR forecasts and can replace the auxiliary test when only ES forecasts are reported.","The misspecification-robust asymptotic theory covers nonlinear models, which was not previously available for joint VaR–ES M-estimation and supports flexible parametric forecast combination."],"supporting_citations":[{"why":"Supplies the strictly consistent joint loss function for the VaR–ES pair that makes ES forecast evaluation and M-estimation possible.","marker":"Fissler and Ziegel (2016)"},{"why":"Provides the dynamic semiparametric VaR–ES models and the M-estimation approach whose asymptotic theory this paper extends, and the 1F/2F GAS models are used as simulation DGPs.","marker":"Patton et al. (2019)"},{"why":"Establishes the joint quantile–ES regression framework and the covariance estimators (scl-sp and density-quantile) that the encompassing tests rely on.","marker":"Dimitriadis and Bayer (2019)"},{"why":"Provides misspecification-robust asymptotic theory for linear ES regressions, the direct predecessor of the strict test's robustness argument.","marker":"Bayer and Dimitriadis (2020)"},{"why":"Defines forecast encompassing for quantiles, supplies the VaR encompassing test used as benchmark, and motivates the forecast-selection decision rule.","marker":"Giacomini and Komunjer (2005)"},{"why":"Introduces GAS models with time-varying parameters, used as an additional DGP that generates misspecification for the strict ES test.","marker":"Creal et al. (2013)"},{"why":"Introduces the ES-CAViaR dynamic models used both as a misspecified DGP and as empirical forecasting models.","marker":"Taylor (2019)"},{"why":"Supplies the nid estimator used to estimate the density-quantile term in the asymptotic covariance matrix.","marker":"Hendricks and Koenker (1992)"}],"fun_headline_variants":["New ES encompassing tests work without VaR forecasts","Misspecification-robust tests for Expected Shortfall forecasts","Test Expected Shortfall forecasts without VaR input","Forecast encompassing tests for ES, no VaR needed","Robust ES forecast encompassing tests skip VaR forecasts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that putting ES forecasts where the quantile forecasts belong in the strict test does not materially shift the estimated combination weights away from $(1,0)$ when one forecast truly encompasses the other; the paper argues this effect is negligible rather than proving it, and its own VaR/ES GAS simulation still rejects about 13.5% of the time at a 10% nominal level when $n=5000$.","fun_headline_variants_meta":{"raw":{"variants":["New ES encompassing tests work without VaR forecasts","Misspecification-robust tests for Expected Shortfall forecasts","Test Expected Shortfall forecasts without VaR input","Forecast encompassing tests for ES, no VaR needed","Robust ES forecast encompassing tests skip VaR forecasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000576,"raw_usage":{"total_tokens":2679,"prompt_tokens":870,"completion_tokens":1809,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":1742}},"tokens_in":486,"tokens_out":1809,"duration_ms":15219,"temperature":1.0,"reasoning_tokens":1742,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:37:37.192339+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a data-generating process with pronounced time-varying skewness so the ratio of ES to VaR moves substantially, generate two forecasts where one truly encompasses the other, run the strict ES test at $n=5000$, and check the empirical rejection rate; if it stays far above the 10% nominal level or the estimated weight on the encompassing forecast is not near one, the paper's negligibility argument fails.","supporting_citations":[{"cited_title":"and Bayer, S","cited_arxiv_id":null,"evidence_quote":"Establishes the joint quantile–ES regression framework and the covariance estimators (scl-sp and density-quantile) that the encompassing tests rely on."},{"cited_title":"and Komunjer, I","cited_arxiv_id":null,"evidence_quote":"Defines forecast encompassing for quantiles, supplies the VaR encompassing test used as benchmark, and motivates the forecast-selection decision rule."},{"cited_title":"J., and Lucas, A","cited_arxiv_id":null,"evidence_quote":"Introduces GAS models with time-varying parameters, used as an additional DGP that generates misspecification for the strict ES test."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the ES-CAViaR dynamic models used both as a misspecified DGP and as empirical forecasting models."}],"review_version":1}