Pith. sign in

REVIEW 4 major objections 7 minor 3 references

Conditional Method Confidence Set

T0 review · 4 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Running the Model Confidence Set state by state yields per-regime forecast-model sets with an asymptotic coverage guarantee.

desk verdict A sensible statewise extension of Hansen-Lunde-Nason with an informative simulation study, but the central CLT and bootstrap arguments for non-contiguous statewise subsamples are asserted rather than proved, so the asymptotic claims need careful refereeing. read the letter →

arxiv 2505.21278 v1 pith:WQCSQB53 submitted 2025-05-27 econ.EM

classification econ.EM MSC 62P2062M1062F40
keywords conditionalforecastevaluationModelConfidenceSetExpectedShortfallValue-at-Riskstresstestingmultiplefamilywiseerrorratestate-dependentloss
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper extends the Model Confidence Set — the set of forecasting methods whose predictive ability is statistically indistinguishable from the best — from unconditional to conditional evaluation. The proposed Conditional Method Confidence Set runs the same selection logic separately for each observed economic regime, so the winner can differ across states. The authors prove that, statewise, the test statistic converges to the maximum of a normal vector under the null and diverges under the alternative, and that the resulting set contains every conditionally superior method with asymptotic probability at least $1-\alpha$. If correct, this gives regulators and risk managers a tool that controls the per-state familywise error rate while allowing the best forecast model to be different in calm and crisis periods. The paper demonstrates the tool on Expected Shortfall stress testing across liquidity horizons.

What carries the argument

The load-bearing device is the statewise application of the MCS closure principle. The paper splits the full out-of-sample loss series into state-specific subsequences $L^l_{i,\tau}$ by conditioning on a categorical state variable $S_t$ at the forecast origin, computes per-state loss differentials $d^l_{ij,\tau}$ and standardized excess losses $t^l_{i\cdot}$, and applies the elimination rule $e^l_{\max,M}=\arg\max_{i\in M} t^l_{i\cdot}$ whenever $T^l_{\max,M}$ exceeds a bootstrap critical value. A coherency condition ties the test to the elimination rule so that no conditionally superior method is removed with probability exceeding $\alpha$. The circular block bootstrap supplies the critical values, implicitly accounting for the correlation among loss differentials and for the temporal dependence in the state subsamples.

What would settle it

Take a state with at least two competing methods and simulate a DGP in which the conditional expected loss differential changes sign within that state, for example with $E[d]=+\delta$ in the first half of the state subsample and $-\delta$ in the second half. If the CMCS is applied and its coverage is computed over many repeated samples, coverage below $1-\alpha$ for the true statewise superior set would directly contradict Theorem A.1, and the failure should become visible when the sign flip occurs near the middle of the subsample with $n_l$ large.

Watch

Extended reading notes

Core claim

The central claim is that forecast evaluation can be made conditional on a discrete state variable without losing the set-based guarantees of the Model Confidence Set. For each state $l$, the procedure tests $H^l_{0,M}$: all methods in $M$ have equal conditional expected loss, using statewise $t$-statistics $t^l_{i\cdot}$ and the maximum statistic $T^l_{\max,M}=\max_i t^l_{i\cdot}$. Theorem 2.1 states that under the null $T^l_{\max,M}$ converges in distribution to the maximum of a normal vector with the statewise correlation matrix, while under the alternative it diverges and the elimination rule removes an inferior method. Theorem A.1 then delivers the headline guarantee: the conditional confidence set $\widehat{M}^{l,*}_{1-\alpha}$ contains the set of conditionally superior methods with asymptotic probability at least $1-\alpha$, and eliminates each inferior method with probability tending to one. Thus the CMCS controls the familywise error rate state by state.

Load-bearing premise

The procedure assumes that, within each economic regime, the ranking of the forecasting methods never changes over time; if the best method switches within a regime, the set the test is trying to cover is no longer well defined and the coverage guarantee does not apply.

Editorial extensions

If this is right

  • For each regime, the analyst obtains a set of models whose conditional predictive ability is indistinguishable at level $\alpha$, rather than a single forced winner.
  • A model that looks best unconditionally but is weak in one state is still eliminated in that state, so forecast combinations adapt to the regime.
  • Because the procedure is run separately on disjoint states, false eliminations are controlled within each state and do not contaminate other regimes.
  • In the Expected Shortfall application, the CMCS selects different models across stocks, liquidity horizons, and forecasting horizons, yielding lower and more stable average ES forecasts than the Wald-type dynamic forecast combination benchmark.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The proof relies on discrete states with positive probability; conditioning on continuous or high-dimensional states would require local smoothing and is not covered by the present asymptotics.
  • If within-state rankings can switch, the set of conditionally superior methods may cease to be well defined; an alternative target would be the best method inside rolling sub-periods, replacing Assumption 2.3(i) with piecewise stability.
  • The same statewise closure idea could transfer to other set-valued selection tasks, such as variable selection or nowcast averaging, whenever the loss can be indexed by a regime.
  • One can test the weakest assumption directly: estimate statewise loss differentials on rolling windows and look for sign changes within a state; frequent sign flips would signal that the asymptotic coverage guarantee should not be relied on.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper proposes a Conditional Method Confidence Set (CMCS) that applies the MCS algorithm of Hansen et al. (2011) separately to statewise loss subsequences selected by a discrete state variable S_t. It defines conditionally superior sets M^{l,*}, statewise hypotheses H^l_{0,M}, and a T^l_max test statistic. Theorem 2.1 claims that T^l_max converges to the distribution of the maximum of a normal vector under the null and diverges under the alternative, and Theorem A.1 claims asymptotic coverage at least 1−α for the statewise MCS. The paper also provides simulation evidence comparing CMCS with the unconditional MCS and with the Wald-based dynamic forecast combination of Borup et al. (2024), and an empirical application to Expected Shortfall stress testing across liquidity horizons.

Significance. If the central claims were fully established, the paper would make a useful contribution: a statewise MCS is a natural and practically relevant extension, especially for regulatory stress testing where model rankings may differ across regimes. The simulation design in Section 3.2 is informative about finite-sample power differences between statewise t-tests and the Wald test. The paper is not circular: no fitted parameters are used to make predictions, and the procedure is derived from stated assumptions. However, the proof of the main asymptotic result is incomplete, the bootstrap validity is not established, and the empirical evidence is descriptive and lacks backtesting. The contribution is therefore conditional on closing these gaps.

major comments (4)
  1. [§2.4 and Appendix A.1, Lemma A.1] The statewise CLT is asserted, not proved. Corollary A.2 shows that z_t = h_t d_{ij,t} is α-mixing of the stated size when Assumption 2.1 holds, but the statewise sequence d^l_{ij,n_lτ} is obtained by selecting times {t: S_t = s_l}. Since S_t is F_t-measurable and in the application is a stress indicator (often a 252-day window selected by a risk factor, Section 4.1 and Figure 3), the selected times are neither independent of the losses nor regularly spaced; the statewise sample is a concatenation of possibly separated stress episodes. Mixing of a subsequence selected by a dependent indicator does not follow from mixing of the original sequence. Lemma A.1 simply cites "the CLT for α-mixing processes" after defining X^l_τ from the statewise losses. Without a proof that {X^l_τ} is itself strong mixing, or an alternative CLT for the statewise sample mean, the convergence statement in Theorem 2.1 is unsupported. Please provide the missing argument or add assumptions that make the statewise sequence mixing by construction.
  2. [Appendix A.2, Steps l.1–l.3] The circular block bootstrap is applied to the concatenated statewise loss sequence, but no consistency result is proved for this object. Even if Assumption 2.3(i) holds, blocks drawn from the concatenated sequence can bridge distinct stress episodes and thereby create resampled samples that do not reflect the dependence structure of the original process. Gonçalves and White (2002) is cited for the block bootstrap variance estimator under dependence, but that result concerns the mean of an observed array, not a bootstrap applied to a subsequence defined by a data-dependent state indicator. The bootstrap p-value in Step l.3(c) is therefore not shown to approximate the distribution F^l_{ρ_{n_l}} needed in Theorem 2.1, and the coverage claim in Theorem A.1 is not closed.
  3. [Theorem 2.1 and Lemma A.2] The limit distribution is stated as F^l_{ρ_{n_l}}, which depends on n_l through the correlation matrix ρ_{n_l} implied by Ω_{n_l}. Lemma A.2 states (n_l)^{1/2}(V̄^l − ψ^l) →^d N(0, Ω_{n_l}) without assuming convergence of Ω_{n_l} to a fixed limit. Consequently "T^l_{max,M} →^d F^l_{ρ_{n_l}}" is not a standard convergence-in-distribution statement: the limiting object changes with the sample size. If Ω_{n_l} does not converge, the bootstrap critical values cannot be expected to converge either. Please either impose convergence of Ω_{n_l} (or of ρ_{n_l}) to a fixed limit and prove it under the maintained assumptions, or formulate the result as a uniform approximation over the sequence.
  4. [Section 2.3, Assumption 2.3(i) and Definition 2.1] The paper assumes that sgn(E[d^l_{ij,n_lτ}]) is constant across the within-state index τ, and states that this "implies that the null and the alternative are exhaustive." If this sign varies within a state, the set M^{l,*} defined in Definition 2.1 is not a well-defined target: the conditionally superior set could differ across episodes within the same state. The authors acknowledge this restriction in the text, but it is central to the coverage claim in Theorem A.1 and is not tested or discussed in the application, where stress states are defined by 252-day windows of a risk factor and may contain heterogeneous episodes. Please add a formal discussion of what M^{l,*} means when Assumption 2.3(i) is violated, or provide a diagnostic test for the assumption.
minor comments (7)
  1. [Abstract] The term "CMC S" appears with a spacing typo; it should be "CMCS".
  2. [Section 4.1 and Appendix A.5] There are typos: "defied" should be "defined" in Section 4.1, and "Appedix" should be "Appendix" in the heading of Appendix A.5.
  3. [Table 3 caption] The caption duplicates the sentence "The level of all tests is α = 0.05." Please remove the repetition.
  4. [Appendix A.2, Step l.1(b)(i)] The notation "(v_{b,1},...,v_{b,p}) = (ξ_{b1}, ξ_{b1}+1, ξ_{b1}+p−1)" does not describe a sequence of length p; it should read "(ξ_{b1}, ξ_{b1}+1, ..., ξ_{b1}+p−1) modulo n_l."
  5. [Appendix A.2, Step l.3(a)] The quantity \(\hat{var}(\bar d_{i·})\) is defined as (1/B)Σ_b(ξ*_{b,i}−ξ*_{b,·})^2, where ξ*_{b,·} is the cross-method average for bootstrap b. This is not the usual bootstrap variance estimator for \(\bar d_{i·}\); please clarify whether this is intentional and justify it, since Theorem 2.1 uses t^l_{i·} with \(\hat{var}(\bar d^l_{i·,n_l})\).
  6. [Theorem A.1 proof] The proof refers to "Assumption 4.1" instead of Assumption A.1, and part (ii) states lim inf P(i ∈ \hat M^l_{1−α}) = 0 where the stronger statement lim P = 0 is what is actually proved.
  7. [Section 4] The empirical comparison of CMCS and DFC is descriptive; no backtesting of the resulting ES forecasts or inference on the differences is provided, so the concluding claim that CMCS is "a robust tool" is stronger than what the tables demonstrate.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the CMCS derivation is self-contained and no claim reduces to its own inputs by construction.

full rationale

The paper does not fit any parameter and then rename it a prediction. The CMCS test statistic is built from statewise loss differentials under explicit Assumptions 2.1–2.4, and Theorem 2.1 follows from Lemma A.2, which invokes a CLT for alpha-mixing processes. The bootstrap implementation is explicitly the circular block bootstrap of Politis and Romano (1991) and the consistency argument cites Gonçalves and White (2002), both external to the authors. Hansen et al. (2011) is used as a benchmark and template, not as a self-citation, and no uniqueness theorem from the authors' own prior work is imported. The empirical application compares CMCS with the DFC procedure of Borup et al. (2024) on out-of-sample ES forecasts, but the comparison is descriptive and not a fitted-input-called-prediction loop. There are open technical questions about the step from mixing of z_t to mixing of the non-contiguous statewise subsequence {d^l_ij,nlτ}, and about the sample-size-dependent limit F^l_{rho_nl} in Theorem 2.1; however, those are concerns about proof correctness and asymptotic statement rigor, not about circularity. No load-bearing step reduces, by the paper's own equations, to its own inputs.

Assumptions & free parameters 3 free parameters · 8 assumptions · 0 invented entities

The central claim rests on standard mixing and CLT assumptions, a restrictive sign-constancy condition on within-state rankings, and an unproved assertion that non-contiguous statewise subsamples inherit the mixing and bootstrap properties needed.

free parameters (3)
  • Bootstrap block length p = Not reported
    Chosen by the researcher in Step l.1(a) of the CMCS bootstrap; the empirical section does not report the block length used, so results depend on an unspecified tuning parameter.
  • Number of bootstrap replications B = 100
    Set to B=100 in Section 4.2; this is small for stable p-value estimation, and no sensitivity analysis is provided.
  • MCS significance level alpha = 0.05
    Standard choice, not fitted. Listed because it is a free tuning constant of the procedure, though conventional.
assumptions (8)
  • domain assumption The process {W_t} and test functions {h_t} are alpha-mixing of order -r/(r-2) for some r>2 and gamma>0 (Assumption 2.1).
    Invoked at Section 2.4 to apply CLTs and the bootstrap theory for weakly dependent data.
  • domain assumption Finite (r+gamma)-th moments of conditional loss differentials (Assumption 2.2).
    Needed for the CLT and convergence of sample means.
  • ad hoc to paper The sign of E[d^l_{ij,nlτ}] is constant in τ, with positive variance and uniformly positive variance of normalized partial sums (Assumption 2.3).
    This sign-constancy is the paper's additional restriction beyond Giacomini-White and Borup et al.; it ensures the conditional ranking is time-invariant within a state and is essential for the definition of M^{l,*} and the validity of the statewise tests.
  • domain assumption Each state l has probability at least epsilon>0, so n_l goes to infinity (Assumption 2.4).
    Required for the statewise sample size to diverge.
  • ad hoc to paper The statewise loss subsequences are alpha-mixing of at most order -r/(r-2) (Corollary A.2).
    This is asserted, but the proof misapplies the finite-lag mixing lemma to non-contiguous subsamples selected by the state variable; the random sampling times can destroy mixing, so this is not established.
  • ad hoc to paper The circular block bootstrap applied to statewise losses consistently estimates the distribution of the statewise mean (Section 2.5).
    No theorem is provided for bootstrap consistency when the resampled series is a non-contiguous subsample; the procedure treats the subsample as if it were a contiguous stationary time series.
  • standard math A CLT for heterogeneous mixing arrays applies to the statewise means (Lemma A.1).
    The paper cites White (2001) Th. 5.20; application requires the variance conditions in Assumption 2.3, so it is conditionally justified.
  • domain assumption The loss function is strictly consistent for the target functional (Gneiting-Raftery).
    Ensures that minimizing expected loss corresponds to the true functional; used in the ES application.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Conditional Method Confidence Set." pith.science (2026). https://pith.science/paper/WQCSQB53

@misc{pith2026250521278,
  author       = {Pith},
  title        = {Pith review of: Conditional Method Confidence Set},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WQCSQB53}},
  note         = {Machine review of arXiv:2505.21278}
}
read the original abstract

This paper proposes a Conditional Method Confidence Set (CMCS) which allows to select the best subset of forecasting methods with equal predictive ability conditional on a specific economic regime. The test resembles the Model Confidence Set by Hansen et al. (2011) and is adapted for conditional forecast evaluation. We show the asymptotic validity of the proposed test and illustrate its properties in a simulation study. The proposed testing procedure is particularly suitable for stress-testing of financial risk models required by the regulators. We showcase the empirical relevance of the CMCS using the stress-testing scenario of Expected Shortfall. The empirical evidence suggests that the proposed CMCS procedure can be used as a robust tool for forecast evaluation of market risk models for different economic regimes.

Figures

Figures reproduced from arXiv: 2505.21278 by the authors.

Figure 1
Figure 1. Power properties 0.1 0.2 0.3 0.4 0.5 µ 1.0 2.5 5.0 7.5 10.0 n = 150 MCS CMCS 0.1 0.2 0.3 0.4 0.5 µ n = 500 MCS CMCS 0.1 0.2 0.3 0.4 0.5 µ n = 1000 MCS CMCS The figure displays the simulated power of the tests defined as the number of models in the confidence set averaged over 5000 simulation iterations, of the CMCS procedure (in dashed-dotted line) and the unconditional MCS procedure (in solid line). The x-axis show… view at source ↗
Figure 3
Figure 3. Illustration of stress periods.                            [PITH_FULL_IMAGE:figures/full_fig_p024_3.png] view at source ↗
Figure 5
Figure 5. Time series of 10-day ahead ES forecasts for GE: CMCS. [PITH_FULL_IMAGE:figures/full_fig_p027_5.png] view at source ↗
Figures from the paper (9 more)
Figure 7
Figure 7. Figure 7: Model selection for 10 days ahead ES forecasts for differ [PITH_FULL_IMAGE:figures/full_fig_p028_7.png]
Figure 9
Figure 9. Figure 9: Time series of 10-day ahead ES forecasts for BAC: DFC wit [PITH_FULL_IMAGE:figures/full_fig_p030_9.png]
Figure 10
Figure 10. Figure 10: Method selection for 1 day ahead ES forecasts for differ [PITH_FULL_IMAGE:figures/full_fig_p054_10.png]
Figure 11
Figure 11. Figure 11: Method selection for 1 day ahead ES forecasts for differ [PITH_FULL_IMAGE:figures/full_fig_p054_11.png]
Figure 12
Figure 12. Figure 12: Method selection for 1 day ahead ES forecasts for differ [PITH_FULL_IMAGE:figures/full_fig_p055_12.png]
Figure 13
Figure 13. Figure 13: Time series of 10 days ahead ES forecasts for BAC: DFC w [PITH_FULL_IMAGE:figures/full_fig_p060_13.png]
Figure 14
Figure 14. Figure 14: Time series of 10 days ahead ES forecasts for BAC: DFC w [PITH_FULL_IMAGE:figures/full_fig_p060_14.png]
Figure 15
Figure 15. Figure 15: Method selection for 10 day ahead ES forecasts for diffe [PITH_FULL_IMAGE:figures/full_fig_p061_15.png]
Figure 16
Figure 16. Figure 16: Method selection for 10 day ahead ES forecasts for diffe [PITH_FULL_IMAGE:figures/full_fig_p061_16.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

3 extracted references · 3 canonical work pages

  1. [1]

    Regarding notation, δM = 0 refers to the case when H l 0,M cannot be rejected, and δM = 1 refers to the case when it is rejected

    For δM and ǫM, we suppress ( ·)l for ease of notation, though the pairs of equivalence test and elimination rule could differ per state. Regarding notation, δM = 0 refers to the case when H l 0,M cannot be rejected, and δM = 1 refers to the case when it is rejected. Note that the former als o implies that the testing procedure is halted. In the latter case...

  2. [2]

    − (p∆ 1 + (1 −p)∆ 2)p∆ 1 =pσ2 +p(1 −p)(∆ 2 1 − ∆ 1∆ 2) Σ 22 =Var (Dt · /BD St=1) =E[(Dt · /BD St=1)2] −E[Dt · /BD St=1]2 =E[D2 t |St = 1] ·E[/BD St=1] −E[Dt|St = 1] 2 ·E[/BD St=1]2 =p(σ2 + ∆ 2

  3. [3]

    standard

    −p2∆ 2 1 =pσ2 +p(1 −p)∆ 2 1 Given that it exists, the inverse of the 2 × 2 matrix Σ is given by Σ −1 = 1 det(Σ)   Σ 22 −Σ 12 −Σ 12 Σ 11  , 42 with det(Σ) = (1 −p)pσ2( (1 −p)∆ 2 1 +p∆ 2 2 +σ2) . (13) Plugging into T h Thus, as factors in D,E,F in the expression of T h in Equation 12 above we obtain Σ −1 11 + 2Σ −1 12 + Σ −1 22 = 1 det(Σ) × ( pσ2 +p(1 −...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.