{"id":"ce5e2945-0818-4717-91aa-218585598b03","arxiv_id":"2507.14408","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A Bayesian time-varying parameter model with Markov switching and dynamic shrinkage is applied to economic exchange-rate models, reporting out-of-sample predictive gains over a random walk with stochastic volatility.","lead":"The authors introduce a Markov-switching version of the dynamic shrinkage process (MSDSP) that lets regression coefficients switch between zero, constant, and drifting regimes. Applied to classical exchange-rate models, the method reports out-of-sample predictive gains over a random walk with stochastic volatility, offering a new angle on the Meese-Rogoff puzzle.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The out-of-sample comparison in §4 is only valid if standardization is recursive; §4.3 does not state this, and full-sample standardization would leak future information into the supposed out-of-sample forecasts.","rationale":"The reader's weakest assumption is exactly the issue I would flag: the out-of-sample evaluation can be invalidated by full-sample standardization, and §4.3 does not rule this out. I agree with that identification. I also considered other potential concerns — the absence of uncertainty quantification for the LPDR, the insignificant individual Diebold-Mariano tests, and the lack of code or data — but none has the same capacity to overturn the central claim by itself. The DM tests are supplementary; the LPDR, CRPS, and MCS already carry the argument. The standardization ambiguity, by contrast, sits at the entrance of the empirical analysis: if the transformation uses future information, every LPDR and RMSFE in Tables 3-5 is contaminated. The phrase 'expanding' in §4.1 refers to model estimation, not to the data transformation, and no appendix passage clarifies the preprocessing. Because the concern is addressable and the paper's contribution is otherwise coherent, I recommend keeping the reader's conditional verdict. The recursive-standardization re-run would settle the matter; if the results survive, the MSDSP contribution stands as a well-motivated extension of the dynamic shrinkage literature.","tokens_in":24062,"tokens_out":6491,"duration_ms":84200,"concrete_test":"Recompute Tables 3-5 and the MCS in Section 4 with recursive standardization: at each forecast origin t, standardize each variable using only observations through t (or, for the h-step direct regression pairs, through t-h), then re-estimate all models and re-evaluate the predictive distributions. If the MSDSP LPDRs remain above 3 and RMSFE ratios remain below 1, the concern is resolved; if the improvements shrink or vanish, the headline claim is an artifact of data leakage. Requiring the authors to release the preprocessing code would also settle the ambiguity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim — that MSDSP-augmented economic models outperform the random walk out of sample — depends entirely on the integrity of the out-of-sample protocol used in Section 4. Section 4.3 describes the data only as: 'All variables are standardized to have a mean of 0 and a variance of 1.' It does not specify whether the standardization is recursive (moments computed from data available at the forecast origin) or uses the full sample. If full-sample moments are used, the realized target y_{t+h} is normalized using its own future mean and variance, and the predictor x_t is normalized using future values of x, including future exchange-rate levels through the PPP regressor (p_t - p*_t - e_t). These leaked moments can differentially improve the predictive likelihood, CRPS, and RMSFE of the economic models relative to the random walk, because the economic models use covariates while the random walk does not. Tables 3-5 report LPDRs around 2-6 and RMSFE ratios near 0.98, magnitudes that could plausibly be generated by such leakage. The sentence in §4.1 that predictions are 'out-of-sample by using expanding' describes parameter estimation, not the standardization transformation. Because the paper provides no code and no further detail on the preprocessing, this ambiguity is the most load-bearing unresolved assumption: if full-sample standardization was used, the headline comparison is not a genuine out-of-sample exercise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Markov Switching Dynamic Shrinkage Process (MSDSP) that extends the Dynamic Shrinkage Process of Kowal et al. (2019) by attaching independent two-state Markov processes to each coefficient, allowing a time-varying parameter to switch between exact zero (sparsity) and a DSP-driven path that permits shrinkage toward constancy as well as gradual and abrupt change; the measurement equation is a linear regression with stochastic volatility. The authors provide an MCMC scheme based on Polya-gamma and normal-mixture augmentation and report two simulation designs in which the MSDSP recovers zero and non-zero regimes with tighter credible intervals than the DSP. The empirical application revisits the Meese-Rogoff puzzle: MSDSP and DSP versions of seven economic models (IRP, Taylor rule, PPP, monetary model, and oil/gold/copper commodity models) are evaluated against random-walk benchmarks with stochastic volatility for GBP/USD over 1990-2017, with out-of-sample evaluation from Nov 2000 to Jun 2017. The paper reports that MSDSP versions achieve positive log predictive density ratios (about 2.2 to 6.0), lower CRPS, RMSFE ratios near 0.98, and top ranks in model confidence sets, while linear and DSP versions generally underperform the random walk. A two-stage model-assembly extension applies MSDSP to the weights of a Bayesian predictive synthesis and is reported to dominate dynamic Bayesian predictive synthesis.","tokens_in":24380,"tokens_out":18972,"duration_ms":154519,"significance":"The contribution is potentially important. The MSDSP is a natural and coherent unification of sparsity, shrinkage, and structural change in a single state-space prior, and it nests the DSP and standard TVP specifications; the simulation results provide honest evidence that the model can separate zero regimes from active regimes, which the DSP alone cannot do. The empirical exercise follows the Rossi (2013) benchmark taxonomy rather than a bespoke design, and the evaluation uses five complementary criteria (LPDR, CRPS, RMSFE, tail coverage, and model confidence sets). Credit is also due for reporting the unfavorable Diebold-Mariano results explicitly rather than suppressing them, for making the direct h-step-ahead scheme clear, and for the scalable two-stage assembly method. However, the headline empirical finding is only as good as the out-of-sample protocol, and the current manuscript does not specify whether the standardization in Section 4.3 is recursive or full-sample; this is the main reason the result cannot be accepted as it stands.","major_comments":[{"comment":"Section 4.3 states 'All variables are standardized to have a mean of 0 and a variance of 1' without saying whether the standardization is recursive (using only data available up to the forecast origin) or full-sample. If full-sample moments are used, the realized target y_{t+h} is normalized using its own future mean and variance, and the covariates (including the PPP regressor p_t - p*_t - e_t) are normalized using values observed after time t; this mechanically contaminates the out-of-sample predictive likelihoods, CRPS, coverage rates, and RMSFE in Tables 3-6 in a way that can differentially favor the covariate-bearing economic models over the random walk. The sentence in Section 4.1 that predictions are 'out-of-sample by using expanding' describes the parameter-estimation window, not the transformation, and Section 4.4, which is cited for details, contains only results rather than the protocol. The authors must state explicitly whether standardization is recursive and, if it was full-sample, redo the entire evaluation with recursive moments; the paper currently provides no code with which to resolve this ambiguity.","section":"§4.3 (Data Description); §4.1 (Competing Economic Models)"},{"comment":"The point-forecast claim is weaker than the text suggests. The RMSFE ratios in Table 5 are all close to unity (roughly 0.97-0.99 for MSDSP), and the authors report that pairwise Diebold-Mariano tests 'did not achieve many significant result'. The joint Wilcoxon test pools 84 statistics (7 models times 12 horizons) that are highly dependent across horizons and models, so the reported p-value of 0.0000 is not a reliable basis for the conclusion that the MSDSP is 'systematically' superior in point forecasts. The authors should either temper the RMSFE-based conclusions, provide a multiple-testing correction, or report the distribution of the individual DM statistics so that the reader can judge the strength of the point-forecast evidence.","section":"§4.4.2 (Point Forecasts), Table 5"},{"comment":"The text claims that 'the point forecast (RMSFE) from MSDSP is better than the random walk when h > 9', but Table 8 shows the opposite at h=9 (1.0104 versus 0.9902) and at h=12 (0.9886 versus 0.9848); only h=10 and h=11 are better. Moreover, the MSDSP assembly has worse LPL and CRPS than the RW-SV at every horizon (for example, LPL -142.4 versus -140.0 at h=1), so the summary sentence 'slightly worse ... at short horizons but tends to improve at longer horizons' misrepresents the reported results. The claim that the MSDSP assembly 'consistently beat' DBPS is supported by Table 8, but the comparison against the random walk should be restated to match the numbers.","section":"§5.4 (Model Assembly Results), Table 8"}],"minor_comments":[{"comment":"The manuscript contains numerous typos and grammatical slips that should be corrected: 'resprentation' in §2.3, 'daws' in §2.4, 'Markow' in §2.2, 'As in Case 2' should be 'As in Case 1' in §3.2, 'the cases ofthe Canadian dollar' in §4.3, 'benchmerk' in §4.4.2, and 'eliminated form the confidence set' in §4.4.3; Section 5.3 also contains the unfinished sentence 'and thus the structure of the model weights ωt It cannot be inferred without input from the base models'.","section":"§2.1, §2.3, §2.4, §3.2, §4.3, §4.4.2, §4.4.3 (typos/presentation)"},{"comment":"The left-hand side 'log(p(Y_{T0+h:T}|I_{T0}, M))' suggests conditioning only on the initial information set I_{T0}, while the definition is a sum of log predictive densities that condition on I_t at each t; please align the notation with the recursive updating scheme.","section":"§4.2, Eq. (15)"},{"comment":"The Z-distribution shape parameters α_h and β_h are not listed in Table 1 and no fixed values or priors are given in the main text; because the shrinkage behavior of the DSP depends on these shapes, the authors should state how they are set in both the simulations and the application.","section":"§2.1, Eq. (8), Table 1"},{"comment":"The tail-coverage comparisons use 200 out-of-sample observations, so the standard error of a nominal 2.5% coverage rate is roughly one percentage point; differences such as 1.5% versus 2.5% are within sampling noise, and the text should not describe such differences as decisive without uncertainty quantification.","section":"§4.4.1, Table 6"},{"comment":"The comparison between the MSDSP assembly, which inputs a single point prediction per model, and the full DBPS, which inputs predictive distributions, is not entirely apples-to-apples; please acknowledge in the text that the two procedures use different information inputs, and optionally add a point-input DBPS variant for a cleaner comparison.","section":"§5.1, §5.2 (Model Assembly)"},{"comment":"For the assembly application (p = 8), the assumption that all Markov switching processes are independent is substantive; at least a brief discussion of this restriction or a sensitivity exercise with a common state process would be informative.","section":"§2.1, Eq. (9)"},{"comment":"No data availability statement or replication code is provided; given that the empirical claim is the centerpiece of the paper, the authors should make the code and the precise recursive standardization and estimation protocol available as part of the revision.","section":"General (reproducibility)"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the key question is the standardization protocol in Section 4.3. If the authors used full-sample standardization, the out-of-sample comparison is invalid and the headline result may be an artifact; if it was recursive, the result is genuine. Since no code is provided, I could not verify this externally. I recommend asking for the replication code and a precise protocol description as part of the revision, with re-estimation under recursive standardization if necessary. There is also a smaller reporting inconsistency in Section 5.4 versus Table 8 that should be corrected. The methodological contribution (the MSDSP model and its sampler) is sound and publishable independently of the empirical outcome, so this is a strong candidate for major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the MSDSP process is the real deal as a modeling contribution, and the simulations back it up. But the Meese-Rogoff headline should not be trusted until the authors clarify how they standardized the data.\n\nThe new piece is the Markov switching layer on top of the DSP of Kowal et al. (2019). That lets each coefficient switch between an exact zero state and a DSP state, which covers sparsity, constancy, and gradual drift in one process. The model nests DSP and standard TVP, and the simulation study shows the posterior correctly finds the designed zero and non-zero episodes. That is a legitimate extension of a well-known prior, and the paper is honest about the related work (Rockova-McAlinn, Bernardi et al., Uribe-Lopes, etc.). The BPS application is a nice extra, and the writing is clear.\n\nThe soft spot is exactly where the reader put it. Section 4.3 says only 'All variables are standardized to have a mean of 0 and a variance of 1.' No recursive or expanding-window caveat. Given that the out-of-sample predictions start in 2000 and the sample ends in 2017, full-sample standardization would use future information to scale both the response and the covariates (the PPP regressor contains e_t). That can differentially lift the economic models' density and point scores relative to the RW-SV, which uses no covariates. The magnitudes in Tables 3-5 are small enough that such leakage could plausibly account for them. The statement in Section 4.1 about 'expanding' refers to parameter estimation, not to the transformation. So this is a load-bearing ambiguity, not a footnote.\n\nThere are also smaller issues: individual DM tests are mostly insignificant, and the significance comes from a joint Wilcoxon test; predictive metrics have no uncertainty quantification; and no code or data are given. None of these are fatal, but they make it hard to verify the main result.\n\nMy recommendation: send it to a serious referee. The MSDSP contribution is worth publishing, but the authors must either confirm that standardization was recursive (or fixed on the initial training window) and give code/data, or the empirical claim should be downgraded to an illustration rather than a resolution of the Meese-Rogoff puzzle.","headline":"The MSDSP model is a real methodological contribution, but the paper's headline empirical claim rests on an unstated standardization choice that could leak future information.","tokens_in":24887,"tokens_out":4452,"would_cite":true,"duration_ms":48052,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes the Markov Switching Dynamic Shrinkage Process (MSDSP), a time-varying-parameter prior that lets each coefficient switch exactly to zero or shrink to a constant, and reports that economic models equipped with it beat…","keywords":["Bayesian Econometrics","Shrinkage Methods","Sparsity","Model Combination","Variable Selection","Exchange Rate Prediction","Markov Switching","Time-varying parameters"],"falsifier":"Recompute the out-of-sample LPDR, CRPS, and RMSFE for the GBP/USD application after standardizing each predictor recursively with moments computed only up to the forecast origin, and check whether every MSDSP model still beats the random walk with stochastic volatility at all horizons; if the original tables were built with full-sample standardization, the reported dominance may shrink or disappear. Inspecting the replication code for whether the standardization step uses full-sample moments would settle the question directly.","tokens_in":23883,"feed_emoji":"💱","tokens_out":7570,"duration_ms":74435,"temperature":0.7,"pith_summary":"The paper proposes a new prior process for time-varying parameter models, the Markov Switching Dynamic Shrinkage Process (MSDSP), and uses it to argue that the Meese–Rogoff puzzle is not a verdict on economic fundamentals but on model rigidity. In the MSDSP, each regression coefficient is either switched exactly to zero by a two-state Markov chain or follows a dynamic shrinkage process that can hold it constant over stretches of time. Applied to seven standard economic models for monthly GBP/USD exchange rates, the MSDSP versions outperform the random walk with stochastic volatility on density and point forecast metrics at horizons of one to twelve months, where constant-coefficient versions of the same models lose. The paper also builds an MSDSP-based model combination scheme and reports that it beats dynamic Bayesian predictive synthesis benchmarks. A sympathetic reader would take the paper to establish that allowing sparse, occasionally constant, occasionally shifting coefficients is what lets fundamentals-based models compete with the random walk.","feed_headline":"Switch-off coefficients let economic models beat the FX random walk","feed_subtitle":"Each predictor can switch to zero, and economic models then beat the random walk on density and point forecasts","key_machinery":"The load-bearing object is the MSDSP equation system: for each coefficient $i$, $\\beta_{it}=s_{it}\\tilde\\beta_{it}$ with $s_{it}\\in\\{0,1\\}$ following an independent two-state Markov chain, and the shadow coefficient $\\tilde\\beta_{it}$ following a random walk $\\tilde\\beta_{it}=\\tilde\\beta_{i,t-1}+\\omega_{it}$ with $\\omega_{it}\\sim N(0,\\exp(h_{it}))$. The log-volatility $h_{it}$ follows an AR(1) process with innovations from the Z-distribution, whose heavy negative tail drives $\\exp(h_{it})$ to zero so that the shadow coefficient is 'shrunken' to a constant; the Markov chain is what gives the exact zero state that DSP alone lacks. This combination lets one data-driven mechanism deliver sparsity, dynamic shrinkage, and structural change simultaneously. The empirical application uses a direct $h$-step-ahead version of the model, replacing $y_t$ by $y_{t+h}$ with covariates $x_t$, so predictive distributions are built without forecasting the predictors.","core_discovery":"On the paper's own terms, the central discovery is that the Meese–Rogoff puzzle can be overturned by giving each predictor in an economic exchange-rate model its own 'off switch.' The MSDSP nests the Dynamic Shrinkage Process of Kowal et al. (2019) by adding a two-state Markov chain per coefficient: in state 0 the coefficient is exactly zero (sparsity), and in state 1 it evolves as a random walk whose innovation variance can shrink to near zero, so the coefficient can be constant, gradually changing, or abruptly shifting. In the GBP/USD out-of-sample evaluation from November 2000 to June 2017, every economic model equipped with MSDSP beats the stochastic-volatility random walk benchmark on log predictive density ratio, CRPS, and RMSFE, while the same models with constant coefficients or with DSP alone do not; the MSDSP-PPP model is reported to strongly dominate the random walk at all horizons. In model assembly, treating each model's predictive mean as data and letting MSDSP govern the combination weights yields better log predictive likelihood, CRPS, and RMSFE than dynamic Bayesian predictive synthesis.","pith_inferences":["Inference: the same switch-on/switch-off mechanism should transfer to other macroeconomic forecasting settings where predictors matter only in some episodes, such as inflation or output-gap forecasting; the paper does not test this.","Inference: if the standardization used full-sample moments, the headline comparison could be optimistic; a recursive-standardization robustness check would cleanly separate the MSDSP mechanism from any look-ahead leakage.","Inference: one could test whether the MSDSP's advantage comes mainly from the exact zero state or from the dynamic shrinkage by comparing it with a model that uses an independent Bernoulli switch rather than a Markov chain, isolating the role of regime persistence.","Inference: an extension the paper leaves open is feeding full predictive densities, not just predictive means, into the MSDSP assembly; if that works, the two-stage cost saving could be combined with density information."],"forward_implications":["If the paper's central claim is right, the Meese–Rogoff puzzle should be restated: constant-coefficient economic models fail against the random walk, but the same fundamentals models with sparse, time-varying coefficients can win.","The MSDSP specification makes the 'off switch' state a structural feature, so forecasts come with an interpretable record of when each fundamental mattered; this could guide which economic relationships are alive at a given date.","The success of MSDSP-PPP in particular implies that purchasing-power-parity deviations have predictive content for the exchange rate once the relationship is allowed to be intermittently active.","Because the MSDSP assembly outperforms standard dynamic Bayesian predictive synthesis while using only point predictions, cheap two-stage sparse combination could be applied to larger model pools without heavy distributional inputs.","The same qualitative conclusions are reported to extend to the Canadian dollar and Japanese yen exchange rates in the paper's appendix, suggesting the result is not specific to GBP/USD."],"supporting_citations":[{"why":"Establishes the puzzle that structural exchange-rate models cannot beat the random walk out of sample, the benchmark the paper aims to overturn.","marker":"(Meese and Rogoff, 1983a)"},{"why":"Provides the exhaustive survey confirming the puzzle and identifies the no-drift random walk with stochastic volatility as the toughest benchmark.","marker":"(Rossi, 2013)"},{"why":"Supplies the Dynamic Shrinkage Process that the MSDSP nests and extends with Markov switching.","marker":"(Kowal et al., 2019)"},{"why":"Provides dynamic Bayesian predictive synthesis, the model assembly framework the MSDSP assembly is built on and compared with.","marker":"(McAlinn and West, 2019)"},{"why":"Introduces the Z-distribution whose extreme negative innovations drive the shrinkage of coefficient volatility in the MSDSP.","marker":"(Barndorff-Nielsen et al., 1982)"},{"why":"Motivates the commodity-price (oil, gold, copper) exchange-rate models included in the forecast comparison.","marker":"(Ferraro et al., 2015)"}],"fun_headline_variants":["Switch-off coefficients let economic FX models beat the random walk","Markov switching sparsity overturns the Meese-Rogoff puzzle","Economic models beat FX random walk when coefficients can vanish","Each predictor off-switch flips Meese-Rogoff: economic FX wins","Sparse dynamics help economic FX models outforecast random walk"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The empirical comparison assumes that the standardization of all predictor variables to mean 0 and variance 1 uses only information available at each forecast origin; the paper does not state whether this is recursive or full-sample, and full-sample standardization would let future data leak into the 'out-of-sample' predictions.","fun_headline_variants_meta":{"raw":{"variants":["Switch-off coefficients let economic FX models beat the random walk","Markov switching sparsity overturns the Meese-Rogoff puzzle","Economic models beat FX random walk when coefficients can vanish","Each predictor off-switch flips Meese-Rogoff: economic FX wins","Sparse dynamics help economic FX models outforecast random walk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000385,"raw_usage":{"total_tokens":2050,"prompt_tokens":974,"completion_tokens":1076,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":987}},"tokens_in":590,"tokens_out":1076,"duration_ms":11029,"temperature":1.0,"reasoning_tokens":987,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:57:02.480584+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the out-of-sample LPDR, CRPS, and RMSFE for the GBP/USD application after standardizing each predictor recursively with moments computed only up to the forecast origin, and check whether every MSDSP model still beats the random walk with stochastic volatility at all horizons; if the original tables were built with full-sample standardization, the reported dominance may shrink or disappear. Inspecting the replication code for whether the standardization step uses full-sample moments would settle the question directly.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the exhaustive survey confirming the puzzle and identifies the no-drift random walk with stochastic volatility as the toughest benchmark."},{"cited_title":"and West, M","cited_arxiv_id":null,"evidence_quote":"Provides dynamic Bayesian predictive synthesis, the model assembly framework the MSDSP assembly is built on and compared with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Z-distribution whose extreme negative innovations drive the shrinkage of coefficient volatility in the MSDSP."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the commodity-price (oil, gold, copper) exchange-rate models included in the forecast comparison."}],"review_version":1}