{"id":"c2da6c0a-5dfc-4c95-ad70-dbc3695aeeb8","arxiv_id":"2505.05334","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":7,"one_line_summary":"Simple Bayesian shrinkage regressions with up to 56 predictors often improve one-step-ahead Thai inflation forecasts, but the paper's evidence is internally inconsistent and does not support the broader claims.","lead":"This paper compares Bayesian shrinkage priors for forecasting Thai inflation with a univariate regression, and reports that wide predictor sets without stochastic volatility beat benchmark and SV-augmented models. The finding is a single-country empirical result, and the paper's own tables frequently contradict its headline claims.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central SV-vs-non-SV comparison is unauditable because the SV-augmented models and the DL prior are never defined, and the abstract's claim about LASSO large is contradicted by its own Table 2.","rationale":"The reader's REJECT verdict is correct and well-supported. The load-bearing concern is not merely missing code but the complete absence of definitions for two pillars of the headline result: the DL prior and all SV-augmented models. Without these, the SV-vs-non-SV comparison, which is the main novel claim, cannot be audited or reproduced. The paper's own Table 2 also directly contradicts the abstract's claim that LASSO large is superior across multiple horizons. The extreme LPL values further indicate a likely implementation error. Section 6's causal interpretation of shrinkage coefficients is an additional overreach. These are correctness risks that invalidate the central claims, not disagreements over modeling choices. The paper could be salvageable with complete model specifications, corrected tables, significance testing, and reproducible code, but in its current form it should be rejected.","tokens_in":21911,"tokens_out":1935,"duration_ms":20691,"concrete_test":"Obtain the complete model definitions and replication code from the authors. The decisive check is to rerun the estimation with the DL prior and the three SV-augmented models (HS SV, RIDGE SV, LASSO SV) using the stated design matrix and rolling window, and verify that Table 2's relative RMSE entries, especially LASSO large at h=4, 8, 12 (1.05, 1.54, 2.69), and the LPL trajectories in Figures 1-3 are reproduced. If the code produces different numbers, or if the DL/SV specifications cannot be written down, the central claims collapse.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim that HS, DL, and LASSO in the large predictor setting without SV are superior, and that SV 'increases estimation noise,' rests on models that are never specified. Section 2 defines the noninformative, Ridge, Adaptive Lasso, Spike-and-Slab, Horseshoe, and Horseshoe+ priors, but the Dirichlet-Laplace (DL) prior, which appears in all results tables (Tables 2-5) and the abstract, is never defined by equations, likelihood, or sampler. Likewise, the SV-augmented implementations ('HS SV', 'RIDGE SV', 'LASSO SV' in Figures 1-3) have no stated state equation, prior on volatility parameters, estimation algorithm, or convergence diagnostics. The LPL numbers quoted in Section 5, such as HS SV at -1070.0754 versus HS at 1.6045, are so extreme that they suggest a computational failure or a different scoring convention, but the text treats them as meaningful evidence. Moreover, the abstract's claim that 'HS, DL and LASSO in large-sized model setting without SV exhibit superior performance across multiple horizons' is contradicted by Table 2: LASSO large has relative RMSE 1.05, 1.54, and 2.69 at h=4, 8, and 12 in the full period, all worse than the benchmark. These are not stylistic issues; if the DL prior or SV models are misspecified or incorrectly estimated, the headline ranking is an artifact, not a finding about Thai inflation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares Bayesian shrinkage priors (noninformative, Ridge, Adaptive Lasso, Spike-and-Slab, Horseshoe, Horseshoe+, and Dirichlet–Laplace) for univariate out-of-sample forecasting of Thai CPI inflation. Forecasts are produced at horizons h = 1, 4, 8, 12 in three predictor settings (AR(2) only, moderate with 20 predictors, large with 56 predictors), evaluated by relative RMSE, quantile-weighted CRPS (uniform, tails, right, left), and cumulative log predictive likelihood. The authors compare these models with and without stochastic volatility and conclude that SV-augmented models underperform non-SV counterparts, that HS, DL and LASSO in the large predictor setting are superior across multiple horizons, and that left-tail (deflationary) risks are better captured than right-tail (inflationary) risks. A final section interprets Horseshoe shrinkage weights (kappa) as identifying supply-side drivers of Thai inflation.","tokens_in":22283,"tokens_out":4907,"duration_ms":44532,"significance":"If the empirical claims were fully supported, the paper would provide a useful case study of high-dimensional univariate Bayesian forecasting for an emerging economy, and a cautionary note on the value of stochastic volatility in this context. The design has components that are appropriate for this purpose: a monthly Thai dataset with a broad set of predictors, a rolling-window out-of-sample scheme, multiple forecast horizons, and a range of proper scoring rules. The paper also engages with an interesting substantive question about sparse versus dense predictor sets and tail risks in inflation. However, in its current form the significance cannot be assessed because the central results are internally inconsistent and key model components are never specified. The paper does not ship reproducible code or data, and several numerical claims in the text do not match the tables.","major_comments":[{"comment":"The Dirichlet–Laplace (DL) prior, which appears in the abstract, in all results tables (Tables 2–5), and in the conclusions, is never defined. Section 2 defines the noninformative, Ridge, Adaptive Lasso, Spike-and-Slab, Horseshoe, and Horseshoe+ priors (Eqs. 4–11), but DL is absent from the methodology. Without the prior distribution, its hyperparameters, the likelihood, or the sampling algorithm, the DL results cannot be audited or reproduced. This is load-bearing because the abstract's headline claim ('HS, DL and LASSO in large-sized model setting without SV exhibit superior performance') directly depends on a model that is never specified.","section":"Section 2"},{"comment":"The SV-augmented models ('HS SV', 'RIDGE SV', 'LASSO SV', and their moderate and large versions) are not specified. The text refers to Stock and Watson (2007) but gives no state equation, no prior distribution for the volatility parameters, no estimation algorithm, and no convergence diagnostics for the SV extensions of the Bayesian regression in Eq. (1). The central conclusion that SV 'appears to increase estimation noise rather than improving forecast accuracy' rests entirely on these undefined implementations. As presented, the comparison between SV and non-SV models cannot be audited or reproduced, and the possibility of a mis-specified or incorrectly estimated SV model cannot be ruled out.","section":"Section 5, Figures 1–3"},{"comment":"The abstract and conclusion claim that LASSO in the large-sized model without SV exhibits superior performance across multiple horizons, but Table 2 contradicts this. For the full period, LASSO large has relative RMSE 1.05 at h=4, 1.54 at h=8, and 2.69 at h=12, all worse than the benchmark; in the pandemic period, the relative RMSE is 1.08 at h=4 and 1.77 at h=12. Section 4 itself states that LASSO shows 'substantial degradation at h=8 and h=12'. The conclusion statement that 'the LASSO large model without SV consistently provides superior performance across both short- and long-term horizons' is therefore inconsistent with the reported results, and this inconsistency affects the main policy message of the paper.","section":"Table 2 and Abstract/Conclusion"},{"comment":"The cumulative log predictive likelihood (LPL) values in Section 5 (e.g., HS 1.6045 vs. HS SV -1070.0754 at h=1; HS moderate 35.2046 vs. HS moderate SV -1549.2438 at h=4; HS large -6.0832 vs. HS large SV -1116.3535 at h=12) are not defined and appear implausible. No formula for cumulative LPL is given, no table reports these values, and no explanation is provided for the order of magnitude. A value of -1070 for a cumulative monthly log score over roughly 40 evaluation periods would imply an average log likelihood near -26, which is not credible for a well-calibrated predictive density. These numbers are the primary evidence for the SV-versus-non-SV conclusion, but as they stand they suggest a computational failure or a scoring convention that is not described, rather than a reliable empirical result.","section":"Section 5, LPL numbers"},{"comment":"Several prose numbers in Section 4.1 do not appear in the tables. The text claims 'Ridge ... an extreme qwCRPS score of almost 200%', but Table 3 (tails) reports RIDGE large full h=1 = 1.03, not approximately 2. The text claims 'DL suffered a catastrophic breakdown (7.65)' in the right-weighted CRPS discussion, but Table 4 (right) shows DL large h=4 = 0.93 (full) and 0.80 (pandemic), and no value near 7.65 appears in any table. The same section refers to DL 'extreme inefficiencies (6.51 in full range, 3.50 in pandemic)', which also do not appear in Tables 3–5. These discrepancies mean the density-forecast results, which support the tail-risk asymmetry claim, cannot be reliably interpreted from the manuscript as written.","section":"Section 4.1"}],"minor_comments":[{"comment":"The text says 'we investigate the forecasting performance of six priors' but then lists noninformative, Ridge, Lasso, Horseshoe, Horseshoe+, and Spike-and-Slab plus, later, DL in the results; the count and the list should be reconciled. Also, the intro sentence about frequentist approaches is an incomplete fragment.","section":"Section 2, introductory paragraph"},{"comment":"The Adaptive Lasso notation is confusing: lambda is used as the Laplace rate on the left-hand side and as the local regularization parameter on the right-hand side, and the 'epsilon' in the denominator is undefined. Clarifying these definitions would help reproducibility.","section":"Equation (6)"},{"comment":"The driver narrative in Section 6 interprets the Horseshoe shrinkage weight kappa from the in-sample full-period fit as evidence of causal importance for Thai inflation. This is not supported by an out-of-sample validation or a formal test, and the reader should be told that these are descriptive posterior summaries rather than identified causal effects. This is not central to the forecasting comparison, but it is presented as a substantive finding.","section":"Section 6"},{"comment":"The sentence 'In my opinion, the results highlight an important trade-off...' uses a first-person phrase that is unusual in a co-authored research paper and should be reworded.","section":"Section 5, Ridge discussion"},{"comment":"The data description says 'the first 20 series are included in moderate-sized model setup', but the list in Table 6 includes 56 series and the first series is CPI; it is unclear whether the moderate model includes CPI plus 19 predictors or 20 predictors including CPI. Please clarify the exact variable sets.","section":"Table 6 / Section 4"},{"comment":"Several references and terms have typographical issues (e.g., 'Hern' andez-Lobato' with a stray space, 'Ba' nbura', 'sucha as', 'a the AR(2)'). A careful proofread would be needed before resubmission. The paper also does not state whether data and code are available for reproducibility.","section":"Throughout"}],"recommendation":"reject","confidential_remarks":"Given the direct contradiction between the abstract/conclusion and Table 2, the undefined DL prior and SV-augmented models, and the mismatch between the Section 5 LPL numbers and any plausible scoring convention, I do not see how a revision within the normal scope could repair the central claims. The manuscript would need a substantial redo of the empirical analysis and a transparent specification of all models, so I recommend rejection rather than major revision. There is no indication of academic misconduct; the issues are internal inconsistencies and incomplete model descriptions, but they are severe enough to make the reported rankings unauditable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a genuine, good-faith empirical exercise, but the manuscript is not ready for circulation as a research paper. The new thing is the application: comparing Horseshoe, Horseshoe+, Lasso, Ridge, Spike-and-Slab, and noninformative priors for forecasting Thai headline inflation in univariate regressions with 2, 20, and 56 predictors, using rolling direct forecasts and three scoring rules. That is a reasonable and useful exercise for an emerging-market central bank, and the paper does some things well: the non-SV priors are mostly defined, the forecast evaluation is out-of-sample, and the scoring metrics are standard.\n\nThe soft spots are serious. The Dirichlet-Laplace prior, which appears in the abstract and all the results tables, is never defined anywhere. The SV-augmented models ('HS SV', 'RIDGE SV', 'LASSO SV') that the entire Section 5 comparison rests on have no equations, no priors on volatility, no sampler details, no convergence checks. That means the central claim—that SV adds noise rather than accuracy—cannot be audited. The numbers in Section 5 reinforce this: HS at 1.6045 versus HS SV at -1070.0754 is not a believable model comparison; it looks like a computational failure or a different scoring convention being described as if it were the same metric.\n\nThe abstract's summary sentence is also contradicted by Table 2. It says LASSO large without SV is superior across multiple horizons, but the table shows relative RMSE of 1.05, 1.54, and 2.69 at h=4, 8, and 12—worse than the benchmark. Several prose numbers in Section 4 (Ridge qwCRPS 'almost 200%', DL '7.65') do not appear in any table. Section 6's driver narrative uses the same fitted Horseshoe model to 'discover' which predictors matter and then tells an economic story, without any out-of-sample check or identification; that is speculation, not evidence.\n\nNone of these flaws are fatal to the underlying question. The data could support a good paper, and the non-SV comparison among the defined priors is at least partially interpretable. But as it stands, the headline findings rest on undefined models and internally inconsistent tables. I would not cite it, and I would not bring it to a reading group as a substantive result. That said, I would send it to peer review rather than desk-reject: the topic is legitimate, the authors are engaging the right literature, and a referee report that demands complete model definitions, corrected tables, and reproducibility would either fix the paper or end it. That is what the process is for.","headline":"A real empirical exercise on Bayesian shrinkage priors for Thai inflation, but the headline SV comparison is unauditable and the abstract's own claim is contradicted by Table 2.","tokens_in":22826,"tokens_out":2500,"would_cite":false,"duration_ms":25628,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62J07","62M20","91B84"],"pacs":[],"model":"deepseek-v4-flash","headline":"For Thai inflation, Bayesian shrinkage with many predictors and no stochastic volatility outperforms SV-augmented models on RMSE, qwCRPS, and log predictive likelihood.","keywords":["Thai inflation","Bayesian shrinkage priors","Horseshoe prior","Dirichlet-Laplace prior","stochastic volatility","direct multi-step forecasting","quantile-weighted CRPS","inflation forecasting"],"falsifier":"Run the same rolling direct-forecast exercise with an explicit stochastic-volatility likelihood, say an AR(2) mean equation with a log-volatility random walk, and an explicit Dirichlet-Laplace prior, then compare cumulative log predictive likelihood at h=8 and h=12; if any SV-augmented model beats its non-SV counterpart at those horizons, or if the large Lasso without SV loses to a correctly specified benchmark, the paper's central claim is not supported.","tokens_in":21665,"feed_emoji":"📈","tokens_out":7073,"duration_ms":69106,"temperature":0.7,"pith_summary":"The paper tries to establish that, for Thai headline inflation, a univariate Bayesian linear regression with a wide panel of predictors and a shrinkage prior forecasts better when it omits stochastic volatility (SV) than when it builds time-varying volatility in. Across rolling direct forecasts from 2006 to 2024 and on RMSE, quantile-weighted CRPS, and log predictive likelihood, the Horseshoe, Dirichlet-Laplace, and Lasso priors in the large 56-predictor specification rank at or near the top, while SV-augmented versions underperform most clearly in high-dimensional settings. The paper also argues that the same horseshoe machinery isolates cost-push drivers, especially food and fuel components and Bangkok alien work permits, while shrinking demand-side and monetary variables to zero. A sympathetic reader would care because the claim challenges the default use of stochastic volatility in inflation forecasting and points to simpler, broad-data models as the more reliable tool for an emerging-market central bank.","feed_headline":"Stochastic volatility degrades Thai inflation forecasts, study finds","feed_subtitle":"A 56-predictor horseshoe/Lasso model without time-varying variance beats SV-augmented rivals, especially at the one-month horizon.","key_machinery":"The engine of the paper is a univariate Bayesian linear regression $y = X\\beta + \\epsilon$, $\\epsilon \\sim N(0,\\sigma^2 I)$, with priors on $\\beta$: the noninformative benchmark $N(0, 10^4 \\sigma^2 I)$, Ridge, adaptive Lasso, Spike-and-Slab, Horseshoe, and Horseshoe+, plus a Dirichlet-Laplace prior that is named in the results but not defined in the methodology. The Horseshoe family is the central object: it is a global-local shrinkage prior with a heavy-tailed local parameter $\\lambda_j$ and global $\\tau$, which shrinks noise toward zero while leaving large signals almost untouched, and the paper uses its shrinkage factor $\\kappa$ (values near 1 meaning 'unshrunk') to rank predictors. The forecasting machinery is rolling direct multi-step estimation at horizons $h=1,4,8,12$ over a 340-month sample with 20- and 56-variable predictor sets and an AR(2) benchmark, scored by RMSE, quantile-weighted CRPS with tail, center, left, and right weight functions, and cumulative log predictive likelihood. LPL carries the SV-versus-no-SV comparison, while qwCRPS carries the tail-risk asymmetry.","core_discovery":"On its own terms, the paper's discovery is an empirical regularity: in a direct multi-step forecasting exercise for Thai headline CPI, models without stochastic volatility dominate their SV-augmented counterparts, and the gap grows with the predictor set. The large 56-predictor Horseshoe, Horseshoe+, Dirichlet-Laplace, and Lasso regressions produce relative RMSE values well below 1 at the one-month horizon (for example 0.15–0.16 for HS and HS+, 0.21 for DL) and relative tail-weighted CRPS values around 0.58–0.62, while the SV versions, especially HS with SV, show severely negative cumulative log predictive likelihoods. The paper also reports an asymmetry: left-tail (deflationary) risk is captured well by the heavy-tailed shrinkage priors, whereas right-tail (inflationary-spike) forecasting is unstable, with some priors breaking down under stress. Finally, the Horseshoe's shrinkage factor $\\kappa$ identifies a small set of supply-side CPI components plus the number of alien work permits in Bangkok as the predictors that consistently resist shrinkage, which the authors read as the cost-push signature of a credible inflation-targeting regime.","pith_inferences":["If the result holds beyond Thailand, it suggests that for small open inflation-targeting economies with relatively stable volatility, the default assumption that SV improves density forecasts should be reversed; a testable extension is to rerun the same direct-forecast design on Indonesia, the Philippines, or other emerging economies.","The right-tail weakness may be partly an artifact of the chosen quantile weights; a natural test is to tilt qwCRPS weights even more heavily toward the upper 5% quantile and see whether the large shrinkage priors recover their rank or whether the failure is intrinsic.","The Horseshoe-$\\kappa$ ranking could be turned into a real-time monitoring dashboard: tracking whether supply-side CPI components move above $\\kappa = 0.6$ would give a data-driven early signal of cost-push pressure before inflation materializes.","Because the SV implementations are not specified in the paper, a task for future work is to re-estimate with a standard SV likelihood, such as an AR(2) mean equation with a log-volatility random walk, and check whether the ordering is an artifact of estimation noise rather than a property of SV itself."],"forward_implications":["Policymakers at the Bank of Thailand can treat the 56-predictor Horseshoe, Dirichlet-Laplace, or Lasso regression without SV as a better short-run inflation forecast than SV-augmented alternatives, especially at the one-month horizon.","In high-dimensional settings, adding stochastic volatility is predicted to reduce rather than improve forecast accuracy, so model builders should decide on the predictor set before adding SV.","Persistent high-$\\kappa$ variables such as Eggs & Dairy, Electricity/Fuel/Water, and Bangkok alien work permits give a short list of supply-side indicators to watch for cost-push inflation under the targeting regime.","Left-tail deflationary forecasts are more trustworthy than right-tail inflationary-surge forecasts, so tail-risk reporting should report left and right quantile scores separately rather than one average.","The COVID-19 sub-period results suggest the no-SV large models also hold up in crisis episodes, implying the finding is not limited to calm periods."],"supporting_citations":[{"why":"Introduces the Horseshoe prior whose global-local shrinkage structure is the paper's primary forecasting prior.","marker":"Carvalho et al. (2010)"},{"why":"Provides the Gibbs sampler the paper uses to draw from the Horseshoe posterior.","marker":"Makalic and Schmidt (2015)"},{"why":"Introduces Horseshoe+, the ultra-sparse heavy-tailed prior evaluated throughout the paper.","marker":"Bhadra et al. (2017)"},{"why":"Introduces the Dirichlet-Laplace prior that appears among the top-performing large models.","marker":"Bhattacharya et al. (2015)"},{"why":"Defines the Lasso penalty that the adaptive Bayesian Lasso prior builds on.","marker":"Tibshirani (1996)"},{"why":"Defines adaptive Lasso weights used in the localized Lasso prior specification.","marker":"Zou (2006)"},{"why":"Provides the UC-SV iterated benchmark and the motivation for comparing stochastic-volatility variants.","marker":"Stock and Watson (2007)"},{"why":"Defines CRPS, the density scoring rule underlying qwCRPS.","marker":"Gneiting and Raftery (2007)"},{"why":"Introduces quantile- and threshold-weighted scoring rules used for tail-risk evaluation.","marker":"Gneiting and Ranjan (2011)"},{"why":"Motivates the direct multi-step forecast design by comparing direct and iterated multi-step AR forecasts.","marker":"Marcellino et al. (2006)"}],"fun_headline_variants":["SV models lose to simpler shrinkage in Thai inflation tests","Thai inflation forecast: Skip stochastic volatility, use many predictors","Horseshoe prior without SV dominates Thai inflation forecasts","SV harms Thai CPI forecasts; simpler priors with 56 predictors win"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The head-to-head that supports the title claim assumes the SV-augmented models and the Dirichlet-Laplace prior are implemented correctly, but the paper gives no equations, likelihood, or sampler details for the SV versions and never defines the Dirichlet-Laplace prior, so the reported rankings cannot be audited or reproduced.","fun_headline_variants_meta":{"raw":{"variants":["SV models lose to simpler shrinkage in Thai inflation tests","Thai inflation forecast: Skip stochastic volatility, use many predictors","Horseshoe prior without SV dominates Thai inflation forecasts","SV harms Thai CPI forecasts; simpler priors with 56 predictors win"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000374,"raw_usage":{"total_tokens":2045,"prompt_tokens":1044,"completion_tokens":1001,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":933}},"tokens_in":660,"tokens_out":1001,"duration_ms":8457,"temperature":1.0,"reasoning_tokens":933,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:06:45.318853+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same rolling direct-forecast exercise with an explicit stochastic-volatility likelihood, say an AR(2) mean equation with a log-volatility random walk, and an explicit Dirichlet-Laplace prior, then compare cumulative log predictive likelihood at h=8 and h=12; if any SV-augmented model beats its non-SV counterpart at those horizons, or if the large Lasso without SV loses to a correctly specified benchmark, the paper's central claim is not supported.","supporting_citations":[{"cited_title":"G., and Willard, B","cited_arxiv_id":null,"evidence_quote":"Introduces Horseshoe+, the ultra-sparse heavy-tailed prior evaluated throughout the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the UC-SV iterated benchmark and the motivation for comparing stochastic-volatility variants."},{"cited_title":"H., and Watson, M","cited_arxiv_id":null,"evidence_quote":"Motivates the direct multi-step forecast design by comparing direct and iterated multi-step AR forecasts."}],"review_version":1}