{"id":"0ae2612f-8959-48f2-9cff-aa0242e5cc77","arxiv_id":"2411.16244","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Spike-and-slab selection on intraday FX returns identifies nine US and Australian macro announcements as the only events with over 95% posterior inclusion probability, all linked by the authors to Taylor-rule fundamentals.","lead":"This paper fits a Bayesian stochastic volatility model to 5-minute Australian dollar and Swiss franc returns and uses spike-and-slab priors to select which macroeconomic announcements move exchange rate volatility. It offers a data-driven way to identify the few macro releases that actually matter for currency markets, and it claims practical gains in volatility forecasting and portfolio allocation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Binary announcement indicators ignore surprise magnitude, so the nine 'relevant' events may reflect large realized surprises in 2017-2023 rather than intrinsic importance.","rationale":"The reader's weakest_assumption is the same one I identify: binary release dummies cannot separate event importance from surprise size. This is the most load-bearing concern because the paper's distinctive empirical contribution is the specific list of nine events and the claim that data-driven selection can replace researcher-chosen event lists. If the binary specification conflates event identity with surprise realizations, that list and the Taylor-rule narrative could change once surprises are included. I found no reason to move the verdict beyond the reader's CONDITIONAL: the concern is real but testable, and the paper has partial independent support (forecast horse-race results and the seasonality-volume correlation) that would survive even if the event list shifts. A separate internal issue is that Table 2 shows SSV annual volatility 10.49 below the proposal's 10.50, contradicting the text's claim of 'smallest volatility among all models'; this concerns the portfolio contribution rather than the central event-selection claim, so I treat it as secondary. MCMC diagnostics and prior hyperparameter disclosure are also absent, but those are reproducibility concerns rather than a demonstrated flaw in the argument. The proposed extension with standardized surprises would resolve the main uncertainty without requiring the authors to abandon the model family.","tokens_in":24204,"tokens_out":6229,"duration_ms":68624,"concrete_test":"Re-estimate the model with the event component extended to et = I'_t alpha + I'_t (S_t ⊙ delta), where S_t is the absolute standardized surprise (|actual - consensus| / historical standard deviation) for the event occurring at time t, using Bloomberg consensus data; report posterior inclusion probabilities and effect sizes for alpha and delta for all 117 events. If the same nine events remain above 95% inclusion and the surprise interaction is large for them, the binary-only concern is not the source of the selection. If inclusion probabilities change materially (e.g., events like ADP employment, initial jobless claims, or PMIs enter, or some of the nine drop below 95%), the headline event list is confounded by surprise realizations and the central claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central selection result rests on Equation (5), et = I'_t alpha, where a scheduled release enters as a 0/1 dummy with one coefficient per event type. This treats a release with a negligible surprise exactly like one with a large surprise, even though the FX announcement literature models volatility as a function of standardized surprises. Under surprise-driven volatility, the estimated alpha_i is a pooled average over heterogeneous release sizes, and the posterior inclusion probability measures whether that event type had sufficiently large average surprise-driven volatility in the 2017-2023 sample, not whether the event intrinsically matters. The sample includes the COVID-era inflation surge, so US CPI, nonfarm payrolls and FOMC decisions had unusually large surprises; Australian events like retail sales and employment change also had volatile realizations. The paper neither tests nor acknowledges this. If the concern lands, the headline 'only nine events matter' and the Taylor-rule interpretation are at least partly an artifact of which surprises happened to be large, weakening the claim that data-driven selection can replace hand-picked event lists. The out-of-sample forecast gains and volume-seasonality correlation are independent support for the overall model, but they do not validate the identity of the selected events.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a Bayesian stochastic volatility model for 5-minute Australian Dollar returns that jointly captures volatility persistence, intraday seasonality, and the effects of a large set of scheduled macroeconomic announcements from Australia and the US. Event effects are modeled with spike-and-slab priors, which lets the data select a sparse set of relevant announcements. The authors report that only nine events have posterior inclusion probability above 95%, that the seasonal component follows a W-shaped pattern correlated with trading volume, and that the model outperforms standard SV and GARCH alternatives in out-of-sample realized volatility forecasting and in a global minimum variance portfolio allocation with the Swiss Franc.","tokens_in":24399,"tokens_out":6317,"duration_ms":206231,"significance":"If the results are robust, the paper offers a useful data-driven alternative to hand-picked event lists in intraday FX volatility modeling, and it provides concrete out-of-sample forecast gains and portfolio improvements over standard benchmarks. The authors ship a transparent decomposition of log-variance into persistence, seasonality, and announcement components, and they connect the seasonal component to trading volume with a high R-squared. The out-of-sample design and the horse-race comparisons are strengths.","major_comments":[{"comment":"The headline selection result ('only nine events are included more than 95%') is based on pointwise posterior means of the inclusion indicators π_i, but the paper reports no credible intervals for these inclusion probabilities, no MCMC convergence diagnostics, and no sensitivity analysis to the Beta prior on γ or the slab variance σ_α^2. With 702 event-related dummies, the posterior of a single inclusion probability can be highly uncertain; a posterior mean of 0.95 does not by itself establish that the event is selected with high probability. This is load-bearing for the central claim, so the authors should add interval estimates and prior sensitivity checks.","section":"Section 3.1 and Appendix D"},{"comment":"The announcement effect is modeled with binary time-of-release indicators, so a release with a negligible surprise is treated identically to one with a large surprise. The FX announcement literature typically models volatility responses to standardized surprises relative to expectations. Under surprise-driven volatility, the estimated α_i and inclusion probabilities pool over heterogeneous surprise realizations, and the selected list may reflect which events happened to have large surprises in 2017–2023 (e.g., the COVID-era inflation surge) rather than intrinsic importance. The paper neither tests this assumption nor acknowledges it as a limitation. The authors should at least discuss this issue and, ideally, include a robustness check using surprise measures to show the selection is not an artifact of the sample window.","section":"Section 2.2, Equation (5)"},{"comment":"The forecasting exercise does not describe how one-step-ahead volatility forecasts are generated from the SV model. The statement that forecasts are the 'average volatility forecast by our proposal when parameters are fixed at their posterior mean' is insufficient, because the latent SV state x_t must be filtered or integrated out in a nonlinear state space model. Without a precise algorithm (e.g., particle filter or a Kalman filter on the linearized mixture), the out-of-sample comparison is not reproducible and the Diebold-Mariano tests cannot be verified. Please provide the exact forecasting procedure.","section":"Section 4.1 and Appendix C"},{"comment":"The MCMC scheme is described algorithmically but lacks essential implementation details: number of iterations, burn-in, thinning, starting values, and convergence diagnostics are not reported. Since the posterior inclusion probabilities are MCMC estimates, evidence of convergence is needed to trust the 95% threshold claims. This is a reproducibility issue that should be fixed.","section":"Section 2.2 and Appendix C"},{"comment":"The portfolio application reports annualized volatilities and Sharpe ratios without any measure of uncertainty. The difference between the proposal (Sharpe 0.81) and SSV (0.69) may be within sampling variation, yet the paper concludes that the proposed model yields the highest Sharpe ratio. Add confidence intervals or bootstrap standard errors for the performance metrics to support this claim.","section":"Section 4.2, Table 2"}],"minor_comments":[{"comment":"The text refers to 'Sidney' instead of 'Sydney', and the dashed-line colors in the caption (pink, red, light and dark blue, green, orange) do not match the colors described in the main text (red, light and dark blue, green, orange). Please harmonize the figure caption and text.","section":"Figure 3 and text"},{"comment":"The sentence 'Figures 7 and show' is incomplete; it should refer to Figures 7 and 8.","section":"Appendix B"},{"comment":"The symbol s_t is used both for the seasonal component in Equation (4) and for the nominal exchange rate in Equation (8). This notation conflict may confuse readers; please use a different symbol for one of them.","section":"Equations (4) and (8)"},{"comment":"The paper does not cite the classic announcement-surprise literature (e.g., Andersen, Bollerslev, Diebold, Vega 2003; Balduzzi, Elton, Green 2001) or discuss why it departs from that framework by using binary indicators. Adding this context would help position the contribution.","section":"Introduction"},{"comment":"The table layout and the regression equation contain spacing artifacts (e.g., 'P roposal dV olt|t−1' and 'Competitor dV olt|t−1'). These should be typeset cleanly.","section":"Table 1"},{"comment":"The horse-race regression in Equation (9) restricts b1 to [0,1]; the paper does not explain why this restriction is imposed or how the t-statistics are computed. A brief note would improve clarity.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a competent application of existing spike-and-slab variable selection to a new intraday FX setting. The main novelty is the event-selection result, but the absence of uncertainty quantification and prior sensitivity analysis makes the headline claim fragile. The authors should also be encouraged to discuss the sample-period dependence (2017–2023 includes the COVID-19 inflation surge) and to compare against a hand-picked event model to substantiate the claim that data-driven selection beats expert selection. The forecasting and portfolio sections would also benefit from more rigorous uncertainty reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this paper's contribution is a spike-and-slab stochastic volatility model that selects from 702 announcement-lag dummies for 5-minute AUD returns, and that is genuinely new. Prior work picks a handful of events by hand. The paper should get a serious referee, but the \"only nine events matter\" claim is stronger than the evidence supports.\n\nWhat's good: the modeling is sensible and the forecast comparison is genuinely out-of-sample. The Diebold-Mariano results against six competitors are clean, and the horse-race regressions support the model. The seasonality result is a nice separate finding: the estimated daily pattern is W-shaped, and the posterior mean of the seasonal component correlates with average traded volume at R2=0.88 even though only returns were used. That is falsifiable and reproducible. The W-shaped pattern with peaks at market openings extends the old U-shaped literature and is worth discussing.\n\nThe soft spots: first, the binary event indicator assumption, et = I'_t alpha, treats a release with a tiny surprise identically to one with a large surprise. The FX announcement literature mostly models volatility as a function of standardized surprises. If that is right, the posterior inclusion probability measures whether an event type had large realized surprise volatility in this sample, not whether it intrinsically matters. The sample includes the COVID inflation surge, so US CPI, nonfarm payrolls, and FOMC decisions had unusually large surprises. The paper neither tests nor acknowledges this. This does not kill the method; it makes the headline \"only nine events matter\" too strong, and the nine-event list may be sample-specific.\n\nSecond, the Bayesian machinery is under-documented. There are no MCMC convergence diagnostics, no credible intervals for the inclusion probabilities, and no sensitivity analysis for the beta prior on gamma or the slab variance. No code or data are shipped. For a paper whose central result is a set of posterior inclusion probabilities, that is a real gap.\n\nThird, a small but concrete error: Table 2 shows the proposal's annualized volatility at 10.50, with SSV at 10.49. The text says the proposal yields the smallest volatility, which is contradicted by their own table. The Sharpe claim still holds, but the sentence should be corrected.\n\nWho this is for: researchers working on intraday FX volatility, macro announcement effects, and Bayesian variable selection in SV models. The method is plausible and the empirical findings are interesting enough to deserve referee time. A good referee should push for disclosure of MCMC details, a surprise-size sensitivity check, and a more honest statement about what the selected event list does and does not show.","headline":"Data-driven event selection is a real step forward, but the paper overclaims the identity of the nine events and under-reports MCMC validation.","tokens_in":24936,"tokens_out":2861,"would_cite":true,"duration_ms":41719,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62M10","91B84","62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A spike-and-slab stochastic volatility model applied to five-minute Australian dollar returns finds that only nine of 117 U.S.","keywords":["stochastic volatility","exchange rate volatility","macroeconomic announcements","spike-and-slab priors","intraday seasonality","realized volatility forecasting","portfolio allocation","carry trade"],"falsifier":"Re-estimate the model replacing each release dummy with the standardized surprise, defined as the announced value minus the consensus forecast divided by its standard deviation, and compare the inclusion probabilities and effect sizes; if surprise size rather than mere release is what drives volatility, the binary model's event rankings should shift and its out-of-sample forecast advantage should shrink.","tokens_in":23967,"feed_emoji":"📈","tokens_out":4267,"duration_ms":41251,"temperature":0.7,"pith_summary":"This paper tries to show that the macroeconomic events that truly drive exchange-rate volatility can be discovered from data instead of being chosen from experience. It fits a Bayesian stochastic volatility model to five-minute Australian dollar returns, letting every one of 117 U.S. and Australian announcements enter through spike-and-slab priors that assign each event a posterior probability of inclusion and an effect size. The model reports that only nine events, all plausibly related to Taylor-rule fundamentals, pass the 95% inclusion threshold, and that they raise volatility for up to thirty minutes after release. It also recovers a W-shaped intraday seasonal pattern linked to market openings and trading volume, and shows that including events and seasonality improves realized-volatility forecasts and portfolio allocation relative to standard SV and GARCH benchmarks. If the paper is right, researchers and traders no longer need hand-picked event lists, because the release calendar itself can tell which events matter.","feed_headline":"Nine macro events drive Aussie dollar volatility","feed_subtitle":"A spike-and-slab FX volatility model selects Taylor-rule releases and beats GARCH in forecasts and portfolios.","key_machinery":"The key machinery is a spike-and-slab prior on announcement coefficients inside a multiplicative stochastic volatility model. Log-variance is decomposed into level, persistent stochastic volatility, intraday seasonality, and announcement effects: $h_t = \\mu_h + x_t + s_t + e_t$, with $e_t = I_t'\\alpha$ and $s_t = H_t'\\beta$. Each announcement coefficient comes from either a Dirac spike at zero or a Gaussian slab, so the model learns which events are included and how large their effects are. Estimation uses MCMC, including the seven-component Gaussian mixture approximation of Kim et al. for the latent volatility states and Geweke's method for sampling inclusion indicators despite the Dirac mass.","core_discovery":"The central discovery is that data-driven event selection, rather than prior experience, identifies a small set of announcements that reliably move Australian dollar volatility: FOMC rate decisions, U.S. nonfarm payrolls, U.S. CPI, FOMC meeting minutes, U.S. retail sales, the RBA cash rate target, Australian employment change, Australian GDP, and Australian retail sales. These nine events all connect to the interest-rate rule that central banks follow, linking them to exchange-rate determination through interest differentials. The paper also finds that the estimated intraday seasonal component has a W shape, with peaks at the openings of Asian, European, and U.S. markets, and that this component tracks average traded volume closely, with a simple regression yielding an R-squared of 0.88. Finally, the model out-forecasts standard SV and GARCH specifications in out-of-sample realized-volatility horse races, with Diebold-Mariano p-values near zero and b1 coefficients close to one, and it delivers the smallest portfolio volatility and highest Sharpe ratio among the compared specifications.","pith_inferences":["Because the model represents each announcement as a binary release indicator with a constant multiplicative effect, it implicitly assumes that a scheduled release with a trivial surprise moves volatility as much as one with a large surprise; an extension replacing dummies with standardized surprise magnitudes could change the event ranking and the estimated effect sizes.","The same sparse event-selection approach could be applied to other liquid currencies and asset classes with long event calendars, such as bond yields or equity index volatility, where the combination of release dummies, intraday seasonality, and persistent stochastic volatility could be reused directly.","The W-shaped seasonality and its strong link to trading volume suggest a testable labor-supply story: if market openings drive volatility because traders concentrate work at the start of the day, the seasonal peaks should shift with daylight-saving time changes or holiday schedules that alter opening hours relative to Greenwich Mean Time."],"forward_implications":["Macroeconomic event lists for volatility modeling can be built directly from release calendars, avoiding the cherry-picking of events that plagues hand-selected lists.","The selected events all line up with Taylor-rule fundamentals, giving a coherent economic interpretation: news about interest-rate-setting variables moves exchange rates through expected interest differentials.","Intraday volatility is not simply U-shaped; the estimated W-shaped seasonality reflects the global sequence of market openings and is strongly associated with trading volume, so volume and volatility are linked at high frequency.","Ignoring announcement effects and seasonality costs forecast accuracy: the full model dominates SV, GARCH, HAR, and related benchmarks in out-of-sample realized-volatility prediction, with competitor models adding almost no information once the proposal is included.","The model's volatility forecasts translate into economic gains, producing the lowest variance and highest Sharpe ratio in an Australian dollar/Swiss franc global minimum variance portfolio."],"supporting_citations":[{"why":"Supplies the seven-component Gaussian mixture approximation used to sample the latent stochastic volatility states in the MCMC algorithm.","marker":"Kim et al. [1998]"},{"why":"Provides the method for sampling spike-and-slab inclusion indicators without getting stuck on the Dirac mass at zero.","marker":"Geweke [1996]"},{"why":"Source of the volatility decomposition into level, stochastic volatility, seasonality, and announcement components, as well as the horse-race regression used for forecast comparison.","marker":"Stroud and Johannes [2014]"},{"why":"Defines the Taylor rule that gives the economic interpretation for why the nine selected announcements, all linked to inflation and output-gap variables, should move exchange rate volatility.","marker":"Taylor [1993]"},{"why":"Connects exchange rates to expected interest-rate differentials, which is the channel through which Taylor-rule-related news affects currency volatility.","marker":"Campbell and Clarida [1987]"},{"why":"Establishes the persistence and announcement effects in high-frequency FX volatility that the model is designed to capture.","marker":"Andersen and Bollerslev [1998]"},{"why":"Provides an example of hand-picked event selection that the paper contrasts with its data-driven approach, including events the paper's model does not select.","marker":"Bauwens et al. [2005]"},{"why":"Supplies the volume-volatility relationship that the paper extends to high-frequency currency data and to the seasonal component specifically.","marker":"Abanto-Valle et al. [2010]"},{"why":"Justifies the choice of Australian dollar and Swiss franc as investment and funding currencies in the portfolio application.","marker":"Lustig et al. [2011]"}],"fun_headline_variants":["Nine macro events move the Aussie dollar, model shows","Data-driven method pins FX swings on nine key releases","Taylor-rule releases predict currency volatility better","Sparse model reveals which macro news drive the Aussie","Aussie FX volatility: nine events that matter most"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"A scheduled announcement moves volatility just by being released at a set time, with the same multiplicative effect every time it occurs, regardless of whether the number announced matches what markets expected.","fun_headline_variants_meta":{"raw":{"variants":["Nine macro events move the Aussie dollar, model shows","Data-driven method pins FX swings on nine key releases","Taylor-rule releases predict currency volatility better","Sparse model reveals which macro news drive the Aussie","Aussie FX volatility: nine events that matter most"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000351,"raw_usage":{"total_tokens":1902,"prompt_tokens":924,"completion_tokens":978,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":915}},"tokens_in":540,"tokens_out":978,"duration_ms":9805,"temperature":1.0,"reasoning_tokens":915,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:20:53.071986+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate the model replacing each release dummy with the standardized surprise, defined as the announced value minus the consensus forecast divided by its standard deviation, and compare the inclusion probabilities and effect sizes; if surprise size rather than mere release is what drives volatility, the binary model's event rankings should shift and its out-of-sample forecast advantage should shrink.","supporting_citations":[],"review_version":1}