Walk-forward out-of-sample SHAP shows a random forest's PCE inflation importance swings from 0.778 on revised data to 0.039 on real-time data, while real-time/revised point-forecast accuracy differences are mostly insignificant.
Forecasting and Explaining the Phillips Curve: A SHAP-Based Comparison of Machine Learning and Traditional Time-Series Models for Canadian Unemployment and Inflation
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This study evaluates the out-of-sample forecasting ability of six model types: ARIMA, VAR, Random Forest, XGBoost, LSTM, and GRU, for monthly Canadian inflation from January 2012 to April 2026 (n = 172). The evaluation employs expanding-window walk-forward validation across 1-, 3-, 6-, and 12-month horizons. Results reveal a horizon-dependent shift: ARIMA significantly outperforms all machine learning and deep learning models at the one-month horizon (Diebold-Mariano p < 0.05). However, Random Forest and XGBoost become notably superior at six and twelve months, reducing RMSE by approximately 30-75 percent compared to ARIMA and VAR. LSTM and GRU perform well only at the shortest horizon, likely due to overfitting given the limited data. Analyzing four macroeconomic sub-periods shows that no single model consistently dominates. SHAP analysis of the top-performing XGBoost model indicates that lagged inflation is more influential than unemployment, which only becomes significantly impactful during the pandemic tail. The findings clarify when machine learning methods can surpass traditional benchmarks.
fields
stat.AP 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Forecasting in the Fog: Real-Time versus Revised-Data Evidence on Machine Learning's Edge over the Phillips Curve
Walk-forward out-of-sample SHAP shows a random forest's PCE inflation importance swings from 0.778 on revised data to 0.039 on real-time data, while real-time/revised point-forecast accuracy differences are mostly insignificant.