Random splits of panel data into training and test sets cause temporal and cross-sectional leakage that overstates out-of-sample performance; the correct split depends on whether the task is cross-sectional prediction or sequential forecasting.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
econ.EM 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
On the (Mis)Use of Machine Learning with Panel Data
Random splits of panel data into training and test sets cause temporal and cross-sectional leakage that overstates out-of-sample performance; the correct split depends on whether the task is cross-sectional prediction or sequential forecasting.