A best-practices workflow for unsupervised scientific discovery, illustrated by a stability- and generalizability-driven clustering case study of Milky Way globular clusters using APOGEE data.
Rashomon effect in Educational Research: Why More is Better Than One for Measuring the Importance of the Variables?
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This study explores how the Rashomon effect influences variable importance in the context of student demographics used for academic outcomes prediction. Our research follows the way machine learning algorithms are employed in Educational Data Mining, focusing on highlighting the so-called Rashomon effect. The study uses the Rashomon set of simple-yet-accurate models trained using decision trees, random forests, light GBM, and XGBoost algorithms with the Open University Learning Analytics Dataset. We found that the Rashomon set improves the predictive accuracy by 2-6%. Variable importance analysis revealed more consistent and reliable results for binary classification than multiclass classification, highlighting the complexity of predicting multiple outcomes. Key demographic variables imd_band and highest_education were identified as vital, but their importance varied across courses, especially in course DDD. These findings underscore the importance of model choice and the need for caution in generalizing results, as different models can lead to different variable importance rankings. The codes for reproducing the experiments are available in the repository: https://anonymous.4open.science/r/JEDM_paper-DE9D.
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Unsupervised Machine Learning for Scientific Discovery: Workflow and Best Practices
A best-practices workflow for unsupervised scientific discovery, illustrated by a stability- and generalizability-driven clustering case study of Milky Way globular clusters using APOGEE data.