REVIEW 5 cited by
OpenML Benchmarking Suites
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Machine learning research depends on objectively interpretable, comparable, and reproducible algorithm benchmarks. We advocate the use of curated, comprehensive suites of machine learning tasks to standardize the setup, execution, and reporting of benchmarks. We enable this through software tools that help to create and leverage these benchmarking suites. These are seamlessly integrated into the OpenML platform, and accessible through interfaces in Python, Java, and R. OpenML benchmarking suites (a) are easy to use through standardized data formats, APIs, and client libraries; (b) come with extensive meta-information on the included datasets; and (c) allow benchmarks to be shared and reused in future studies. We then present a first, carefully curated and practical benchmarking suite for classification: the OpenML Curated Classification benchmarking suite 2018 (OpenML-CC18). Finally, we discuss use cases and applications which demonstrate the usefulness of OpenML benchmarking suites and the OpenML-CC18 in particular.
Forward citations
Cited by 5 Pith papers
-
PIPES: A Meta-dataset of Machine Learning Pipelines
PIPES is a new meta-dataset containing results of 9,408 combinations of imputation, encoding, scaling, feature preprocessing, and classification techniques on 280 successfully processed datasets.
-
Meta-learning ecological priors from large language models explains human learning and decision making
A meta-learned transformer trained on LLM-generated tasks (ERMI) outperforms classical cognitive models in predicting human choices across function learning, category learning, and decision making.
-
TabFlex: Scaling Tabular Learning to Millions with Linear Attention
Linear attention lets a TabPFN-style model process millions of tabular samples in seconds with near-identical accuracy on small datasets.
-
Towards Benchmarking Foundation Models for Tabular Data With Text
A new 13-dataset benchmark shows that adding text embeddings to tabular models usually improves accuracy, but no embedding or downsampling strategy dominates.
-
Divide, Specialize, and Route: A New Approach to Efficient Ensemble Learning
A difficulty-based ensemble that routes instances to specialized models is proposed, but reported gains are not tested against standard ensemble baselines and are filtered to favorable cases.
Discussion (0). Sign in to comment.