Pith. sign in

REVIEW 5 cited by

OpenML Benchmarking Suites

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1708.03731 v3 pith:ZA3EPOEV submitted 2017-08-11 stat.ML cs.LG

classification stat.MLcs.LG
keywords benchmarkingopenmlsuitesbenchmarkscuratedclassificationlearningmachine
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning research depends on objectively interpretable, comparable, and reproducible algorithm benchmarks. We advocate the use of curated, comprehensive suites of machine learning tasks to standardize the setup, execution, and reporting of benchmarks. We enable this through software tools that help to create and leverage these benchmarking suites. These are seamlessly integrated into the OpenML platform, and accessible through interfaces in Python, Java, and R. OpenML benchmarking suites (a) are easy to use through standardized data formats, APIs, and client libraries; (b) come with extensive meta-information on the included datasets; and (c) allow benchmarks to be shared and reused in future studies. We then present a first, carefully curated and practical benchmarking suite for classification: the OpenML Curated Classification benchmarking suite 2018 (OpenML-CC18). Finally, we discuss use cases and applications which demonstrate the usefulness of OpenML benchmarking suites and the OpenML-CC18 in particular.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PIPES: A Meta-dataset of Machine Learning Pipelines

    cs.LG 2025-09 conditional novelty 6.0 of 10

    PIPES is a new meta-dataset containing results of 9,408 combinations of imputation, encoding, scaling, feature preprocessing, and classification techniques on 280 successfully processed datasets.

  2. Meta-learning ecological priors from large language models explains human learning and decision making

    q-bio.NC 2025-08 conditional novelty 6.0 of 10

    A meta-learned transformer trained on LLM-generated tasks (ERMI) outperforms classical cognitive models in predicting human choices across function learning, category learning, and decision making.

  3. TabFlex: Scaling Tabular Learning to Millions with Linear Attention

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Linear attention lets a TabPFN-style model process millions of tabular samples in seconds with near-identical accuracy on small datasets.

  4. Towards Benchmarking Foundation Models for Tabular Data With Text

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A new 13-dataset benchmark shows that adding text embeddings to tabular models usually improves accuracy, but no embedding or downsampling strategy dominates.

  5. Divide, Specialize, and Route: A New Approach to Efficient Ensemble Learning

    cs.LG 2025-06 reject novelty 4.0 of 10

    A difficulty-based ensemble that routes instances to specialized models is proposed, but reported gains are not tested against standard ensemble baselines and are filtered to favorable cases.

Pith tools