Pith. sign in

REVIEW 1 cited by

AutoML using Metadata Language Embeddings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.03698 v1 pith:XDDWYYAG submitted 2019-10-08 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords automlembeddingsperformancealgorithmscomputationdatasetframeworkimproves
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As a human choosing a supervised learning algorithm, it is natural to begin by reading a text description of the dataset and documentation for the algorithms you might use. We demonstrate that the same idea improves the performance of automated machine learning methods. We use language embeddings from modern NLP to improve state-of-the-art AutoML systems by augmenting their recommendations with vector embeddings of datasets and of algorithms. We use these embeddings in a neural architecture to learn the distance between best-performing pipelines. The resulting (meta-)AutoML framework improves on the performance of existing AutoML frameworks. Our zero-shot AutoML system using dataset metadata embeddings provides good solutions instantaneously, running in under one second of computation. Performance is competitive with AutoML systems OBOE, AutoSklearn, AlphaD3M, and TPOT when each framework is allocated a minute of computation. We make our data, models, and code publicly available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MetaRank: Task-Aware Metric Selection for Model Transferability Estimation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A meta-learner ranks MTE metrics for a target dataset from text descriptions, improving average rank over fixed metrics on 11 datasets.

Pith tools