Pith. sign in

REVIEW 6 cited by

Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.16243 v1 pith:YPLBYTYS submitted 2024-12-19 cs.LG

Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data

classification cs.LG
keywords datamultimodalautomltabulartextbenchmarkcombinationsdatasets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper studies the best practices for automatic machine learning (AutoML). While previous AutoML efforts have predominantly focused on unimodal data, the multimodal aspect remains under-explored. Our study delves into classification and regression problems involving flexible combinations of image, text, and tabular data. We curate a benchmark comprising 22 multimodal datasets from diverse real-world applications, encompassing all 4 combinations of the 3 modalities. Across this benchmark, we scrutinize design choices related to multimodal fusion strategies, multimodal data augmentation, converting tabular data into text, cross-modal alignment, and handling missing modalities. Through extensive experimentation and analysis, we distill a collection of effective strategies and consolidate them into a unified pipeline, achieving robust performance on diverse datasets.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond IID: How General Are Tabular Foundation Models, Really?

    cs.LG 2026-06 unverdicted novelty 7.0

    Tabular foundation models excel on tiny- to medium-sized IID data but are outperformed by traditional tree-based and deep learning models on non-IID, large, and high-dimensional datasets, based on evaluations across 1...

  2. MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

    cs.LG 2026-05 unverdicted novelty 7.0

    MulTaBench is a new collection of 40 image-tabular and text-tabular datasets designed to test target-aware representation tuning in multimodal tabular models.

  3. VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning

    cs.CV 2026-05 unverdicted novelty 7.0

    VT-Bench is the first unified benchmark aggregating 14 visual-tabular datasets with over 756K samples and evaluating 23 models to expose challenges in this multi-modal area.

  4. VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning

    cs.CV 2026-05 unverdicted novelty 7.0

    VT-Bench aggregates 14 datasets totaling over 756K samples across 9 domains and evaluates 23 models to establish a unified testbed for visual-tabular multi-modal discriminative and generative tasks.

  5. VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning

    cs.CV 2026-05 unverdicted novelty 7.0

    VT-Bench aggregates 14 datasets from 9 domains and evaluates 23 models to standardize visual-tabular discriminative and generative tasks.

  6. MultiModalPFN: Extending Prior-Data Fitted Networks for Multimodal Tabular Learning

    cs.LG 2026-02 unverdicted novelty 6.0

    MultiModalPFN extends TabPFN with modality projectors, a multi-head gated MLP, and cross-attention pooler to unify tabular and non-tabular inputs, outperforming prior methods on medical and general multimodal datasets.