Pith. sign in

REVIEW 12 cited by

TIME: TabPFN-Integrated Multimodal Engine for Robust Tabular-Image Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.00813 v1 pith:EALDPAJD submitted 2025-06-01 cs.CV cs.LG

TIME: TabPFN-Integrated Multimodal Engine for Robust Tabular-Image Learning

classification cs.CV cs.LG
keywords tabularmultimodaldatatimelearningmedicaldatasetsengine
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Tabular-image multimodal learning, which integrates structured tabular data with imaging data, holds great promise for a variety of tasks, especially in medical applications. Yet, two key challenges remain: (1) the lack of a standardized, pretrained representation for tabular data, as is commonly available in vision and language domains; and (2) the difficulty of handling missing values in the tabular modality, which are common in real-world medical datasets. To address these issues, we propose the TabPFN-Integrated Multimodal Engine (TIME), a novel multimodal framework that builds on the recently introduced tabular foundation model, TabPFN. TIME leverages TabPFN as a frozen tabular encoder to generate robust, strong embeddings that are naturally resilient to missing data, and combines them with image features from pretrained vision backbones. We explore a range of fusion strategies and tabular encoders, and evaluate our approach on both natural and medical datasets. Extensive experiments demonstrate that TIME consistently outperforms competitive baselines across both complete and incomplete tabular inputs, underscoring its practical value in real-world multimodal learning scenarios.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond IID: How General Are Tabular Foundation Models, Really?

    cs.LG 2026-06 unverdicted novelty 7.0

    Tabular foundation models excel on tiny- to medium-sized IID data but are outperformed by traditional tree-based and deep learning models on non-IID, large, and high-dimensional datasets, based on evaluations across 1...

  2. MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

    cs.LG 2026-05 unverdicted novelty 7.0

    MulTaBench is a new collection of 40 image-tabular and text-tabular datasets designed to test target-aware representation tuning in multimodal tabular models.

  3. Disentangled Fine-Grained Prototype Learning for Incomplete Image-Tabular Classification

    cs.CV 2026-06 unverdicted novelty 6.0

    DFPL introduces prototype-based disentanglement and alignment modules to preserve fine-grained consistency across heterogeneous modalities for better performance under missing data conditions.

  4. TabPFN-3: Technical Report

    cs.LG 2026-05 unverdicted novelty 6.0

    TabPFN-3 scales tabular foundation models to 1M rows with synthetic pretraining, test-time compute, and benchmark-leading performance on tabular, relational, and tabular-text tasks while being up to 20x faster than Ta...

  5. TabPFN-3: Technical Report

    cs.LG 2026-05 unverdicted novelty 6.0

    TabPFN-3 delivers state-of-the-art tabular prediction performance on benchmarks up to 1M rows, is up to 20x faster than prior versions, and introduces test-time scaling that beats non-TabPFN models by hundreds of Elo points.

  6. MultiModalPFN: Extending Prior-Data Fitted Networks for Multimodal Tabular Learning

    cs.LG 2026-02 unverdicted novelty 6.0

    MultiModalPFN extends TabPFN with modality projectors, a multi-head gated MLP, and cross-attention pooler to unify tabular and non-tabular inputs, outperforming prior methods on medical and general multimodal datasets.

  7. TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models

    cs.LG 2025-11 unverdicted novelty 6.0

    TabPFN-2.5 scales tabular foundation models to 20x larger datasets, outperforms tuned tree models on TabArena, achieves near-perfect win rates against default XGBoost, and adds a distillation engine for fast productio...

  8. Parameter-Efficient Adapter Tuning for Tabular-Image Multimodal Learning

    cs.CV 2026-06 unverdicted novelty 5.0

    TI-Adapter applies embedding-level and bottleneck adapters to achieve competitive or better performance than full fine-tuning on 20 tabular-image datasets while training far fewer parameters.

  9. Modular Multimodal Classification Without Fine-Tuning: A Simple Compositional Approach

    cs.LG 2026-05 unverdicted novelty 5.0

    CoMET achieves strong multimodal classification performance by composing frozen modality encoders, PCA compression, and tabular foundation models without any training, reaching state-of-the-art on diverse benchmarks i...

  10. Brain Vascular Age Prediction Using Cerebral Blood Flow Velocity and Machine Learning Algorithms

    cs.AI 2026-05 conditional novelty 5.0

    TCD-derived MOCAIP and HRV features predict chronological age in healthy adults (MAE ~3.7 y), and diseased cohorts show larger errors that the authors interpret as accelerated cerebrovascular aging.

  11. Adaptation and Fine-tuning with TabPFN for Travelling Salesman Problem

    cs.LG 2025-11 conditional novelty 5.0

    TabPFN-v2, fine-tuned on one 500-node TSP sample in about two minutes, constructs TSP tours reaching 2-5% of Concorde optimality after 2-opt post-processing, across instance sizes 50 to 1000.

  12. Brain Vascular Age Prediction Using Cerebral Blood Flow Velocity and Machine Learning Algorithms

    cs.AI 2026-05 unverdicted novelty 4.0

    Machine learning models trained on TCD-derived features from healthy subjects predict brain vascular age and indicate accelerated cerebrovascular aging in subjects with stroke, Alzheimer's, and other conditions.