Pith. sign in

REVIEW 4 cited by

Tabular Data Augmentation for Machine Learning: Progress and Prospects of Embracing Generative AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.21523 v1 pith:YHHKER3Z submitted 2024-07-31 cs.LG cs.AIcs.DB

classification cs.LGcs.AIcs.DB
keywords datatableaugmentationgenerativemethodstabularcurrentfuture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning (ML) on tabular data is ubiquitous, yet obtaining abundant high-quality tabular data for model training remains a significant obstacle. Numerous works have focused on tabular data augmentation (TDA) to enhance the original table with additional data, thereby improving downstream ML tasks. Recently, there has been a growing interest in leveraging the capabilities of generative AI for TDA. Therefore, we believe it is time to provide a comprehensive review of the progress and future prospects of TDA, with a particular emphasis on the trending generative AI. Specifically, we present an architectural view of the TDA pipeline, comprising three main procedures: pre-augmentation, augmentation, and post-augmentation. Pre-augmentation encompasses preparation tasks that facilitate subsequent TDA, including error handling, table annotation, table simplification, table representation, table indexing, table navigation, schema matching, and entity matching. Augmentation systematically analyzes current TDA methods, categorized into retrieval-based methods, which retrieve external data, and generation-based methods, which generate synthetic data. We further subdivide these methods based on the granularity of the augmentation process at the row, column, cell, and table levels. Post-augmentation focuses on the datasets, evaluation and optimization aspects of TDA. We also summarize current trends and future directions for TDA, highlighting promising opportunities in the era of generative AI. In addition, the accompanying papers and related resources are continuously updated and maintained in the GitHub repository at https://github.com/SuDIS-ZJU/awesome-tabular-data-augmentation to reflect ongoing advancements in the field.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ensembling Membership Inference Attacks Against Tabular Generative Models

    cs.CR 2025-09 conditional novelty 6.0 of 10

    No single membership inference attack dominates across tabular generative models, and unsupervised ensembles of attacks achieve better average rankings.

  2. TableCopilot: A Table Assistant Empowered by Natural Language Conditional Table Discovery

    cs.DB 2025-07 conditional novelty 5.0 of 10

    A new task and system for finding unionable or joinable tables that satisfy a user's natural language condition, evaluated on a new benchmark.

  3. Template-Based Schema Matching of Multi-Layout Tenancy Schedules:A Comparative Study of a Template-Based Hybrid Matcher and the ALITE Full Disjunction Model

    cs.DB 2025-07 conditional novelty 4.0 of 10

    A template-based hybrid schema matcher aligns multi-layout tenancy schedules to a fixed target schema and reports an F1 of 0.881, but the score is obtained by grid search on the evaluation ground truth.

  4. Fabrication of nano-diamonds with a single NV center: Towards matter-wave interferometry with massive objects

    quant-ph 2025-08 unverdicted novelty 3.0 of 10

    The authors describe design considerations and fabrication of 40x65x80 nm nanodiamond pillars with single NV centers as a step toward matter-wave interferometry of massive objects.

Pith tools