Pith. sign in

REVIEW 4 cited by

Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.12112 v3 pith:FRQ3SXZ6 submitted 2023-12-19 cs.LG cs.AI

classification cs.LGcs.AI
keywords datalow-dataaugmentationcurationdatasetsllmsaugmentedcllm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine Learning (ML) in low-data settings remains an underappreciated yet crucial problem. Hence, data augmentation methods to increase the sample size of datasets needed for ML are key to unlocking the transformative potential of ML in data-deprived regions and domains. Unfortunately, the limited training set constrains traditional tabular synthetic data generators in their ability to generate a large and diverse augmented dataset needed for ML tasks. To address this challenge, we introduce CLLM, which leverages the prior knowledge of Large Language Models (LLMs) for data augmentation in the low-data regime. However, not all the data generated by LLMs will improve downstream utility, as for any generative model. Consequently, we introduce a principled curation mechanism, leveraging learning dynamics, coupled with confidence and uncertainty metrics, to obtain a high-quality dataset. Empirically, on multiple real-world datasets, we demonstrate the superior performance of CLLM in the low-data regime compared to conventional generators. Additionally, we provide insights into the LLM generation and curation mechanism, shedding light on the features that enable them to output high-quality augmented datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes

    cs.LG 2025-09 conditional novelty 6.0 of 10

    ReFine combines rule-guided prompting and dual-granularity filtering to improve LLM-based tabular data generation when only 30 to 90 labeled rows exist, achieving top average rank over baselines.

  2. Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Fine-tuning on data aligned with an LLM's prior knowledge induces overconfidence, and CogCalib mitigates this by gating a calibration loss to known data.

  3. A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.

  4. Few-shot LLM Synthetic Data with Distribution Matching

    cs.CL 2025-02 conditional novelty 5.0 of 10

    SynAlign generates LLM synthetic text from diversity-guided demonstrations, then reweights it by MMD-based distribution matching, improving downstream classification accuracy.

Pith tools