Pith. sign in

REVIEW 4 cited by

Towards Better Serialization of Tabular Data for Few-shot Classification with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.12464 v2 pith:5YOAFYQP submitted 2023-12-18 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords serializationdatamethodmodelsclassificationefficiencyincludinglanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a study on the integration of Large Language Models (LLMs) in tabular data classification, emphasizing an efficient framework. Building upon existing work done in TabLLM (arXiv:2210.10723), we introduce three novel serialization techniques, including the standout LaTeX serialization method. This method significantly boosts the performance of LLMs in processing domain-specific datasets, Our method stands out for its memory efficiency and ability to fully utilize complex data structures. Through extensive experimentation, including various serialization approaches like feature combination and importance, we demonstrate our work's superiority in accuracy and efficiency over traditional models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Serialization format and in-context examples change both accuracy and gender fairness of LLM loan approvals, with finance-tuned models often showing larger disparities.

  2. Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A 22-dataset benchmark compares existing multimodal AutoML tricks, and an automatic ensemble of those tricks achieves the most robust performance.

  3. Knowledge prompt chaining for semantic modeling

    cs.CL 2025-01 reject novelty 4.0 of 10

    A two-stage prompt-chaining framework with JSON serialization and graph pruning improves LLM-based semantic modeling of structured data over prior systems on three benchmarks.

  4. Improving LLM Group Fairness on Tabular Data via In-Context Learning

    cs.LG 2024-12

Pith tools