REVIEW 8 cited by
TabLLM: Few-shot Classification of Tabular Data with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
TabLLM: Few-shot Classification of Tabular Data with Large Language Models
read the original abstract
We study the application of large language models to zero-shot and few-shot classification of tabular data. We prompt the large language model with a serialization of the tabular data to a natural-language string, together with a short description of the classification problem. In the few-shot setting, we fine-tune the large language model using some labeled examples. We evaluate several serialization methods including templates, table-to-text models, and large language models. Despite its simplicity, we find that this technique outperforms prior deep-learning-based tabular classification methods on several benchmark datasets. In most cases, even zero-shot classification obtains non-trivial performance, illustrating the method's ability to exploit prior knowledge encoded in large language models. Unlike many deep learning methods for tabular datasets, this approach is also competitive with strong traditional baselines like gradient-boosted trees, especially in the very-few-shot setting.
Forward citations
Cited by 8 Pith papers
-
Flow Map Denoisers: Traversing the Distortion-Perception Plane for Inverse Problems
Flow map denoisers use a lookahead parameter t to span the distortion-perception frontier, proven optimal for Gaussian targets and effective for natural images and inverse problems.
-
Reciprocal Co-Training (RCT): Coupling Gradient-Based and Non-Differentiable Models via Reinforcement Learning
RCT couples an LLM and Random Forest via RL feedback so each augments the other's features and rewards, producing consistent gains on three medical datasets.
-
Controllability-Aware Adversarial Examples Against LLM-Based Network Traffic Classifiers
Under DC-only transfer attacks, LLM IDS vulnerability is substantial but dataset- and comparator-dependent, with gradient/score attacks transferring better than greedy ones.
-
Collaborative Large and Small Language Models for Accurate and Scalable Data Repair
LasRepair++ pairs an LLM instructor with an SLM corrector, refines context via EM, and down-weights uncertain repairs using column-calibrated confidence, reporting 18.1% average F1 gain over baselines on data repair tasks.
-
Learning Normalized Energy Models for Linear Inverse Problems
Energy-based model with covariance regularization computes normalized posteriors for linear inverse problems without retraining, enabling adaptive sampling and blind estimation on image datasets.
-
ReSS: Learning Reasoning Models for Tabular Data Prediction via Symbolic Scaffold
ReSS uses decision-tree scaffolds to fine-tune LLMs for faithful tabular reasoning, reporting up to 10% gains over baselines on medical and financial data.
-
ReSS: Learning Reasoning Models for Tabular Data Prediction via Symbolic Scaffold
ReSS extracts decision paths from trees as scaffolds to guide LLM reasoning generation, fine-tunes the LLM on the resulting dataset with scaffold-invariant augmentation, and reports up to 10% gains on medical and fina...
-
Reciprocal Co-Training (RCT): Coupling Gradient-Based and Non-Differentiable Models via Reinforcement Learning
Reciprocal co-training links an LM and a random forest via RL so LM embeddings enrich the forest and forest probabilities reward LM updates, yielding gains on three medical datasets.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.