LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers

Nikhil Abhyankar, Parshin Shojaee, Chandan K Reddy · 2025 · cs.LG · arXiv 2503.14434

9 Pith papers cite this work. Polarity classification is still indexing.

9 Pith papers citing it

open full Pith review browse 9 citing papers arXiv PDF

abstract

Automated feature engineering plays a critical role in improving predictive model performance for tabular learning tasks. Traditional automated feature engineering methods are limited by their reliance on pre-defined transformations within fixed, manually designed search spaces, often neglecting domain knowledge. Recent advances using Large Language Models (LLMs) have enabled the integration of domain knowledge into the feature engineering process. However, existing LLM-based approaches use direct prompting or rely solely on validation scores for feature selection, failing to leverage insights from prior feature discovery experiments or establish meaningful reasoning between feature generation and data-driven performance. To address these challenges, we propose LLM-FE, a novel framework that combines evolutionary search with the domain knowledge and reasoning capabilities of LLMs to automatically discover effective features for tabular learning tasks. LLM-FE formulates feature engineering as a program search problem, where LLMs propose new feature transformation programs iteratively, and data-driven feedback guides the search process. Our results demonstrate that LLM-FE consistently outperforms state-of-the-art baselines, significantly enhancing the performance of tabular prediction models across diverse classification and regression benchmarks. The code is available at: https://github.com/nikhilsab/LLMFE

citation-role summary

background 2

citation-polarity summary

background 2

representative citing papers

LATTEArena: An Evaluation Framework for LLM-powered Tabular Feature Engineering (Extended Version)

cs.AI · 2026-06-08 · accept · novelty 7.0

LATTEArena introduces a 6-dimensional taxonomy and modular evaluation framework for LLM-powered tabular feature engineering, benchmarking 24 configurations to produce 17 findings on cost-effectiveness trade-offs with public artifacts.

TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks

cs.LG · 2026-06-01 · unverdicted · novelty 7.0

TabPrep is a new feature engineering pipeline that targets three data patterns and improves performance of tree-based, neural, linear, and foundation models on tabular benchmarks, often more than model architecture changes.

Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data

cs.AI · 2026-04-22 · unverdicted · novelty 7.0

MALMAS is a memory-augmented multi-agent LLM system that generates diverse, high-quality features for tabular data via agent decomposition, routing, and iterative memory-guided refinement.

CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association

cs.LG · 2026-06-26 · unverdicted · novelty 6.0

CPAgents is a three-agent framework that generates interpretable composite phenotypes from cardiac imaging features and shows improved disease discrimination over baselines in a population cohort.

Bridging Expert Knowledge and Automated Feature Engineering via Self-Evolution

cs.AI · 2026-06-07 · unverdicted · novelty 6.0

FEST uses self-evolving trees to produce expert-aligned, auditable features from unstructured data and outperforms baselines on brand, authenticity, and stress tasks while releasing the BrandGuide dataset.

BoostLLM: Boosting-inspired LLM Fine-tuning for Few-shot Tabular Classification

cs.LG · 2026-05-07 · unverdicted · novelty 6.0 · 2 refs

BoostLLM trains sequential PEFT adapters in a boosting framework with tree path inputs to improve LLM performance on few-shot tabular classification, matching or exceeding XGBoost.

FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data

cs.AI · 2025-10-29 · unverdicted · novelty 6.0

FELA deploys specialized LLM agents in an evolutionary framework to generate, validate, and refine explainable features from heterogeneous industrial event logs, improving downstream model performance.

RelAgent: LLM Agents as Data Scientists for Relational Learning

cs.LG · 2026-05-08 · unverdicted · novelty 5.0

RelAgent uses an LLM agent to autonomously generate SQL feature programs paired with classical models for interpretable relational learning predictions that execute efficiently on standard databases.

TriAlignGR: Triangular Multitask Alignment with Multimodal Deep Interest Mining for Generative Recommendation

cs.IR · 2026-05-05 · unverdicted · novelty 5.0 · 3 refs

TriAlignGR introduces cross-modal alignment, deep interest mining via CoT, and triangular multitask training to fix semantic degradation and opacity in SID-based generative recommendation.

citing papers explorer

Showing 8 of 8 citing papers after filters.

LATTEArena: An Evaluation Framework for LLM-powered Tabular Feature Engineering (Extended Version) cs.AI · 2026-06-08 · accept · none · ref 1 · internal anchor
LATTEArena introduces a 6-dimensional taxonomy and modular evaluation framework for LLM-powered tabular feature engineering, benchmarking 24 configurations to produce 17 findings on cost-effectiveness trade-offs with public artifacts.
TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks cs.LG · 2026-06-01 · unverdicted · none · ref 63 · internal anchor
TabPrep is a new feature engineering pipeline that targets three data patterns and improves performance of tree-based, neural, linear, and foundation models on tabular benchmarks, often more than model architecture changes.
Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data cs.AI · 2026-04-22 · unverdicted · none · ref 57 · internal anchor
MALMAS is a memory-augmented multi-agent LLM system that generates diverse, high-quality features for tabular data via agent decomposition, routing, and iterative memory-guided refinement.
CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association cs.LG · 2026-06-26 · unverdicted · none · ref 14 · internal anchor
CPAgents is a three-agent framework that generates interpretable composite phenotypes from cardiac imaging features and shows improved disease discrimination over baselines in a population cohort.
Bridging Expert Knowledge and Automated Feature Engineering via Self-Evolution cs.AI · 2026-06-07 · unverdicted · none · ref 1 · internal anchor
FEST uses self-evolving trees to produce expert-aligned, auditable features from unstructured data and outperforms baselines on brand, authenticity, and stress tasks while releasing the BrandGuide dataset.
BoostLLM: Boosting-inspired LLM Fine-tuning for Few-shot Tabular Classification cs.LG · 2026-05-07 · unverdicted · none · ref 1 · 2 links · internal anchor
BoostLLM trains sequential PEFT adapters in a boosting framework with tree path inputs to improve LLM performance on few-shot tabular classification, matching or exceeding XGBoost.
RelAgent: LLM Agents as Data Scientists for Relational Learning cs.LG · 2026-05-08 · unverdicted · none · ref 44 · internal anchor
RelAgent uses an LLM agent to autonomously generate SQL feature programs paired with classical models for interpretable relational learning predictions that execute efficiently on standard databases.
TriAlignGR: Triangular Multitask Alignment with Multimodal Deep Interest Mining for Generative Recommendation cs.IR · 2026-05-05 · unverdicted · none · ref 30 · 3 links · internal anchor
TriAlignGR introduces cross-modal alignment, deep interest mining via CoT, and triangular multitask training to fix semantic degradation and opacity in SID-based generative recommendation.

LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer