Pith. sign in

REVIEW 3 cited by

Evolutionary Large Language Model for Automated Feature Transformation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.16203 v2 pith:E2IMWXDK submitted 2024-05-25 cs.LG

classification cs.LG
keywords featureevolutionarytransformationdatabasespaceautomateddownstreamfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Feature transformation aims to reconstruct the feature space of raw features to enhance the performance of downstream models. However, the exponential growth in the combinations of features and operations poses a challenge, making it difficult for existing methods to efficiently explore a wide space. Additionally, their optimization is solely driven by the accuracy of downstream models in specific domains, neglecting the acquisition of general feature knowledge. To fill this research gap, we propose an evolutionary LLM framework for automated feature transformation. This framework consists of two parts: 1) constructing a multi-population database through an RL data collector while utilizing evolutionary algorithm strategies for database maintenance, and 2) utilizing the ability of Large Language Model (LLM) in sequence understanding, we employ few-shot prompts to guide LLM in generating superior samples based on feature transformation sequence distinction. Leveraging the multi-population database initially provides a wide search scope to discover excellent populations. Through culling and evolution, the high-quality populations are afforded greater opportunities, thereby furthering the pursuit of optimal individuals. Through the integration of LLMs with evolutionary algorithms, we achieve efficient exploration within a vast space, while harnessing feature knowledge to propel optimization, thus realizing a more adaptable search paradigm. Finally, we empirically demonstrate the effectiveness and generality of our proposed method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DELTA: Variational Disentangled Learning for Privacy-Preserving Data Reprogramming

    cs.LG 2025-08 reject novelty 6.0 of 10

    DELTA uses reinforcement learning to find useful feature transformations, then a disentangled variational autoencoder to generate transformed features that keep task utility while reducing sensitive-attribute predicti...

  2. ELATE: Evolutionary Language model for Automated Time-series Engineering

    cs.LG 2025-08 conditional novelty 5.0 of 10

    An LLM-guided evolutionary feature engineering method for time-series forecasting reduces RMSE by 8.4% on average across seven datasets.

  3. A Survey on Data-Centric AI: Tabular Learning from Reinforcement Learning and Generative AI Perspective

    cs.LG 2025-02 conditional novelty 1.0 of 10

    A review that organizes RL-based and generative methods for tabular feature selection and generation into a taxonomy, compares their strengths and limitations, and outlines research challenges.

Pith tools