Pith. sign in

REVIEW 4 major objections 4 minor 9 cited by

Towards Data-Centric AI: A Comprehensive Survey of Traditional, Reinforcement, and Generative Approaches for Tabular Data Transformation

T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper argues that tabular data-centric AI reduces to two data transformations — feature selection and feature generation — and that one taxonomy can organize traditional, reinforcement-learning, and generative methods for both.

desk verdict A useful and current survey frame for tabular feature selection and generation, but the broken citation-reference machinery undercuts its reliability as a reference work and needs a serious pass before it can serve its purpose. read the letter →

arxiv 2501.10555 v1 pith:2JIFUQNR submitted 2025-01-17 cs.LG cs.AI

classification cs.LGcs.AI
keywords data-centricAItabulardatafeatureselectiongenerationreinforcementlearninggenerativeautomatedengineeringsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey claims that the practical route to better AI on tabular data runs through the data itself, not just the model. It argues that tabular data transformation has two core tasks — feature selection, which keeps useful columns and drops redundant ones, and feature generation, which creates new columns from existing ones — and that all current methods fit into a single taxonomy. The survey places classical filter, wrapper, and embedded selection alongside multi-view variants, then shows how reinforcement learning and generative AI recast both tasks as automated search problems. It closes with a comparison table and practical guidelines. If the synthesis is right, a practitioner can look up any tabular-data pain point and find the family of solutions designed for it.

What carries the argument

The load-bearing object is the survey's taxonomy, summarized in Figure 2, which partitions tabular data-centric AI into feature selection and feature generation, then under each into traditional methods (filter, wrapper, embedded, hybrid, and multi-view) and advanced methods (reinforcement learning and generative AI). Two formal mechanisms carry the advanced half: the Markov decision process formulation, whose state is the current feature set (plus transformation history for generation), whose actions are selecting or deselecting a feature or applying a mathematical operation, and whose reward balances downstream performance against set size or complexity; and the encoder-decoder-evaluator formulation, in which transformation sequences are encoded into a continuous embedding, decoded back to sequences, and evaluated by predicted performance, enabling a gradient-based search for better transformations.

What would settle it

A reader can test the survey's utility claim directly by resolving every citation marker in Section 3.1.1 and Figure 2 against the reference list; broken markers — such as the maximum-relevance minimum-redundancy method cited only as '[? ]' and figure entries whose numbers do not correspond to the listed references — would falsify the claim that this is a dependable map of the field.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that feature selection and feature generation are the two essential operations of tabular data-centric AI, and that prior surveys covered only part of this picture. It claims to bridge those gaps by integrating traditional methods, reinforcement learning, and generative AI for both operations under one taxonomy. Feature selection is formalized as choosing a subset $S \subseteq \{1,\dots,d\}$ of the $d$ original features, and feature generation as producing a transformed dataset through mathematical operations. Advanced methods are then formulated as Markov decision processes for reinforcement learning — the state is the current feature set, the action selects or deselects a feature or applies a transformation, and the reward balances downstream performance against penalties — and as an encoder-decoder-evaluator architecture for generative AI, where transformation sequences are mapped to a continuous embedding space, reconstructed, and scored so that better transformations can be found by gradient-based search. The paper also identifies open challenges, including scalability, interpretability, privacy-preserving feature engineering, and LLM-based generation.

Load-bearing premise

The survey's usefulness rests on its descriptions and citations of prior work being accurate and correctly attributed, because a reader who cannot trust the mapping from method to citation cannot act on the taxonomy.

Editorial extensions

If this is right

  • A practitioner facing a high-dimensional table can use the taxonomy to place their problem: cheap statistical screening first, wrapper or embedded selection when feature interactions matter, reinforcement learning for sequential or dynamic feature decisions, and generative models for transferring feature knowledge across tasks.
  • Feature selection and feature generation, usually treated as separate pipelines, can be viewed as one data-space refinement problem with a shared optimization language.
  • Reinforcement-learning-generated feature transformation records can be compressed into an embedding space, meaning experience from many datasets can be reused instead of rediscovered.
  • LLM-based feature generation can operate without a separate learned predictor, using in-context learning and retrieval-augmented generation from external knowledge to create explainable features.
  • The comparison of traditional and advanced methods gives concrete guidance: prefer traditional methods for small, static, interpretability-critical datasets and advanced methods for high-dimensional, dynamic, or multimodal ones.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the unified formulations are taken seriously, a natural next step the paper does not take is a shared benchmark that scores feature-selection and feature-generation methods on identical tabular datasets, which would make the taxonomy testable rather than descriptive.
  • The encoder-decoder-evaluator pattern suggests a testable extension: pretraining the embedding space on synthetic transformation records from many datasets should improve downstream feature quality on a held-out tabular task, and failure would indicate the embedding is not capturing transferable knowledge.
  • The survey's emphasis on interpretability points to an under-explored hybrid: generative methods that produce candidate features but require a domain expert's validation before entering the model, potentially combining scalability with trust.
  • The LLM-based directions imply that text-informed feature generation will matter most in domains where tabular rows come with side text, such as medical records or product descriptions; a controlled comparison against text-free generation on such data would quantify the gain.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper surveys tabular data-centric AI, focusing on feature selection and feature generation. It organizes feature selection into single-view (filter, wrapper, embedded, hybrid) and multi-view (supervised, semi-supervised, unsupervised) methods, and feature generation into human-driven and automated techniques. It then reviews advanced methods based on reinforcement learning and generative AI, provides an MDP formulation for RL-based feature engineering, an encoder-decoder-evaluator formulation for generative approaches, a comparative analysis with practical guidelines, and a discussion of open challenges such as AutoML, LLMs, and federated learning.

Significance. The topic is timely and the paper's architecture is a real strength: it unifies feature selection and generation under a single taxonomy, distinguishes traditional, RL, and generative approaches, and gives explicit formalizations in Section 5. The comparative discussion in Section 6 and the practical guidelines in Section 6.2 are useful and broadly accurate at a high level. If the citation apparatus were reliable, this would be a valuable reference for researchers entering tabular data-centric AI. The manuscript makes no novel algorithmic claims, so there are no derivations or fitted parameters to verify; its evidentiary burden is entirely attributional, which is exactly where the current version fails. The breadth of coverage is genuine, but the broken citations mean the paper cannot currently serve as the trustworthy comprehensive reference it sets out to be.

major comments (4)
  1. [§3.1.1 and Figure 2] The citation apparatus for the filter-based methods is internally inconsistent and contains a placeholder. The text cites the chi-square test as [38, 99], ANOVA as [77, 132, 155], Pearson correlation as [193], and mutual information as [5, 193], while Figure 2 attributes the same methods to [45, 46], [47, 48, 49], [50], and [50, 51], respectively. In addition, the sentence "Methods like mRMR (Maximum Relevance Minimum Redundancy) [? ]" contains a literal unresolved placeholder. Because a survey's primary function is to let readers locate and verify the methods it describes, these inconsistencies make the central claim of a reliable comprehensive survey unverifiable as submitted. The authors should audit every in-text citation against the final reference list.
  2. [References [32]/[33], [54]/[55], [91]/[92]] The reference list contains at least three duplicate entries: [32] and [33] are the same Fan et al. paper, [54] and [55] are the same Huang et al. paper, and [91] and [92] are the same Liu et al. paper. These duplicates are not harmless: Section 5.1.1 cites [92] for the Monte Carlo early-stopping method while [91] is the identical paper, and Section 3.1.4 cites [55] for the genetic-algorithm-plus-mutual-information hybrid while [54] is the same work. The list must be deduplicated and all downstream numbering recomputed.
  3. [§5.1.1 and Figure 2] The advanced RL feature-selection entries disagree between text and taxonomy figure. The text attributes the multi-agent framework to Liu et al. [89], the GCN-based collaboration to [90], the group-wise method to Fan et al. [34], IRFS to [35], and decision-tree feedback to [33], whereas Figure 2 lists `Group-Wise Method [8]`, `Advanced Statistical Summaries & GCNs State Representation [39]`, `Enhanced Reward Scheme [173]`, `Interactive Reinforced Feature Selection [7]`, `Pre-Filtering And Iterative RL [177]`, and `Decision Tree Based [178]`. Some of these entries point to papers unrelated to the cited method; for example, [8] is a pattern-recognition textbook and [39] is a forest-optimization paper. The same problem affects the LLM branch, where the text cites [189] and [188] but Figure 2 cites [190] and [191]. This makes it impossible to verify the survey's coverage of advanced methods.
  4. [Figure 2 vs Section 4] Figure 2 promises a `Latent Representation Learning` category under traditional feature transformation, with subcategories for autoencoders, categorical embeddings, and PCA, but Section 4, the corresponding textual treatment, contains no such subsection. PCA appears only later as a dimensionality-reduction suggestion in Section 6.2, and autoencoders appear only in the feature-selection discussion, e.g., SDAE-LSTM in Section 3.1.3. Either a substantive subsection is missing or the taxonomy overstates the coverage; the paper should align the figure with the text before claiming comprehensiveness.
minor comments (4)
  1. [Section 1] The phrase `F or feature selection` contains a stray capitalization artifact and should read `For feature selection`.
  2. [Figure 2 and Section 4] The figure labels a whole branch `Feature Transformation`, while the text consistently uses `Feature Generation`; please unify the terminology.
  3. [Section 5.1] The reward formula for feature selection is split across a line break in a way that could confuse readers; the expression `R(t) = Perf(...) - lambda * |F(t+1)|` should be set on one line with proper punctuation.
  4. [Reference list] Several entries, including [41], [56], [154], [174], [188], and [189], are arXiv preprints; please mark them clearly as preprints and include access dates.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a descriptive survey with no fitted parameters, no derivation chain, and no prediction that reduces to its own inputs; citation-reference mismatches are a verifiability issue, not circularity.

full rationale

This manuscript is a survey, not a derivation or an empirical study. It organizes existing work on tabular data-centric AI into a taxonomy, describes methods from the literature, and provides comparative guidance. There is no fitted parameter that is later renamed as a prediction, no result claimed to follow from first principles, and no equation in the paper is shown to be equivalent to an input by construction. The formal definitions in Section 2 and the RL/generative AI formulations in Section 5 are general mathematical descriptions of problem settings, not derivations of new results. The self-citations to the authors' prior works (e.g., [147], [148], [165], [175], [177], [178]) are used to describe those works as part of the surveyed landscape, not as the sole justification for a new central claim, so they are not load-bearing in a circularity sense. The significant citation-reference mismatches, duplicate references, and the literal '[? ]' placeholder for mRMR that the reviewer identified are real correctness and verifiability defects, but they are not instances of the seven circularity patterns: they do not make any claimed result equivalent to its own input. Under the hard rules, this non-finding is the appropriate outcome, and the circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

A survey introduces no free parameters or invented entities. It relies on the accuracy of its citations and the adequacy of its taxonomy as domain assumptions.

assumptions (2)
  • domain assumption The survey's characterizations of cited methods are faithful to the original publications.
    The entire utility of the survey rests on accurate representation of prior work; any mischaracterization would mislead readers.
  • domain assumption The taxonomy categories (filter/wrapper/embedded, single/multi-view, RL/generative) are exhaustive and non-overlapping.
    The claim of comprehensiveness depends on the taxonomy covering all relevant methods without omission.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Data-Centric AI: A Comprehensive Survey of Traditional, Reinforcement, and Generative Approaches for Tabular Data Transformation." pith.science (2026). https://pith.science/paper/2JIFUQNR

@misc{pith2026250110555,
  author       = {Pith},
  title        = {Pith review of: Towards Data-Centric AI: A Comprehensive Survey of Traditional, Reinforcement, and Generative Approaches for Tabular Data Transformation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2JIFUQNR}},
  note         = {Machine review of arXiv:2501.10555}
}
read the original abstract

Tabular data is one of the most widely used formats across industries, driving critical applications in areas such as finance, healthcare, and marketing. In the era of data-centric AI, improving data quality and representation has become essential for enhancing model performance, particularly in applications centered around tabular data. This survey examines the key aspects of tabular data-centric AI, emphasizing feature selection and feature generation as essential techniques for data space refinement. We provide a systematic review of feature selection methods, which identify and retain the most relevant data attributes, and feature generation approaches, which create new features to simplify the capture of complex data patterns. This survey offers a comprehensive overview of current methodologies through an analysis of recent advancements, practical applications, and the strengths and limitations of these techniques. Finally, we outline open challenges and suggest future perspectives to inspire continued innovation in this field.

Figures

Figures reproduced from arXiv: 2501.10555 by the authors.

Figure 1
Figure 1. Real-world applications often generate vast amounts of tabular data, making Data-Centric AI essential for optimizing [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. An overview of the taxonomy for existing tabular data-centric AI techniques. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Feature selection is framed as a RL problem, wherein the agent’s actions correspond to the selection of individual features. [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Feature generation is formulated as a RL problem, where the agent’s actions involve selecting individual features and applying [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]
Figure 5
Figure 5. Figure 5: RL extensively explores the feature space, generating abundant feature learning (feature selection/generation) knowledge. [PITH_FULL_IMAGE:figures/full_fig_p022_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging the Domain Gap in Equation Distillation with Reinforcement Feedback

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Reinforcement learning fine-tuning with numerical fitness rewards improves equation discovery accuracy and noise robustness of a pretrained symbolic regression transformer.

  2. Brownian Bridge Augmented Surrogate Simulation and Injection Planning for Geological CO$_2$ Storage

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A Brownian bridge augmented framework improves surrogate simulation accuracy and injection plan quality on synthetic CO2 storage datasets compared with established baselines.

  3. Sculpting Features from Noise: Reward-Guided Hierarchical Diffusion for Task-Optimal Feature Transformation

    cs.LG 2025-05 conditional novelty 6.0 of 10

    DIFFT generates task-optimal feature transformations via reward-guided latent diffusion with a semi-autoregressive decoder, outperforming ten baselines on 14 tabular datasets.

  4. Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning

    cs.AI 2025-04 conditional novelty 6.0 of 10

    HRLFS combines LLM semantic feature states with Gaussian mixture distributions and hierarchical multi-agent reinforcement learning to select feature subsets, reporting improved downstream performance and reduced agent...

  5. Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives

    cs.LG 2025-06 reject novelty 5.0 of 10

    TimesCLIP aligns image-based and text-based views of the same time series via contrastive learning to improve forecasting accuracy on several benchmarks, but the full multimodal model is not used on two of the six lon...

  6. Agentic Feature Augmentation: Unifying Selection and Generation with Teaming, Planning, and Memories

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A router-selector-generator LLM agent team with offline PPO and dual memories unifies feature selection and generation, reporting improved downstream performance on six tabular datasets.

  7. Unsupervised Feature Transformation via In-context Generation, Generator-critic LLM Agents, and Duet-play Teaming

    cs.LG 2025-04 conditional novelty 5.0 of 10

    A generator-critic pair of LLM agents creates new tabular features from feature names, data summaries, and task descriptions, and is reported to improve Random Forest accuracy more than supervised feature engineering ...

  8. Collaborative Multi-Agent Reinforcement Learning for Automated Feature Transformation with Graph-Driven Path Optimization

    cs.LG 2025-04 conditional novelty 5.0 of 10

    TCTO, a graph-based multi-agent RL method that prunes and backtracks over feature transformation paths, outperforms existing automated feature engineering methods on most of 25 tabular datasets.

  9. LLM-ML Teaming: Integrated Symbolic Decoding and Gradient Search for Valid and Stable Generative Feature Transformation

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A product-of-experts decoder that blends a fine-tuned LLM's token probabilities with a gradient-searched sequence decoder produces more valid and stable feature transformations than either alone.

Reference graph

Works this paper leans on

197 extracted references · 61 canonical work pages · cited by 9 Pith papers

  1. [33]

    Wei Fan, Kunpeng Liu, Hao Liu, Yong Ge, Hui Xiong, and Yanjie Fu. 2021. Interactive reinforcement learning for feature selection with decision tree in the loop. IEEE Transactions on Knowledge and Data Engineering 35, 2 (2021), 1624–1636

  2. [55]

    Jinjie Huang, Yunze Cai, and Xiaoming Xu. 2007. A hybrid genetic algorithm for feature selection wrapper based on mutual information. Pattern recognition letters 28, 13 (2007), 1825–1844

  3. [91]

    Kunpeng Liu, Pengfei Wang, Dongjie Wang, Wan Du, Dapeng Oliver Wu, and Yanjie Fu. 2021. Efficient Reinforced Feature Selection via Early Stopping Traverse Strategy. In 2021 IEEE International Conference on Data Mining (ICDM) . IEEE, 399–408

  4. [92]

    Kunpeng Liu, Pengfei Wang, Dongjie Wang, Wan Du, Dapeng Oliver Wu, and Yanjie Fu. 2021. Efficient reinforced feature selection via early stopping traverse strategy. In 2021 IEEE International Conference on Data Mining (ICDM) . IEEE, 399–408

  5. [193]

    Hongfang Zhou, Xiqian Wang, and Rourou Zhu. 2022. Feature selection based on mutual information with correlation coefficient. Applied intelligence 52, 5 (2022), 5457–5474

  6. [50]

    Franziska Horn, Robert Pack, and Michael Rieger. 2019. The autofeat python library for automated feature engineering and selection. arXiv preprint arXiv:1901.07329 (2019)

  7. [89]

    Kunpeng Liu, Yanjie Fu, Pengfei Wang, Le Wu, Rui Bo, and Xiaolin Li. 2019. Automating feature subspace exploration via multi-agent reinforcement learning. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 207–215

  8. [90]

    Kunpeng Liu, Yanjie Fu, Le Wu, Xiaolin Li, Charu Aggarwal, and Hui Xiong. 2021. Automated feature selection: A reinforcement learning perspective. IEEE Transactions on Knowledge and Data Engineering 35, 3 (2021), 2272–2284

  9. [34]

    Wei Fan, Kunpeng Liu, Hao Liu, Ahmad Hariri, Dejing Dou, and Yanjie Fu. 2021. Autogfs: Automated group-based feature selection via interactive reinforcement learning. In Proceedings of the 2021 SIAM International Conference on Data Mining (SDM) . SIAM, 342–350

  10. [35]

    Wei Fan, Kunpeng Liu, Hao Liu, Pengyang Wang, Yong Ge, and Yanjie Fu. 2020. Autofs: Automated feature selection via diversity-aware interactive reinforcement learning. In 2020 IEEE International Conference on Data Mining (ICDM) . IEEE, 1008–1013

  11. [8]

    Christopher M Bishop and Nasser M Nasrabadi. 2006. Pattern recognition and machine learning . Vol. 4. Springer

  12. [39]

    Manizheh Ghaemi and Mohammad-Reza Feizi-Derakhshi. 2016. Feature selection using forest optimization algorithm. Pattern Recognition 60 (2016), 121–129. Manuscript submitted to ACM Towards Data-Centric AI: A Comprehensive Survey of Traditional, Reinforcement, and Generative Approaches for Tabular Data Transformation 31

  13. [173]

    2011.ℓ2,1-norm regularized discriminative feature selection for unsupervised learning

    Yi Yang, Heng Tao Shen, Zhigang Ma, Zi Huang, and Xiaofang Zhou. 2011.ℓ2,1-norm regularized discriminative feature selection for unsupervised learning. In Proceedings of International Joint Conference on Artificial Intelligence . 1589–1594

  14. [7]

    Jacek Biesiada and Wlodzisław Duch. 2008. Feature selection for high-dimensional data—a Pearson redundancy based filter. InComputer recognition systems 2. Springer, 242–249

  15. [177]

    Aggarwal, and Yanjie Fu

    Wangyang Ying, Dongjie Wang, Xuanming Hu, Yuanchun Zhou, Charu C. Aggarwal, and Yanjie Fu. 2024. Unsupervised Generative Feature Transformation via Graph Contrastive Pre-training and Multi-objective Fine-tuning. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Barcelona, Spain) (KDD ’24). Association for Computing M...

  16. [178]

    Wangyang Ying, Dongjie Wang, Kunpeng Liu, Leilei Sun, and Yanjie Fu. 2023. Self-optimizing feature generation via categorical hashing representation and hierarchical reinforcement crossing. In 2023 IEEE International Conference on Data Mining (ICDM) . IEEE, 748–757

  17. [189]

    Xinhao Zhang, Jinghan Zhang, Banafsheh Rekabdar, Yuanchun Zhou, Pengfei Wang, and Kunpeng Liu. 2024. Dynamic and Adaptive Feature Generation with LLM. arXiv preprint arXiv:2406.03505 (2024)

  18. [188]

    Xinhao Zhang, Jinghan Zhang, Fengran Mo, Yuzhong Chen, and Kunpeng Liu. 2024. TIFG: Text-Informed Feature Generation with Large Language Models. arXiv preprint arXiv:2406.11177 (2024)

  19. [190]

    Jing Zhao, Xijiong Xie, Xin Xu, and Shiliang Sun. 2017. Multi-View Learning Overview: Recent Progress and New Challenges. Information Fusion 38 (2017), 43–54

  20. [191]

    Xiaosa Zhao, Kunpeng Liu, Wei Fan, Lu Jiang, Xiaowei Zhao, Minghao Yin, and Yanjie Fu. 2020. Simplifying reinforced feature selection via restructured choice strategy of single agent. In 2020 IEEE International conference on data mining (ICDM) . IEEE, 871–880

Show all 197 references
  1. [1]

    Mohammed Ghaith Altarabichi, Sławomir Nowaczyk, Sepideh Pashami, and Peyman Sheikholharam Mashhadi. 2023. Fast Genetic Algorithm for feature selection—A qualitative approximation approach. Expert systems with applications 211 (2023), 118528

  2. [2]

    Sven Apel, Sergiy Kolesnikov, Norbert Siegmund, Christian Kästner, and Brady Garvin. 2013. Exploring feature interactions in the wild: the new feature-interaction challenge. In Proceedings of the 5th international workshop on feature-oriented software development . 1–8

  3. [3]

    Xiangpin Bai, Lei Zhu, Cheng Liang, Jingjing Li, Xiushan Nie, and Xiaojun Chang. 2020. Multi-view feature selection via nonnegative structured graph learning. Neurocomputing 387 (2020), 110–122

  4. [4]

    William H Baltosser. 1996. Biostatistical analysis. Ecology 77, 7 (1996), 2266–2268

  5. [5]

    Roberto Battiti. 1994. Using mutual information for selecting features in supervised neural net learning. IEEE Transactions on neural networks 5, 4 (1994), 537–550

  6. [6]

    James C Bezdek, Robert Ehrlich, and William Full. 1984. FCM: The fuzzy c-means clustering algorithm. Computers & geosciences 10, 2-3 (1984), 191–203

  7. [9]

    Muffy Calder, Mario Kolberg, Evan H Magill, and Stephan Reiff-Marganiec. 2003. Feature interaction: a critical review and considered forecast. Computer Networks 41, 1 (2003), 115–141

  8. [10]

    E Jane Cameron, Nancy Griffeth, Y-J Lin, Margaret E Nilson, William K Schnure, and Hugo Velthuijsen. 1993. A feature-interaction benchmark for IN and beyond. IEEE Communications Magazine 31, 3 (1993), 64–69. Manuscript submitted to ACM 30 Wang et al

  9. [11]

    Zhiwen Cao, Xijiong Xie, and Yuqi Li. 2024. Multi-view unsupervised feature selection with consensus partition and diverse graph. Information Sciences 661 (2024), 120178

  10. [12]

    Zhiwen Cao, Xijiong Xie, Feixiang Sun, and Jiabei Qian. 2023. Consensus cluster structure guided multi-view unsupervised feature selection. Knowledge-Based Systems 271 (2023), 110578

  11. [13]

    Girish Chandrashekar and Ferat Sahin. 2014. A survey on feature selection methods. Computers & electrical engineering 40, 1 (2014), 16–28

  12. [14]

    Hong Chen, Feiping Nie, Rong Wang, and Xuelong Li. 2022. Unsupervised feature selection with flexible optimal graph. IEEE Transactions on Neural Networks and Learning Systems 35, 2 (2022), 2014–2027

  13. [15]

    Xiangning Chen, Qingwei Lin, Chuan Luo, Xudong Li, Hongyu Zhang, Yong Xu, Yingnong Dang, Kaixin Sui, Xu Zhang, Bo Qiao, et al. 2019. Neural feature search: A neural architecture for automated feature engineering. In2019 IEEE International Conference on Data Mining (ICDM) . IEEE, 71–80

  14. [16]

    Yuehui Chen, Ajith Abraham, and Bo Yang. 2006. Feature selection and classification using flexible neural tree. Neurocomputing 70, 1-3 (2006), 305–313

  15. [17]

    Yumin Chen, Duoqian Miao, Ruizhi Wang, and Keshou Wu. 2011. A rough set approach to feature selection based on power set tree.Knowledge-Based Systems 24, 2 (2011), 275–281

  16. [18]

    Yi-Wei Chen, Qingquan Song, and Xia Hu. 2021. Techniques for automated machine learning. ACM SIGKDD Explorations Newsletter 22, 2 (2021), 35–50

  17. [19]

    Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al. 2016. Wide & deep learning for recommender systems. In Proceedings of the 1st workshop on deep learning for recommender systems . 7–10

  18. [20]

    Xiaohui Cheng, Yonghua Zhu, Jingkuan Song, Guoqiu Wen, and Wei He. 2017. A novel low-rank hypergraph feature selection for multi-view classification. Neurocomputing 253 (2017), 115–121

  19. [21]

    Grigorios G Chrysos, Stylianos Moschoglou, Giorgos Bouritsas, Jiankang Deng, Yannis Panagakis, and Stefanos Zafeiriou. 2021. Deep polynomial neural networks. IEEE transactions on pattern analysis and machine intelligence 44, 8 (2021), 4021–4034

  20. [22]

    Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh. 2009. Learning non-linear combinations of kernels. Advances in neural information processing systems 22 (2009)

  21. [23]

    Menghan Cui, Kaixiang Wang, Xiaojian Ding, Zihan Xu, Xin Wang, and Pengcheng Shi. 2024. Multi-view Stable Feature Selection with Adaptive Optimization of View Weights. Knowledge-Based Systems 299 (2024), 111970

  22. [24]

    Nattane Luíza da Costa, Márcio Dias de Lima, and Rommel Barbosa. 2021. Evaluation of feature selection methods based on artificial neural network weights. Expert Systems with Applications 168 (2021), 114312

  23. [25]

    Houtao Deng and George Runger. 2012. Feature selection via regularized trees. In Proceedings International Joint Conference on Neural Networks . IEEE, 1–8

  24. [26]

    Hui Ding, Peng-Mian Feng, Wei Chen, and Hao Lin. 2014. Identification of bacteriophage virion proteins by the ANOVA feature selection and analysis. Molecular BioSystems 10, 8 (2014), 2229–2235

  25. [27]

    Ofer Dor and Yoram Reich. 2012. Strengthening learning algorithms by feature discovery. Information Sciences 189 (2012), 176–190

  26. [28]

    Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. 2006. Calibrating noise to sensitivity in private data analysis. In Theory of Cryptography: Third Theory of Cryptography Conference, TCC 2006, New York, NY, USA, March 4-7, 2006. Proceedings 3 . Springer, 265–284

  27. [29]

    Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. 2019. Neural architecture search: A survey. The Journal of Machine Learning Research 20, 1 (2019), 1997–2017

  28. [30]

    Frank J Fabozzi, Harry M Markowitz, and Francis Gupta. 2011. Portfolio selection. The Theory and Practice of Investment Management (2011), 45–78

  29. [31]

    Peter S Fader and Bruce GS Hardie. 2005. A note on deriving the Pareto/NBD model and related expressions

  30. [36]

    Si-Guo Fang, Dong Huang, Chang-Dong Wang, and Yong Tang. 2023. Joint multi-view unsupervised feature selection and graph learning. IEEE Transactions on Emerging Topics in Computational Intelligence 8, 1 (2023), 16–31

  31. [37]

    Ronald C Fisher. 2022. State and local public finance . Routledge

  32. [38]

    Todd Michael Franke, Timothy Ho, and Christina A Christie. 2012. The chi-square test: Often used and more often misinterpreted. American journal of evaluation 33, 3 (2012), 448–458

  33. [40]

    Abdullah Saeed Ghareb, Azuraliza Abu Bakar, and Abdul Razak Hamdan. 2016. Hybrid feature selection based on enhanced genetic algorithm for text categorization. Expert Systems with Applications 49 (2016), 31–47

  34. [41]

    Nanxu Gong, Wangyang Ying, Dongjie Wang, and Yanjie Fu. 2024. Neuro-Symbolic Embedding for Short and Effective Feature Selection via Autoregressive Generation. arXiv:2404.17157 [cs.LG] https://arxiv.org/abs/2404.17157

  35. [42]

    Pablo M Granitto, Cesare Furlanello, Franco Biasioli, and Flavia Gasperi. 2006. Recursive feature elimination with random forest for PTR-MS analysis of agroindustrial products. Chemometrics and intelligent laboratory systems 83, 2 (2006), 83–90

  36. [43]

    Isabelle Guyon, Jason Weston, Stephen Barnhill, and Vladimir Vapnik. 2002. Gene selection for cancer classification using support vector machines. Machine learning 46 (2002), 389–422

  37. [44]

    Pingting Hao, Kunpeng Liu, and Wanfu Gao. 2024. Anchor-guided global view reconstruction for multi-view multi-label feature selection. Information Sciences 679 (2024), 121124

  38. [45]

    Pingting Hao, Kunpeng Liu, and Wanfu Gao. 2024. Double-Layer Hybrid-Label Identification Feature Selection for Multi-View Multi-Label Learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 12295–12303

  39. [46]

    Mohammad Nazmul Haque, Nasimul Noman, Regina Berretta, and Pablo Moscato. 2016. Heterogeneous ensemble combination search using genetic algorithm for class imbalanced data classification. PloS one 11, 1 (2016), e0146116

  40. [47]

    Frank E Harrell et al. 2001. Regression modeling strategies: with applications to linear models, logistic regression, and survival analysis . Vol. 608. Springer

  41. [48]

    Xiangnan He and Tat-Seng Chua. 2017. Neural factorization machines for sparse predictive analytics. In Proceedings of the 40th International ACM SIGIR conference on Research and Development in Information Retrieval . 355–364

  42. [49]

    Xin He, Kaiyong Zhao, and Xiaowen Chu. 2021. AutoML: A survey of the state-of-the-art. Knowledge-Based Systems 212 (2021), 106622

  43. [51]

    Chenping Hou, Changshui Zhang, Yi Wu, and Feiping Nie. 2010. Multiple view semi-supervised dimensionality reduction. Pattern Recognition 43, 3 (2010), 720–730

  44. [52]

    Hui-Huang Hsu, Cheng-Wei Hsieh, and Ming-Da Lu. 2011. Hybrid feature selection by combining filters and wrappers. Expert Systems with Applications 38, 7 (2011), 8144–8150

  45. [53]

    Xuanming Hu, Dongjie Wang, Wangyang Ying, and Yanjie Fu. 2024. Reinforcement Feature Transformation for Polymer Property Performance Prediction. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management . 4538–4545

  46. [56]

    Xiaohan Huang, Dongjie Wang, Zhiyuan Ning, Ziyue Qiao, Qingqing Long, Haowei Zhu, Min Wu, Yuanchun Zhou, and Meng Xiao. 2024. Enhancing Tabular Data Optimization with a Flexible Graph-based Reinforced Exploration Strategy. arXiv preprint arXiv:2406.07404 (2024)

  47. [57]

    Yanyong Huang, Zongxin Shen, Yuxin Cai, Xiuwen Yi, Dongjie Wang, Fengmao Lv, and Tianrui Li. 2023. C2IMUFS: Complementary and Consensus Learning-Based Incomplete Multi-View Unsupervised Feature Selection. IEEE Transactions on Knowledge and Data Engineering 35, 10 (2023), 10681–10694

  48. [58]

    RJ Hyndman. 2018. Forecasting: principles and practice . OTexts

  49. [59]

    OECD Indicators and OECD Hagvísar. 2019. Health at a glance 2019: OECD indicators . Paris: OECD Publishing

  50. [60]

    Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, and Ondrej Chum. 2019. Label propagation for deep semi-supervised learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5070–5079

  51. [61]

    Xiaodong Jia, Xiao-Yuan Jing, Xiaoke Zhu, Songcan Chen, Bo Du, Ziyun Cai, Zhenyu He, and Dong Yue. 2020. Semi-supervised multi-view deep discriminant representation learning. IEEE transactions on pattern analysis and machine intelligence 43, 7 (2020), 2496–2509

  52. [62]

    Bingbing Jiang, Xingyu Wu, Xiren Zhou, Yi Liu, Anthony G Cohn, Weiguo Sheng, and Huanhuan Chen. 2022. Semi-supervised multiview feature selection with adaptive graph learning. IEEE Transactions on Neural Networks and Learning Systems 35, 3 (2022), 3615–3629

  53. [63]

    Qi Jiang, Guoxu Zhou, and Qibin Zhao. 2024. Semi-supervised multi-view concept decomposition. Expert Systems with Applications 241 (2024), 122572

  54. [64]

    Xiao-Yuan Jing, Rui-Min Hu, Yang-Ping Zhu, Shan-Shan Wu, Chao Liang, and Jing-Yu Yang. 2014. Intra-view and inter-view supervised correlation analysis for multi-view feature learning. In Proceedings of the AAAI Conference on Artificial Intelligence . AAAI, 1882–1889

  55. [65]

    James Max Kanter and Kalyan Veeramachaneni. 2015. Deep feature synthesis: Towards automating data science endeavors. In2015 IEEE international conference on data science and advanced analytics (DSAA) . IEEE, 1–10

  56. [66]

    Shubhra Kanti Karmaker, Md Mahadi Hassan, Micah J Smith, Lei Xu, Chengxiang Zhai, and Kalyan Veeramachaneni. 2021. Automl to date and beyond: Challenges and opportunities. ACM Computing Surveys (CSUR) 54, 8 (2021), 1–36

  57. [67]

    Gilad Katz, Eui Chul Richard Shin, and Dawn Song. 2016. Explorekit: Automatic feature generation and selection. In 2016 IEEE 16th International Conference on Data Mining (ICDM) . IEEE, 979–984

  58. [68]

    Ambika Kaul, Saket Maheshwary, and Vikram Pudi. 2017. Autolearn—automated feature generation and selection. In 2017 IEEE International Conference on data mining (ICDM) . IEEE, 217–226. Manuscript submitted to ACM 32 Wang et al

  59. [69]

    Udayan Khurana, Horst Samulowitz, and Deepak Turaga. 2018. Feature engineering for predictive modeling using reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32

  60. [70]

    Udayan Khurana, Deepak Turaga, Horst Samulowitz, and Srinivasan Parthasrathy. 2016. Cognito: Automated feature engineering for supervised learning. In 2016 IEEE 16th International Conference on Data Mining Workshops (ICDMW) . IEEE, 1304–1307

  61. [71]

    Kazuki Koyama, Keisuke Kiritoshi, Tomomi Okawachi, and Tomonori Izumitani. 2022. Effective nonlinear feature selection method based on hsic lasso and with variational inference. In International Conference on Artificial Intelligence and Statistics . PMLR, 10407–10421

  62. [72]

    Krzysztof Krawiec. 2002. Genetic programming-based construction of features for machine learning and knowledge discovery tasks. Genetic Programming and Evolvable Machines 3 (2002), 329–343

  63. [73]

    Atsutoshi Kumagai, Tomoharu Iwata, Yasutoshi Ida, and Yasuhiro Fujiwara. 2022. Few-shot Learning for Feature Selection with Hilbert-Schmidt Independence Criterion. Advances in Neural Information Processing Systems 35 (2022), 9577–9590

  64. [74]

    Andrew Kusiak. 2001. Feature transformation methods in data mining. IEEE Transactions on Electronics packaging manufacturing 24, 3 (2001), 214–221

  65. [75]

    Hoang Thanh Lam, Johann-Michael Thiebaut, Mathieu Sinn, Bei Chen, Tiep Mai, and Oznur Alkan. 2017. One button machine for automating feature engineering in relational databases. arXiv preprint arXiv:1706.00327 (2017)

  66. [76]

    Gongmin Lan, Chenping Hou, Feiping Nie, Tingjin Luo, and Dongyun Yi. 2018. Robust feature selection via simultaneous sapped norm and sparse regularizer minimization. Neurocomputing 283 (2018), 228–240

  67. [77]

    Cosmin Lazar, Jonatan Taminau, Stijn Meganck, David Steenhoff, Alain Coletta, Colin Molter, Virginie de Schaetzen, Robin Duque, Hugues Bersini, and Ann Nowe. 2012. A survey on filter techniques for feature selection in gene expression microarray analysis. IEEE/ACM transactions...

  68. [78]

    Riccardo Leardi. 1996. Genetic algorithms in feature selection. In Genetic algorithms in molecular modeling . Elsevier, 67–86

  69. [79]

    Ismael Lemhadri, Feng Ruan, and Rob Tibshirani. 2021. Lassonet: Neural networks with feature sparsity. In International Conference on Artificial Intelligence and Statistics. PMLR, 10–18

  70. [80]

    Jundong Li, Kewei Cheng, Suhang Wang, Fred Morstatter, Robert P Trevino, Jiliang Tang, and Huan Liu. 2017. Feature selection: A data perspective. ACM computing surveys (CSUR) 50, 6 (2017), 1–45

  71. [81]

    Jundong Li, Kewei Cheng, Suhang Wang, Fred Morstatter, Robert P Trevino, Jiliang Tang, and Huan Liu. 2017. Feature Selection: A Data Perspective. Comput. Surveys 50, 6 (2017), 1–45

  72. [82]

    Yaliang Li, Zhen Wang, Yuexiang Xie, Bolin Ding, Kai Zeng, and Ce Zhang. 2021. Automl: From methodology to application. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management . 4853–4856

  73. [83]

    Cheng Liang, Lianzhi Wang, Li Liu, Huaxiang Zhang, and Fei Guo. 2023. Multi-view unsupervised feature selection with tensor robust principal component analysis and consensus graph learning. Pattern Recognition 141 (2023), 109632

  74. [84]

    Qiang Lin, Min Men, Liran Yang, and Ping Zhong. 2022. A supervised multi-view feature selection method based on locally sparse regularization and block computing. Information Sciences 582 (2022), 146–166

  75. [85]

    Qiang Lin, Yiming Xue, Juan Wen, and Ping Zhong. 2019. A sharing multi-view feature selection method via alternating direction method of multipliers. Neurocomputing 333 (2019), 124–134

  76. [86]

    Qiang Lin, Liran Yang, Ping Zhong, and Hui Zou. 2021. Robust supervised multi-view feature selection with weighted shared loss and maximum margin criterion. Knowledge-Based Systems 229 (2021), 107331

  77. [87]

    Greg Linden, Brent Smith, and Jeremy York. 2003. Amazon. com recommendations: Item-to-item collaborative filtering. IEEE Internet computing 7, 1 (2003), 76–80

  78. [88]

    Hongfu Liu, Ming Shao, and Yun Fu. 2018. Feature selection with unsupervised consensus guidance. IEEE Transactions on Knowledge and Data Engineering 31, 12 (2018), 2319–2331

  79. [93]

    Xiangjie Liu, Hao Zhang, Xiaobing Kong, and Kwang Y Lee. 2020. Wind speed forecasting using deep neural network with feature selection. Neurocomputing 397 (2020), 393–403

  80. [94]

    Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir. 2014. On the computational efficiency of training neural networks.Advances in neural information processing systems 27 (2014)

  81. [95]

    Huijuan Lu, Junying Chen, Ke Yan, Qun Jin, Yu Xue, and Zhigang Gao. 2017. A hybrid feature selection algorithm for gene expression data classification. Neurocomputing 256 (2017), 56–62

  82. [96]

    Lundberg and Su-In Lee

    Scott M. Lundberg and Su-In Lee. 2017. A Unified Approach to Interpreting Model Predictions. In Neural Information Processing Systems . https://api.semanticscholar.org/CorpusID:21889700 Manuscript submitted to ACM Towards Data-Centric AI: A Comprehensive Survey of Traditional,...

  83. [97]

    Yuanfei Luo, Mengshuo Wang, Hao Zhou, Quanming Yao, Wei-Wei Tu, Yuqiang Chen, Wenyuan Dai, and Qiang Yang. 2019. Autocross: Automatic feature crossing for tabular data in real-world applications. InProceedings of the 25th ACM SIGKDD International Conference on Knowledge Discov...

  84. [98]

    Thomas Marill and D Green. 1963. On the effectiveness of receptors in recognition systems. IEEE transactions on Information Theory 9, 1 (1963), 11–17

  85. [99]

    Mary L McHugh. 2013. The chi-square test of independence. Biochemia medica 23, 2 (2013), 143–149

  86. [100]

    Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2017. Communication-efficient learning of deep networks from decentralized data. In Artificial intelligence and statistics. PMLR, 1273–1282

  87. [101]

    Min Men, Ping Zhong, Zhi Wang, and Qiang Lin. 2020. Distributed learning for supervised multiview feature selection. Applied Intelligence 50 (2020), 2749–2769

  88. [102]

    Thaha Muhammed and Riaz Ahmed Shaikh. 2017. An analysis of fault detection strategies in wireless sensor networks. Journal of Network and Computer Applications 78 (2017), 267–287

  89. [103]

    Alhassan Mumuni and Fuseini Mumuni. 2024. Automated data processing and feature engineering for deep learning and big data applications: a survey. Journal of Information and Intelligence (2024)

  90. [104]

    R Muthukrishnan and R Rohini. 2016. LASSO: A feature selection technique in predictive modeling for machine learning. In Proceedings of IEEE International Conference on Advances in Computer Applications . Ieee, 18–20

  91. [105]

    Songyot Nakariyakul and David P Casasent. 2009. An improvement on floating search algorithms for feature subset selection. Pattern Recognition 42, 9 (2009), 1932–1940

  92. [106]

    Nassim Nicholas. 2008. The black swan: the impact of the highly improbable. Journal of the Management Training Institut 36, 3 (2008), 56

  93. [107]

    Feiping Nie, Zheng Wang, Lai Tian, Rong Wang, and Xuelong Li. 2020. Subspace sparse discriminative feature selection. IEEE transactions on Cybernetics 52, 6 (2020), 4221–4233

  94. [108]

    Feiping Nie, Wei Zhu, and Xuelong Li. 2019. Structured graph optimization for unsupervised feature selection. IEEE Transactions on Knowledge and Data Engineering 33, 3 (2019), 1210–1222

  95. [109]

    Vahid Noroozi, Sara Bahaadini, Lei Zheng, Sihong Xie, Weixiang Shao, and S Yu Philip. 2018. Semi-supervised deep representation learning for multi-view problems. In 2018 IEEE international conference on big data (Big Data) . IEEE, 56–64

  96. [110]

    Arie Nugroho, Ahmad Zainul Fanani, and Guruh Fajar Shidik. 2021. Evaluation of feature selection using wrapper for numeric dataset with random forest algorithm. In Proceedings of International Seminar on Application for Technology of Information and Communication . IEEE, 179–183

  97. [111]

    Il-Seok Oh, Jin-Seon Lee, and Byung-Ro Moon. 2004. Hybrid genetic algorithms for feature selection. IEEE Transactions on pattern analysis and machine intelligence 26, 11 (2004), 1424–1437

  98. [112]

    J Osborne. 2003. Notes on the use of data transformations. Practical Assessment, Research, and Evaluations. 2002. Retrieved May 10 (2003), 8

  99. [113]

    Yassine Ouali, Céline Hudelot, and Myriam Tami. 2020. An overview of deep semi-supervised learning. arXiv preprint arXiv:2006.05278 (2020)

  100. [114]

    Tianji Pang, Feiping Nie, Junwei Han, and Xuelong Li. 2018. Efficient feature selection viaℓ2,0-norm constrained sparse regression. IEEE Transactions on Knowledge and Data Engineering 31, 5 (2018), 880–893

  101. [115]

    Karl Pearson. 1901. LIII. On lines and planes of closest fit to systems of points in space. The London, Edinburgh, and Dublin philosophical magazine and journal of science 2, 11 (1901), 559–572

  102. [116]

    Hanchuan Peng, Fuhui Long, and Chris Ding. 2005. Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy. IEEE Transactions on pattern analysis and machine intelligence 27, 8 (2005), 1226–1238

  103. [117]

    Barbara Pes, Nicoletta Dessì, and Marta Angioni. 2017. Exploiting the ensemble paradigm for stable feature selection: a case study on high- dimensional genomic data. Information Fusion 35 (2017), 132–147

  104. [118]

    Pavel Pudil, Jana Novovičová, and Josef Kittler. 1994. Floating search methods in feature selection. Pattern recognition letters 15, 11 (1994), 1119–1125

  105. [119]

    Stefan Rahmstorf and Dim Coumou. 2011. Increase of extreme events in a warming world. Proceedings of the National Academy of Sciences 108, 44 (2011), 17905–17909

  106. [120]

    Punch, Erik D Goodman, Leslie A Kuhn, and Anil K Jain

    Michael L Raymer, William F. Punch, Erik D Goodman, Leslie A Kuhn, and Anil K Jain. 2000. Dimensionality reduction using genetic algorithms. IEEE transactions on evolutionary computation 4, 2 (2000), 164–171

  107. [121]

    Why should i trust you?

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " Why should i trust you?" Explaining the predictions of any classifier. InProceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1135–1144

  108. [122]

    Debaditya Roy, K Sri Rama Murty, and C Krishna Mohan. 2015. Feature selection using deep neural networks. In Proceedings of International Joint Conference on Neural Networks . IEEE, 1–6

  109. [123]

    Borja Seijo-Pardo, Verónica Bolón-Canedo, and Amparo Alonso-Betanzos. 2017. Testing different ensemble configurations for feature selection. Neural Processing Letters 46, 3 (2017), 857–880

  110. [124]

    Borja Seijo-Pardo, Verónica Bolón-Canedo, and Amparo Alonso-Betanzos. 2019. On developing an automatic threshold applied to feature selection ensembles. Information Fusion 45 (2019), 227–245

  111. [125]

    William F Sharpe. 1994. The sharpe ratio. Journal of portfolio management 21, 1 (1994), 49–58

  112. [126]

    Caijuan Shi, Gaoyun An, Ruizhen Zhao, Qiuiqi Ruan, and Qi Tian. 2016. Multiview hessian semisupervised sparse feature selection for multimedia analysis. IEEE Transactions on Circuits and Systems for Video Technology 27, 9 (2016), 1947–1961. Manuscript submitted to ACM 34 Wang et al

  113. [127]

    Caijuan Shi, Qiuqi Ruan, Gaoyun An, and Chao Ge. 2015. Semi-supervised sparse feature selection based on multi-view Laplacian regularization. Image and Vision Computing 41 (2015), 1–10

  114. [128]

    Yong Shi, Jianyu Miao, Zhengyu Wang, Peng Zhang, and Lingfeng Niu. 2018. Feature selection withℓ2,1−2 regularization. IEEE Transactions on Neural Networks and Learning Systems 29, 10 (2018), 4967–4982

  115. [129]

    Wojciech Siedlecki and Jack Sklansky. 1989. A note on genetic algorithms for large-scale feature selection. Pattern recognition letters 10, 5 (1989), 335–347

  116. [130]

    Petr Somol, Pavel Pudil, Jana Novovičová, and Pavel Paclık. 1999. Adaptive floating search methods in feature selection. Pattern recognition letters 20, 11-13 (1999), 1157–1163

  117. [131]

    Xian-Fang Song, Yong Zhang, Dun-Wei Gong, and Xiao-Zhi Gao. 2021. A fast hybrid feature selection based on correlation-guided clustering and particle swarm optimization for high-dimensional data. IEEE Transactions on Cybernetics 52, 9 (2021), 9573–9586

  118. [132]

    Lars St, Svante Wold, et al. 1989. Analysis of variance (ANOVA). Chemometrics and intelligent laboratory systems 6, 4 (1989), 259–272

  119. [133]

    Chang Tang, Jiajia Chen, Xinwang Liu, Miaomiao Li, Pichao Wang, Minhui Wang, and Peng Lu. 2018. Consensus learning guided multi-view unsupervised feature selection. Knowledge-Based Systems 160 (2018), 49–60

  120. [134]

    Chang Tang, Xiao Zheng, Xinwang Liu, Wei Zhang, Jing Zhang, Jian Xiong, and Lizhe Wang. 2021. Cross-view locality preserved diversity and consensus learning for multi-view unsupervised feature selection. IEEE Transactions on Knowledge and Data Engineering 34, 10 (2021), 4705–4716

  121. [135]

    Chang Tang, Xinzhong Zhu, Xinwang Liu, and Lizhe Wang. 2019. Cross-view local structure preserved diversity and consensus learning for multi-view unsupervised feature selection. In Proceedings of the AAAI Conference on Artificial Intelligence . 5101–5108

  122. [136]

    Jiliang Tang, Salem Alelyani, and Huan Liu. 2014. Feature selection for classification: A review. Data classification: Algorithms and applications (2014), 37

  123. [137]

    Robert Tibshirani. 1996. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society Series B: Statistical Methodology 58, 1 (1996), 267–288

  124. [138]

    Jingwei Too and Abdul Rahim Abdullah. 2021. A new and fast rival genetic algorithm for feature selection. The Journal of Supercomputing 77, 3 (2021), 2844–2874

  125. [139]

    Binh Tran, Bing Xue, and Mengjie Zhang. 2016. Genetic programming for feature construction and selection in classification on high-dimensional data. Memetic Computing 8, 1 (2016), 3–15

  126. [140]

    Chih-Fong Tsai, William Eberle, and Chi-Yuan Chu. 2013. Genetic algorithms in feature and instance selection. Knowledge-Based Systems 39 (2013), 240–247

  127. [141]

    Ruey S Tsay. 2005. Analysis of financial time series . John wiley & sons

  128. [142]

    John W Tukey. 1977. Exploratory data analysis. Reading/Addison-Wesley (1977)

  129. [143]

    Alper Kursat Uysal and Serkan Gunal. 2014. Text classification using genetic algorithm oriented latent semantic features. Expert Systems with Applications 41, 13 (2014), 5938–5947

  130. [144]

    Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of machine learning research 9, 11 (2008)

  131. [145]

    Antanas Verikas and Marija Bacauskiene. 2002. Feature selection with neural networks. Pattern Recognition Letters 23, 11 (2002), 1323–1335

  132. [146]

    Dimitrios Ververidis and Constantine Kotropoulos. 2005. Sequential forward feature selection with low computational cost. In 2005 13th European Signal Processing Conference. IEEE, 1–4

  133. [147]

    Dongjie Wang, Yanjie Fu, Kunpeng Liu, Xiaolin Li, and Yan Solihin. 2022. Group-wise Reinforcement Feature Generation for Optimal and Explainable Representation Space Reconstruction. Proceedings of the 28th ACM SIGKDD international conference on Knowledge discovery and data min...

  134. [148]

    Dongjie Wang, Meng Xiao, Min Wu, Pengfei Wang, Yuanchun Zhou, and Yanjie Fu. 2024. Reinforcement-enhanced autoregressive feature transformation: gradient-steered search in continuous space for postfix expressions. In Proceedings of the 37th International Conference on Neural I...

  135. [149]

    Hua Wang, Feiping Nie, Heng Huang, and Chris Ding. 2013. Heterogeneous visual features fusion via sparse multimodal machine. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 3097–3102

  136. [150]

    Jian Wang, Huaqing Zhang, Junze Wang, Yifei Pu, and Nikhil R Pal. 2020. Feature selection using a neural network with group lasso regularization and controlled redundancy. IEEE Transactions on Neural Networks and Learning Systems 32, 3 (2020), 1110–1123

  137. [151]

    Rong Wang, Jintang Bian, Feiping Nie, and Xuelong Li. 2022. Nonlinear feature selection neural network via structured sparse regularization. IEEE Transactions on Neural Networks and Learning Systems 34, 11 (2022), 9493–9505

  138. [152]

    Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & cross network for ad click predictions. In Proceedings of the ADKDD’17 . 1–7

  139. [153]

    Rong Wang, Canyu Zhang, Jintang Bian, Zheng Wang, Feiping Nie, and Xuelong Li. 2022. Sparse and flexible projections for unsupervised feature selection. IEEE Transactions on Knowledge and Data Engineering 35, 6 (2022), 6362–6375

  140. [154]

    Xinyuan Wang, Dongjie Wang, Wangyang Ying, Rui Xie, Haifeng Chen, and Yanjie Fu. 2024. Knockoff-Guided Feature Selection via A Single Pre-trained Reinforced Agent. arXiv preprint arXiv:2403.04015 (2024)

  141. [155]

    Yu-Hao Wang, Yu-Fei Zhang, Ying Zhang, Zhi-Feng Gu, Zhao-Yue Zhang, Hao Lin, and Ke-Jun Deng. 2022. Identification of adaptor proteins using the ANOVA feature selection technique. Methods 208 (2022), 42–47

  142. [156]

    Zheng Wang, Feiping Nie, Canyu Zhang, Rong Wang, and Xuelong Li. 2021. Joint nonlinear feature selection and continuous values regression network. Pattern Recognition Letters 150 (2021), 197–206. Manuscript submitted to ACM Towards Data-Centric AI: A Comprehensive Survey of Tr...

  143. [157]

    Melvyn Weeks. 2002. Introductory econometrics: a modern approach

  144. [158]

    A Wayne Whitney. 1971. A direct method of nonparametric measurement selection. IEEE transactions on computers 100, 9 (1971), 1100–1103

  145. [159]

    Jian-Sheng Wu, Jun-Xiao Gong, Jing-Xin Liu, and Weidong Min. 2023. Multi-level correlation learning for multi-view unsupervised feature selection. Knowledge-Based Systems 281 (2023), 111073

  146. [160]

    Le Wu, Xiangnan He, Xiang Wang, Kun Zhang, and Meng Wang. 2022. A survey on accuracy-oriented neural recommendation: From collaborative filtering to information-rich recommendation. IEEE Transactions on Knowledge and Data Engineering 35, 5 (2022), 4425–4445

  147. [161]

    Zhihao Wu, Xincan Lin, Zhenghong Lin, Zhaoliang Chen, Yang Bai, and Shiping Wang. 2023. Interpretable graph convolutional network for multi-view semi-supervised learning. IEEE Transactions on Multimedia 25 (2023), 8593–8606

  148. [162]

    Lihu Xiao, Zhenan Sun, Ran He, and Tieniu Tan. 2013. Coupled feature selection for cross-sensor iris recognition. In 2013 IEEE Sixth International Conference on Biometrics: Theory, Applications and Systems (BTAS) . IEEE, 1–6

  149. [163]

    Meng Xiao, Dongjie Wang, Min Wu, Kunpeng Liu, Hui Xiong, Yuanchun Zhou, and Yanjie Fu. 2024. Traceable group-wise self-optimizing feature transformation learning: A dual optimization perspective. ACM Transactions on Knowledge Discovery from Data 18, 4 (2024), 1–22

  150. [164]

    Meng Xiao, Dongjie Wang, Min Wu, Ziyue Qiao, Pengfei Wang, Kunpeng Liu, Yuanchun Zhou, and Yanjie Fu. 2023. Traceable Automatic Feature Transformation via Cascading Actor-Critic Agents. In Proceedings of the 2023 SIAM International Conference on Data Mining (SDM) . SIAM, 775–783

  151. [165]

    Meng Xiao, Dongjie Wang, Min Wu, Pengfei Wang, Yuanchun Zhou, and Yanjie Fu. 2023. Beyond Discrete Selection: Continuous Embedding Space Optimization for Generative Feature Selection . In 2023 IEEE International Conference on Data Mining (ICDM) . IEEE Computer Society, Los Ala...

  152. [166]

    Meng Xiao, Weiliang Zhang, Xiaohan Huang, Hengshu Zhu, Min Wu, Xiaoli Li, and Yuanchun Zhou. 2025. Knowledge-Guided Biomarker Identification for Label-Free Single-Cell RNA-Seq Data: A Reinforcement Learning Perspective. (2025). arXiv:2501.04718 [q-bio.GN] https: //arxiv.org/ab...

  153. [167]

    Jinglin Xu, Junwei Han, Feiping Nie, and Xuelong Li. 2020. Multi-View Scaling Support Vector Machines for Classification and Feature Selection. IEEE Transactions on Knowledge and Data Engineering 32, 7 (2020), 1419–1430

  154. [168]

    Zhaozhao Xu, Fangyuan Yang, Chaosheng Tang, Hong Wang, Shuihua Wang, Junding Sun, and Yudong Zhang. 2024. FG-HFS: A feature filter and group evolution hybrid feature selection algorithm for high-dimensional gene expression data. Expert Systems with Applications 245 (2024), 123069

  155. [169]

    Zihao Xu, Chenglong Zhang, Zhaolong Ling, Peng Zhou, Yan Zhong, Li Li, Han Zhang, Weiguo Sheng, and Bingbing Jiang. 2024. Multi-View Semi-Supervised Feature Selection with Graph Convolutional Networks. In 2024 International Joint Conference on Neural Networks (IJCNN) . IEEE, 1–8

  156. [170]

    Ke Yan and David Zhang. 2015. Feature selection and analysis on correlated gas sensor data with recursive feature elimination. Sensors and Actuators B: Chemical 212 (2015), 353–363

  157. [171]

    Jian-Bo Yang, Kai-Quan Shen, Chong-Jin Ong, and Xiao-Ping Li. 2009. Feature selection for MLP neural network: The use of random permutation of probabilistic outputs. IEEE Transactions on Neural Networks 20, 12 (2009), 1911–1922

  158. [172]

    Xiangli Yang, Zixing Song, Irwin King, and Zenglin Xu. 2022. A survey on deep semi-supervised learning. IEEE Transactions on Knowledge and Data Engineering 35, 9 (2022), 8934–8954

  159. [174]

    Wangyang Ying, Haoyue Bai, Kunpeng Liu, and Yanjie Fu. 2024. Topology-aware Reinforcement Feature Space Reconstruction for Graph Data. arXiv preprint arXiv:2411.05742 (2024)

  160. [175]

    Wangyang Ying, Dongjie Wang, Haifeng Chen, and Yanjie Fu. 2024. Feature Selection as Deep Sequential Generative Learning. ACM Trans. Knowl. Discov. Data 18, 9, Article 221 (Oct. 2024), 21 pages. doi:10.1145/3687485

  161. [176]

    Wangyang Ying, Dongjie Wang, Xuanming Hu, Ji Qiu, Jin Park, and Yanjie Fu. 2024. Revolutionizing Biomarker Discovery: Leveraging Generative AI for Bio-Knowledge-Embedded Continuous Space Exploration. InProceedings of the 33rd ACM International Conference on Information and Kno...

  162. [179]

    Sejong Yoon and Saejoon Kim. 2009. Mutual information-based SVM-RFE for diagnostic classification of digitized mammograms.Pattern Recognition Letters 30, 16 (2009), 1489–1495

  163. [180]

    Wenjie You, Zijiang Yang, and Guoli Ji. 2014. Feature selection for high-dimensional multi-category data using PLS-based local recursive feature elimination. Expert Systems with Applications 41, 4 (2014), 1463–1475

  164. [181]

    Haoliang Yuan, Junyu Li, Yong Liang, and Yuan Yan Tang. 2022. Multi-view unsupervised feature selection with tensor low-rank minimization. Neurocomputing 487 (2022), 75–85

  165. [182]

    Chenglong Zhang, Yang Fang, Xinyan Liang, Xingyu Wu, Bingbing Jiang, et al. 2024. Efficient Multi-view Unsupervised Feature Selection with Adaptive Structure Learning and Inference. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (I...

  166. [183]

    Chenglong Zhang, Bingbing Jiang, Zidong Wang, Jie Yang, Yangfeng Lu, Xingyu Wu, and Weiguo Sheng. 2023. Efficient multi-view semi-supervised feature selection. Information Sciences 649 (2023), 119675

  167. [184]

    Chenglong Zhang, Xinyan Liang, Peng Zhou, Zhaolong Ling, Yingwei Zhang, Xingyu Wu, Weiguo Sheng, and Bingbing Jiang. 2024. Scalable Multi-view Unsupervised Feature Selection with Structure Learning and Fusion. In Proceedings of the 32nd ACM International Conference on Multimed...

  168. [185]

    Huaqing Zhang, Jian Wang, Zhanquan Sun, Jacek M Zurada, and Nikhil R Pal. 2019. Feature selection for neural networks using group lasso regularization. IEEE Transactions on Knowledge and Data Engineering 32, 4 (2019), 659–673

  169. [186]

    Rui Zhang, Feiping Nie, Xuelong Li, and Xian Wei. 2019. Feature Selection with Multi-View Data: A Survey. Information Fusion 50 (2019), 158–167

  170. [187]

    Xinhao Zhang, Zaitian Wang, Lu Jiang, Wanfu Gao, Pengfei Wang, and Kunpeng Liu. 2024. TFWT: Tabular Feature Weighting with Transformer. arXiv preprint arXiv:2405.08403 (2024)

  171. [192]

    Changkang Zhong, Yu Chen, and Jian Peng. 2020. Feature selection based on a novel improved tree growth algorithm. International Journal of Computational Intelligence Systems 13, 1 (2020), 247–258

  172. [194]

    HongFang Zhou, JiaWei Zhang, YueQing Zhou, XiaoJie Guo, and YiMing Ma. 2021. A feature selection algorithm of decision tree based on feature weight. Expert Systems with Applications 164 (2021), 113842

  173. [195]

    Shixuan Zhou and Peng Song. 2024. Consistency–exclusivity guided unsupervised multi-view feature selection. Neurocomputing 569 (2024), 127119

  174. [196]

    Guanghui Zhu, Shen Jiang, Xu Guo, Chunfeng Yuan, and Yihua Huang. 2022. Evolutionary Automated Feature Engineering. In PRICAI 2022: Trends in Artificial Intelligence: 19th Pacific Rim International Conference on Artificial Intelligence, PRICAI 2022, Shanghai, China, November 1...

  175. [197]

    Guanghui Zhu, Zhuoer Xu, Chunfeng Yuan, and Yihua Huang. 2022. DIFER: differentiable automated feature engineering. In International Conference on Automated Machine Learning . PMLR, 17–1

  176. [198]

    Jianyong Zhu, Jingwei Chen, Bin Xu, Hui Yang, and Feiping Nie. 2023. Fast orthogonal locality-preserving projections for unsupervised feature selection. Neurocomputing 531 (2023), 100–113

  177. [199]

    Yongbin Zhu, Wenshan Li, and Tao Li. 2023. A hybrid artificial immune optimization for high-dimensional feature selection. Knowledge-Based Systems 260 (2023), 110111. Manuscript submitted to ACM

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.