{"id":"f11c5714-ae10-4985-b925-466cb2601206","arxiv_id":"2502.08828","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review that organizes RL-based and generative methods for tabular feature selection and generation into a taxonomy, compares their strengths and limitations, and outlines research challenges.","lead":"Surveys how reinforcement learning and generative AI are used to automate feature selection and feature generation in tabular data. A useful map of an emerging data-centric AI area for practitioners choosing between these approaches.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey's central map rests on an unverified, self-heavy citation base; without a systematic selection protocol, the taxonomy's completeness is not established.","rationale":"The reader's conditional verdict is appropriate. The paper's central contribution is a map of RL and generative tabular feature engineering; for that map to hold, the papers used to draw it must be representative. Examining Sections 3–4 and Figure 1: of the 25+ method papers discussed, roughly 20 are from the same core group or close collaborators, and the manuscript provides no search protocol, inclusion/exclusion criteria, or independent sampling. This makes the taxonomy's completeness unfalsifiable from the text alone. That is a coverage-bias risk, not an allegation of bad faith. The concrete test—a systematic literature search compared against the reference list—would settle whether the map is complete. If independent works are missing, the survey should either broaden its coverage or explicitly reframe itself as a review of the authors' own line of work. Until that check is run, rejecting the paper would be too harsh: the taxonomy is still a useful starting point, and most cited works are real, peer-reviewed contributions. Hence the reader's 'conditional' verdict stands unchanged.","tokens_in":11460,"tokens_out":4977,"duration_ms":48361,"concrete_test":"Run a pre-registered literature search on DBLP/Google Scholar/Semantic Scholar with queries such as 'reinforcement learning feature selection', 'generative feature generation tabular', and 'automated feature engineering RL', restricted to 2015–2025; screen titles/abstracts with explicit inclusion criteria (tabular data, feature selection or generation, RL or generative methodology); and compare the resulting paper set against the survey's Sections 3–4 references. If more than five independent groups' qualifying papers are absent, the taxonomy's completeness and the comparative conclusions are not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's contribution is the taxonomy and comparative analysis (Figure 1, Sections 3–5), which claims to 'systematically review' RL and generative methods for tabular feature selection and generation. The load-bearing condition is that the cited papers are representative of the field. Of the roughly 25 method papers discussed in Sections 3–4, about 20 are from the same author group or immediate collaborators (e.g., Liu et al. 2019/2021; Fan et al. 2020/2021; Ying et al. 2023/2024a–d; Wang et al. 2022/2024a–b; Gong et al. 2024a–b). The manuscript gives no search strategy, inclusion/exclusion criteria, or quality screen, so a reader cannot tell whether Figure 1 was derived from the broader literature or from a single research trajectory. If independent, top-venue work in these areas is omitted, the taxonomy's categories and the Section 5/7 comparative guidance may be overfit to one family of methods. This is a structural limitation of the survey's central argument, not an allegation of misconduct.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews reinforcement learning (RL) and generative approaches for feature selection and feature generation in tabular data-centric AI. It proposes a taxonomy in Figure 1, reviews methods in Sections 3 and 4, compares the two families qualitatively in Section 5, offers practical strategies in Section 6, gives selection guidance in Section 7, and discusses challenges and future directions in Section 8. The paper's central claim is that RL-based and generative techniques can be organized into the proposed taxonomy and usefully compared in terms of performance, interpretability, adaptability, and automation.","tokens_in":11662,"tokens_out":5011,"duration_ms":49358,"significance":"If the taxonomy and comparative conclusions were established, the survey would provide a useful structured map of an emerging area and would help practitioners choose between RL-based and generative feature engineering. The paper offers a clear conceptual framing, especially the contrast between RL's discrete sequential search and generative methods' continuous embedding-space optimization. However, the survey's evidentiary base is currently too narrow and self-referential to support the claim of a systematic review. The load-bearing assumption that the cited papers are representative of the field is not justified by any stated selection protocol, and the qualitative comparisons are not grounded in empirical evidence. The manuscript would be a valuable contribution after a major revision that addresses these structural issues.","major_comments":[{"comment":"The survey's core claim is that it 'systematically reviews existing feature selection and generation techniques' (Section 1), but the manuscript gives no search strategy, database list, time span, inclusion/exclusion criteria, or screening procedure. The method papers discussed in Sections 3–4 are overwhelmingly from the same author group or immediate collaborators: of the roughly 25 works reviewed, about 20 are by Liu, Fan, Wang, Ying, Gong, Xiao, and coauthors. As a result, a reader cannot verify that the taxonomy in Figure 1 and the qualitative comparisons in Section 5 are representative of the broader field rather than a description of one research trajectory. This is a structural limitation of the central argument, not an allegation of misconduct; it is fixable by adding a transparent literature-selection protocol and expanding coverage to independent work.","section":"§1, Figure 1; Sections 3–4"},{"comment":"The proposed taxonomy is not consistently defined. Figure 1 places methods under branches such as 'Two-Step Generative Approaches', 'Cascading Frameworks', and 'Hybrid And Specialized RL Approaches', but the text never defines 'two-step', and several methods appear to fit more than one branch: Xiao et al. [2023] is presented in Section 4.1 as a generative method but 'leverages reinforcement learning' for data collection; Wang et al. [2024a] is classified under generative feature generation yet its title is 'Reinforcement-Enhanced Autoregressive Feature Transformation'. Because the central contribution is the map itself, overlapping categories and unstated criteria for branch assignment need to be clarified.","section":"Figure 1; Sections 3.1–4.2"},{"comment":"The comparative claims in Section 5 and the guidance in Section 7 are stated as established findings, but no empirical evidence, benchmark table, or cited source is provided for statements such as 'RL-based methods: RL-based methods offer better interpretability', 'Generative-based methods: ... more stable than RL in some cases', and 'Generative models ... automation level is higher than RL'. These are plausible hypotheses, but the paper does not distinguish them from documented results, and some claims are internally qualified later (e.g., deep RL models are admitted to become harder to interpret). The authors should either ground the comparisons in a systematic synthesis of reported experimental results or explicitly label Section 5 as design considerations/opinion.","section":"§5; §7"}],"minor_comments":[{"comment":"The sentence 'improving model performance, efficiency, and interoperability' appears to use 'interoperability' where 'interpretability' is meant; this should be corrected.","section":"§2"},{"comment":"The citation 'Bai et al.' in the discussion of differential privacy has no year or venue and is not listed in the references; the bibliographic entry should be completed.","section":"§6"},{"comment":"The phrase 'This chapter explores' should be 'This section explores', since the manuscript is organized into sections, not chapters.","section":"§8"},{"comment":"The feature generation example '[f1, f2] → [f1/f2, f1 − f2, f1+f2/f1]' is ambiguous; parentheses such as '(f1+f2)/f1' would avoid implying f1 + (f2/f1).","section":"§2"},{"comment":"Several reference entries are incomplete or inconsistently formatted: 'Sutton [2018]' is listed as a book without the full title formatting, and 'Kamatchi and Uma [2025]' has inconsistent capitalization; a careful reference-checking pass is needed.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The self-citation density is high enough that the editor may wish to require a transparent statement of scope and coverage. As it stands, the paper reads more like a research-group retrospective than a field survey. If the journal expects broad literature coverage, the authors should either substantially broaden the cited base or explicitly narrow the paper's scope to their own research program."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does one useful thing well: it organizes the small, recent literature on RL-based and generative feature selection/generation for tabular data into a clean two-by-two taxonomy. Sections 6 and 7 contain practical, level-headed advice that a practitioner new to this area would find helpful. The writing is clear, and the framing of feature engineering as a data-centric optimization problem is sensible.\n\nThat said, the central claim of the survey—that it systematically reviews the field—does not survive contact with the reference list. The large majority of method papers discussed are from the same research group or immediate collaborators (Ying, Wang, Gong, Fan, Liu, Xiao, Fu, and so on). There is no search strategy, no inclusion or exclusion criteria, and no attempt to account for work that might fall outside this orbit. The effect is that the taxonomy in Figure 1 looks more like a map of one laboratory's research program than a map of the field.\n\nThe qualitative comparisons in Section 5 (performance, interpretability, adaptability) are asserted without any benchmark evidence or systematic analysis. They may be plausible, but they are not demonstrated. I also noticed a few citation oddities, like unrelated biomedical/dental papers cited as examples of high-stakes domains in Section 8; that undercuts the impression of careful scholarship.\n\nThe stress-test note is right: the representativeness problem is structural, not a matter of misconduct. But it is also fixable. The authors could either broaden coverage with a transparent literature selection protocol or, failing that, reframe the survey as a focused review of a particular family of methods—their own—rather than a comprehensive review. Either move would make the paper honest about its scope.\n\nFor who it helps: someone looking for a quick orientation to RL/generative feature engineering and who understands they are getting one group's perspective. It does not replace a broader survey.\n\nRecommendation: send to peer review, but clearly ask for major revision: clarify inclusion criteria, broaden or reposition coverage, and either substantiate or soften the comparative claims in Section 5. The taxonomy itself is worth preserving.","headline":"A readable, well-organized survey of a narrow subfield, but the review's coverage is heavily tilted toward the authors' own research line and its comparative claims rest on assertion rather than evidence.","tokens_in":12142,"tokens_out":1914,"would_cite":false,"duration_ms":19389,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey organizes RL-based and generative approaches to tabular feature selection and generation into one taxonomy and argues they have complementary trade-offs.","keywords":["tabular data","data-centric AI","feature selection","feature generation","reinforcement learning","generative AI","automated feature engineering","large language models"],"falsifier":"A systematic literature search followed by a benchmark would settle it: if methods outside the surveyed set populate categories the taxonomy lacks, or if on a fixed collection of tabular datasets the RL-versus-generative performance ranking contradicts the survey's stated trade-offs, the central claim is falsified.","tokens_in":11294,"feed_emoji":"📊","tokens_out":5877,"duration_ms":51214,"temperature":0.7,"pith_summary":"This survey argues that the many scattered methods for automatically improving tabular data can be seen as two coherent families: those that use reinforcement learning to search over discrete feature choices, and those that use generative models to optimize in a continuous embedding space. It organizes these families into a taxonomy covering both feature selection and feature generation, and it compares their strengths and weaknesses across performance, interpretability, adaptability, and automation. A reader would care because automated feature engineering is the main route to better machine learning on tabular data without manual effort; a clear map tells practitioners when to use RL-based methods, when to use generative methods, and when to combine them. If the taxonomy holds, it gives researchers a shared language for positioning new methods and identifies concrete open problems, from privacy-preserving feature engineering to LLM-based and multimodal feature generation.","feed_headline":"Survey maps RL and generative AI for tabular feature engineering","feed_subtitle":"Taxonomy tells when reinforcement learning beats generative models for tabular feature engineering.","key_machinery":"The central object is the taxonomy itself, grounded in two named mechanisms. Reinforcement learning treats feature selection and generation as a Markov decision process: an agent selects features or applies transformation operators, receives a reward from the downstream model, and iteratively refines its policy. Generative models use an embedding-optimization-generation loop: observed feature sets are encoded into a continuous latent space, the space is searched by gradient-based optimization, and new feature decisions are decoded from the optimized embedding. The encoder-decoder-evaluator architecture appears repeatedly as the concrete implementation of the generative paradigm, with long-range dependencies captured by transformer-based variational autoencoders and redundancy controlled by orthogonality constraints.","core_discovery":"The paper's central claim is that the state of data-centric AI for tabular learning is best understood through a two-axis map: one axis is the task (feature selection versus feature generation), and the other is the optimization paradigm (reinforcement learning versus generative modeling). On the RL side, feature selection and generation are cast as sequential decision processes driven by reward signals; on the generative side, features are embedded into a continuous space where selection or construction is found by gradient-based search and then decoded. The survey further claims that these two paradigms have complementary trade-offs: RL offers traceable decision paths and adaptivity to streaming data but suffers from high computational cost and sensitivity to reward design, while generative methods enable smoother high-dimensional search and higher automation but bring black-box interpretability and dependence on training data quality. It completes the picture with practical selection criteria, hybrid strategies, and a list of open challenges that follow from the comparison.","pith_inferences":["The RL-versus-generative divide in this survey looks like a special case of a broader spectrum between discrete combinatorial search and continuous relaxation; the same trade-off likely applies to other data-centric tasks such as data cleaning, imputation, and augmentation.","The emphasis on the embedding-optimization-generation paradigm suggests a testable extension: applying the same continuous-space approach to feature selection in non-tabular modalities, such as graph or time-series data, might outperform RL baselines on tasks with high-dimensional feature spaces.","A practical benchmark could decide the comparative claims: on a fixed set of public tabular datasets, measure RL-based versus generative feature engineering under a fixed compute budget; if their relative performance reverses between tasks, the survey's guidance would need to be conditioned on more than data dynamics and dimensionality.","The survey implies that interpretability is a key differentiator, but post-hoc interpretability tools could narrow that gap; a hybrid pipeline that uses RL for selectivity and generative models with surrogate explanations could serve both goals."],"forward_implications":["If the taxonomy is correct, new RL-based or generative feature-engineering methods can be positioned by which cell they fill, and practitioners can choose approaches by task type and data characteristics.","The comparative analysis implies that RL-based methods should be preferred for dynamic, streaming, or sequentially changing data, and generative methods for static high-dimensional datasets with ample unlabeled structure.","The survey's hybrid scenario suggests a concrete architecture: a generative model proposes a wide pool of candidate features, and an RL agent selects and refines them, balancing exploration with long-term rewards.","The stated future directions indicate that LLM-based feature generation and multimodal integration are the next frontier, with open questions about tabular encoding and cross-modal alignment.","If the identified limitations are taken seriously, research priority should shift to reward design for RL and to making generative feature engineering interpretable and privacy-preserving."],"supporting_citations":[{"why":"Establishes that feature engineering significantly impacts predictive modeling, motivating the survey's focus.","marker":"Heaton [2016]"},{"why":"Foundational multi-agent RL framework for feature subspace exploration, defining the RL feature selection category.","marker":"Liu et al. [2019]"},{"why":"Introduces cascaded group-wise RL feature generation, a central example of the RL feature generation category.","marker":"Wang et al. [2022]"},{"why":"Proposes the encoder-decoder-evaluator paradigm for generative feature selection, anchoring the generative category.","marker":"Xiao et al. [2023]"},{"why":"Redefines feature selection as sequential token generation with a transformer-based VAE, a key generative feature selection method.","marker":"Ying et al. [2024b]"},{"why":"Presents an unsupervised generative feature transformation framework using graph contrastive learning, a central generative feature generation method.","marker":"Ying et al. [2024d]"},{"why":"Provides a differentiable automated feature engineering baseline that the generative feature generation discussion extends.","marker":"Zhu et al. [2022]"},{"why":"Deep feature synthesis is an earlier automated feature generation approach against which newer generative and RL methods are positioned.","marker":"Kanter and Veeramachaneni [2015]"}],"fun_headline_variants":["Tabular feature engineering: RL vs generative trade-offs mapped","When to pick RL or generative for tabular features","Two-axis survey reveals RL and generative strengths for tabular","Data-centric AI for tabular: the RL-generative decision guide","RL vs generative: choosing the right tool for tabular features"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the papers selected for review—many written by the same author group—are representative of the broader field of RL- and generative-based tabular feature engineering; if the selection is one-sided, the taxonomy and the strengths-and-limitations comparison could be distorted.","fun_headline_variants_meta":{"raw":{"variants":["Tabular feature engineering: RL vs generative trade-offs mapped","When to pick RL or generative for tabular features","Two-axis survey reveals RL and generative strengths for tabular","Data-centric AI for tabular: the RL-generative decision guide","RL vs generative: choosing the right tool for tabular features"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000775,"raw_usage":{"total_tokens":3394,"prompt_tokens":874,"completion_tokens":2520,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":2437}},"tokens_in":490,"tokens_out":2520,"duration_ms":16924,"temperature":1.0,"reasoning_tokens":2437,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T23:32:41.681739+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic literature search followed by a benchmark would settle it: if methods outside the surveyed set populate categories the taxonomy lacks, or if on a fixed collection of tabular datasets the RL-versus-generative performance ranking contradicts the survey's stated trade-offs, the central claim is falsified.","supporting_citations":[{"cited_title":"An empirical analysis of feature engineering for predictive modeling","cited_arxiv_id":null,"evidence_quote":"Establishes that feature engineering significantly impacts predictive modeling, motivating the survey's focus."},{"cited_title":"Beyond Discrete Se- lection: Continuous Embedding Space Optimization for Generative Feature Selection","cited_arxiv_id":null,"evidence_quote":"Proposes the encoder-decoder-evaluator paradigm for generative feature selection, anchoring the generative category."},{"cited_title":"Difer: differentiable automated feature engineer- ing","cited_arxiv_id":null,"evidence_quote":"Provides a differentiable automated feature engineering baseline that the generative feature generation discussion extends."},{"cited_title":"Deep fea- ture synthesis: Towards automating data science endeav- ors","cited_arxiv_id":null,"evidence_quote":"Deep feature synthesis is an earlier automated feature generation approach against which newer generative and RL methods are positioned."}],"review_version":1}