{"id":"a6c4f7de-99f0-448a-bfdf-08b5d95b7033","arxiv_id":"1908.08006","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":0.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A book chapter that reviews evolutionary algorithms and feature extraction techniques for high-dimensional data, but delivers no new results and does not fulfill its promised formal definition of the curse of dimensionality.","lead":"This chapter surveys evolutionary and nature-inspired algorithms for feature selection and optimization in data science, including GA, ABC, PSO, ACO, GWO, and COA. It is an expository review with no new experiments or formal derivations.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The promised 'clear and formal definition' of the curse of dimensionality is absent from Section 2.3; the section is informal prose with no mathematical characterization, so the central claim of the chapter is unsupported.","rationale":"The reader's verdict is REJECT, and my independent reading agrees that the paper does not meet the standard for acceptance. The reader's weakest_assumption focused on the fidelity of the algorithm pseudocode, which is a real concern given the corrupted placeholders in Algorithms 4 and 8 and the missing update equations in Algorithm 3. However, the most load-bearing gap is even more direct: the chapter's headline promise of a 'clear and formal definition' of the curse of dimensionality is never fulfilled. Section 2.3 contains only informal prose and references to other papers. This is not a disagreement with the scientific consensus about CoD; it is an internal failure to deliver the stated contribution. A survey can be informal, but it cannot claim to provide a formal definition and then omit one. The malformed pseudocode is a secondary but compounding issue: it makes the overview of evolutionary algorithms unreliable, yet that part could in principle be repaired. The formal-definition absence is the central claim's weakest point. Since the reader already rejected the paper, my analysis does not change the verdict; it sharpens the primary reason for rejection.","tokens_in":17551,"tokens_out":2491,"duration_ms":27150,"concrete_test":"Run a systematic extraction of all formal mathematical objects from Section 2.3 (and any later section that revisits CoD): numbered equations, displayed formulas, 'Definition' environments, theorem/proposition statements, or explicit set-theoretic/statistical/complexity-theoretic characterizations. If the extraction yields zero such objects, the abstract's promise of a 'clear and formal definition of the CoD problem' is not met by the manuscript.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and introduction promise that the chapter will 'develop a clear and formal definition of the CoD problem.' The central value of the chapter depends on that definition existing somewhere in the text. Section 2.3, titled 'Curse of Dimensionality,' contains only informal statements such as 'Curse of dimensionality is related to the fact that the input data is too huge that no human being can analyze it' and 'This overabundance of data is called the curse of dimensionality.' No equation, formal definition, theorem, or precise complexity-theoretic or statistical statement is provided. The section is a short prose discussion with references to Altman and Krzywinski [7] and Guo et al. [8], but those citations do not supply the promised definition within this chapter. Because the primary advertised contribution is this formal definition, its absence is a load-bearing failure: even if every algorithm overview were corrected, the central claim as stated would still be unfulfilled. The same section also contains unsupported and incorrect claims, such as the assertion that CoD is 'due to the large amount of generated/sensed/collected data' rather than to the geometry and sampling properties of high-dimensional spaces, which further undermines the survey's reliability. The pseudocode listings in Section 4 contain corrupted placeholder tokens (e.g., '/v.alt', '/afii10069.ital', '/y.alt') and syntactically malformed lines, especially in Algorithms 4 and 8, but the formal-definition gap is more directly connected to the chapter's stated main contribution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey chapter on evolutionary algorithms (EAs) and other metaheuristics for data science, focusing on the curse of dimensionality (CoD), feature extraction and selection, and dimension reduction. The abstract and introduction promise a 'clear and formal definition' of CoD, followed by a survey of feature extraction techniques and an overview of EA families including GA, GP, EP, ABC, PSO, ACO, GWO, COA, CSO, and FSA, with pseudocode for feature selection. The paper also claims to show how nature-inspired algorithms can address CoD in large-scale engineering and science problems, and it references several applications including the authors' own prior work on image steganalysis. The manuscript appears to be an early version of a book chapter, referencing a companion chapter [67] for further details.","tokens_in":17879,"tokens_out":7367,"duration_ms":68392,"significance":"If the survey were accurate and self-contained, it could serve as a useful introductory reference for practitioners applying evolutionary algorithms to feature selection and dimension reduction in high-dimensional data. The authors have assembled a broad bibliography and cover a wide range of algorithms, which is potentially helpful for readers seeking an entry point. However, the paper's central advertised contribution—the formal definition of CoD—is not delivered, and the manuscript contains numerous factual errors, corrupted pseudocode, and incomplete placeholder citations. These issues are not local presentation problems; they undermine the reliability of the survey as a reference. The paper offers no new experimental results or reproducible code, and its educational value is currently limited because readers cannot trust the algorithm descriptions or the historical and technical claims.","major_comments":[{"comment":"The abstract promises 'a clear and formal definition of the CoD problem,' but Section 2.3 provides only informal prose, describing CoD as 'related to the fact that the input data is too huge that no human being can analyze it' and equating it with an 'overabundance of data.' No mathematical definition, complexity statement, or statistical characterization is given. Because the formal definition is the paper's primary advertised contribution, its absence is a load-bearing failure.","section":"Section 2.3 / Abstract"},{"comment":"The pseudocode listings contain corrupted placeholder tokens from a broken font encoding, such as '/v.alt', '/afii10069.ital', and '/y.alt', as well as malformed equations, for example Algorithm 4's input line 'K}≥ 1', Algorithm 6's line 'usin/afii10069.ital Pe = 0.005· N 2c', and Algorithm 8's lines 7 and 17. These listings cannot be executed and are not faithful representations of the cited algorithms, so the overview misleads readers who rely on the pseudocode to understand or implement the methods.","section":"Section 4, Algorithms 2, 4, 6, 8"},{"comment":"The claim that 'Evolutionary algorithms (EAs) is invented not more than 28 years' is historically inaccurate and is contradicted by the paper's own references, which include Fogel's 1965 work on evolutionary programming [39], Koza's 1992 genetic programming book [37], and the historical treatment in [12]. Additionally, Table 2 defines 'SVD' as 'singular value dimension' instead of 'singular value decomposition.' These factual errors reduce the survey's credibility.","section":"Section 2.6 / Table 2"},{"comment":"The manuscript contains unfinished placeholder citations, notably 'cite all papers from 2.4 here [9, 1]', 'cite all papers from 2.5 here [10, 11]', and 'cite all papers from 2.6 here [12, 13, 14, 15, 13]'. This is direct evidence that the manuscript is not complete and has not been prepared for formal review. A published survey must not include author instructions to itself.","section":"Sections 2.2 and 2.4"}],"minor_comments":[{"comment":"The statement that CoD is 'due to the large amount of generated/sensed/collected data' is misleading; the curse of dimensionality concerns the difficulties of estimation, sampling, and optimization in high-dimensional spaces, not merely data volume.","section":"Section 2.3"},{"comment":"The example expression '4 ∗ tan(x) +/y.alt2' contains an encoding artifact and should be written with proper mathematical notation, e.g., 4*tan(x) + y^2.","section":"Section 4.2.2"},{"comment":"The sentence 'PCA completely is used to generate a new dimension using a certain formula' is vague, and the claim that PCA fails for 'circle-based and sine or cosine-based distribution of instances' needs a concrete example or a citation to be informative.","section":"Section 3"},{"comment":"The text says 'most of research studies are accomplished using SSA' but the intended abbreviation is likely SSGA; SSA is never defined, and the surrounding discussion refers to steady-state genetic algorithms.","section":"Section 2.6.1"}],"recommendation":"reject","confidential_remarks":"This manuscript appears to be an unfinished draft of a book chapter: it contains placeholder citations, encoding artifacts, and numerous factual errors, and it fails to deliver the formal definition of CoD promised in the abstract. In my view, these problems go beyond what a standard revision can address; the chapter needs a substantial rewrite and careful technical checking before it could be considered for publication. I would encourage the authors to remove or fulfill the promise of a formal CoD definition, correct the pseudocode and factual errors, and complete the citations before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zach/Tara — quick take on 1908.08006. It is a review chapter, not a research paper: no new method, theorem, dataset, or experiment. The abstract promises “a clear and formal definition” of the curse of dimensionality, and that promise is the closest thing to a contribution. Section 2.3 does not deliver it — no equation, no theorem, no precise statistical statement. It says CoD is “related to the fact that the input data is too huge that no human being can analyze it” and “due to the large amount of generated/sensed/collected data.” That is not the curse of dimensionality; it is a hand-wavy description of big data. This is a load-bearing gap, not a missing flourish.\n\nWhat the chapter does reasonably: it organizes feature extraction into auto-encoder, feature selection, and feature generation, and it walks through ABC, PSO, ACO, GWO, COA, CSO, FSA with pseudocode. That structure is fine for a beginner, and the reference list draws on the right literature, including the authors’ own prior work on image steganalysis and their companion chapter. But the execution undermines the survey. The pseudocode in Algorithms 4 and 8 contains literal placeholder tokens (“/v.alt”, “/afii10069.ital”, “/y.alt”) and malformed lines; Algorithm 1 has an unexecutable output line; the abbreviations table defines SVD as “singular value dimension”; and the text claims EAs were invented “not more than 28 years” ago, which is wrong by decades. Those are not style quibbles. A reader trying to implement or teach from the pseudocode will be misled.\n\nThe stress-test note is correct: the formal-definition gap is the central issue. Everything else — the placeholders, the factual errors — are symptoms of the same lack of polish. There is no new result to salvage, and the survey would need substantial rewriting to be trustworthy. I would not cite it for the CoD definition, and I would not send it to a serious referee in this state. A desk reject with an invitation to resubmit a thoroughly revised version, or a clear rejection, is the right call.","headline":"A review chapter that promises a formal definition of the curse of dimensionality and never delivers; the placeholder-laden pseudocode and factual errors make it unsuitable even as an educational resource.","tokens_in":18347,"tokens_out":2751,"would_cite":false,"duration_ms":24822,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A review chapter argues that the curse of dimensionality is a feature-space failure and that evolutionary algorithms, embedded in feature extraction, are the practical way to escape it.","keywords":["evolutionary algorithms","curse of dimensionality","feature extraction","feature selection","metaheuristics","dimension reduction","data science","swarm intelligence"],"falsifier":"Re-implement each presented algorithm, especially the ACO and FSA listings whose equations contain placeholder symbols, from the pseudocode alone and run them on a benchmark dataset with a known feature-selection result; if any algorithm cannot be executed or fails to reproduce the cited method's reported behavior, the chapter's claim to provide an accurate overview is falsified.","tokens_in":1671,"feed_emoji":"🧬","tokens_out":2859,"duration_ms":75832,"temperature":0.7,"pith_summary":"This paper is a tutorial-style argument that the curse of dimensionality, the performance and time-complexity collapse caused by too many features in high-dimensional data, can be attacked inside the preprocessing step by evolutionary algorithms. It defines CoD as a feature-space disease marked by data sparsity, multiple testing, overfitting, and high time complexity, then separates feature extraction into auto-encoder dimension reduction, feature selection, and feature generation. The chapter's practical claim is that nature-inspired metaheuristics such as genetic algorithms, particle swarm optimization, artificial bee colony, ant colony optimization, grey wolf optimizer, and coyote optimization algorithm can find near-optimal feature subsets when classical gradient-based learners would settle in local optima. A sympathetic reader takes away a map of the field and a set of pseudocode templates for embedding these algorithms into a classifier pipeline.","feed_headline":"Evolutionary algorithms take on the curse of dimensionality","feed_subtitle":"A survey maps high-dimensional data problems onto swarm and evolutionary search for feature extraction.","key_machinery":"The central machinery is the generic evolutionary-algorithm loop, presented as a reusable skeleton: represent a candidate feature subset, initialize a population, evaluate a fitness function, select parents, apply crossover and mutation, replace or update the population, and stop on a stall condition. Around this loop the chapter organizes its survey, with the same skeleton appearing through species-specific update rules in genetic algorithms, artificial bee colony, particle swarm optimization, ant colony optimization, grey wolf optimizer, and coyote optimization algorithm, each supplied with pseudocode for feature selection. The other load-bearing object is the three-way taxonomy of feature extraction, namely auto-encoder dimension reduction, feature selection with filter, wrapper, and embedded variants, and feature generation, because it determines where in the pipeline the evolutionary search is inserted.","core_discovery":"On the paper's own terms, the central claim is organizational and prescriptive: the curse of dimensionality in data science is best understood as a feature-extraction failure, and evolutionary algorithms are the class of optimizers suited to repair it. The chapter maintains that once data is represented by a large set of extracted features, CoD manifests as data sparsity, multiple testing, overfitting, and prohibitive time complexity; classical feature extraction such as PCA is insufficient because it assumes linear correlation and can destroy information on non-linear structures. Feature selection that keeps original feature values is therefore presented as the corrective, and wrapper-based selection in particular is said to outperform filter-based selection at the cost of time. The chapter further claims that evolutionary algorithms are adopted precisely when a problem suffers from placement in local optima rather than global ones, and that their stochastic nature is managed by running them twenty to thirty times and reporting the mean, yielding a stable near-optimal solution.","pith_inferences":["The chapter does not provide detailed experimental comparisons among the surveyed algorithms, so a benchmark with identical feature-selection representation would directly test which method is best for CoD.","The chapter's emphasis on running each algorithm twenty to thirty times and averaging implies that evolutionary results should be treated as random variables; reporting variance and confidence intervals alongside the mean is a testable extension.","The pseudocode contains placeholder symbols and malformed equations in some listings, so a reader who wants a reliable implementation should consult the original papers, and a clean re-derivation of each algorithm from the cited sources is a natural companion."],"forward_implications":["If CoD is a feature-space problem, then preprocessing, not the classifier, is the right place to apply optimization, and embedding evolutionary algorithms into feature extraction should lower time complexity while preserving or improving classifier accuracy.","If wrapper-based feature selection consistently outperforms filter-based selection, then accuracy-critical systems should accept the higher runtime and use classifier-driven fitness evaluation, with evolutionary algorithms searching the feature-subset space.","If the generic evolutionary-algorithm loop is reliable across species, then a practitioner can port the same representation, selection, crossover, and mutation logic to new problems by swapping the fitness function and the update rule.","If evolutionary algorithms are adopted only when local optima are the obstruction, then their value for a given dataset is diagnostic: a problem that is not locally trapping will not benefit from evolutionary search.","If the presented algorithms are domain-independent, then the same feature-selection machinery applies across engineering, medicine, network analysis, and image classification without redesigning the optimizer."],"supporting_citations":[{"why":"Supplies the curse-of-dimensionality concept and the framing that more data is beneficial until dimensionality becomes the obstacle.","marker":"[7]"},{"why":"Supports the claim that CoD brings multiple-testing problems and offers an approach to large-scale hypothesis testing.","marker":"[8]"},{"why":"Provides the categorization of nature-inspired computation that the chapter reorganizes into metaheuristic and evolutionary branches.","marker":"[9]"},{"why":"Supplies the coyote optimization algorithm, including its pack-based update and birth-and-death mechanism, that the chapter presents as a recent evolutionary method.","marker":"[13]"},{"why":"Provides an artificial bee colony based feature selection method used to show how ABC is adapted to discrete feature-selection problems.","marker":"[14]"},{"why":"Supports the discussion of big data classification and classifies particle swarm optimization into classical, scale-free, and binary versions.","marker":"[23]"},{"why":"Supplies the survey of artificial bee colony algorithms and their applications that grounds the chapter's claim that ABC is the most widely used bee-inspired method.","marker":"[40]"},{"why":"Gives the original particle swarm optimization algorithm that the chapter summarizes and positions among population-based metaheuristics.","marker":"[48]"},{"why":"Gives the original ant colony optimization method and grounds the chapter's description of pheromone-based stochastic search.","marker":"[56]"},{"why":"Supplies the grey wolf optimizer and the claim that it outperforms other evolutionary algorithms on large-scale engineering problems.","marker":"[59]"}],"fun_headline_variants":["Evolutionary search tames high-dimensional data","When local optima strike, evolution answers","Curse of dimensionality? Let evolution attack it","Feature selection gets an evolutionary edge"],"cache_read_input_tokens":20480,"weakest_assumption_plain":"The whole overview stands or falls on the assumption that the pseudocode and parameter descriptions in Section 4 faithfully represent the cited algorithms, because that is what a reader would use to implement them.","fun_headline_variants_meta":{"raw":{"variants":["Evolutionary search tames high-dimensional data","When local optima strike, evolution answers","Curse of dimensionality? Let evolution attack it","Feature selection gets an evolutionary edge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000152,"raw_usage":{"total_tokens":1212,"prompt_tokens":963,"completion_tokens":249,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":195}},"tokens_in":579,"tokens_out":249,"duration_ms":2718,"temperature":1.0,"reasoning_tokens":195,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:56:02.257151+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-implement each presented algorithm, especially the ACO and FSA listings whose equations contain placeholder symbols, from the pseudocode alone and run them on a benchmark dataset with a known feature-selection result; if any algorithm cannot be executed or fails to reproduce the cited method's reported behavior, the chapter's claim to provide an accurate overview is falsified.","supporting_citations":[{"cited_title":"The curse (s) of dimensionality","cited_arxiv_id":null,"evidence_quote":"Supplies the curse-of-dimensionality concept and the framing that more data is beneficial until dimensionality becomes the obstacle."},{"cited_title":"A reminiscent study of nature inspired computation","cited_arxiv_id":null,"evidence_quote":"Provides the categorization of nature-inspired computation that the chapter reorganizes into metaheuristic and evolutionary branches."},{"cited_title":"Coyote optimization algorithm: a new metaheuristic for global optimization problems","cited_arxiv_id":null,"evidence_quote":"Supplies the coyote optimization algorithm, including its pack-based update and birth-and-death mechanism, that the chapter presents as a recent evolutionary method."},{"cited_title":"Big data classiﬁcation using scale-free binary particle swarm optimization","cited_arxiv_id":null,"evidence_quote":"Supports the discussion of big data classification and classifies particle swarm optimization into classical, scale-free, and binary versions."},{"cited_title":"A comprehen- sive survey: artiﬁcial bee colony (abc) algorithm and applications","cited_arxiv_id":null,"evidence_quote":"Supplies the survey of artificial bee colony algorithms and their applications that grounds the chapter's claim that ABC is the most widely used bee-inspired method."},{"cited_title":"Eberhart","cited_arxiv_id":null,"evidence_quote":"Gives the original particle swarm optimization algorithm that the chapter summarizes and positions among population-based metaheuristics."},{"cited_title":"Ant colony optimization: a new meta-heuristic","cited_arxiv_id":null,"evidence_quote":"Gives the original ant colony optimization method and grounds the chapter's description of pheromone-based stochastic search."},{"cited_title":"Grey wolf optimizer","cited_arxiv_id":null,"evidence_quote":"Supplies the grey wolf optimizer and the claim that it outperforms other evolutionary algorithms on large-scale engineering problems."}],"review_version":1}