{"id":"b5ced652-74f0-4880-9141-d5993953b558","arxiv_id":"2604.18227","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":3.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"FSEVAL is a unified toolbox and dashboard for comprehensively evaluating and visualizing feature selection algorithms across supervised and unsupervised settings.","lead":"The paper introduces FSEVAL, a software toolbox and visualization dashboard for evaluating feature selection algorithms in machine learning. It provides a standardized framework so researchers can compare feature selection methods more easily and comprehensively.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"No significant objection identified: this is a software toolbox paper reviewed at abstract-only level; the central claim is descriptive, not quantitative, and cannot be meaningfully stress-tested without the full text and code repository.","rationale":"The reader's verdict of UNVERDICTED with LOW confidence is correct for an abstract-only review of a software toolbox paper. The abstract contains no quantitative claims, theorems, or empirical results to stress-test. The reader correctly identified the load-bearing premise (coverage breadth and metric correctness) and correctly noted that these cannot be assessed without the full text and code. I agree with this assessment. No adjustment to the verdict is warranted. The only additional observation I would make is that the abstract's failure to enumerate supported algorithms, metrics, or comparisons with existing toolboxes is a presentation gap that, if reflected in the full paper, would weaken the contribution's novelty claim. But this is speculative without the full text. The concrete test I propose — verifying metric implementations against ground-truth synthetic data and checking whether the toolbox adds evaluation infrastructure beyond scikit-learn's existing feature_selection module — is the single most informative check once the full text and code are available.","tokens_in":1416,"tokens_out":1049,"duration_ms":37625,"concrete_test":"Obtain the full text and the code repository. Verify: (1) that at least one supervised and one unsupervised feature selection metric is implemented and matches a reference computation on a small synthetic dataset with known ground truth; (2) that the supported algorithm list is not merely a re-export of sklearn.feature_selection without added evaluation infrastructure. If both checks pass, the 'standardized, comprehensive' claim has初步 support; if either fails, the value proposition weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader correctly identifies that the load-bearing premise is whether FSEVAL's metric implementations are correct and its algorithm/metric coverage is broad enough to justify 'comprehensive' and 'standardized.' However, this is fundamentally an abstract-only review of a software tool paper. The abstract makes no falsifiable quantitative claim — it describes a toolbox and its intended purpose. There is no equation, theorem, or empirical result to scrutinize. The honest assessment is that no load-bearing concern can be identified from the abstract alone. The real risks (metric computation bugs, narrow algorithm coverage, thin wrapper over scikit-learn, lack of unit tests, no comparison with existing FS evaluation frameworks) are all verifiable only by inspecting the full text and the shipped code. The reader's UNVERDICTED verdict with LOW confidence is the appropriate posture given the available information. The only observation worth noting is that the abstract is notably vague: it does not enumerate which algorithms, metrics, or data regimes are supported, nor does it position FSEVAL against existing toolboxes. This vagueness is a presentation weakness rather than a logical flaw in any argument.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"The manuscript proposes FSEVAL, a feature selection evaluation toolbox with an accompanying visualization dashboard. The stated goal is to provide a standardized, unified framework for comprehensively evaluating feature selection algorithms across both supervised and unsupervised settings. The paper is positioned as a software tool paper for researchers in the field. Only the abstract was available for review; the full text and code repository were not provided. Consequently, this report is based solely on the abstract and cannot assess the correctness of metric implementations, the breadth of algorithm and metric coverage, the quality of the visualization dashboard, or the software engineering practices (testing, documentation, reproducibility).","tokens_in":1783,"tokens_out":670,"duration_ms":73284,"significance":"If the toolbox delivers on its stated goals, it could be a useful contribution to the feature selection community by lowering the barrier to rigorous, standardized evaluation. However, the abstract makes no falsifiable quantitative claim and provides no evidence of the toolbox's scope, correctness, or adoption. The significance of a software tool paper rests entirely on execution details that are not available for assessment. No machine-checked proofs, reproducible code artifacts, parameter-free derivations, or falsifiable predictions are evident from the abstract. The central value proposition — that FSEVAL enables 'comprehensive' and 'standardized' evaluation — is a claim that requires substantiation through the full paper and code repository.","major_comments":[{"comment":"The full text of the manuscript was not available for review. Only the abstract was provided. For a software toolbox paper, the load-bearing claims concern metric implementation correctness, breadth of algorithm and metric coverage, usability of the dashboard, comparison with existing FS evaluation frameworks, and software engineering quality (unit tests, documentation, reproducibility). None of these can be assessed from the abstract alone. The manuscript cannot be meaningfully evaluated without the full text and access to the code repository. This is the primary obstacle to providing a verdict.","section":null},{"comment":"Even at the abstract level, the claims of 'comprehensive' and 'standardized' evaluation are unsubstantiated. The abstract does not enumerate which feature selection algorithms are supported, which evaluation metrics are included, which data regimes are covered, or how FSEVAL compares to existing toolboxes (e.g., scikit-feature, ASU's feature selection repository). For a tool paper, these details are essential to justify the central value proposition. The abstract should at minimum specify the scope of coverage and position the tool against prior work.","section":null}],"minor_comments":[{"comment":"The abstract would benefit from concrete specifics: the number of supported algorithms, the number of evaluation metrics, the supported data types, and a brief comparison statement against existing feature selection repositories.","section":null},{"comment":"The phrase 'involved with discriminating redundant features from informative ones' is slightly awkward; consider rephrasing for clarity (e.g., 'concerned with distinguishing redundant features from informative ones').","section":null}],"recommendation":"uncertain","confidential_remarks":"The manuscript was submitted as an abstract-only submission; the full text and code repository were not available. I cannot provide a substantive review of a software tool paper without access to the full text and the code. If the editor can provide the full manuscript and a link to the code repository, I can complete a proper review. As things stand, any verdict other than 'uncertain' would be unjustified."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for the careful reading and for clearly identifying the core issue: the review was conducted on the abstract alone, with the full text and code repository unavailable. We agree that a software tool paper cannot be meaningfully evaluated without access to the full manuscript and code. We address each major comment below.","responses":[{"response":"The referee is entirely correct that the full text and code repository are essential for evaluating a software tool paper, and we agree that the abstract alone is insufficient. We believe this was a submission or review system issue rather than a deficiency in the manuscript itself: the full paper, supplementary materials, and the public code repository (with documentation, unit tests, and usage examples) were submitted as part of the package. We respectfully request that the full manuscript and repository be made available for review so that the referee can assess the specific execution details listed — metric implementation correctness, algorithm and metric coverage, dashboard usability, comparison with existing frameworks, and software engineering practices. We are confident that these materials substantiate the claims made in the abstract. That said, we acknowledge the referee's point that the abstract could be more informative even in isolation, and we will revise it accordingly (see our response to the second major comment).","revision_made":"partial","referee_comment":"The full text of the manuscript was not available for review. Only the abstract was provided. For a software toolbox paper, the load-bearing claims concern metric implementation correctness, breadth of algorithm and metric coverage, usability of the dashboard, comparison with existing FS evaluation frameworks, and software engineering quality (unit tests, documentation, reproducibility). None of these can be assessed from the abstract alone. The manuscript cannot be meaningfully evaluated without the full text and access to the code repository. This is the primary obstacle to providing a verdict."},{"response":"This is a fair criticism. Even granting that the full text was unavailable, the abstract should stand on its own and provide enough detail for a reader to understand the tool's scope and positioning. We will revise the abstract to (1) enumerate the categories of feature selection algorithms supported (e.g., filter, wrapper, embedded, and hybrid methods across supervised and unsupervised settings), (2) list the key evaluation metrics included (e.g., stability measures, clustering quality metrics, downstream classifier accuracy, redundancy analysis), (3) specify the data regimes covered, and (4) explicitly position FSEVAL against existing toolboxes such as scikit-feature and the ASU repository, noting what FSEVAL adds — namely, a unified dashboard for visualization and a standardized evaluation pipeline across both supervised and unsupervised settings. We agree that without these details the terms 'comprehensive' and 'standardized' are unsupported at the abstract level.","revision_made":"yes","referee_comment":"Even at the abstract level, the claims of 'comprehensive' and 'standardized' evaluation are unsubstantiated. The abstract does not enumerate which feature selection algorithms are supported, which evaluation metrics are included, which data regimes are covered, or how FSEVAL compares to existing toolboxes (e.g., scikit-feature, ASU's feature selection repository). For a tool paper, these details are essential to justify the central value proposition. The abstract should at minimum specify the scope of coverage and position the tool against prior work."}],"tokens_in":1150,"tokens_out":774,"duration_ms":59551,"standing_objections":["The referee's primary concern — that the full text and code were unavailable — cannot be resolved through revision alone. We believe the full manuscript and repository were submitted but were not accessible to the referee, likely due to a system or handling issue. We respectfully request that the complete submission package be made available for re-review. We cannot manufacture a response to substantive critiques of implementation correctness, coverage breadth, or software engineering quality when those critiques have not yet been raised against the actual content."]},"desk_editor":{"model":"glm-5.2","letter":"Here's the honest situation: I only have the abstract for this paper, and the abstract is thin. FSEVAL is a software toolbox for evaluating feature selection algorithms, with a visualization dashboard, covering both supervised and unsupervised settings. That's the entire substantive content I can see. There is no full text, no code link, no list of supported algorithms or metrics, no comparison to existing toolboxes, and no quantitative claim of any kind. So this letter is necessarily limited — I'm telling you what I can and cannot infer, not giving you a verdict on the work itself. What the paper appears to offer: a unified evaluation framework for feature selection that bundles multiple metrics and algorithms under one interface with a dashboard. If the implementation is solid, broad, and well-tested, this is genuinely useful infrastructure. Feature selection is a fragmented subfield where people routinely hand-roll their own evaluation pipelines, and a well-maintained, open-source toolbox with correct metric implementations would save real time and improve comparability across papers. That is a legitimate contribution for a tool paper. The problem is that the abstract gives me no way to assess whether any of that holds. It does not enumerate which algorithms or metrics are supported. It does not mention unit tests, benchmark datasets, or comparison with scikit-feature, scikit-learn's feature_selection module, or other existing frameworks. It does not say whether the code is open-source or where to find it. The word 'comprehensive' appears twice but is not substantiated. For a tool paper, the abstract needs to at minimum signal coverage breadth and code availability — this one does neither. The reader's UNVERDICTED verdict with low confidence is the right call given what we have. I disagree with nothing in the stress-test note; it correctly identifies that the real risks (metric bugs, narrow coverage, thin wrapper over existing libraries) are all uncheckable from the abstract. This is not a flaw in the paper's logic — there is no argument to stress-test — but it is a presentation weakness. A tool paper that does not foreground its code repository and coverage in the abstract is making it hard for reviewers to do their job. My recommendation: this deserves a serious referee only if the full text ships open code with documented algorithm/metric coverage and evidence of correct implementation (unit tests, benchmark comparisons, or at least a clear API specification). Without that, there is nothing to review. If the full text does include those things, I would want a referee who can actually run the toolbox and check metric correctness against known baselines. Set would_accept_peer_review=true — the paper's value hinges entirely on details I cannot see, and a referee with full access is the right person to make that call.","headline":"Abstract-only review of a feature selection evaluation toolbox; cannot assess implementation quality or novelty from available material","tokens_in":2010,"tokens_out":647,"would_cite":false,"duration_ms":43689,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Toolbox standardizes how feature selection algorithms are judged","keywords":["feature selection","evaluation toolbox","machine learning","supervised learning","unsupervised learning","visualization dashboard","reproducibility"],"falsifier":"If independent users find that FSEVAL's metric coverage omits widely used evaluation criteria, or that its results disagree with manually computed benchmarks on standard datasets, the claim of standardized comprehensive evaluation would not hold.","tokens_in":1506,"feed_emoji":"🔧","tokens_out":793,"duration_ms":21473,"temperature":0.7,"pith_summary":"The paper introduces FSEVAL, a software toolbox and visualization dashboard designed to standardize the evaluation of feature selection algorithms. Feature selection — the process of identifying which input variables carry useful signal while discarding redundant ones — is used in both supervised and unsupervised machine learning, but practitioners lack a unified way to compare methods. FSEVAL bundles multiple evaluation metrics and a visual interface so that researchers can run comprehensive, side-by-side assessments of different feature selection algorithms without each team reimplementing its own evaluation pipeline. The central claim is that a single, well-designed toolbox can make feature selection evaluation more consistent, reproducible, and accessible across the field.","feed_headline":"One toolbox to standardize feature-selection evaluation","feed_subtitle":"FSEVAL bundles supervised and unsupervised metrics into a single dashboard so researchers stop reinventing evaluation pipelines.","key_machinery":"FSEVAL is a software toolbox paired with a visualization dashboard. It integrates multiple evaluation metrics applicable to both supervised and unservised feature selection, and presents results through a unified interface so that different algorithms can be compared on common ground.","core_discovery":"The paper's central contribution is the FSEVAL toolbox itself: a unified evaluation and visualization platform that covers both supervised and unsupervised feature selection, providing standardized metrics and a dashboard so that comparisons between algorithms are no longer ad hoc. The load-bearing premise is that the set of metrics, algorithms, and data handling routines bundled into FSEVAL is broad and correct enough to justify calling the evaluation comprehensive and standardized.","pith_inferences":["The value of FSEVAL depends on community adoption: a standardization tool only standardizes if enough researchers use it, so the paper's impact is partly sociological rather than purely technical.","If the toolbox is open and extensible, the most useful long-term contribution may be the evaluation protocol itself rather than any particular implementation, since the protocol could be re-implemented in other frameworks.","Comparing supervised and unsupervised feature selection within one toolbox raises the question of whether there exist universal quality measures for selected feature subsets that are independent of downstream task type — a question the toolbox could empirically surface even if it does not resolve it."],"forward_implications":["If adopted, FSEVAL could reduce the reproducibility gap in feature selection research by giving every group the same evaluation baseline.","A standardized evaluation toolbox makes it easier to detect when a newly proposed feature selection method is merely re-deriving an existing approach under different packaging.","The dashboard lowers the barrier for practitioners in applied domains who need to choose a feature selection method but lack the expertise to assemble a full evaluation suite.","Unified metrics across supervised and unsupervised settings could reveal whether algorithms that perform well in one regime transfer to the other."],"fun_headline_variants":["FSEVAL bundles feature selection evaluation into one dashboard","A single dashboard for evaluating feature selection algorithms","Unifying supervised and unsupervised feature selection evaluation","FSEVAL standardizes feature selection algorithm comparisons","One dashboard to standardize feature selection evaluation"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The paper assumes that one toolbox can meaningfully cover the diverse landscape of feature selection algorithms, metrics, and data regimes well enough to justify calling its evaluation comprehensive and standardized. If the coverage is too narrow or the metrics are implemented incorrectly, the central value proposition weakens.","fun_headline_variants_meta":{"raw":{"variants":["FSEVAL bundles feature selection evaluation into one dashboard","A single dashboard for evaluating feature selection algorithms","Unifying supervised and unsupervised feature selection evaluation","FSEVAL standardizes feature selection algorithm comparisons","One dashboard to standardize feature selection evaluation"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":908,"prompt_tokens":381,"completion_tokens":527,"prompt_tokens_details":null},"tokens_in":381,"tokens_out":527,"duration_ms":14636,"temperature":1.0,"reasoning_tokens":562,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-05T12:36:32.069274+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If independent users find that FSEVAL's metric coverage omits widely used evaluation criteria, or that its results disagree with manually computed benchmarks on standard datasets, the claim of standardized comprehensive evaluation would not hold.","supporting_citations":[],"review_version":2}