{"id":"faa0ddd0-a269-42f2-85f7-1089df9ee55d","arxiv_id":"2412.11386","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A descriptor-based model combining category-specific persistent homology with gradient boosting reports higher R2 than transformer baselines on eight MOF gas selectivity datasets, but the comparison uses non-identical datasets.","lead":"Metal-organic framework properties are predicted by a machine learning method that turns atomic structures into topological barcodes grouped by chemical element type. The method, called category-specific topological learning, is claimed to beat large transformer models on eight gas selectivity tasks while staying interpretable.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SOTA claim rests on uncontrolled comparisons: dataset sizes and splits differ from baselines (Table 4) and no uncertainty intervals are reported, so the reported R² improvements may be within noise.","rationale":"The reader's weakest assumption—that reported comparisons are fair—is exactly the load-bearing point. The paper's own Table 4 and Appendix B concede that dataset sizes differ across methods and that exact data details for MOFTransformer and PMTransformer were not provided. Because the reported R² gaps are small, an uncontrolled comparison cannot support the 'outperforming all previous results' claim. A concrete fix is to retrain or evaluate all baselines on identical data and splits and report uncertainty. The paper has real strengths: a clearly described geometric pipeline, a universal hyperparameter set, a 100-run repeatability protocol, and interpretability analysis (t-SNE, feature importance), which credit the method's plausibility. However, those strengths do not fix the comparison protocol. I agree with the conditional verdict: the paper should be published only if the controlled benchmark and data/code release substantiate the SOTA claim. If the controlled comparison shows overlapping intervals, the headline should be moderated to competitive accuracy rather than SOTA.","tokens_in":16614,"tokens_out":9637,"duration_ms":88733,"concrete_test":"Run the released descriptor-based model of Orhan et al. [27] and the pretrained MOFTransformer and PMTransformer on CSTL's exact filtered datasets from the source repository, using the same 80:10:10 random splits with the same 10 split seeds (23-32) and 10 model seeds (13-22) for a total of 100 runs per method. Compare the per-run R², MAE, and RMSE distributions. If the CSTL mean exceeds each baseline mean by more than the pooled standard error (or the 95% CI of the difference excludes zero), the outperformance claim is supported; if intervals overlap, the claim should be revised to parity under identical conditions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that CSTL outperforms all previous results depends on the fairness of the comparisons in Table 2. Appendix B and Table 4 show the compared datasets are not identical: CSTL uses, for example, 5132 samples for N2 uptake while MOFTransformer and PMTransformer report 5286, and the exact splits used by those baselines are not reproduced. Since the reported R² gaps are small (0.79 vs 0.78 for N2 uptake; 0.85 vs 0.83 for O2 uptake), differences in dataset filtering, sample composition, or random splits could account for the margin. The paper also reports only averaged metrics over 100 models without standard deviations or confidence intervals, so it is unknown whether the margin is statistically significant. This is load-bearing because the abstract's headline claim of outperforming all previous results is not established without a controlled benchmark on identical data and splits. The authors acknowledge the dataset discrepancy in Appendix B but do not remediate it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces a category-specific topological learning (CSTL) pipeline for predicting eight N2/O2-related properties of metal-organic frameworks. The method constructs alpha complexes from MOF crystal structures, groups elements into eight chemically motivated categories (C0-C7) plus an all-atom set, computes persistent homology in dimensions 0-2 for each category, bins the resulting barcodes on a 0-25 Å grid with 0.1 Å resolution, and concatenates the bin counts into a 6750-dimensional feature vector. A gradient boosting tree regressor with fixed hyperparameters is trained on 80% of each dataset, and metrics are averaged over 100 train/model repetitions. The authors report R2 values between 0.79 and 0.85 across the eight datasets and claim that CSTL outperforms a descriptor-based model, MOFTransformer, and PMTransformer. The paper also presents t-SNE visualizations and tree-based feature importance analyses to support interpretability.","tokens_in":16797,"tokens_out":7030,"duration_ms":61921,"significance":"If the comparison to prior models were controlled, CSTL would be a valuable contribution: it is a compact, interpretable, descriptor-based alternative to large pretrained transformers, it requires no external pretraining corpus, and its descriptors are derived transparently from crystal structure topology rather than from learned representations. The element-category construction is chemically meaningful, and the feature importance analysis gives plausible physical interpretation (e.g., carbon-based cavities for uptake, overall cycles/cavities for diffusivity). I also found no target leakage in the descriptor construction: the element categories are based on elemental occurrence frequencies, not on the labels. However, the headline claim of state-of-the-art performance currently rests on comparisons with datasets of different sizes/filtering and without uncertainty quantification, so the significance of the empirical advantage is not yet established.","major_comments":[{"comment":"The central claim that CSTL 'outperforms all previous results' is not established because the benchmark comparisons are not controlled. The CSTL datasets are smaller than those used by MOFTransformer and PMTransformer (e.g., 5132 vs 5286 samples for N2 uptake and 5241 vs 5286 for O2 uptake), the outlier filtering differs, and the exact train/validation/test splits of the baselines are not reproduced or re-run on the CSTL splits. Since the reported R2 advantages are small (0.79 vs 0.78 for N2 uptake; 0.85 vs 0.83 for O2 uptake), differences in dataset composition or split could account for the margin. The authors acknowledge the size discrepancy in Appendix B but do not remediate it. I request a controlled evaluation on identical data and splits, or at least results on both dataset versions and an explicit discussion of how filtering affects each model.","section":"Section 3.1, Appendix B, Table 4"},{"comment":"The text states that CSTL 'consistently outperforms these models across all datasets,' but for four of the eight datasets (Henry's constants for N2/O2 and self-diffusivities at infinite dilution for N2/O2) the MOFTransformer and PMTransformer columns are blank, so there is no comparison to support the claim on those datasets. Either fill in these entries with values from the original papers or by re-running the baselines, or restrict the claim to the datasets for which comparisons exist.","section":"Table 2, Section 2.2"},{"comment":"The reported metrics are means over 100 models, but no standard deviations, confidence intervals, or significance tests are reported. The heatmaps in Figures S2-S4 are presented instead of quantitative uncertainty estimates. Given the small R2 differences in Table 2 and the fact that the models are trained on the same data with different seeds, the paper should report the spread of the 100 metrics (e.g., mean ± std or bootstrap CIs) to show whether the improvement over baselines is statistically meaningful.","section":"Appendix C, Table 2"}],"minor_comments":[{"comment":"The sentence 'The properties of interest, such as O2 and N2 selectivity, were simulated in earlier studies [25]' cites the CoRE MOF database reference [25]; the simulations are from Orhan et al. [27], which is cited later in the same paragraph. Please correct the citation.","section":"Section 3.1"},{"comment":"The phrase 'topological representations are constructed to capture the interactions among atoms across different categories' is imprecise, because each C_i sub-complex contains atoms of a single category; cross-category interactions enter only through the C_all complex. Please rephrase.","section":"Section 3.2"},{"comment":"The table mixes available metrics across methods (CSTL has r2/MAE/RMSE; Descriptor-based has r2/RMSE; MOFTransformer has r2/MAE; PMTransformer has only MAE) and leaves blanks. Clarify which metrics are available and move the full comparison to a supplementary table if needed.","section":"Table 2"},{"comment":"The column header 'CSTL(80% training, 20% test)' conflicts with the 80:10:10 split described in Section 3.3; if the second column is the 80% training / 20% test+validation holdout, say so explicitly.","section":"Appendix E, Table 5"},{"comment":"The data availability section points to the source repository for the datasets but does not provide code or exact preprocessing scripts for the CSTL descriptors; providing these would improve reproducibility.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and the method is well motivated. The empirical claim is currently stronger than the evidence; the controlled-benchmark request in the major comments is the main obstacle. If the authors re-run the baselines on identical data/splits and report uncertainty intervals, I would be supportive. There is no indication of target leakage or circularity in the descriptor construction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is sound and worth engaging: category-specific persistent homology, with elements grouped by chemical role and frequency, gives a modest but real extension of element-specific PH for MOF property prediction. The paper does a clean, systematic job on eight gas selectivity datasets, uses a single hyperparameter set, and reports averages over 100 models. The feature importance analysis is a nice touch—it gives some interpretability that transformers and graph networks often lack. I'd call this a solid incremental contribution.\n\nThe soft spot is the headline claim. The abstract says CSTL \"outperforms all previous results,\" but the evidence for that is not controlled. In Appendix B the authors themselves show that their dataset sizes differ from MOFTransformer and PMTransformer (e.g., 5132 vs 5286 for N2 uptake), and they do not reproduce the exact splits used by those baselines. The R2 gaps are small—0.79 vs 0.78, 0.85 vs 0.83—so without matching data or at least confidence intervals, the margin could easily be noise. They average over 100 models but report only means; the heatmaps suggest some spread, but no standard deviations are given. This is load-bearing because the paper's main pitch is state-of-the-art performance.\n\nA couple of minor issues: there's a typo in the abstract (\"topological leaning\"), and Section 3.1 refers to \"Table S1\" where it should point to Table 4 in Appendix B. Also, no code or data are released, which hurts reproducibility, though the datasets are public.\n\nThe categorization itself is hand-defined—C0–C7 based on valence electrons and frequency. That's reasonable, but there's no sensitivity analysis to show how much the choices matter. That's a minor concern, not a fatal one.\n\nOverall, the method is plausible and the study is clearly presented. The stress-test concern about uncontrolled comparisons holds up: the authors acknowledge the dataset discrepancy but do not remediate it. I think this deserves a serious referee, but only with major revision. The authors should either run the baselines on identical splits, or soften the \"outperforms all previous results\" claim to \"competitive,\" and report uncertainty intervals. If they do that, the paper becomes a useful contribution. If they don't, the headline is not supported.\n\nI'd send it to review—it's the kind of work that can be fixed with a proper re-benchmark.","headline":"A useful descriptor extension, but the SOTA claim rests on datasets and splits that don't match the baselines; needs a controlled re-benchmark and released code before I'd trust the headline.","tokens_in":17300,"tokens_out":2058,"would_cite":true,"duration_ms":20289,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["55N31","68T05"],"pacs":[],"model":"deepseek-v4-flash","headline":"By grouping atoms into chemical categories before computing persistent homology, CSTL predicts eight MOF gas properties with R2 up to 0.85, beating transformer models trained on millions of structures.","keywords":["metal-organic frameworks","persistent homology","topological data analysis","category-specific descriptors","gas selectivity","machine learning property prediction","gradient boosting","alpha complex"],"falsifier":"Retrain the descriptor-based model, MOFTransformer, and PMTransformer on the exact eight filtered datasets and 80/10/10 splits used for CSTL; if any baseline matches or exceeds its R2 on the same test folds, the central outperformance claim is refuted.","tokens_in":16420,"feed_emoji":"🧪","tokens_out":5372,"duration_ms":45201,"temperature":0.7,"pith_summary":"This paper introduces category-specific topological learning (CSTL), a descriptor method that predicts eight N2/O2-related properties of metal-organic frameworks from crystal structure alone. It represents each MOF as a simplicial complex, splits atoms into eight chemically meaningful element categories plus an all-atom category, and runs persistent homology on each category to produce barcode-derived feature vectors. A gradient boosting model on these 6750-dimensional descriptors reaches R2 values between 0.79 and 0.85 across the eight datasets, exceeding the descriptor-based, MOFTransformer, and PMTransformer baselines. The authors argue the method is interpretable: feature importance traces gas selectivity to carbon cavities and metal-node influence. If correct, CSTL offers a lighter, explainable route to accurate MOF property screening without pretraining on millions of structures.","feed_headline":"Element-aware topology beats transformers for MOF gas properties","feed_subtitle":"Persistent homology grouped by chemical role predicts N2/O2 selectivity with R2 up to 0.85.","key_machinery":"The central object is category-specific persistent homology on alpha complexes. An alpha complex is a simplicial complex grown from the Delaunay triangulation of atomic positions as a radius parameter increases; persistent homology tracks when connected components, loops, and cavities appear and disappear, summarized as barcodes. The paper's twist is to compute these barcodes separately for eight element categories (alkali and other metals C0, transition/lanthanide/actinide metals C1, metalloids C2, halogens C3, hydrogen C4, carbon C5, nitrogen/phosphorus C6, oxygen/sulfur/selenium C7) and for all atoms together. Binning each barcode over 0 to 25 angstroms with 0.1 angstrom steps produces a length-fixed vector, and concatenation gives a 6750-dimensional descriptor that a gradient boosting tree model maps to each target property. The categories make the topology chemically aware and keep rare elements from being swamped by abundant carbon and oxygen.","core_discovery":"The central claim is that category-specific persistent homology, rather than more data or deeper models, carries the predictive signal for MOF gas properties. For each elemental category C0 through C7 and the whole structure Call, the method builds an alpha complex filtration and records Betti numbers of H0, H1, and H2 on a 0 to 25 angstrom grid, yielding 750 features per category and 6750 features in total. Feeding these descriptors to a gradient boosting regressor gives test R2 values between 0.79 and 0.85 on eight datasets covering Henry constants, uptake, and self-diffusivity of N2 and O2, with lower MAE and RMSE than all compared prior models, using one fixed hyperparameter set and 100 repeated splits. The paper also shows t-SNE plots where CSTL features separate MOFs with extreme property values, and tree-based feature importance that assigns physical meaning to the categories.","pith_inferences":["If the element-category assignment is what supplies the chemical inductive bias, then coarsening the categories or randomizing the grouping should measurably lower R2; the paper does not run this ablation, but it is directly testable.","Because CSTL needs no pretraining corpus, it could serve as a computationally cheap screening baseline for hypothetical MOF libraries, where transformer models trained on existing structures may transfer poorly.","The comparison with prior models rests on datasets of slightly different sizes and unreported splits; a strict head-to-head with baselines retrained on CSTL's exact splits would reveal how much of the gap is the descriptor rather than data handling.","The fixed 0.1 angstrom binning and 25 angstrom cutoff assume structure-property signal lives in that range; properties governed by finer or longer-range features could escape the descriptor."],"forward_implications":["On the eight N2/O2 datasets, CSTL reports higher R2 and lower MAE/RMSE than the descriptor-based model, MOFTransformer, and PMTransformer, so it would become the benchmark to beat for these properties.","Because the descriptors come only from CIF structure files and a fixed hyperparameter set, CSTL can be applied to new MOF datasets without pretraining or fine-tuning.","The feature-importance analysis links specific categories to physics: carbon cycles and oxygen/sulfur spacing drive gas selectivity, while loops and cavities in the all-atom complex drive diffusivity.","The reported stability across 100 random splits and a 20 percent holdout suggests the accuracy gain is not an artifact of one lucky train/test partition."],"supporting_citations":[{"why":"Supplies the CoRE MOF 2019 database from which all eight datasets are drawn.","marker":"[25]"},{"why":"Provides the O2/N2 property datasets, outlier filtering thresholds, and the descriptor-based baseline.","marker":"[27]"},{"why":"Transformer baseline that CSTL outperforms and source of the 80:10:10 split convention.","marker":"[32]"},{"why":"Second transformer baseline that CSTL outperforms.","marker":"[33]"},{"why":"Foundational method for computing persistent homology used in the barcode pipeline.","marker":"[38]"},{"why":"Introduces element-specific persistent homology, the precursor idea CSTL generalizes to categories.","marker":"[39]"},{"why":"Defines the alpha complex filtration used to build category-specific topological representations.","marker":"[52]"},{"why":"Gradient boosting regressor implementation used for prediction.","marker":"[54]"}],"fun_headline_variants":["Element-aware topology beats deep baselines for MOF gas predictions","Category-specific persistent homology outperforms prior MOF models","Topological features grouped by element nail MOF gas property tests","CSTL: chemistry-informed topology tops MOF prediction benchmarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim of outperforming prior models assumes the comparison is fair: the baselines must be evaluated on the same filtered data and same train/test splits, otherwise differences in dataset size or composition, not the descriptors, could explain the accuracy gap.","fun_headline_variants_meta":{"raw":{"variants":["Element-aware topology beats deep baselines for MOF gas predictions","Category-specific persistent homology outperforms prior MOF models","Topological features grouped by element nail MOF gas property tests","CSTL: chemistry-informed topology tops MOF prediction benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000531,"raw_usage":{"total_tokens":2534,"prompt_tokens":902,"completion_tokens":1632,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":1563}},"tokens_in":518,"tokens_out":1632,"duration_ms":11441,"temperature":1.0,"reasoning_tokens":1563,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:58:36.463114+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the descriptor-based model, MOFTransformer, and PMTransformer on the exact eight filtered datasets and 80/10/10 splits used for CSTL; if any baseline matches or exceeds its R2 on the same test folds, the central outperformance claim is refuted.","supporting_citations":[{"cited_title":"Advances, updates, and analytics for the computation-ready, experimental metal– organic framework database: Core mof 2019","cited_arxiv_id":null,"evidence_quote":"Supplies the CoRE MOF 2019 database from which all eight datasets are drawn."},{"cited_title":"Prediction of o2/n2 selectivity in metal–organic frameworks via high-throughput computational screening and machine learning","cited_arxiv_id":null,"evidence_quote":"Provides the O2/N2 property datasets, outlier filtering thresholds, and the descriptor-based baseline."},{"cited_title":"A multi-modal pre-training transformer for universal transfer learning in metal–organic frameworks","cited_arxiv_id":null,"evidence_quote":"Transformer baseline that CSTL outperforms and source of the 80:10:10 split convention."},{"cited_title":"Enhancing structure–property relationships in porous materials through transfer learning and cross-material few-shot learning","cited_arxiv_id":null,"evidence_quote":"Second transformer baseline that CSTL outperforms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces element-specific persistent homology, the precursor idea CSTL generalizes to categories."},{"cited_title":"Smooth surfaces for multi-scale shape representation","cited_arxiv_id":null,"evidence_quote":"Defines the alpha complex filtration used to build category-specific topological representations."},{"cited_title":"Scikit- learn: Machine learning in python","cited_arxiv_id":null,"evidence_quote":"Gradient boosting regressor implementation used for prediction."}],"review_version":1}