REVIEW 4 major objections 6 minor 52 references
Adding aligned multi-resolution cluster assignments to any predictor yields consistent gains, raising a hard microbial regression from R² 0.6 to 0.98.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
2026-08-04 08:42 UTC pith:CVD5ZVN7
load-bearing objection Potentially useful plug-in feature augmentation with striking DSNI gains, but currently unverifiable: no code/data, unspecified image embeddings, and an absent formal characterization. the 4 major comments →
Mixing Configurations for Downstream Prediction
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Unsupervised multi-resolution community detection on a kNN graph of any frozen embedding yields a small set of stable configurations—clusterings at multiple valid resolutions that capture hierarchical structure discarded by single-resolution clustering. GraMixC extracts these configurations, aligns train- and test-time partitions with a Reverse Merge/Split (RMS) procedure that treats differing cluster counts as merges and splits, and fuses the aligned tokens via attention heads before passing them to any downstream predictor. The paper shows this consistently improves every baseline it tests, raising DSNI-pH R² from about 0.6 to 0.98 for the best configuration, and bringing a non-graph model
What carries the argument
Configurations: the finite set of stable hierarchical clusterings obtained by varying the resolution parameter in modularity-based community detection (BlueRed/parallel-DT) on a kNN graph. The argument is carried by three pieces: (1) multi-gamma community detection that automatically discovers all valid resolutions without ad-hoc parameter choice; (2) Reverse Merge & Split (RMS) alignment, which matches train and test partitions of different sizes using a two-walk Laplacian of a confusion matrix; and (3) attention heads that learn per-sample, per-task mixing weights over the aligned configuration tokens before concatenation with raw features.
Load-bearing premise
The kNN graph built on the given representation must encode task-relevant grouping: if the chosen features are unrelated to the target, the configurations derived from them add no signal. For CIFAR-10, the paper never specifies which embedding was used to build the graph, so the largest reported gains rest on an unstated feature choice.
What would settle it
Remove the cluster–task correlation: build configurations on a shuffled or random representation (or on features known to be irrelevant to the target) while keeping the same pipeline; if GraMixC still produces large gains, the improvements come from something other than the configurations themselves. A cheaper check: on a dataset with known cluster structure orthogonal to the label, GraMixC should show no improvement over baseline.
If this is right
- Any baseline predictor—from random forests to tabular transformers—improves when aligned multi-resolution configurations are concatenated (GC), and improves further when mixed via attention (GMC).
- On the DSNI 16S rRNA cultivation-media task, GraMixC raises R² for pH from roughly 0.6 to 0.98 for the best configuration, setting a new state of the art for growth-media prediction.
- Incrementally adding more configurations decreases error until a plateau, confirming that the finite set of resolutions carries complementary information rather than redundant noise.
- Configurations are less redundant than vision-transformer register tokens and track task changes in attention maps without retraining, suggesting they can serve as a general-purpose, interpretable feature augmentation.
- GraMixC is plug-and-play: it requires no retraining of the base predictor and works with frozen embeddings, so gains transfer across model architectures.
Where Pith is reading between the lines
- Beyond the paper: if the gains hold under proper ablation, the biggest practical windfall is in label-scarce scientific settings like microbial growth prediction, where 0.1% anchors suffice to align test partitions—pointing toward a cheap semi-supervised recipe.
- Beyond the paper: the CIFAR-10 gains (e.g., TabTransformer 46.3% to 87.6%) depend on the representation used to build the kNN graph, which the paper never states; a fair test would vary that embedding (random, pixel-level, supervised) and measure how much of the gain survives.
- Beyond the paper: because RMS alignment is done at inference with a small anchor set, the method could be stress-tested under domain shift—if train and test lie on different manifolds, the alignment may fail, making this a concrete testable extension.
- Beyond the paper: the attention weights over configurations may serve as an interpretability tool, since per-sample mixing reveals which resolution level matters for a given prediction and could expose dataset hierarchies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes GraMixC, a plug-and-play module that builds a kNN graph on an input representation, extracts multi-resolution community-detection partitions ('configurations') using BlueRed/parallel-DT, aligns train and test configurations with a proposed Reverse Merge/Split (RMS) procedure, and fuses them with attention heads before a downstream predictor. The authors report consistent improvements from adding static configurations (GC) and attention-mixed configurations (GMC) across tabular, molecular, vision, and text benchmarks, with the headline result being a large R² increase on a 16S rRNA cultivation-media prediction task. The paper also claims a formal characterization of configurations and draws analogies to ViT register tokens.
Significance. If the empirical pattern holds, the idea of using structurally stable multi-resolution clusterings as plug-and-play features is simple, novel, and potentially valuable for low-data biological and tabular problems. The DSNI-pH improvement (base R² around 0.6 to >0.95 for several predictors) is striking and the repeated-seed reporting in Table 1 is a strength. However, as submitted, the paper cannot be fully checked: there is no code/data release, the vision-benchmark graph-construction input is not specified, Table 2 lacks error bars, and the claimed formal characterization does not appear in the paper. The central claim of consistency is also contradicted by entries in the paper's own Table 2. These issues are fixable but require substantive revision.
major comments (4)
- [§4.2, Table 2] The claim that 'adding GC yields consistent gains' is contradicted by the paper's own Table 2. On CIFAR-10, TabTransformer+GC degrades relative to TabTransformer (CE 1.028→1.049; Acc 0.706→0.704) and FT-Transformer+GC degrades (Acc 0.874→0.870). On Boston Housing, TabTransformer+GMC drops from R² 0.811 to 0.671. Since Table 2 reports no repeated runs or error bars, it is impossible to tell whether these are noise or systematic exceptions. Please report seed statistics for all Table 2 entries and either revise the 'consistent gains' statement or explain these exceptions quantitatively.
- [§3.1, §4.1, Table 2] The input matrix X used to build the kNN graph is never specified for MNIST and CIFAR-10. Section 3.1 says only 'Given an input matrix X∈R^{N×d}', and §4.1 does not state whether X is raw pixels, a pretrained embedding, or another representation. This is not a cosmetic detail: Table 2's largest gains (e.g., TabN+GC on CIFAR-10 from 0.463 to 0.876) depend entirely on whether the graph carries task-relevant signal. If a strong frozen embedding (e.g., DINOv2/ResNet) is used while baselines consume raw pixels, the comparison is confounded; if raw pixels are used, the magnitude of the gains must be explained. Please specify the embedding for every benchmark and ideally run an ablation with and without the same embedding for the baselines.
- [Abstract; Section 2] The abstract claims that 'we formally characterize these configurations', but the paper contains no formal characterization, theorem, or precise definition of the three claimed properties beyond qualitative observations (attention maps, histograms in Figs. 2–3). The phrase 'register token can be considered as a latent configuration' is an analogy, not a characterization. Either provide precise formal statements with rigorous definitions/proofs or remove the 'formally characterize' claim from the abstract and introduction.
- [Table 2; §4.1] Table 2 reports a single run per configuration, unlike Table 1 which gives mean±std. Given that the paper's central claim is consistency across tasks and architectures, and that the one explicitly acknowledged counterexample involves a large drop (TabT+GMC on Boston: R² 0.811→0.671), the absence of error bars and seed information for Table 2 is a load-bearing reproducibility gap. Please report at least 5 seeds per entry (or explain computational constraints) and state whether the Boston drop is statistically significant.
minor comments (6)
- [Abstract; §4.2] The abstract says GraMixC improves R² 'from 0.6 to 0.9', but Table 1 shows a range of base R² values and best R² reaching 0.984. Please state the exact improvements more precisely.
- [§3.2] The description 'we carry a small portion (0.1%) of train samples as anchors during inference' is confusing. Clarify whether anchors are a fixed subset of the training set, how they are selected, and whether this introduces any test-time dependence.
- [§3.1, §3.2] Hyperparameters θ=0.1, λ=15, k=log10 N, and anchor fraction 0.1% are introduced without sensitivity analysis. The paper relies on 'previous work' for λ but not for θ or the anchor fraction; please provide at least a brief sensitivity study or cite a precedent.
- [§4.3, Figure 8] The 'same embedding budget' in the comparison of PCA/UMAP/AE versus GC/GMC is not precisely defined. Please state the number of embedding dimensions and all configuration levels used so the comparison is reproducible.
- [Table 2] The table caption and §E.1 state that the sole exception is TabTransformer on Boston, but CIFAR-10 entries for TabT+GC and FTT+GC also show small degradations. Please correct the summary text.
- [General] No code or data link is provided. Given that the method involves several nonstandard components (BlueRed/parallel-DT, RMS alignment, SG-t-SNE reweighting), a public implementation would greatly aid verification.
Circularity Check
No significant circularity: configurations are label-free inputs; labels only train the attention/predictor, and the reported gains are empirical, not fitted targets renamed as predictions.
full rationale
The claimed derivation chain is not circular in the sense defined here. Configurations are produced from the input matrix X by kNN graph construction, SG-t-SNE reweighting, and modularity/BlueRed multi-gamma community detection, none of which uses the downstream labels; the labels enter only in the supervised training of the attention heads and the final predictor, which is standard supervised learning and not an equivalent-by-construction prediction. The method does not fit a parameter to a subset of the target and then report a closely related quantity as a prediction; the baseline/+GC/+GMC comparisons are honest ablations on held-out test sets. No load-bearing self-citation appears: the citations to Pitsianis et al., Floros et al., and Darcet et al. are external prior work, and no uniqueness theorem from the present authors is invoked to force the choice of configurations. The qualitative Section 2 attention-map demonstration is partly by construction—attention heads are trained on the labels before showing task-dependent attention—but it is a motivating visualization, not one of the reported prediction results, so it does not affect the circularity score. The principal weaknesses are descriptive/reproducibility issues (e.g., the embedding X used to build the kNN graph for MNIST/CIFAR-10 is not specified, leaving a possible confound between the embedding and the configuration module), not circularity.
Axiom & Free-Parameter Ledger
free parameters (4)
- kNN neighborhood k=log10 N =
log10 N
- SG-t-SNE reweighting λ =
15
- RMS balance θ =
0.1
- anchor fraction =
0.1%
axioms (4)
- domain assumption BlueRed/parallel-DT returns a finite set of valid configurations Γ and that these configurations are the relevant hierarchical clusterings
- domain assumption Euclidean kNN graph preserves the latent manifold structure of the embedding
- ad hoc to paper RMS alignment using 0.1% anchors and the two-walk Laplacian yields semantically correct train/test cluster correspondence
- domain assumption 7-mer count vectors are a sufficient representation of 16S rRNA for growth-condition prediction
Cite this review
Pith. "Pith review of Mixing Configurations for Downstream Prediction." pith.science (2026). https://pith.science/paper/CVD5ZVN7
@misc{pith2026251019248,
author = {Pith},
title = {Pith review of: Mixing Configurations for Downstream Prediction},
year = {2026},
howpublished = {\url{https://pith.science/paper/CVD5ZVN7}},
note = {Machine review of arXiv:2510.19248}
}
read the original abstract
Clustering-based features are widely used in machine learning, but most methods must choose a resolution -- a choice that is global, fixed, and ad hoc. Recent work shows that varying the resolution parameter produces only a finite set of structurally stable partitions, known as configurations. Based on this, we introduce Configuration-Mixed Prediction (CMP), a setting where models learn to adaptively weight these configurations per sample for downstream prediction. We propose MixConfig, a plug-and-play feature augmentation module that extracts configurations from any frozen embedding and learns energy-aware mixing weights via a novel selector that jointly reasons about sample context, cluster assignments, and stability statistics. Experiments across tabular, molecular, vision, and text domains demonstrate consistent improvements over single-resolution and static baselines across diverse predictor architectures, with gains particularly pronounced in low-data regimes.
Figures
Reference graph
Works this paper leans on
-
[1]
Some Methods for Classification and Analysis of Multivariate Observations
J. MacQueen. “Some Methods for Classification and Analysis of Multivariate Observations”. In:Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Statistics. V ol. 5.1. University of California Press, Jan. 1, 1967, pp. 281–298
1967
-
[2]
Normalized Cuts and Image Segmentation
Jianbo Shi and J. Malik. “Normalized Cuts and Image Segmentation”. In:IEEE Trans. Pattern Anal. Machine Intell.22.8 (Aug. 2000), pp. 888–905
2000
-
[3]
On Spectral Clustering: Analysis and an Algorithm
A. Ng, M. Jordan, and Y . Weiss. “On Spectral Clustering: Analysis and an Algorithm”. In:Advances in Neural Information Processing Systems. V ol. 14. MIT Press, 2001
2001
-
[4]
Perceptual Cues That Permit Categorical Differentiation of Animal Species by Infants
P. C. Quinn and P. D. Eimas. “Perceptual Cues That Permit Categorical Differentiation of Animal Species by Infants”. In:J Exp Child Psychol63.1 (Oct. 1996), pp. 189–211. PMID:8812045
1996
-
[5]
Infant Object Categorization Transcends Diverse Object- Context Relations
M. H. Bornstein, M. E. Arterberry, and C. Mash. “Infant Object Categorization Transcends Diverse Object- Context Relations”. In:Infant Behav Dev33.1 (Feb. 2010), pp. 7–15. PMID:20031232
2010
-
[6]
Lessons from Infant Learning for Unsupervised Machine Learning
L. Zaadnoordijk, T. R. Besold, and R. Cusack. “Lessons from Infant Learning for Unsupervised Machine Learning”. In:Nat Mach Intell4.6 (June 2022), pp. 510–520
2022
-
[7]
L. Muttenthaler, K. Greff, F. Born, B. Spitzer, S. Kornblith, M. C. Mozer, K.-R. Müller, T. Unterthiner, and A. K. Lampinen.Aligning Machine and Human Visual Representations across Abstraction Levels. Oct. 29, 2024. arXiv: 2409.06509 [cs].URL:http://arxiv.org/abs/2409.06509(visited on 05/11/2025). Pre-published
Pith/arXiv arXiv 2024
-
[8]
Parallel Clustering with Resolution Variation
N. Pitsianis, D. Floros, T. Liu, and X. Sun. “Parallel Clustering with Resolution Variation”. In:2023 IEEE High Performance Extreme Computing Conference (HPEC). 2023 IEEE High Performance Extreme Computing Conference (HPEC). Boston, MA, USA: IEEE, Sept. 25, 2023, pp. 1–8
2023
-
[9]
Krizhevsky.Learning Multiple Layers of Features from Tiny Images
A. Krizhevsky.Learning Multiple Layers of Features from Tiny Images. Toronto, ON, Canada, 2009, pp. 32–33. 9 Juntang Wang, Hao Wu et al
2009
-
[10]
16S rRNA Gene Sequencing for Bacterial Identification in the Diagnostic Laboratory: Pluses, Perils, and Pitfalls
J. M. Janda and S. L. Abbott. “16S rRNA Gene Sequencing for Bacterial Identification in the Diagnostic Laboratory: Pluses, Perils, and Pitfalls”. In:J Clin Microbiol45.9 (Sept. 2007), pp. 2761–2764. PMID:17626177
2007
-
[11]
Naive Bayesian Classifier for Rapid Assignment of rRNA Sequences into the New Bacterial Taxonomy
Q. Wang, G. M. Garrity, J. M. Tiedje, and J. R. Cole. “Naive Bayesian Classifier for Rapid Assignment of rRNA Sequences into the New Bacterial Taxonomy”. In:Appl Environ Microbiol73.16 (Aug. 2007), pp. 5261–5267. PMID:17586664
2007
-
[12]
The Active Microbial Community More Accurately Reflects the Anaerobic Digestion Process: 16S rRNA (Gene) Sequencing as a Predictive Tool
J. De Vrieze, A. J. Pinto, W. T. Sloan, and U. Z. Ijaz. “The Active Microbial Community More Accurately Reflects the Anaerobic Digestion Process: 16S rRNA (Gene) Sequencing as a Predictive Tool”. In:Microbiome 6.1 (Apr. 2, 2018), p. 63
2018
-
[13]
Evaluation of 16S rRNA Gene Sequencing for Species and Strain-Level Microbiome Analysis
J. S. Johnson, D. J. Spakowicz, B.-Y . Hong, L. M. Petersen, P. Demkowicz, L. Chen, S. R. Leopold, B. M. Hanson, H. O. Agresta, M. Gerstein, E. Sodergren, and G. M. Weinstock. “Evaluation of 16S rRNA Gene Sequencing for Species and Strain-Level Microbiome Analysis”. In:Nat Commun10.1 (Nov. 6, 2019), p. 5029
2019
-
[14]
M. Caron, P. Bojanowski, A. Joulin, and M. Douze.Deep Clustering for Unsupervised Learning of Visual Features. Version 2. Mar. 18, 2019. arXiv:1807.05520 [cs].URL: http://arxiv.org/abs/1807.05520 (visited on 05/13/2025). Pre-published
Pith/arXiv arXiv 2019
-
[15]
Y . Yang, Z. Guan, Z. Wang, W. Zhao, C. Xu, W. Lu, and J. Huang.Self-Supervised Heterogeneous Graph Pre-training Based on Structural Clustering. Apr. 12, 2023. arXiv:2210.10462 [cs].URL: http://arxiv. org/abs/2210.10462(visited on 05/13/2025). Pre-published
Pith/arXiv arXiv 2023
-
[16]
T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski.Vision Transformers Need Registers. Apr. 12, 2024. arXiv: 2309.16588 [cs].URL:http://arxiv.org/abs/2309.16588(visited on 03/25/2025). Pre-published
Pith/arXiv arXiv 2024
-
[17]
F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf.Object-Centric Learning with Slot Attention. Oct. 14, 2020. arXiv: 2006.15055 [cs] .URL: http: //arxiv.org/abs/2006.15055(visited on 05/11/2025). Pre-published
Pith/arXiv arXiv 2020
-
[18]
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin.Emerging Properties in Self-Supervised Vision Transformers. May 24, 2021. arXiv: 2104.14294 [cs] .URL: http://arxiv.org/ abs/2104.14294(visited on 03/25/2025). Pre-published
Pith/arXiv arXiv 2021
-
[19]
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski.DINOv2: Learning Robust Visual Features without Superv...
Pith/arXiv arXiv 2024
-
[20]
O. Siméoni, G. Puy, H. V . V o, S. Roburin, S. Gidaris, A. Bursuc, P. Pérez, R. Marlet, and J. Ponce.Localizing Objects with Self-Supervised Transformers and No Labels. Sept. 29, 2021. arXiv: 2109.14279 [cs] .URL: http://arxiv.org/abs/2109.14279(visited on 03/25/2025). Pre-published
Pith/arXiv arXiv 2021
-
[21]
H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. M. Ni, and H. -Y . Shum.DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. July 11, 2022. arXiv: 2203.03605 [cs] .URL: http://arxiv.org/abs/2203.03605(visited on 03/25/2025). Pre-published
Pith/arXiv arXiv 2022
-
[22]
Attention Is All You Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. ukasz Kaiser, and I. Polosukhin. “Attention Is All You Need”. In:Advances in Neural Information Processing Systems. V ol. 30. Curran Associates, Inc., 2017
2017
-
[23]
Digraph Clustering by the BlueRed Method
T. Liu, D. Floros, N. Pitsianis, and X. Sun. “Digraph Clustering by the BlueRed Method”. In:2021 IEEE High Performance Extreme Computing Conference (HPEC). 2021 IEEE High Performance Extreme Computing Conference (HPEC). Sept. 2021, pp. 1–7
2021
-
[24]
A Global Geometric Framework for Nonlinear Dimensionality Reduction
J. B. Tenenbaum, V . de Silva, and J. C. Langford. “A Global Geometric Framework for Nonlinear Dimensionality Reduction”. In:Science290.5500 (Dec. 22, 2000), pp. 2319–2323
2000
-
[25]
Spaceland Embedding of Sparse Stochastic Graphs
N. Pitsianis, A.-S. Iliopoulos, D. Floros, and X. Sun. “Spaceland Embedding of Sparse Stochastic Graphs”. In: 2019 IEEE High Performance Extreme Computing Conference (HPEC). 2019 IEEE High Performance Extreme Computing Conference (HPEC). Sept. 2019, pp. 1–8
2019
-
[26]
Visualizing Data Using T-SNE
L. Van der Maaten and G. Hinton. “Visualizing Data Using T-SNE.” In:Journal of machine learning research 9.11 (2008), pp. 2579–2605
2008
-
[27]
From Louvain to Leiden: Guaranteeing Well-Connected Communi- ties
V . A. Traag, L. Waltman, and N. J. Van Eck. “From Louvain to Leiden: Guaranteeing Well-Connected Communi- ties”. In:Sci. Rep.9.1 (Mar. 26, 2019), p. 5233
2019
-
[28]
Automated Selection of Results in Hierarchical Segmentations of Remotely Sensed Hyperspectral Images
A. Plaza and J. Tilton. “Automated Selection of Results in Hierarchical Segmentations of Remotely Sensed Hyperspectral Images”. In:Proceedings. 2005 IEEE International Geoscience and Remote Sensing Symposium,
2005
-
[29]
Comparing Partitions
L. J. Hubert and P. Arabie. “Comparing Partitions”. In:Journal of Classification2.2–3 (1985), pp. 193–218. 10 Juntang Wang, Hao Wu et al
1985
-
[30]
The Hungarian Method for the Assignment Problem
H. W. Kuhn. “The Hungarian Method for the Assignment Problem”. In:Nav. Res. Logist. Q.2.1–2 (Mar. 1955), pp. 83–97
1955
-
[31]
A Shortest Augmenting Path Algorithm for Dense and Sparse Linear Assignment Problems
R. Jonker and A. V olgenant. “A Shortest Augmenting Path Algorithm for Dense and Sparse Linear Assignment Problems”. In:Computing38.4 (Dec. 1, 1987), pp. 325–340
1987
-
[32]
Auction Algorithms for Network Flow Problems: A Tutorial Introduction
D. P. Bertsekas. “Auction Algorithms for Network Flow Problems: A Tutorial Introduction”. In:Comput Optim Applic1.1 (Oct. 1992), pp. 7–66
1992
-
[33]
Algebraic Connectivity of Graphs
M. Fiedler. “Algebraic Connectivity of Graphs”. In:Czech. Math. J.23.2 (1973), pp. 298–305
1973
-
[34]
Algebraic Vertex Ordering of a Sparse Graph for Adjacency Access Locality and Graph Compression
D. Floros, N. Pitsianis, and X. Sun. “Algebraic Vertex Ordering of a Sparse Graph for Adjacency Access Locality and Graph Compression”. In:2024 IEEE High Performance Extreme Computing Conference (HPEC). 2024 IEEE High Performance Extreme Computing Conference (HPEC). Sept. 2024, pp. 1–7
2024
-
[35]
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba. “Adam: A Method for Stochastic Optimization”. In:arXiv:1412.6980 [cs.LG](Jan. 30, 2017). arXiv:1412.6980 [cs.LG]
Pith/arXiv arXiv 2017
-
[36]
German Collection of Microorganisms and Cell Cultures GmbH.DSMZ – German Collection of Microorganisms and Cell Cultures. 2025
2025
-
[37]
National Institutes of Health (NIH).URL: https://www.nih.gov/ (visited on 02/04/2025)
National Institutes of Health (NIH). National Institutes of Health (NIH).URL: https://www.nih.gov/ (visited on 02/04/2025)
2025
-
[38]
Revisiting K-mer Profile for Effective and Scalable Genome Representation Learning
A. Çelikkanat, A. R. Masegosa, and T. D. Nielsen. “Revisiting K-mer Profile for Effective and Scalable Genome Representation Learning”. In: (). [39]Kraken: Ultrafast Metagenomic Sequence Classification Using Exact Alignments | Genome Biology | Full Text. URL: https://genomebiology.biomedcentral.com/articles/10.1186/gb-2014-15-3-r46 (visited on 02/04/2025)
-
[40]
com/articles/nbt.2023(visited on 02/04/2025)
How to Apply de Bruijn Graphs to Genome Assembly | Nature Biotechnology.URL: https://www.nature. com/articles/nbt.2023(visited on 02/04/2025)
2023
-
[41]
Quantum Chemistry Structures and Properties of 134 Kilo Molecules
R. Ramakrishnan, P. O. Dral, M. Rupp, and O. A. von Lilienfeld. “Quantum Chemistry Structures and Properties of 134 Kilo Molecules”. In:Sci Data1.1 (Aug. 5, 2014), p. 140022
2014
-
[42]
Hedonic Housing Prices and the Demand for Clean Air
D. Harrison and D. L. Rubinfeld. “Hedonic Housing Prices and the Demand for Clean Air”. In:Journal of Environmental Economics and Management5.1 (Mar. 1, 1978), pp. 81–102
1978
-
[43]
Gradient-Based Learning Applied to Document Recognition
Y . Lecun, L. Bottou, Y . Bengio, and P. Haffner. “Gradient-Based Learning Applied to Document Recognition”. In:Proceedings of the IEEE86.11 (Nov. 1998), pp. 2278–2324
1998
-
[44]
Random Forests
L. Breiman. “Random Forests”. In:Machine Learning45.1 (2001), pp. 5–32
2001
-
[45]
XGBoost: A Scalable Tree Boosting System
T. Chen and C. Guestrin. “XGBoost: A Scalable Tree Boosting System”. In:Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Aug. 13, 2016, pp. 785–794. arXiv:1603.02754 [cs.LG]
Pith/arXiv arXiv 2016
-
[46]
CatBoost: Unbiased Boosting with Categorical Features
L. Prokhorenkova, G. Gusev, A. V orobev, A. V . Dorogush, and A. Gulin. “CatBoost: Unbiased Boosting with Categorical Features”. In:Advances in Neural Information Processing Systems. V ol. 31. Curran Associates, Inc., 2018
2018
-
[47]
S. O. Arik and T. Pfister.TabNet: Attentive Interpretable Tabular Learning. Dec. 9, 2020. arXiv: 1908.07442 [cs].URL:http://arxiv.org/abs/1908.07442(visited on 05/14/2025). Pre-published
Pith/arXiv arXiv 2020
-
[48]
X. Huang, A. Khetan, M. Cvitkovic, and Z. Karnin.TabTransformer: Tabular Data Modeling Using Contextual Embeddings. Dec. 11, 2020. arXiv: 2012.06678 [cs].URL: http://arxiv.org/abs/2012.06678 (visited on 05/05/2025). Pre-published
Pith/arXiv arXiv 2020
-
[49]
Y . Gorishniy, I. Rubachev, V . Khrulkov, and A. Babenko.Revisiting Deep Learning Models for Tabular Data. Oct. 26, 2023. arXiv:2106.11959 [cs].URL: http://arxiv.org/abs/2106.11959 (visited on 05/14/2025). Pre-published
Pith/arXiv arXiv 2023
-
[50]
Predicting the Optimal Growth Temperatures of Prokaryotes Using Only Genome Derived Features
D. B. Sauer and D.-N. Wang. “Predicting the Optimal Growth Temperatures of Prokaryotes Using Only Genome Derived Features”. In:Bioinformatics35.18 (Sept. 15, 2019), pp. 3224–3231
2019
-
[51]
Geometry-Enhanced Molecular Representation Learning for Property Prediction
X. Fang, L. Liu, J. Lei, D. He, S. Zhang, J. Zhou, F. Wang, H. Wu, and H. Wang. “Geometry-Enhanced Molecular Representation Learning for Property Prediction”. In:Nat Mach Intell4.2 (Feb. 2022), pp. 127–134. [52]Principal Component Analysis. Springer Series in Statistics. New York: Springer-Verlag, 2002
2022
-
[53]
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
L. McInnes, J. Healy, and J. Melville. “UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction”. In:arXiv:1802.03426 [stat.ML](Sept. 18, 2020). arXiv:1802.03426 [stat.ML]. 11 Juntang Wang, Hao Wu et al. A An Intuitive Example of Configuration Mixing To illustrate the necessity of fusing valid clusterings across resolution scales, we u...
Pith/arXiv arXiv 2020
-
[2005]
2005 IEEE International Geoscience and Remote Sensing Symposium, 2005
IGARSS ’05.. 2005 IEEE International Geoscience and Remote Sensing Symposium, 2005. IGARSS ’05. V ol. 7. July 2005, pp. 4946–4949
2005
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.