REVIEW 3 major objections 5 minor 1 cited by
Convolutional Vision Transformer for Cosmology Parameter Inference
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A hybrid convolutional vision transformer yields tighter constraints on matter density and clustering amplitude than pure CNN or plain transformer baselines on simulated cosmic fields.
desk verdict Clean, honest first CvT benchmark for cosmology parameter inference, but the architecture-level claims rest on capacity-mismatched single runs and the transfer-learning claim is contradicted by the paper's own ViT numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the CvT architecture: a multi-stage vision transformer in which convolutional token embedding layers progressively reduce spatial resolution while increasing feature width, and depth-wise separable convolutions replace the linear query, key, and value projections of a standard transformer. This design captures local structure through convolutions and global relationships through self-attention in every block, and it eliminates the need for positional encoding. The paper's argument attributes CvT's tighter constraints and its successful dark-matter-to-halo transfer to this combined inductive bias. A supporting mechanism is the moment-matching loss, which trains the network to output both a posterior mean and a standard deviation for each parameter, allowing accuracy and uncertainty to be evaluated on the test set.
What would settle it
Retrain CvT, ViT, and the CNN with matched numbers of parameters, several random initializations, and hyperparameter optimization on the same dark matter and halo maps; if CvT's RMSE advantage disappears or reverses, the claim that the hybrid architecture itself is better is refuted.
Extended reading notes
Core claim
In the paper's own terms, the central discovery is that the CvT design, replacing a vision transformer's linear projections with convolutional projections and convolutional token embeddings, gives tighter likelihood-free constraints on $\Omega_m$ and $\sigma_8$ from both dark matter and halo density maps. On the test set, CvT achieves lower RMSE than ViT for both parameters in the dark-matter pretraining and halo-transfer settings, and it achieves lower RMSE than the CNN in every setting except $\sigma_8$ from dark matter, where the CNN is closer but CvT is still best. The paper also reports that initializing from a dark-matter pretrained model before finetuning on halo maps lowers RMSE and produces more representative error bars than training on halo maps from scratch, and that this benefit appears for CvT but not for ViT or the CNN. The authors interpret this as CvT leveraging the shared large-scale structure of dark matter and halos, while noting that more detailed tests are needed to confirm that mechanism.
Load-bearing premise
The comparison assumes that the performance gap reflects architectural merit even though CvT is far larger than the baselines and no hyperparameter tuning was performed.
Editorial extensions
If this is right
- If robust, the comparison suggests hybrid convolution-transformer architectures should be the default choice for simulation-based cosmological parameter inference.
- Dark-matter pretraining followed by halo finetuning gives a recipe for working with expensive, sparsely sampled tracers such as galaxies, where large simulation suites are often unavailable.
- Because CvT does not require positional encoding, the same trained architecture can be applied to maps of different resolutions without modification, which is relevant for survey data with varying depth.
- The transfer-learning result indicates that features learned from dark matter simulations carry over to biased tracers, so pretrained models could in principle be reused across multiple observables.
Reading between the lines
- Inference: the architecture comparison is not capacity-matched, since CvT has 17.6 million parameters versus 1.6 million for the plain ViT and a five-layer CNN, and no hyperparameter tuning was performed, so part of the reported gap may reflect model size and training budget rather than the CvT design itself.
- Inference: a direct test of the transfer claim would give ViT and CNN the same pretraining data and budget as CvT; the paper reports no transfer benefit for them, but their smaller capacity may be the reason, and a capacity-matched comparison would clarify whether the benefit is specific to the hybrid design.
- Inference: since the paper reports underestimated error bars on halo fields, especially for $\Omega_m$, a practical next step is to recalibrate predicted variances on a validation set before quoting $1\sigma$ uncertainties.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper applies the Convolutional vision Transformer (CvT) to likelihood-free inference of the cosmological parameters Ωm and σ8 from 2D projected dark-matter and halo maps of the public QUIJOTE simulations. The authors train a moment-network regression head that predicts marginal posterior means and variances, pretrain CvT on dark-matter maps, and then finetune on halo maps, comparing against a vision transformer (ViT) and a small CNN. The headline claims are that CvT constrains both parameters better than ViT and CNN on both field types, that dark-matter pretraining helps CvT for halo-field inference, and that CvT is more computationally efficient than ViT. The paper reports RMSE and averaged error-bar metrics on a held-out test set and provides public code.
Significance. If the empirical claims survive robustness checks, the result would be practically useful: a hybrid convolutional-transformer architecture may combine the local feature extraction of CNNs with the global context of attention for field-level cosmological inference, and the transfer-learning finding could reduce the cost of training on expensive halo or galaxy simulations. The paper has clear strengths: it uses the public QUIJOTE simulations, a simulation-level train/validation/test split, a held-out test set, and released code, and there is no circularity in the evaluation because all numbers are computed on data not used in training. However, the current evidence base is a set of single-run point estimates from models of very different capacities, which is not yet sufficient to establish an architecture-level ranking or a robust transfer-learning conclusion.
major comments (3)
- [§3 'Comparison with traditional ViT' and 'Comparison with CNN'; §2 'Training details'] The headline claim that CvT constrains Ωm and σ8 better than ViT and CNN is not yet supported as an architecture-level statement, because the comparison does not isolate architecture: CvT-13 has 17.6M parameters, the ViT baseline has 1.6M parameters, and the CNN has only five convolutional layers, and each model is trained once with no hyperparameter tuning and no seed averaging. Reported RMSE differences (e.g., dark-matter pretraining Ωm RMSE 0.059 for CvT versus 0.066 for ViT versus 0.073 for CNN) are therefore as consistent with capacity or run-to-run scatter as with the hybrid design. Please provide capacity-matched baselines or, at minimum, multiple seeded runs with error bars on every reported metric, and temper the abstract and conclusions if such a comparison is not performed.
- [Abstract; §4 Conclusion; §3 'Comparison with traditional ViT'] The transfer-learning claim is internally inconsistent. The abstract and conclusion state that ViT and CNN do not benefit from dark-matter pretraining, but the ViT numbers in §3 show the opposite: halo training from scratch has Ωm RMSE 0.074 and σ8 RMSE 0.112, while pretraining then finetuning gives 0.068 and 0.106, which is an improvement for ViT. Only the CNN results fail to benefit. The abstract and conclusion should be corrected to state that the pretraining benefit was not found for CNN and that ViT did show an improvement in these single runs, or the conclusion should be recast with appropriate uncertainty.
- [§3 'Comparison with CNN'] The CNN transfer-learning result appears pathological: for halo after dark-matter pretraining, Ωm RMSE is 0.21 and σ8 RMSE is 0.151 with ¯σ=0 for both parameters. The paper does not report whether the same head-reinitialization and unfrozen finetuning protocol was used for the CNN and ViT baselines, or whether the CNN failed to converge. As presented, this result is not interpretable as a property of CNN transfer learning; please specify the exact protocol for all models and either diagnose the CNN failure or remove it from the transfer-learning comparison.
minor comments (5)
- [§2 'Data'] The preprocessing sentence 'The overdensities are first calculated (ρ/ρ)' is incomplete or typographically wrong; it should read δ = ρ/ρ̄ − 1, and the subsequent transformation should be written consistently as log10(1 + δ).
- [§3 'Results'] The sentence 'The constraints in (c) are worse than in (b) as shown by the lower values of RMSE and ¯σ' appears to say the opposite of what is meant; worse constraints should be reflected by higher RMSE values, so the phrasing should be corrected.
- [Abstract] The abstract uses 'Convolution vision Transformer' while the body uses 'Convolutional vision Transformer'; please unify the terminology.
- [Fig. 2 caption and §3 'Experimental details and evaluation'] Please clarify whether the scatter points represent per-simulation averages over the 30 2D maps and whether the reported RMSE and ¯σ are computed on those averaged predictions or on all individual maps; the definitions in §3 refer to N test examples without specifying the unit.
- [§3 'Experimental details and evaluation'] The statement 'We chose not to freeze any weights, as doing so only provided marginally better constraints...' is ambiguous about which model and finetuning configuration are being described; please make the subject explicit.
Circularity Check
No significant circularity; all reported predictions are evaluated on held-out QUIJOTE test simulations.
full rationale
The paper makes no derivation-from-first-principles claim. Its central empirical claims (CvT RMSE and error-bar comparisons, and the transfer-learning benefit on halo fields) rest on test-set evaluations from the public QUIJOTE simulations split at the simulation level: Section 2 states 'we divide it into training, validation, and testing sets using an 80-10-10% split. Since the splitting is performed at the simulation level, all data from a simulation are either in the training, validation, or testing set.' The loss in Eq. (1) trains the predicted means and variances against batch true values, but the reported RMSE is computed on test examples never used for fitting or for model selection, since 'the model weights corresponding to the lowest validation loss are used for inference.' The transfer-learning conclusion is a direct empirical contrast between a finetuned model and a from-scratch model on the same halo data, not a quantity forced by the loss or preprocessing. The architecture ranking is an experimental observation; the fairness caveat (ViT has 1.6M parameters versus 17.6M for CvT, no seed averaging, and 'No specific hyperparameter tuning has been performed') is an experimental-control concern, not circular reasoning. No load-bearing self-citations appear: reference [26] supplies the CvT architecture as external prior work, and reference [4] is a separate application of CvT to galaxy morphology. One internal inconsistency exists—the Abstract claims 'ViT and CNN do not show these benefits' while Section 3 reports that ViT halo constraints improve with DM pretraining (Omega_m RMSE 0.074 to 0.068; sigma_8 RMSE 0.112 to 0.106)—but an internal inconsistency is a correctness issue, not a circularity. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- Model architecture choices =
CvT-13 (17.6M params), ViT patch size 8 (1.6M params), CNN with 5 convolutional layers
- Training schedule =
Batch size 16, AdamW lr 5e-6, weight decay 1e-5, LR decay factor 0.3 after 5 epochs, 30 epochs
- Map preprocessing =
256^3 CIC grid, log10(1+delta), 10 random slices per projection direction, slice thickness ~3.9 h^-1 Mpc
- Evaluation aggregation =
30 maps per simulation, averaged into one point in figures; RMSE over test examples
assumptions (5)
- domain assumption QUIJOTE simulations accurately represent dark matter and halo fields for Omega_m and sigma_8 inference at z=0.
- domain assumption Projected 2D maps contain sufficient information for parameter inference.
- domain assumption The moment network loss (Eq. 1) yields calibrated marginal posterior means and variances.
- ad hoc to paper A single training run with no hyperparameter tuning gives representative performance rankings.
- domain assumption FoF halo catalogs with no mass cut are valid tracers for the comparison.
Cite this review
Pith. "Pith review of Convolutional Vision Transformer for Cosmology Parameter Inference." pith.science (2026). https://pith.science/paper/LJHOFPZP
@misc{pith2026241114392,
author = {Pith},
title = {Pith review of: Convolutional Vision Transformer for Cosmology Parameter Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/LJHOFPZP}},
note = {Machine review of arXiv:2411.14392}
}
abstract
Parameter inference is a crucial task in modern cosmology that requires accurate and fast computational methods to handle the high precision and volume of observational datasets. In this study, we explore a hybrid vision transformer, the Convolution vision Transformer (CvT), which combines the benefits of vision transformers (ViTs) and convolutional neural networks (CNNs). We use this approach to infer the $\Omega_m$ and $\sigma_8$ cosmological parameters from simulated dark matter and halo fields. Our experiments indicate that the constraints on $\Omega_m$ and $\sigma_8$ obtained using CvT are better than ViT and CNN, using either dark matter or halo fields. For CvT, pretraining on dark matter fields proves advantageous for improving constraints using halo fields compared to training a model from the beginning. However, ViT and CNN do not show these benefits. The CvT is more efficient than ViT since, despite having more parameters, it requires a training time similar to that of ViT and has similar inference times. The code is available at \url{https://github.com/Yash-10/cvt-cosmo-inference/}.
Figures
Forward citations
Cited by 1 Pith paper
-
ViT-based Local Volume dwarf galaxy Identificationin (VIDA) in the CSST survey
A Vision Transformer trained on simulated CSST images identifies Local Volume dwarf galaxies with 85% true positives at 0.1% false positives, reaching M_V = -7 within 10 Mpc.
Reference graph
Works this paper leans on
-
[1]
Bayesian field-level inference of primordial non-Gaussianity using next-generation galaxy surveys
Adam Andrews, Jens Jasche, Guilhem Lavaux, and Fabian Schmidt. Bayesian field-level inference of primordial non-Gaussianity using next-generation galaxy surveys. Monthly Notices of the RAS , 520(4): 5746–5763, April 2023. doi: 10.1093/mnras/stad432
-
[2]
Better plain ViT baselines for ImageNet-1k
Lucas Beyer, Xiaohua Zhai, and Alexander Kolesnikov. Better plain ViT baselines for ImageNet-1k. arXiv e-prints, art. arXiv:2205.01580, May 2022. doi: 10.48550/arXiv.2205.01580
-
[3]
Experiment tracking with weights and biases, 2020
Lukas Biewald. Experiment tracking with weights and biases, 2020. URL https://www.wandb.com/. Software available from wandb.com
2020
-
[4]
Galaxy morphology classification based on Convolutional vision Transformer (CvT)
Jie Cao, Tingting Xu, Yuhe Deng, Linhua Deng, Mingcun Yang, Zhijing Liu, and Weihong Zhou. Galaxy morphology classification based on Convolutional vision Transformer (CvT). Astronomy and Astrophysics, 683:A42, March 2024. doi: 10.1051/0004-6361/202348544
-
[5]
A new approach to observational cosmology using the scattering transform
Sihao Cheng, Yuan-Sen Ting, Brice Ménard, and Joan Bruna. A new approach to observational cosmology using the scattering transform. Monthly Notices of the RAS , 499(4):5902–5914, December 2020. doi: 10.1093/mnras/staa3165
-
[6]
The frontier of simulation-based inference
Kyle Cranmer, Johann Brehmer, and Gilles Louppe. The frontier of simulation-based inference. Pro- ceedings of the National Academy of Science , 117(48):30055–30062, December 2020. doi: 10.1073/pnas. 1912789117
doi:10.1073/pnas 2020
-
[7]
Natalí S. M. de Santi, Helen Shao, Francisco Villaescusa-Navarro, L. Raul Abramo, Romain Teyssier, Pablo Villanueva-Domingo, Yueying Ni, Daniel Anglés-Alcázar, Shy Genel, Elena Hernández-Martínez, Ulrich P. Steinwandel, Christopher C. Lovell, Klaus Dolag, Tiago Castro, and Mark V ogelsberger. Robust Field-level Likelihood-free Inference with Galaxies. Ast...
-
[8]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv e-prints, art. arXiv:2010.11929, October 2020. doi: 10.48550/arXiv...
Show all 27 references
-
[9]
Strong Gravitational Lensing Parameter Estimation with Vision Transformer
Kuan-Wei Huang, Geoff Chih-Fan Chen, Po-Wen Chang, Sheng-Chieh Lin, Chia-Jung Hsu, Vishal Thengane, and Joshua Yao-Yu Lin. Strong Gravitational Lensing Parameter Estimation with Vision Transformer. arXiv e-prints, art. arXiv:2210.04143, October 2022. doi: 10.48550/arXiv.2210.04143
-
[10]
Sabiu, Inkyu Park, and Sungwook E
Se Yeon Hwang, Cristiano G. Sabiu, Inkyu Park, and Sungwook E. Hong. The universe is worth 64 3 pixels: convolution neural network and vision transformers for cosmology. Journal of Cosmology and Astroparticle Physics, 2023(11):075, November 2023. doi: 10.1088/1475-7516/2023/11/075
2023 doi
-
[11]
Željko Ivezi´c, Steven M. Kahn, J. Anthony Tyson, Bob Abel, Emily Acosta, Robyn Allsman, David Alonso, Yusra AlSayyad, Scott F. Anderson, John Andrew, James Roger P. Angel, George Z. Angeli, Reza Ansari, Pierre Antilogus, Constanza Araujo, Robert Armstrong, Kirk T. Arndt, Pier...
2019
- [12]
-
[13]
Laureijs, J
R. Laureijs, J. Amiaux, S. Arduini, J. L. Auguères, J. Brinchmann, R. Cole, M. Cropper, C. Dabin, L. Duvet, A. Ealet, B. Garilli, P. Gondoin, L. Guzzo, J. Hoar, H. Hoekstra, R. Holmes, T. Kitching, T. Maciaszek, Y . Mellier, F. Pasian, W. Percival, J. Rhodes, G. Saavedra Criad...
-
[14]
Extracting cosmological parameters from N-body simulations using machine learning techniques
Andrei Lazanu. Extracting cosmological parameters from N-body simulations using machine learning techniques. Journal of Cosmology and Astroparticle Physics , 2021(9):039, September 2021. doi: 10.1088/ 1475-7516/2021/09/039
2021
-
[15]
On the accuracy and precision of correlation functions and field- level inference in cosmology
Florent Leclercq and Alan Heavens. On the accuracy and precision of correlation functions and field- level inference in cosmology. Monthly Notices of the RAS , 506(1):L85–L90, September 2021. doi: 10.1093/mnrasl/slab081
2021 doi
-
[16]
Parker, ChangHoon Hahn, Shirley Ho, Michael Eickenberg, Jiamin Hou, Elena Massara, Chirag Modi, Azadeh Moradinezhad Dizgah, Bruno Régaldo-Saint Blancard, and David Spergel
Pablo Lemos, Liam H. Parker, ChangHoon Hahn, Shirley Ho, Michael Eickenberg, Jiamin Hou, Elena Massara, Chirag Modi, Azadeh Moradinezhad Dizgah, Bruno Régaldo-Saint Blancard, and David Spergel. SimBIG: Field-level Simulation-based Inference of Large-scale Structure. In Machine...
- [17]
-
[18]
Eisenstein, Sihan Yuan, and Lehman H
Michelle Ntampaka, Daniel J. Eisenstein, Sihan Yuan, and Lehman H. Garrison. A Hybrid Deep Learning Approach to Cosmological Constraints from Galaxy Redshift Surveys. Astrophysical Journal, 889(2):151, February 2020. doi: 10.3847/1538-4357/ab5f5e
2020 doi
-
[19]
Sabiu, ZhiGang Li, HaiTao Miao, and Xiao-Dong Li
ShuYang Pan, MiaoXin Liu, Jaime Forero-Romero, Cristiano G. Sabiu, ZhiGang Li, HaiTao Miao, and Xiao-Dong Li. Cosmological parameter estimation from large-scale structure deep learning. Science China Physics, Mechanics, and Astronomy, 63(11):110412, November 2020. doi: 10.1007...
2020 doi
-
[20]
Price, Shirley Ho, Jeff Schneider, and Barnabas Poczos
Siamak Ravanbakhsh, Junier Oliva, Sebastien Fromenteau, Layne C. Price, Shirley Ho, Jeff Schneider, and Barnabas Poczos. Estimating Cosmological Parameters from the Dark Matter Distribution. arXiv e-prints, art. arXiv:1711.02033, November 2017. doi: 10.48550/arXiv.1711.02033
-
[21]
An improved cosmological parameter infer- ence scheme motivated by deep learning
Dezs˝o Ribli, Bálint Ármin Pataki, and István Csabai. An improved cosmological parameter infer- ence scheme motivated by deep learning. Nature Astronomy, 3:93–98, January 2019. doi: 10.1038/ s41550-018-0596-8
2019
-
[22]
Weak lensing cosmology with convolutional neural networks on noisy data
Dezs˝o Ribli, Bálint Ármin Pataki, José Manuel Zorrilla Matilla, Daniel Hsu, Zoltán Haiman, and István Csabai. Weak lensing cosmology with convolutional neural networks on noisy data. Monthly Notices of the RAS, 490(2):1843–1860, December 2019. doi: 10.1093/mnras/stz2610
2019 doi
-
[23]
Kreisch, Andrina Nicola, Justin Alsing, Roman Scoccimarro, Licia Verde, Matteo Viel, Shirley Ho, Stephane Mallat, Benjamin Wandelt, and David N
Francisco Villaescusa-Navarro, ChangHoon Hahn, Elena Massara, Arka Banerjee, Ana Maria Delgado, Doogesh Kodi Ramanah, Tom Charnock, Elena Giusarma, Yin Li, Erwan Allys, Antoine Brochard, Cora Uhlemann, Chi-Ting Chiang, Siyu He, Alice Pisani, Andrej Obuljen, Yu Feng, Emanuele C...
2020
-
[24]
Spergel, Yin Li, Benjamin Wandelt, Andrina Nicola, Leander Thiele, Sultan Hassan, Jose Manuel Zorrilla Matilla, Desika Narayanan, Romeel Dave, and Mark V ogelsberger
Francisco Villaescusa-Navarro, Daniel Anglés-Alcázar, Shy Genel, David N. Spergel, Yin Li, Benjamin Wandelt, Andrina Nicola, Leander Thiele, Sultan Hassan, Jose Manuel Zorrilla Matilla, Desika Narayanan, Romeel Dave, and Mark V ogelsberger. Multifield Cosmology with Artificial...
-
[25]
Spergel, Rachel S
Francisco Villaescusa-Navarro, Shy Genel, Daniel Anglés-Alcázar, Leander Thiele, Romeel Dave, Desika Narayanan, Andrina Nicola, Yin Li, Pablo Villanueva-Domingo, Benjamin Wandelt, David N. Spergel, Rachel S. Somerville, Jose Manuel Zorrilla Matilla, Faizan G. Mohammad, Sultan ...
2022
- [26]
-
[27]
Training details
Dominik Zürcher, Janis Fluri, Raphael Sgier, Tomasz Kacprzak, and Alexandre Refregier. Cosmological forecast for non-Gaussian statistics in large-scale weak lensing surveys. Journal of Cosmology and Astroparticle Physics, 2021(1):028, January 2021. doi: 10.1088/1475-7516/2021/...
2021 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.