REVIEW 4 major objections 6 minor 56 references
Vision Transformers for Weakly-Supervised Microorganism Enumeration
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Transformer vision models can count microorganisms from whole-image labels, though ResNets still lead.
desk verdict Solid, reproducible benchmark; the headline ResNet-over-ViT claim is protocol-dependent and needs qualifiers or pretrained baselines. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the backbone-plus-regression-head pipeline for weakly-supervised counting: a feature extractor (ResNet, CNN, or vision transformer) maps the image to embeddings, and a single fully connected layer regresses those embeddings to a total count. For ViTs, images are divided into patches ($16\times16$ or $32\times32$) and processed by multi-head self-attention; CrossViT adds a dual-branch design with cross-attention that fuses features from two patch scales. The authors also constructed a synthetic dataset by placing Gaussian-ellipse fluorescent bacteria onto real backgrounds, yielding 12,000 images with counts up to 1,855, to cover dense and homogeneously sparse scenes. Training every architecture from scratch with near-identical hyperparameters is what lets the paper attribute observed differences to architecture choice.
What would settle it
Run the same four datasets with an ImageNet-pretrained CrossViT and an ImageNet-pretrained ResNet50 under identical fine-tuning schedules: if the pretrained CrossViT still fails to beat ResNet50 on the artificial bacteria dataset, the paper's case for ViT competitiveness would be weakened; if it does beat it, the from-scratch protocol would be the reason for the reported gap.
Extended reading notes
Core claim
On all four datasets, ResNet backbones reach the lowest mean absolute error, with ResNet101 best on fluorescent neurons (MAE 1.400) and ResNet50 best on VGG-cells (MAE 1.225), human cancer cells (MAE 27.206), and the artificial bacteria dataset (MAE 22.030). The ViT family is not far behind in most settings: CrossViT captures the best overall result on the artificial bacteria dataset (MAE 20.011) with the lowest transformer FLOPS ($50.91\times 10^8$), TransCrowd-Token matches ResNet-level accuracy on fluorescent neurons and artificial bacteria, DeepViT is the best ViT on fluorescent neurons, and vanilla ViT is best on VGG-cells. The paper reads this as evidence that ViTs are viable feature extractors for weakly-supervised microorganism enumeration, especially on homogeneous data, and that their gap to ResNets may shrink with pretraining and fine-tuning.
Load-bearing premise
The comparison assumes that training all architectures from scratch with nearly identical hyperparameters is a fair way to judge them, even though vision transformers are known to need pretraining and large datasets to reach their potential.
Editorial extensions
If this is right
- If ViTs can count microorganisms from global counts, laboratories could skip instance-level annotation and still automate enumeration for contamination monitoring and health-standard checks.
- CrossViT's result on the artificial bacteria dataset suggests multi-scale patch fusion is a promising direction for dense, uniformly distributed microorganism images.
- The from-scratch protocol leaves open that pretrained ViTs, fine-tuned in the usual way, could close or reverse the gap to ResNets.
- The released artificial dataset generator gives other researchers a controlled benchmark for weakly-supervised counting with known density.
Reading between the lines
- Because ViTs are known to depend on pretraining, the paper's from-scratch comparison likely understates their ceiling; a fairer test would compare ImageNet-pretrained ViTs against from-scratch ResNets.
- The CrossViT advantage on homogeneous data may extend to other uniform-texture counting problems, such as colony counting on agar plates or particle counting in environmental samples.
- A testable follow-up is to vary patch size and fusion strategy on the artificial dataset to isolate whether multi-scale cross-attention or parameter efficiency drives CrossViT's win.
- Weakly-supervised counting with global labels could combine with active learning to further reduce real-world labeling effort.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a comparative empirical study of vision transformer (ViT) backbones versus traditional CNN and ResNet architectures for weakly-supervised microorganism counting. Four microscopy datasets are used, including a newly created synthetic fluorescent bacteria dataset. All models are trained from scratch with near-identical hyperparameters, and performance is measured by MAE, RMSE, and FLOPs. The central claim is that, while ResNets achieve the best overall accuracy, ViTs—especially CrossViT—are competitive and occasionally superior on homogeneously distributed data. The manuscript also reports that the TransCrowd model, a state-of-the-art ViT-based counting method, underperforms when trained from scratch, in contrast to its pretrained performance reported in the literature.
Significance. The study addresses a practical and under-explored problem: using weakly-supervised counting for microorganism enumeration, where only global counts are available as labels. Its main strengths are the breadth of architectures compared, the introduction of a synthetic fluorescent bacteria dataset with released generation code, and the clear statement of the from-scratch training protocol, which supports reproducibility. If the comparison were properly controlled, the result that from-scratch ViTs are generally worse than ResNets for this task would be a useful negative result. However, because the protocol deliberately excludes pretraining—which is known to be crucial for ViTs—and because the paper itself concedes this limitation in Section VI, the current evidence does not support the unqualified conclusion stated in the abstract and conclusion. The absence of variance estimates further weakens the significance of the reported ranking. The paper is a plausible starting point for a more rigorous benchmarking study, but as it stands the central claim is not fully supported.
major comments (4)
- [Section IV-A and Table III] The central claim that ResNets outperform ViTs is drawn from a protocol where all models are trained from scratch and the paper itself acknowledges in Section VI that 'most studies use pre-trained weights and special fine-tuning for ViT-based approaches' and that TransCrowd, after pretraining, shows state-of-the-art performance 'contrary to the results of our study.' This self-acknowledged limitation means the MAE/RMSE gap in Table III may reflect the training budget rather than architectural merit. To support the abstract's unqualified statement, the authors should either add pretrained ViT baselines (at least for one or two representative architectures) or explicitly restrict the conclusion to the from-scratch setting in both the abstract and the conclusion.
- [Table III and Section IV-A] No standard deviations, confidence intervals, or significance tests are reported for any MAE/RMSE value, despite the statement in Section IV-A that 'Experiments were run with different randomization seeds.' With multiple seeds, variance estimates are cheap and necessary to determine whether gaps such as Vanilla ViT's MAE of 1.886 versus ResNet50's 1.225 on VGG-Cells are meaningful. Without such estimates, the performance ranking in Table III is not statistically grounded.
- [Section IV-A] The training protocol is not actually 'nearly identical' because the batch size is confounded with the architecture: 128 for CNNs, 64 for Vanilla ViT, CrossViT, TransCrowd, and ResNets, and 32 for Parallel ViT, DeepViT, and XCiT. Batch size affects optimization dynamics, regularization, and final accuracy, so the comparison is not controlled across architectures. The authors should either use a common batch size (if memory permits) or justify the choice and analyze its effect; otherwise, any performance difference cannot be attributed solely to the architecture.
- [Section IV-B] The train/validation split procedure is not described. For each dataset, it is not stated how the augmented images were partitioned, what ratio was used, whether the split was stratified by count, or whether the same split was used for all models. This is especially important for the Human Cancer Cells dataset, which has only 1463 augmented images and where results are likely sensitive to the split. The authors should specify the split procedure and ideally run multiple splits or cross-validation to ensure the ranking is robust.
minor comments (6)
- [Section II-C] The citation for CCTrans is incorrect: the text says 'named CCTrans [16]', but reference [16] is the CounTR paper; CCTrans should cite reference [40] (Tian et al., CCTrans). Please correct the citation.
- [Table I] The column header 'MLP Dim.' for CNNs is described as 'CONVOLUTIONAL OUTPUT'; this is confusing because MLP dimension is not a standard term for CNN channels. Please rename the column, e.g., 'Feature channels' for CNNs and keep 'MLP dim' for ViTs.
- [Section IV-A] The paper reports 'average FLOPS (floating point operations per second)' in Table III and Section V-C, but the units (10^8) indicate a count of floating-point operations per inference, not operations per second. If throughput is intended, please report inference time or FPS; if it is FLOPs, please correct the terminology.
- [Table II] The 'Data augmentation' column states 'yes' for three datasets but does not describe what augmentation was applied. Please specify the augmentation operations (e.g., random crops, flips, rotations) in the text or in a footnote.
- [Section I] The phrase 'ViTs performance demonstrates competent results' should be 'ViTs' performance demonstrates competent results' (add apostrophe). In addition, the sentence is grammatically awkward; consider rephrasing.
- [Section III-B-2] The text states 'To achieve higher model complexity without compromising parameter and compute neutrality, Parallel ViT [50] proposes parallelizing the MHSA and feed-forward blocks.' The term 'compute neutrality' is unclear; please rephrase to 'without increasing the number of parameters or FLOPs' for precision.
Circularity Check
No significant circularity: this is an empirical benchmark whose conclusions rest on measured MAE and RMSE values, not on quantities defined in terms of the claims they support.
full rationale
This paper is an empirical comparative study, not a derivation from first principles. The central claims—ResNets outperform ViTs overall and CrossViT is competitive on homogeneous data—are supported by measured MAE, RMSE, and FLOPS on four datasets under a common from-scratch training protocol. No prediction is constructed from a fitted parameter that is then renamed, no architecture or loss is defined in terms of the target result, and no load-bearing premise is imported from the authors' own prior work; in fact, the paper contains no self-citations. The artificial fluorescent bacteria dataset is synthetic and its counts are known by construction, but it serves only as benchmark input rather than as evidence for the comparative conclusion, so the evaluation is independent of the claim being tested. The paper's own Section VI caveat—that ViT approaches in prior work typically use pre-trained weights and fine-tuning, and that TransCrowd's original state-of-the-art result is reversed here—is an acknowledged limitation of the training protocol and a threat to external validity, but it is not circularity. There is no equation where an output reduces to an input, and no fitted value is presented as a prediction of the data from which it was derived. The absence of variance estimates and the potential unfairness of the from-scratch protocol are correctness concerns, not circularity concerns, and the manuscript itself concedes the former limitation explicitly. Accordingly, the circularity score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption Training all models from scratch with roughly identical hyperparameters yields a fair comparison.
- ad hoc to paper The artificial fluorescent bacteria dataset, generated by placing Gaussian ellipses on real backgrounds, is representative of real fluorescent microscopy images.
Cite this review
Pith. "Pith review of Vision Transformers for Weakly-Supervised Microorganism Enumeration." pith.science (2026). https://pith.science/paper/IX6I3EAC
@misc{pith2026241202250,
author = {Pith},
title = {Pith review of: Vision Transformers for Weakly-Supervised Microorganism Enumeration},
year = {2026},
howpublished = {\url{https://pith.science/paper/IX6I3EAC}},
note = {Machine review of arXiv:2412.02250}
}
read the original abstract
Microorganism enumeration is an essential task in many applications, such as assessing contamination levels or ensuring health standards when evaluating surface cleanliness. However, it's traditionally performed by human-supervised methods that often require manual counting, making it tedious and time-consuming. Previous research suggests automating this task using computer vision and machine learning methods, primarily through instance segmentation or density estimation techniques. This study conducts a comparative analysis of vision transformers (ViTs) for weakly-supervised counting in microorganism enumeration, contrasting them with traditional architectures such as ResNet and investigating ViT-based models such as TransCrowd. We trained different versions of ViTs as the architectural backbone for feature extraction using four microbiology datasets to determine potential new approaches for total microorganism enumeration in images. Results indicate that while ResNets perform better overall, ViTs performance demonstrates competent results across all datasets, opening up promising lines of research in microorganism enumeration. This comparative study contributes to the field of microbial image analysis by presenting innovative approaches to the recurring challenge of microorganism enumeration and by highlighting the capabilities of ViTs in the task of regression counting.
Figures
Reference graph
Works this paper leans on
-
[1]
X. Liu, S. Wang, L. Sendi, and M. J. Caulfield, “High-throughput imaging of bacterial colonies grown on filter plates with application to serum bactericidal assays,” Journal of immunological methods, vol. 292, no. 1-2, pp. 187–193, 2004
work page 2004
-
[2]
M. Riepl, S. Schauer, S. Knetsch, E. Holzhammer, A. H. Farnleitner, R. Sommer, and A. K. Kirschner, “Applicability of solid-phase cytometry and epifluorescence microscopy for rapid assessment of the microbio- logical quality of dialysis water,” Nephrology Dialysis Transplantation, vol. 26, no. 11, pp. 3640–3645, 2011
work page 2011
-
[3]
R. L. Kepner Jr and J. R. Pratt, “Use of fluorochromes for direct enumeration of total bacteria in environmental samples: past and present,” Microbiological reviews, vol. 58, no. 4, pp. 603–615, 1994
work page 1994
-
[4]
R. A. Herbert, 1 Methods for Enumerating Microorganisms and De- termining Biomass in Natural Environments , vol. 22 of Techniques in Microbial Ecology, p. 1–39. Academic Press, Jan. 1990
work page 1990
-
[5]
Horwitz, Official methods of analysis of the Association of Official Analytical Chemists
W. Horwitz, Official methods of analysis of the Association of Official Analytical Chemists . Washington, DC: The Association of Official Analytical Chemists, 1970
work page 1970
-
[6]
Absher, CHAPTER 1 - Hemocytometer Counting , p
M. Absher, CHAPTER 1 - Hemocytometer Counting , p. 395–397. Academic Press, Jan. 1973
work page 1973
-
[7]
Chapter 2 - Techniques for Oral Microbiology , p. 15–40. Oxford: Academic Press, Jan. 2015
work page 2015
-
[8]
J. Zhang, C. Li, M. M. Rahaman, Y . Yao, P. Ma, J. Zhang, X. Zhao, T. Jiang, and M. Grzegorzek, “A comprehensive review of image analysis methods for microorganism counting: from classical image processing to deep learning approaches,” Artificial Intelligence Review , vol. 55, pp. 2875–2944, Apr. 2022
work page 2022
Show all 56 references
-
[9]
DeepBacs: Bacterial image analysis using open-source deep learning approaches,
C. Spahn, R. F. Laine, P. M. Pereira, E. Gómez-de Mariscal, L. V on Chamier, M. Conduit, M. G. De Pinho, G. Jacquemet, S. Holden, M. Heilemann, and R. Henriques, “DeepBacs: Bacterial image analysis using open-source deep learning approaches,” preprint, Microbiology, Nov. 2021. 8
2021
-
[10]
Democratising deep learning for microscopy with ZeroCostDL4Mic,
L. V on Chamier, R. F. Laine, J. Jukkala, C. Spahn, D. Krentzel, E. Nehme, M. Lerche, S. Hernández-Pérez, P. K. Mattila, E. Karinou, S. Holden, A. C. Solak, A. Krull, T.-O. Buchholz, M. L. Jones, L. A. Royer, C. Leterrier, Y . Shechtman, F. Jug, M. Heilemann, G. Jacquemet, and...
2021
-
[11]
Automatic bacillus anthracis bacteria detection and segmentation in microscopic images using unet++,
F. Hoorali, H. Khosravi, and B. Moradi, “Automatic bacillus anthracis bacteria detection and segmentation in microscopic images using unet++,” Journal of Microbiological Methods , vol. 177, p. 106056, Oct. 2020
2020
-
[12]
Deeply-Supervised Density Regression for Automatic Cell Counting in Microscopy Images,
S. He, K. T. Minn, L. Solnica-Krezel, M. A. Anastasio, and H. Li, “Deeply-Supervised Density Regression for Automatic Cell Counting in Microscopy Images,” Nov. 2020. arXiv:2011.03683 [cs, eess]
2020 arXiv
-
[13]
Evaluation of two methods for monitoring surface cleanliness-atp bioluminescence and traditional hygiene swabbing.,
C. A. Davidson, C. J. Griffith, A. C. Peters, and L. Fielding, “Evaluation of two methods for monitoring surface cleanliness-atp bioluminescence and traditional hygiene swabbing.,” Luminescence : the journal of biological and chemical luminescence , vol. 14 1, pp. 33–8, 1999
1999
-
[14]
Cell Counting by Regression Using Convolutional Neural Network,
Y . Xue, N. Ray, J. Hugh, and G. Bigras, “Cell Counting by Regression Using Convolutional Neural Network,” in Computer Vision – ECCV 2016 Workshops (G. Hua and H. Jégou, eds.), vol. 9913, pp. 274–290, Cham: Springer International Publishing, 2016. Series Title: Lecture Notes i...
2016
-
[15]
Classification Beats Regression: Counting of Cells from Greyscale Microscopic Images based on Annotation-free Training Samples,
X. Ding, Q. Zhang, and W. J. Welch, “Classification Beats Regression: Counting of Cells from Greyscale Microscopic Images based on Annotation-free Training Samples,” Oct. 2020. arXiv:2010.14782 [cs, eess]
2020 arXiv
-
[16]
Countr: Transformer-based generalised visual counting,
C. Liu, Y . Zhong, A. Zisserman, and W. Xie, “Countr: Transformer-based generalised visual counting,” June 2023. arXiv:2208.13721 [cs]
2023 arXiv
-
[17]
Transcrowd: weakly- supervised crowd counting with transformers,
D. Liang, X. Chen, W. Xu, Y . Zhou, and X. Bai, “Transcrowd: weakly- supervised crowd counting with transformers,” Science China Information Sciences, vol. 65, p. 160104, June 2021. arXiv:2104.09116 [cs]
2021 arXiv
-
[18]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” 2021
2021
-
[19]
Automating cell counting in fluorescent microscopy through deep learning with c-resunet,
R. Morelli, L. Clissa, R. Amici, M. Cerri, T. Hitrec, M. Luppi, L. Rinaldi, F. Squarcio, and A. Zoccoli, “Automating cell counting in fluorescent microscopy through deep learning with c-resunet,” Scientific Reports , vol. 11, p. 22920, 11 2021
2021
-
[20]
Learning to count objects in images,
V . Lempitsky and A. Zisserman, “Learning to count objects in images,” in Advances in Neural Information Processing Systems (J. Lafferty, C. Williams, J. Shawe-Taylor, R. Zemel, and A. Culotta, eds.), vol. 23, Curran Associates, Inc., 2010
2010
-
[21]
Microscope images of human cancer cell lines (u2os and hl-60),
F. Lavitt, D. J. Rijlaarsdam, D. v. d. Linden, E. Weglarz-Tomczak, and J. M. Tomczak, “Microscope images of human cancer cell lines (u2os and hl-60),” Jan. 2021
2021
-
[22]
Transformers in Vision: A Survey,
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in Vision: A Survey,” ACM Computing Surveys , vol. 54, pp. 1–41, Jan. 2022. arXiv:2101.01169 [cs]
2022 arXiv
-
[23]
Monitoring micorbial morphogenetic changes in a fermentation process by a self- tuning vision system (stvs),
Y . Shabtai, M. Ronen, I. Mukmenev, and H. Guterman, “Monitoring micorbial morphogenetic changes in a fermentation process by a self- tuning vision system (stvs),” Computers & Chemical Engineering, vol. 20, p. S321–S326, Jan. 1996
1996
-
[24]
U-net: deep learning for cell counting, detection, and morphometry,
T. Falk, D. Mai, R. Bensch, O. Çiçek, A. Abdulkadir, Y . Marrakchi, A. Böhm, J. Deubner, Z. Jäckel, K. Seiwald, A. Dovzhenko, O. Tietz, C. Dal Bosco, S. Walsh, D. Saltukoglu, T. L. Tay, M. Prinz, K. Palme, M. Simons, I. Diester, T. Brox, and O. Ronneberger, “U-net: deep learni...
2019
-
[25]
Automating cell counting in fluorescent microscopy through deep learning with c-resunet,
R. Morelli, L. Clissa, R. Amici, M. Cerri, T. Hitrec, M. Luppi, L. Rinaldi, F. Squarcio, and A. Zoccoli, “Automating cell counting in fluorescent microscopy through deep learning with c-resunet,” Scientific Reports , vol. 11, p. 22920, Nov. 2021
2021
-
[26]
U-net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” May 2015. arXiv:1505.04597 [cs]
2015 arXiv
-
[27]
Microscopy cell counting and detection with fully convolutional regression networks,
W. Xie, J. A. Noble, and A. Zisserman, “Microscopy cell counting and detection with fully convolutional regression networks,” Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization, vol. 6, pp. 283–292, May 2018
2018
-
[28]
Automatic microscopic cell counting by use of deeply-supervised density regression model,
S. He, K. T. Minn, L. Solnica-Krezel, M. Anastasio, and H. Li, “Automatic microscopic cell counting by use of deeply-supervised density regression model,” in Medical Imaging 2019: Digital Pathology , p. 19, Mar. 2019. arXiv:1903.01084 [cs]
2019 arXiv
-
[29]
Efficient and robust cell detection: A structured regression approach,
Y . Xie, F. Xing, X. Shi, X. Kong, H. Su, and L. Yang, “Efficient and robust cell detection: A structured regression approach,” Medical Image Analysis, vol. 44, p. 245–254, Feb. 2018
2018
-
[30]
Weakly supervised learning for cell recognition in immunohistochemical cytoplasm staining images,
S. Zhang, C. Zhu, H. Li, J. Cai, and L. Yang, “Weakly supervised learning for cell recognition in immunohistochemical cytoplasm staining images,” Feb. 2022. arXiv:2202.13372 [cs, eess]
2022 arXiv
-
[31]
Deep convolutional neural networks for human embryonic cell counting,
A. Khan, S. Gould, and M. Salzmann, “Deep convolutional neural networks for human embryonic cell counting,” in Computer Vision – ECCV 2016 Workshops (G. Hua and H. Jégou, eds.), Lecture Notes in Computer Science, (Cham), p. 339–348, Springer International Publishing, 2016
2016
-
[32]
Crowdclip: Unsupervised crowd counting via vision-language model,
D. Liang, J. Xie, Z. Zou, X. Ye, W. Xu, and X. Bai, “Crowdclip: Unsupervised crowd counting via vision-language model,” Apr. 2023. arXiv:2304.04231 [cs]
2023 arXiv
-
[33]
Deep learning and transfer learning for automatic cell counting in microscope images of human cancer cell lines,
F. Lavitt, D. J. Rijlaarsdam, D. van der Linden, E. Weglarz-Tomczak, and J. M. Tomczak, “Deep learning and transfer learning for automatic cell counting in microscope images of human cancer cell lines,” Applied Sciences, vol. 11, p. 4912, Jan. 2021
2021
-
[34]
Context-aware crowd counting,
W. Liu, M. Salzmann, and P. Fua, “Context-aware crowd counting,” Apr
-
[35]
Crowdformer: Weakly-supervised crowd counting with improved generalizability,
S. S. Savner and V . Kanhangad, “Crowdformer: Weakly-supervised crowd counting with improved generalizability,” Mar. 2022. arXiv:2203.03768 [cs]
2022 arXiv
-
[36]
Dtcc: Multi- level dilated convolution with transformer for weakly-supervised crowd counting,
Z. Miao, Y . Zhang, Y . Peng, H. Peng, and B. Yin, “Dtcc: Multi- level dilated convolution with transformer for weakly-supervised crowd counting,” Computational Visual Media , Apr. 2023
2023
-
[37]
Transformer-based visual segmentation: A survey,
X. Li, H. Ding, H. Yuan, W. Zhang, J. Pang, G. Cheng, K. Chen, Z. Liu, and C. C. Loy, “Transformer-based visual segmentation: A survey,” Dec
-
[38]
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y . Wang, Y . Fu, J. Feng, T. Xiang, P. H. S. Torr, and L. Zhang, “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” July 2021. arXiv:2012.15840 [cs]
2021 arXiv
-
[39]
Boosting crowd counting with transformers,
G. Sun, Y . Liu, T. Probst, D. P. Paudel, N. Popovic, and L. Van Gool, “Boosting crowd counting with transformers,” May 2021. arXiv:2105.10926 [cs]
2021 arXiv
-
[40]
Cctrans: Simplifying and improving crowd counting with transformer,
Y . Tian, X. Chu, and H. Wang, “Cctrans: Simplifying and improving crowd counting with transformer,” Sept. 2021. arXiv:2109.14483 [cs]
2021 arXiv
-
[41]
Joint cnn and transformer network via weakly supervised learning for efficient crowd counting,
F. Wang, K. Liu, F. Long, N. Sang, X. Xia, and J. Sang, “Joint cnn and transformer network via weakly supervised learning for efficient crowd counting,” Mar. 2022. arXiv:2203.06388 [cs]
2022 arXiv
-
[42]
Cctwins: A weakly- supervised transformer-based crowd counting method with adaptive scene consistency attention,
L. Dong, H. Zhang, D. Zhou, J. Shi, and J. Ma, “Cctwins: A weakly- supervised transformer-based crowd counting method with adaptive scene consistency attention,” IEEE Transactions on Consumer Electronics , p. 1–1, 2023
2023
-
[43]
Learning to count anything: Reference- less class-agnostic counting with weak supervision,
M. Hobley and V . Prisacariu, “Learning to count anything: Reference- less class-agnostic counting with weak supervision,” Sept. 2022. arXiv:2205.10203 [cs]
2022 arXiv
-
[44]
Weakly-supervised crowd counting learns from sorting rather than locations,
Y . Yang, G. Li, Z. Wu, L. Su, Q. Huang, and N. Sebe, “Weakly-supervised crowd counting learns from sorting rather than locations,” in Computer Vision – ECCV 2020 (A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, eds.), Lecture Notes in Computer Science, (Cham), p. 1–17, Spri...
2020
-
[45]
Towards using count-level weak supervision for crowd counting,
Y . Lei, Y . Liu, P. Zhang, and L. Liu, “Towards using count-level weak supervision for crowd counting,” 2020
2020
-
[46]
Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes,
Y . Li, X. Zhang, and D. Chen, “Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes,” 2018
2018
-
[47]
Bayesian loss for crowd count estimation with point supervision,
Z. Ma, X. Wei, X. Hong, and Y . Gong, “Bayesian loss for crowd count estimation with point supervision,” 2019
2019
-
[48]
Deepvit: Towards deeper vision transformer,
D. Zhou, B. Kang, X. Jin, L. Yang, X. Lian, Z. Jiang, Q. Hou, and J. Feng, “Deepvit: Towards deeper vision transformer,” 2021
2021
-
[49]
Crossvit: Cross-attention multi-scale vision transformer for image classification,
C.-F. Chen, Q. Fan, and R. Panda, “Crossvit: Cross-attention multi-scale vision transformer for image classification,” 2021
2021
-
[50]
Three things everyone should know about vision transformers,
H. Touvron, M. Cord, A. El-Nouby, J. Verbeek, and H. Jégou, “Three things everyone should know about vision transformers,” 2022
2022
-
[51]
Xcit: Cross-covariance image transformers,
A. El-Nouby, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, I. Laptev, N. Neverova, G. Synnaeve, J. Verbeek, and H. Jegou, “Xcit: Cross-covariance image transformers,” 2021
2021
-
[52]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” Dec. 2015. arXiv:1512.03385 [cs]
2015 arXiv
-
[53]
On layer normalization in the transformer architecture,
R. Xiong, Y . Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y . Lan, L. Wang, and T.-Y . Liu, “On layer normalization in the transformer architecture,” June 2020. arXiv:2002.04745 [cs, stat]
2020 arXiv
-
[54]
Computational framework for simulating fluorescence microscope images with cell populations,
A. Lehmussola, P. Ruusuvuori, J. Selinummi, H. Huttunen, and O. Yli- Harja, “Computational framework for simulating fluorescence microscope images with cell populations,” IEEE transactions on medical imaging , vol. 26, pp. 1010–6, 08 2007
2007
-
[2019]
arXiv:1811.10452 [cs]
-
[2023]
arXiv:2304.09854 [cs]
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.