Pith. sign in

REVIEW 4 major objections 6 minor 56 references

Vision Transformers for Weakly-Supervised Microorganism Enumeration

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Transformer vision models can count microorganisms from whole-image labels, though ResNets still lead.

desk verdict Solid, reproducible benchmark; the headline ResNet-over-ViT claim is protocol-dependent and needs qualifiers or pretrained baselines. read the letter →

arxiv 2412.02250 v1 pith:IX6I3EAC submitted 2024-12-03 cs.CV

classification cs.CV
keywords weakly-supervisedcountingvisiontransformersmicroorganismenumerationResNetCrossViTregressionfluorescentmicroscopycell
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that weakly-supervised counting — regressing a whole-image microorganism count directly from image features, with no localization labels — can be carried out by vision transformer backbones at a level close to that of ResNets. It compares several ViT variants, the TransCrowd counting head, and CNN/ResNet baselines on four microscopy datasets, training all models from scratch under the same schedule. The paper argues that while ResNet50 and ResNet101 achieve the lowest errors overall, CrossViT is competitive, wins on their homogeneous artificial bacteria dataset, and is the most compute-efficient ViT tested. This matters because microorganism enumeration is labor-intensive, and a method that needs only global count labels would cut annotation cost while supporting automation.

What carries the argument

The central object is the backbone-plus-regression-head pipeline for weakly-supervised counting: a feature extractor (ResNet, CNN, or vision transformer) maps the image to embeddings, and a single fully connected layer regresses those embeddings to a total count. For ViTs, images are divided into patches ($16\times16$ or $32\times32$) and processed by multi-head self-attention; CrossViT adds a dual-branch design with cross-attention that fuses features from two patch scales. The authors also constructed a synthetic dataset by placing Gaussian-ellipse fluorescent bacteria onto real backgrounds, yielding 12,000 images with counts up to 1,855, to cover dense and homogeneously sparse scenes. Training every architecture from scratch with near-identical hyperparameters is what lets the paper attribute observed differences to architecture choice.

What would settle it

Run the same four datasets with an ImageNet-pretrained CrossViT and an ImageNet-pretrained ResNet50 under identical fine-tuning schedules: if the pretrained CrossViT still fails to beat ResNet50 on the artificial bacteria dataset, the paper's case for ViT competitiveness would be weakened; if it does beat it, the from-scratch protocol would be the reason for the reported gap.

Watch

Extended reading notes

Core claim

On all four datasets, ResNet backbones reach the lowest mean absolute error, with ResNet101 best on fluorescent neurons (MAE 1.400) and ResNet50 best on VGG-cells (MAE 1.225), human cancer cells (MAE 27.206), and the artificial bacteria dataset (MAE 22.030). The ViT family is not far behind in most settings: CrossViT captures the best overall result on the artificial bacteria dataset (MAE 20.011) with the lowest transformer FLOPS ($50.91\times 10^8$), TransCrowd-Token matches ResNet-level accuracy on fluorescent neurons and artificial bacteria, DeepViT is the best ViT on fluorescent neurons, and vanilla ViT is best on VGG-cells. The paper reads this as evidence that ViTs are viable feature extractors for weakly-supervised microorganism enumeration, especially on homogeneous data, and that their gap to ResNets may shrink with pretraining and fine-tuning.

Load-bearing premise

The comparison assumes that training all architectures from scratch with nearly identical hyperparameters is a fair way to judge them, even though vision transformers are known to need pretraining and large datasets to reach their potential.

Editorial extensions

If this is right

  • If ViTs can count microorganisms from global counts, laboratories could skip instance-level annotation and still automate enumeration for contamination monitoring and health-standard checks.
  • CrossViT's result on the artificial bacteria dataset suggests multi-scale patch fusion is a promising direction for dense, uniformly distributed microorganism images.
  • The from-scratch protocol leaves open that pretrained ViTs, fine-tuned in the usual way, could close or reverse the gap to ResNets.
  • The released artificial dataset generator gives other researchers a controlled benchmark for weakly-supervised counting with known density.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because ViTs are known to depend on pretraining, the paper's from-scratch comparison likely understates their ceiling; a fairer test would compare ImageNet-pretrained ViTs against from-scratch ResNets.
  • The CrossViT advantage on homogeneous data may extend to other uniform-texture counting problems, such as colony counting on agar plates or particle counting in environmental samples.
  • A testable follow-up is to vary patch size and fusion strategy on the artificial dataset to isolate whether multi-scale cross-attention or parameter efficiency drives CrossViT's win.
  • Weakly-supervised counting with global labels could combine with active learning to further reduce real-world labeling effort.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents a comparative empirical study of vision transformer (ViT) backbones versus traditional CNN and ResNet architectures for weakly-supervised microorganism counting. Four microscopy datasets are used, including a newly created synthetic fluorescent bacteria dataset. All models are trained from scratch with near-identical hyperparameters, and performance is measured by MAE, RMSE, and FLOPs. The central claim is that, while ResNets achieve the best overall accuracy, ViTs—especially CrossViT—are competitive and occasionally superior on homogeneously distributed data. The manuscript also reports that the TransCrowd model, a state-of-the-art ViT-based counting method, underperforms when trained from scratch, in contrast to its pretrained performance reported in the literature.

Significance. The study addresses a practical and under-explored problem: using weakly-supervised counting for microorganism enumeration, where only global counts are available as labels. Its main strengths are the breadth of architectures compared, the introduction of a synthetic fluorescent bacteria dataset with released generation code, and the clear statement of the from-scratch training protocol, which supports reproducibility. If the comparison were properly controlled, the result that from-scratch ViTs are generally worse than ResNets for this task would be a useful negative result. However, because the protocol deliberately excludes pretraining—which is known to be crucial for ViTs—and because the paper itself concedes this limitation in Section VI, the current evidence does not support the unqualified conclusion stated in the abstract and conclusion. The absence of variance estimates further weakens the significance of the reported ranking. The paper is a plausible starting point for a more rigorous benchmarking study, but as it stands the central claim is not fully supported.

major comments (4)
  1. [Section IV-A and Table III] The central claim that ResNets outperform ViTs is drawn from a protocol where all models are trained from scratch and the paper itself acknowledges in Section VI that 'most studies use pre-trained weights and special fine-tuning for ViT-based approaches' and that TransCrowd, after pretraining, shows state-of-the-art performance 'contrary to the results of our study.' This self-acknowledged limitation means the MAE/RMSE gap in Table III may reflect the training budget rather than architectural merit. To support the abstract's unqualified statement, the authors should either add pretrained ViT baselines (at least for one or two representative architectures) or explicitly restrict the conclusion to the from-scratch setting in both the abstract and the conclusion.
  2. [Table III and Section IV-A] No standard deviations, confidence intervals, or significance tests are reported for any MAE/RMSE value, despite the statement in Section IV-A that 'Experiments were run with different randomization seeds.' With multiple seeds, variance estimates are cheap and necessary to determine whether gaps such as Vanilla ViT's MAE of 1.886 versus ResNet50's 1.225 on VGG-Cells are meaningful. Without such estimates, the performance ranking in Table III is not statistically grounded.
  3. [Section IV-A] The training protocol is not actually 'nearly identical' because the batch size is confounded with the architecture: 128 for CNNs, 64 for Vanilla ViT, CrossViT, TransCrowd, and ResNets, and 32 for Parallel ViT, DeepViT, and XCiT. Batch size affects optimization dynamics, regularization, and final accuracy, so the comparison is not controlled across architectures. The authors should either use a common batch size (if memory permits) or justify the choice and analyze its effect; otherwise, any performance difference cannot be attributed solely to the architecture.
  4. [Section IV-B] The train/validation split procedure is not described. For each dataset, it is not stated how the augmented images were partitioned, what ratio was used, whether the split was stratified by count, or whether the same split was used for all models. This is especially important for the Human Cancer Cells dataset, which has only 1463 augmented images and where results are likely sensitive to the split. The authors should specify the split procedure and ideally run multiple splits or cross-validation to ensure the ranking is robust.
minor comments (6)
  1. [Section II-C] The citation for CCTrans is incorrect: the text says 'named CCTrans [16]', but reference [16] is the CounTR paper; CCTrans should cite reference [40] (Tian et al., CCTrans). Please correct the citation.
  2. [Table I] The column header 'MLP Dim.' for CNNs is described as 'CONVOLUTIONAL OUTPUT'; this is confusing because MLP dimension is not a standard term for CNN channels. Please rename the column, e.g., 'Feature channels' for CNNs and keep 'MLP dim' for ViTs.
  3. [Section IV-A] The paper reports 'average FLOPS (floating point operations per second)' in Table III and Section V-C, but the units (10^8) indicate a count of floating-point operations per inference, not operations per second. If throughput is intended, please report inference time or FPS; if it is FLOPs, please correct the terminology.
  4. [Table II] The 'Data augmentation' column states 'yes' for three datasets but does not describe what augmentation was applied. Please specify the augmentation operations (e.g., random crops, flips, rotations) in the text or in a footnote.
  5. [Section I] The phrase 'ViTs performance demonstrates competent results' should be 'ViTs' performance demonstrates competent results' (add apostrophe). In addition, the sentence is grammatically awkward; consider rephrasing.
  6. [Section III-B-2] The text states 'To achieve higher model complexity without compromising parameter and compute neutrality, Parallel ViT [50] proposes parallelizing the MHSA and feed-forward blocks.' The term 'compute neutrality' is unclear; please rephrase to 'without increasing the number of parameters or FLOPs' for precision.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: this is an empirical benchmark whose conclusions rest on measured MAE and RMSE values, not on quantities defined in terms of the claims they support.

full rationale

This paper is an empirical comparative study, not a derivation from first principles. The central claims—ResNets outperform ViTs overall and CrossViT is competitive on homogeneous data—are supported by measured MAE, RMSE, and FLOPS on four datasets under a common from-scratch training protocol. No prediction is constructed from a fitted parameter that is then renamed, no architecture or loss is defined in terms of the target result, and no load-bearing premise is imported from the authors' own prior work; in fact, the paper contains no self-citations. The artificial fluorescent bacteria dataset is synthetic and its counts are known by construction, but it serves only as benchmark input rather than as evidence for the comparative conclusion, so the evaluation is independent of the claim being tested. The paper's own Section VI caveat—that ViT approaches in prior work typically use pre-trained weights and fine-tuning, and that TransCrowd's original state-of-the-art result is reversed here—is an acknowledged limitation of the training protocol and a threat to external validity, but it is not circularity. There is no equation where an output reduces to an input, and no fitted value is presented as a prediction of the data from which it was derived. The absence of variance estimates and the potential unfairness of the from-scratch protocol are correctness concerns, not circularity concerns, and the manuscript itself concedes the former limitation explicitly. Accordingly, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper's numerical results rest on two unproved domain assumptions: that from-scratch training is a fair comparison for all architectures, and that the synthetic bacteria dataset captures the relevant statistics of real fluorescence microscopy. Neither is justified with independent evidence, and the first is partially contradicted by the paper's own discussion.

assumptions (2)
  • domain assumption Training all models from scratch with roughly identical hyperparameters yields a fair comparison.
    The conclusion that ResNets outperform ViTs depends on this premise; the paper itself notes in Section VI that most ViT studies use pre-trained weights, so the protocol likely disadvantages ViTs.
  • ad hoc to paper The artificial fluorescent bacteria dataset, generated by placing Gaussian ellipses on real backgrounds, is representative of real fluorescent microscopy images.
    The synthetic dataset is used to benchmark models, but no validation is provided that its difficulty matches real bacteria images.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Vision Transformers for Weakly-Supervised Microorganism Enumeration." pith.science (2026). https://pith.science/paper/IX6I3EAC

@misc{pith2026241202250,
  author       = {Pith},
  title        = {Pith review of: Vision Transformers for Weakly-Supervised Microorganism Enumeration},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IX6I3EAC}},
  note         = {Machine review of arXiv:2412.02250}
}
read the original abstract

Microorganism enumeration is an essential task in many applications, such as assessing contamination levels or ensuring health standards when evaluating surface cleanliness. However, it's traditionally performed by human-supervised methods that often require manual counting, making it tedious and time-consuming. Previous research suggests automating this task using computer vision and machine learning methods, primarily through instance segmentation or density estimation techniques. This study conducts a comparative analysis of vision transformers (ViTs) for weakly-supervised counting in microorganism enumeration, contrasting them with traditional architectures such as ResNet and investigating ViT-based models such as TransCrowd. We trained different versions of ViTs as the architectural backbone for feature extraction using four microbiology datasets to determine potential new approaches for total microorganism enumeration in images. Results indicate that while ResNets perform better overall, ViTs performance demonstrates competent results across all datasets, opening up promising lines of research in microorganism enumeration. This comparative study contributes to the field of microbial image analysis by presenting innovative approaches to the recurring challenge of microorganism enumeration and by highlighting the capabilities of ViTs in the task of regression counting.

Figures

Figures reproduced from arXiv: 2412.02250 by the authors.

Figure 1
Figure 1. Comparison of methodologies in deep learning regarding instance [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. In ViT approaches for WSC, the ViT is used as a backbone, or feature [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Random samples from the four datasets used in this study. Although [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 43 canonical work pages

  1. [1]

    High-throughput imaging of bacterial colonies grown on filter plates with application to serum bactericidal assays,

    X. Liu, S. Wang, L. Sendi, and M. J. Caulfield, “High-throughput imaging of bacterial colonies grown on filter plates with application to serum bactericidal assays,” Journal of immunological methods, vol. 292, no. 1-2, pp. 187–193, 2004

  2. [2]

    Applicability of solid-phase cytometry and epifluorescence microscopy for rapid assessment of the microbio- logical quality of dialysis water,

    M. Riepl, S. Schauer, S. Knetsch, E. Holzhammer, A. H. Farnleitner, R. Sommer, and A. K. Kirschner, “Applicability of solid-phase cytometry and epifluorescence microscopy for rapid assessment of the microbio- logical quality of dialysis water,” Nephrology Dialysis Transplantation, vol. 26, no. 11, pp. 3640–3645, 2011

  3. [3]

    Use of fluorochromes for direct enumeration of total bacteria in environmental samples: past and present,

    R. L. Kepner Jr and J. R. Pratt, “Use of fluorochromes for direct enumeration of total bacteria in environmental samples: past and present,” Microbiological reviews, vol. 58, no. 4, pp. 603–615, 1994

  4. [4]

    R. A. Herbert, 1 Methods for Enumerating Microorganisms and De- termining Biomass in Natural Environments , vol. 22 of Techniques in Microbial Ecology, p. 1–39. Academic Press, Jan. 1990

  5. [5]

    Horwitz, Official methods of analysis of the Association of Official Analytical Chemists

    W. Horwitz, Official methods of analysis of the Association of Official Analytical Chemists . Washington, DC: The Association of Official Analytical Chemists, 1970

  6. [6]

    Absher, CHAPTER 1 - Hemocytometer Counting , p

    M. Absher, CHAPTER 1 - Hemocytometer Counting , p. 395–397. Academic Press, Jan. 1973

  7. [7]

    Chapter 2 - Techniques for Oral Microbiology , p. 15–40. Oxford: Academic Press, Jan. 2015

  8. [8]

    A comprehensive review of image analysis methods for microorganism counting: from classical image processing to deep learning approaches,

    J. Zhang, C. Li, M. M. Rahaman, Y . Yao, P. Ma, J. Zhang, X. Zhao, T. Jiang, and M. Grzegorzek, “A comprehensive review of image analysis methods for microorganism counting: from classical image processing to deep learning approaches,” Artificial Intelligence Review , vol. 55, pp. 2875–2944, Apr. 2022

Show all 56 references
  1. [9]

    DeepBacs: Bacterial image analysis using open-source deep learning approaches,

    C. Spahn, R. F. Laine, P. M. Pereira, E. Gómez-de Mariscal, L. V on Chamier, M. Conduit, M. G. De Pinho, G. Jacquemet, S. Holden, M. Heilemann, and R. Henriques, “DeepBacs: Bacterial image analysis using open-source deep learning approaches,” preprint, Microbiology, Nov. 2021. 8

  2. [10]

    Democratising deep learning for microscopy with ZeroCostDL4Mic,

    L. V on Chamier, R. F. Laine, J. Jukkala, C. Spahn, D. Krentzel, E. Nehme, M. Lerche, S. Hernández-Pérez, P. K. Mattila, E. Karinou, S. Holden, A. C. Solak, A. Krull, T.-O. Buchholz, M. L. Jones, L. A. Royer, C. Leterrier, Y . Shechtman, F. Jug, M. Heilemann, G. Jacquemet, and...

  3. [11]

    Automatic bacillus anthracis bacteria detection and segmentation in microscopic images using unet++,

    F. Hoorali, H. Khosravi, and B. Moradi, “Automatic bacillus anthracis bacteria detection and segmentation in microscopic images using unet++,” Journal of Microbiological Methods , vol. 177, p. 106056, Oct. 2020

  4. [12]

    Deeply-Supervised Density Regression for Automatic Cell Counting in Microscopy Images,

    S. He, K. T. Minn, L. Solnica-Krezel, M. A. Anastasio, and H. Li, “Deeply-Supervised Density Regression for Automatic Cell Counting in Microscopy Images,” Nov. 2020. arXiv:2011.03683 [cs, eess]

  5. [13]

    Evaluation of two methods for monitoring surface cleanliness-atp bioluminescence and traditional hygiene swabbing.,

    C. A. Davidson, C. J. Griffith, A. C. Peters, and L. Fielding, “Evaluation of two methods for monitoring surface cleanliness-atp bioluminescence and traditional hygiene swabbing.,” Luminescence : the journal of biological and chemical luminescence , vol. 14 1, pp. 33–8, 1999

  6. [14]

    Cell Counting by Regression Using Convolutional Neural Network,

    Y . Xue, N. Ray, J. Hugh, and G. Bigras, “Cell Counting by Regression Using Convolutional Neural Network,” in Computer Vision – ECCV 2016 Workshops (G. Hua and H. Jégou, eds.), vol. 9913, pp. 274–290, Cham: Springer International Publishing, 2016. Series Title: Lecture Notes i...

  7. [15]

    Classification Beats Regression: Counting of Cells from Greyscale Microscopic Images based on Annotation-free Training Samples,

    X. Ding, Q. Zhang, and W. J. Welch, “Classification Beats Regression: Counting of Cells from Greyscale Microscopic Images based on Annotation-free Training Samples,” Oct. 2020. arXiv:2010.14782 [cs, eess]

  8. [16]

    Countr: Transformer-based generalised visual counting,

    C. Liu, Y . Zhong, A. Zisserman, and W. Xie, “Countr: Transformer-based generalised visual counting,” June 2023. arXiv:2208.13721 [cs]

  9. [17]

    Transcrowd: weakly- supervised crowd counting with transformers,

    D. Liang, X. Chen, W. Xu, Y . Zhou, and X. Bai, “Transcrowd: weakly- supervised crowd counting with transformers,” Science China Information Sciences, vol. 65, p. 160104, June 2021. arXiv:2104.09116 [cs]

  10. [18]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” 2021

  11. [19]

    Automating cell counting in fluorescent microscopy through deep learning with c-resunet,

    R. Morelli, L. Clissa, R. Amici, M. Cerri, T. Hitrec, M. Luppi, L. Rinaldi, F. Squarcio, and A. Zoccoli, “Automating cell counting in fluorescent microscopy through deep learning with c-resunet,” Scientific Reports , vol. 11, p. 22920, 11 2021

  12. [20]

    Learning to count objects in images,

    V . Lempitsky and A. Zisserman, “Learning to count objects in images,” in Advances in Neural Information Processing Systems (J. Lafferty, C. Williams, J. Shawe-Taylor, R. Zemel, and A. Culotta, eds.), vol. 23, Curran Associates, Inc., 2010

  13. [21]

    Microscope images of human cancer cell lines (u2os and hl-60),

    F. Lavitt, D. J. Rijlaarsdam, D. v. d. Linden, E. Weglarz-Tomczak, and J. M. Tomczak, “Microscope images of human cancer cell lines (u2os and hl-60),” Jan. 2021

  14. [22]

    Transformers in Vision: A Survey,

    S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in Vision: A Survey,” ACM Computing Surveys , vol. 54, pp. 1–41, Jan. 2022. arXiv:2101.01169 [cs]

  15. [23]

    Monitoring micorbial morphogenetic changes in a fermentation process by a self- tuning vision system (stvs),

    Y . Shabtai, M. Ronen, I. Mukmenev, and H. Guterman, “Monitoring micorbial morphogenetic changes in a fermentation process by a self- tuning vision system (stvs),” Computers & Chemical Engineering, vol. 20, p. S321–S326, Jan. 1996

  16. [24]

    U-net: deep learning for cell counting, detection, and morphometry,

    T. Falk, D. Mai, R. Bensch, O. Çiçek, A. Abdulkadir, Y . Marrakchi, A. Böhm, J. Deubner, Z. Jäckel, K. Seiwald, A. Dovzhenko, O. Tietz, C. Dal Bosco, S. Walsh, D. Saltukoglu, T. L. Tay, M. Prinz, K. Palme, M. Simons, I. Diester, T. Brox, and O. Ronneberger, “U-net: deep learni...

  17. [25]

    Automating cell counting in fluorescent microscopy through deep learning with c-resunet,

    R. Morelli, L. Clissa, R. Amici, M. Cerri, T. Hitrec, M. Luppi, L. Rinaldi, F. Squarcio, and A. Zoccoli, “Automating cell counting in fluorescent microscopy through deep learning with c-resunet,” Scientific Reports , vol. 11, p. 22920, Nov. 2021

  18. [26]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” May 2015. arXiv:1505.04597 [cs]

  19. [27]

    Microscopy cell counting and detection with fully convolutional regression networks,

    W. Xie, J. A. Noble, and A. Zisserman, “Microscopy cell counting and detection with fully convolutional regression networks,” Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization, vol. 6, pp. 283–292, May 2018

  20. [28]

    Automatic microscopic cell counting by use of deeply-supervised density regression model,

    S. He, K. T. Minn, L. Solnica-Krezel, M. Anastasio, and H. Li, “Automatic microscopic cell counting by use of deeply-supervised density regression model,” in Medical Imaging 2019: Digital Pathology , p. 19, Mar. 2019. arXiv:1903.01084 [cs]

  21. [29]

    Efficient and robust cell detection: A structured regression approach,

    Y . Xie, F. Xing, X. Shi, X. Kong, H. Su, and L. Yang, “Efficient and robust cell detection: A structured regression approach,” Medical Image Analysis, vol. 44, p. 245–254, Feb. 2018

  22. [30]

    Weakly supervised learning for cell recognition in immunohistochemical cytoplasm staining images,

    S. Zhang, C. Zhu, H. Li, J. Cai, and L. Yang, “Weakly supervised learning for cell recognition in immunohistochemical cytoplasm staining images,” Feb. 2022. arXiv:2202.13372 [cs, eess]

  23. [31]

    Deep convolutional neural networks for human embryonic cell counting,

    A. Khan, S. Gould, and M. Salzmann, “Deep convolutional neural networks for human embryonic cell counting,” in Computer Vision – ECCV 2016 Workshops (G. Hua and H. Jégou, eds.), Lecture Notes in Computer Science, (Cham), p. 339–348, Springer International Publishing, 2016

  24. [32]

    Crowdclip: Unsupervised crowd counting via vision-language model,

    D. Liang, J. Xie, Z. Zou, X. Ye, W. Xu, and X. Bai, “Crowdclip: Unsupervised crowd counting via vision-language model,” Apr. 2023. arXiv:2304.04231 [cs]

  25. [33]

    Deep learning and transfer learning for automatic cell counting in microscope images of human cancer cell lines,

    F. Lavitt, D. J. Rijlaarsdam, D. van der Linden, E. Weglarz-Tomczak, and J. M. Tomczak, “Deep learning and transfer learning for automatic cell counting in microscope images of human cancer cell lines,” Applied Sciences, vol. 11, p. 4912, Jan. 2021

  26. [34]

    Context-aware crowd counting,

    W. Liu, M. Salzmann, and P. Fua, “Context-aware crowd counting,” Apr

  27. [35]

    Crowdformer: Weakly-supervised crowd counting with improved generalizability,

    S. S. Savner and V . Kanhangad, “Crowdformer: Weakly-supervised crowd counting with improved generalizability,” Mar. 2022. arXiv:2203.03768 [cs]

  28. [36]

    Dtcc: Multi- level dilated convolution with transformer for weakly-supervised crowd counting,

    Z. Miao, Y . Zhang, Y . Peng, H. Peng, and B. Yin, “Dtcc: Multi- level dilated convolution with transformer for weakly-supervised crowd counting,” Computational Visual Media , Apr. 2023

  29. [37]

    Transformer-based visual segmentation: A survey,

    X. Li, H. Ding, H. Yuan, W. Zhang, J. Pang, G. Cheng, K. Chen, Z. Liu, and C. C. Loy, “Transformer-based visual segmentation: A survey,” Dec

  30. [38]

    Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,

    S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y . Wang, Y . Fu, J. Feng, T. Xiang, P. H. S. Torr, and L. Zhang, “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” July 2021. arXiv:2012.15840 [cs]

  31. [39]

    Boosting crowd counting with transformers,

    G. Sun, Y . Liu, T. Probst, D. P. Paudel, N. Popovic, and L. Van Gool, “Boosting crowd counting with transformers,” May 2021. arXiv:2105.10926 [cs]

  32. [40]

    Cctrans: Simplifying and improving crowd counting with transformer,

    Y . Tian, X. Chu, and H. Wang, “Cctrans: Simplifying and improving crowd counting with transformer,” Sept. 2021. arXiv:2109.14483 [cs]

  33. [41]

    Joint cnn and transformer network via weakly supervised learning for efficient crowd counting,

    F. Wang, K. Liu, F. Long, N. Sang, X. Xia, and J. Sang, “Joint cnn and transformer network via weakly supervised learning for efficient crowd counting,” Mar. 2022. arXiv:2203.06388 [cs]

  34. [42]

    Cctwins: A weakly- supervised transformer-based crowd counting method with adaptive scene consistency attention,

    L. Dong, H. Zhang, D. Zhou, J. Shi, and J. Ma, “Cctwins: A weakly- supervised transformer-based crowd counting method with adaptive scene consistency attention,” IEEE Transactions on Consumer Electronics , p. 1–1, 2023

  35. [43]

    Learning to count anything: Reference- less class-agnostic counting with weak supervision,

    M. Hobley and V . Prisacariu, “Learning to count anything: Reference- less class-agnostic counting with weak supervision,” Sept. 2022. arXiv:2205.10203 [cs]

  36. [44]

    Weakly-supervised crowd counting learns from sorting rather than locations,

    Y . Yang, G. Li, Z. Wu, L. Su, Q. Huang, and N. Sebe, “Weakly-supervised crowd counting learns from sorting rather than locations,” in Computer Vision – ECCV 2020 (A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, eds.), Lecture Notes in Computer Science, (Cham), p. 1–17, Spri...

  37. [45]

    Towards using count-level weak supervision for crowd counting,

    Y . Lei, Y . Liu, P. Zhang, and L. Liu, “Towards using count-level weak supervision for crowd counting,” 2020

  38. [46]

    Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes,

    Y . Li, X. Zhang, and D. Chen, “Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes,” 2018

  39. [47]

    Bayesian loss for crowd count estimation with point supervision,

    Z. Ma, X. Wei, X. Hong, and Y . Gong, “Bayesian loss for crowd count estimation with point supervision,” 2019

  40. [48]

    Deepvit: Towards deeper vision transformer,

    D. Zhou, B. Kang, X. Jin, L. Yang, X. Lian, Z. Jiang, Q. Hou, and J. Feng, “Deepvit: Towards deeper vision transformer,” 2021

  41. [49]

    Crossvit: Cross-attention multi-scale vision transformer for image classification,

    C.-F. Chen, Q. Fan, and R. Panda, “Crossvit: Cross-attention multi-scale vision transformer for image classification,” 2021

  42. [50]

    Three things everyone should know about vision transformers,

    H. Touvron, M. Cord, A. El-Nouby, J. Verbeek, and H. Jégou, “Three things everyone should know about vision transformers,” 2022

  43. [51]

    Xcit: Cross-covariance image transformers,

    A. El-Nouby, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, I. Laptev, N. Neverova, G. Synnaeve, J. Verbeek, and H. Jegou, “Xcit: Cross-covariance image transformers,” 2021

  44. [52]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” Dec. 2015. arXiv:1512.03385 [cs]

  45. [53]

    On layer normalization in the transformer architecture,

    R. Xiong, Y . Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y . Lan, L. Wang, and T.-Y . Liu, “On layer normalization in the transformer architecture,” June 2020. arXiv:2002.04745 [cs, stat]

  46. [54]

    Computational framework for simulating fluorescence microscope images with cell populations,

    A. Lehmussola, P. Ruusuvuori, J. Selinummi, H. Huttunen, and O. Yli- Harja, “Computational framework for simulating fluorescence microscope images with cell populations,” IEEE transactions on medical imaging , vol. 26, pp. 1010–6, 08 2007

  47. [2019]

    arXiv:1811.10452 [cs]

  48. [2023]

    arXiv:2304.09854 [cs]

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.