REVIEW 4 major objections 7 minor 39 references
IDCIA: Immunocytochemistry Dataset for Cellular Image Analysis
T0 review · 4 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A new 262-image immunolabeled cell dataset shows automated counters still fall short.
desk verdict A useful new multi-antibody cell counting dataset and baseline study; the main caveat is unvalidated manual ground truth. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two things carry the argument. First, the ground-truth format: dot annotations mark each cell, and the paper argues this avoids double counting, is fast to acquire, and supports both detection- and regression-based methods. Second, the evaluation protocol: dot annotations are converted to geometry-adaptive Gaussian density maps for training, and performance is judged by mean absolute error and by a new metric, ACP (Acceptable Error Count Percent), which counts the fraction of images whose predicted count is within 5 percent of the true count. ACP is what converts average error into acceptable to a biologist, and it is the metric on which every tested model fails.
What would settle it
Take a random sample of IDCIA images and have two independent annotators, plus one experienced cell biologist, re-count every cell without seeing the original labels; if inter-annotator agreement is low, or the original counts differ systematically from the expert recount, the conclusion that no automated model is acceptable cannot be separated from errors in the gold standard.
Extended reading notes
Core claim
The central contribution is the dataset itself plus a baseline evaluation. IDCIA contains 262 grayscale 800x600 images of rat adult hippocampal progenitor cells labeled with DAPI and six neural markers (TuJ1, MAP2ab, RIP, GFAP, Nestin, Ki67), with manual dot annotations converted to cell coordinates. Five regression- or density-based deep models were trained and tested with a stratified split, and none met the acceptance standard of a 5 percent error: Count-ception had the lowest mean absolute error (15.47), while CSRNet had the highest acceptable-error percentage (17 percent), and MCNN scored 0 percent on that metric. The paper concludes that existing counting models are not yet accurate enough to replace manual cell counting and that performance varies strongly by antibody stain.
Load-bearing premise
The whole benchmark rests on the manual dot annotations being correct, but the paper reports no check of agreement between annotators or a re-count by an expert, so if those labels are noisy or biased the reported error rates are not a clean measure of model accuracy.
Editorial extensions
If this is right
- The dataset gives the community a public, multi-antibody benchmark for cell counting, with per-image antibody labels that can be used to diagnose where models fail.
- The baseline results establish that density-map methods transfer only partially from crowd counting to immunolabeled stem cells, motivating architectures tuned for antibody-specific appearance.
- Since no model reaches the 5 percent tolerance, the paper's own stated target is a new counting method that does; the dataset is positioned as the training and evaluation ground for that effort.
- The dot annotations are reusable beyond counting, as weak supervision for segmentation or detection, and the antibody label supports an auxiliary classification task.
Reading between the lines
- Beyond the paper: if ACP becomes a standard metric, automated counting systems will be judged by clinical tolerance rather than by mean error, which is a stricter and more decision-relevant bar.
- Beyond the paper: the absence of reported inter-annotator agreement means the benchmark's ceiling could change under a re-annotation study; a second annotation pass on a sample would reveal how much of the 17 percent is label noise.
- Beyond the paper: the strong per-antibody performance differences suggest antibody-aware or stain-conditional models, not a single generic counter, may be the fastest route to expert-acceptable accuracy.
- Beyond the paper: because the images come from a specific electrical-stimulation protocol and scaffold, models trained on IDCIA may need domain adaptation before transferring to other cultures or imaging setups.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces IDCIA, a new annotated fluorescence microscopy dataset of rat adult hippocampal progenitor cells, containing 262 images at 800x600 resolution, stained with DAPI and six primary antibodies (TuJ1, MAP2ab, RIP, GFAP, Nestin, Ki67), with per-image cell counts and dot-annotation coordinates. The authors describe the data collection and annotation pipeline, the dataset structure, suggested metrics (MAE, RMSE, and a new ACP metric based on a 5% acceptable-error threshold), and baseline experiments with five deep-learning counting models (CNN regression, CSRNet, MCNN, Count-ception, FCRN-A) using a stratified 60:20:20 train/validation/test split, grid search for hyperparameters, and five training runs per model. The main claims are that IDCIA covers more staining methods than existing cell-counting datasets and that none of the five evaluated models achieves acceptable counting accuracy, with the best ACP at 17% (CSRNet) and the best MAE at 15.47 (Count-ception). The dataset, code, and trained models are publicly available.
Significance. If the ground-truth annotations are valid, IDCIA is a useful new benchmark for cell counting under diverse immunolabeling conditions, and the paper provides a reproducible evaluation protocol with public data and code. The finding that existing counting models fail the proposed ACP criterion would be a valuable cautionary result for the community. However, the central benchmark conclusion currently rests on an unvalidated manual annotation process: no inter-annotator agreement or expert re-count is reported, and the reported performance metrics lack uncertainty estimates. These issues are addressable and do not negate the dataset's existence, but they must be fixed before the paper's main claims can be accepted.
major comments (4)
- [Section 3.1 / Eq. (3) / Table 5] The ground-truth cell counts are the output of manual dot annotation by a group of undergraduate students, and the manuscript reports no inter-annotator agreement, no expert re-count, and no quantitative quality-control statistic. Because ACP (Eq. 3) is defined as the fraction of predictions within 5% of the manual count, unvalidated manual counts directly undermine the headline result that the best model reaches only 17% ACP and the conclusion that no model can replace manual counting. The statement that imaging and counting were performed blind guards against bias about stimulation condition, but it does not validate counting accuracy. I recommend reporting inter-annotator agreement (e.g., count-level ICC or Cohen's kappa on a subset), an expert re-count of a random sample, or an explicit error-propagation analysis showing how annotation noise affects ACP.
- [Tables 4, 5, and 6] Tables 4 and 6 report only average MAE over five runs without standard deviations or confidence intervals, and Table 5 reports ACP as a single point per model with no uncertainty. With only five runs and small test partitions, the reader cannot assess whether the ranking among models is stable. For example, Table 6's per-antibody MAEs are averaged over at most a handful of test images (the GFAP antibody has 23 images total, so about 4-5 test images under the 60:20:20 split). Please report mean ± std or 95% confidence intervals for all metrics and explicitly state the number of test images per antibody subset.
- [Section 5 / Table 6] The per-antibody baseline results are based on very small test sets: antibodies with 23-25 images contribute only roughly 4-5 test images each. Consequently, the claims that CNN Regression is best on RIP and GFAP and that MCNN is best on Ki67 are not statistically supported; a single image can shift the MAE substantially. I recommend either reporting confidence intervals, aggregating antibodies with small sample sizes, or clearly labeling these results as preliminary observations rather than definitive rankings.
- [Section 4 / Eq. (3)] The 5% acceptable-error threshold used in the ACP metric is described as 'the experts' acceptable error rate' without citation or derivation. Since ACP is the metric used to conclude that all models are unacceptable, this threshold is load-bearing. Please justify the 5% value with a reference or a domain-expert survey, and report sensitivity to the threshold (e.g., ACP at 10% and 15%) so readers can judge how robust the conclusion is.
minor comments (7)
- [Abstract] The phrase 'protein components of immune responses against invaders' is an imprecise definition of antibodies; consider simplifying to 'proteins used to detect specific markers' or similar.
- [Section 1, paragraph 3] The sentence describing scaffolds contains a repetition: '...scaffolds, which are structures providing support for cells to grow within an interdigitated electrode region. Then the voltage is applied to the electrode pads of the scaffold, which are structures providing support for cells to grow.' Please remove the duplicated description.
- [References] References [4] and [26] both cite the same Count-ception paper (Cohen et al., ICCVW 2017); one entry should be removed or the citations merged.
- [Author affiliations] The affiliation line for Surya K. Mallapragada contains an extraneous 'pilcrow' symbol, likely a LaTeX error, and the contact email 'abdu@iastate.com' should probably be 'abdu@iastate.edu'.
- [Figure 2] The caption labels rows inconsistently: 'Row 1 (A-C)' but then describes DAPI in (B and C), while 'Row 2 (D-E)' is used for the dot-annotated images; please clarify the panel layout and the color assignments.
- [Eq. (3)] When the ground-truth count y_i is zero, the ACP condition becomes |prediction| ≤ 0, which counts only exact zero predictions; this edge case should be addressed explicitly since the dataset has images with zero cells.
- [Table 3] The mean cell count for RIP is reported as '49.542' with three decimals while other entries use two; please format consistently.
Circularity Check
No significant circularity: IDCIA is a new annotated dataset and the baseline evaluations are empirical measurements, not derived quantities that reduce to their own inputs.
full rationale
The paper's central claims are that IDCIA is a new annotated immunofluorescence dataset and that five existing counting models do not reach a domain-acceptable error rate on it. Neither claim is circular. The dataset's ground truth comes from manual dot annotations with ImageJ, and the model outputs are evaluated against those annotations using standard MAE and the introduced ACP metric; no fitted parameter is renamed as a prediction and no result is defined in terms of the quantity it is supposed to establish. The ACP definition (Eq. 3) is a thresholded error measure, not a construction that forces the observed 0-17 percent values. The self-citations in the introductory background ([6], [32], [33]) describe earlier electrical-stimulation work from the same labs, but the dataset's novelty and the baseline conclusions do not depend on those citations as evidence; the comparison is against externally published models (CSRNet, MCNN, Count-ception, FCRN-A) using their own code and open protocol. The dataset's practical value and the 'no model is acceptable' conclusion do rest on unvalidated manual annotations, as no inter-annotator agreement or expert re-count is reported, but that is a data-quality limitation rather than a circular-reasoning defect: the benchmark values would remain empirical even if the gold standard were unreliable. No derivation chain in the paper reduces Eq. X to Eq. Y by construction, and no load-bearing conclusion is imported solely from an author self-citation. The honest finding is therefore no significant circularity.
Assumptions & free parameters
free parameters (2)
- Acceptable Error Count Percent threshold =
0.05 (5%)
- Per-model training hyperparameters (learning rate, batch size, epochs) =
See Table 4, e.g., Count-ception: lr=1e-2, batch=2, 1000 epochs
assumptions (4)
- domain assumption Manual dot annotations produced with ImageJ Cell Counter are treated as ground truth cell locations and counts.
- domain assumption The 5 percent acceptable error threshold reflects genuine expert requirements.
- domain assumption Gaussian density maps generated from dot annotations are valid training targets for CSRNet and MCNN.
- standard math Standard deep learning training assumptions, including SGD, L1 loss, and data augmentation, transfer to this microscopy domain.
Cite this review
Pith. "Pith review of IDCIA: Immunocytochemistry Dataset for Cellular Image Analysis." pith.science (2026). https://pith.science/paper/Z74AF7JR
@misc{pith2026241108992,
author = {Pith},
title = {Pith review of: IDCIA: Immunocytochemistry Dataset for Cellular Image Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z74AF7JR}},
note = {Machine review of arXiv:2411.08992}
}
read the original abstract
We present a new annotated microscopic cellular image dataset to improve the effectiveness of machine learning methods for cellular image analysis. Cell counting is an important step in cell analysis. Typically, domain experts manually count cells in a microscopic image. Automated cell counting can potentially eliminate this tedious, time-consuming process. However, a good, labeled dataset is required for training an accurate machine learning model. Our dataset includes microscopic images of cells, and for each image, the cell count and the location of individual cells. The data were collected as part of an ongoing study investigating the potential of electrical stimulation to modulate stem cell differentiation and possible applications for neural repair. Compared to existing publicly available datasets, our dataset has more images of cells stained with more variety of antibodies (protein components of immune responses against invaders) typically used for cell analysis. The experimental results on this dataset indicate that none of the five existing models under this study are able to achieve sufficiently accurate count to replace the manual methods. The dataset is available at https://figshare.com/articles/dataset/Dataset/21970604.
Figures
Reference graph
Works this paper leans on
-
[1]
Yousef Al-Kofahi, Alla Zaltsman, Robert Graves, Will Marshall, and Mirabela Rusu. 2018. A deep learning-based algorithm for 2-D cell segmentation in microscopy images. BMC Bioinformatics 19, 1 (Oct. 2018), 365. https://doi.org/ 10.1186/s12859-018-2375-z
-
[2]
Christopher M Bishop. 2006. Pattern recognition and machine learning . New York : Springer, [2006] ©2006. https://search.library.wisc.edu/catalog/ 9910032530902121
work page 2006
-
[3]
R. A. Bradshaw and P. D. Stahl. 2016. Cell Biology: An Overview. In Refer- ence Module in Biomedical Sciences . Elsevier. https://doi.org/10.1016/B978-0-12- 801238-3.99496-0
-
[4]
Joseph Paul Cohen, Geneviève Boucher, Craig A. Glastonbury, Henry Z. Lo, and Yoshua Bengio. 2017. Count-ception: Counting by Fully Convolutional Redundant Counting. In 2017 IEEE International Conference on Computer Vision Workshops (ICCVW). 18–26. https://doi.org/10.1109/ICCVW.2017.9
-
[5]
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus En- zweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. 2016. The Cityscapes Dataset for Semantic Urban Scene Understanding. https: //doi.org/10.48550/arXiv.1604.01685
-
[6]
Das, Metin Uz, Shaowei Ding, Matthew T
Suprem R. Das, Metin Uz, Shaowei Ding, Matthew T. Lentner, John A. Hondred, Allison A. Cargill, Donald S. Sakaguchi, Surya Mallapragada, and Jonathan C. Claussen. 2017. Electrical Differentiation of Mesenchymal Stem Cells into IDCIA: Immunocytochemistry Dataset for Cellular Image Analysis Schwann-Cell-Like Phenotypes Using Inkjet-Printed Graphene Circuits...
doi:10.1002/adhm 2017
-
[7]
Christoffer Edlund, Timothy R. Jackson, Nabeel Khalid, Nicola Bevan, Timothy Dale, Andreas Dengel, Sheraz Ahmed, Johan Trygg, and Rickard Sjögren. 2021. LIVECell—A large-scale dataset for label-free live cell segmentation. Nature Methods 18, 9 (Sept. 2021), 1038–1045. https://doi.org/10.1038/s41592-021-01249- 6 Number: 9 Publisher: Nature Publishing Group
-
[8]
Everingham, L
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. 2010. The Pascal Visual Object Classes (VOC) Challenge. International Journal of Computer Vision 88, 2 (June 2010), 303–338
2010
Show all 39 references
-
[9]
Seiya Fujita and Xian-Hua Han. 2020. Cell Detection and Segmentation in Microscopy Images with Improved Mask R-CNN. In Proceedings of the Asian Conference on Computer Vision (ACCV) Workshops
2020
-
[10]
Ali Ghaznavi, Renata Rychtarikova, Mohammadmehdi Saberioon, and Dalibor Stys. 2022. Cell segmentation from telecentric bright-field transmitted light microscopy images using a Residual Attention U-Net: a case study on HeLa line. Computers in Biology and Medicine 147 (Aug. 2022...
2022
-
[11]
Yuki Hiramatsu, Kazuhiro Hotta, Ayako Imanishi, Michiyuki Matsuda, and Kenta Terai. 2018. Cell Image Segmentation by Integrating Multiple CNNs. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 2286–22866. https://doi.org/10.1109/CVPRW.2...
2018
-
[12]
Hao Jiang, Sen Li, Weihuang Liu, Hongjin Zheng, Jinghao Liu, and Yang Zhang
-
[13]
Philipp Kainz, Martin Urschler, Samuel Schulter, Paul Wohlhart, and Vincent Lepetit. 2015. You Should Use Regression to Detect Cells. In Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 (Lecture Notes in Computer Science), Nassir Navab, Joachim Hornegge...
2015 doi
-
[14]
Aisha Khan, Stephen Gould, and Mathieu Salzmann. 2016. Deep Convolutional Neural Networks for Human Embryonic Cell Counting. In Computer Vision – ECCV 2016 Workshops (Lecture Notes in Computer Science) , Gang Hua and Hervé Jégou (Eds.). Springer International Publishing, Cham,...
2016 doi
-
[15]
Kirby, Akela A
Elizabeth D. Kirby, Akela A. Kuwahara, Reanna L. Messer, and Tony Wyss- Coray. 2015. Adult hippocampal neural stem and progenitor cells regulate the neurogenic niche by secreting VEGF. Proceedings of the National Academy of Sciences of the United States of America 112, 13 (Mar...
2015 doi
-
[16]
Ambros, Inge M
Florian Kromp, Eva Bozsaky, Fikret Rifatbegovic, Lukas Fischer, Magdalena Ambros, Maria Berneder, Tamara Weiss, Daria Lazic, Wolfgang Dörr, Allan Hanbury, Klaus Beiske, Peter F. Ambros, Inge M. Ambros, and Sabine Taschner- Mandl. 2020. An annotated fluorescence image dataset f...
2020
-
[17]
Rijlaarsdam, Dennet van der Linden, Ewelina Weglarz- Tomczak, and Jakub M
Falko Lavitt, Demi J. Rijlaarsdam, Dennet van der Linden, Ewelina Weglarz- Tomczak, and Jakub M. Tomczak. 2021. Deep Learning and Transfer Learning for Automatic Cell Counting in Microscope Images of Human Cancer Cell Lines. Applied Sciences 11, 11 (Jan. 2021), 4912. https://d...
2021 doi
-
[18]
Victor Lempitsky and Andrew Zisserman. 2010. Learning To Count Ob- jects in Images. In Advances in Neural Information Processing Systems , Vol. 23. Curran Associates, Inc. https://papers.nips.cc/paper/2010/hash/ fe73f687e5bc5280214e0486b273a5f9-Abstract.html
2010
-
[19]
Yuhong Li, Xiaofan Zhang, and Deming Chen. 2018. CSRNet: Dilated Convolu- tional Neural Networks for Understanding the Highly Congested Scenes. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1091–1100. https://doi.org/10.1109/CVPR.2018.00120 ISSN: 2575-7075
2018
-
[20]
Lawrence Zitnick, and Piotr Dollár
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár
-
[21]
Sokolnicki, and Anne E
Vebjorn Ljosa, Katherine L. Sokolnicki, and Anne E. Carpenter. 2012. Annotated high-throughput microscopy image sets for validation. Nature Methods 9, 7 (July 2012), 637–637. https://doi.org/10.1038/nmeth.2083 Number: 7 Publisher: Nature Publishing Group
2012 doi
-
[22]
Fu, Shenghua He, Sabine Dietmann, Steven C
Kyaw Thu Minn, Yuheng C. Fu, Shenghua He, Sabine Dietmann, Steven C. George, Mark A. Anastasio, Samantha A. Morris, and Lilianna Solnica-Krezel
-
[23]
Roberto Morelli, Luca Clissa, Roberto Amici, Matteo Cerri, Timna Hitrec, Marco Luppi, Lorenzo Rinaldi, Fabio Squarcio, and Antonio Zoccoli. 2021. Automating cell counting in fluorescent microscopy through deep learning with c-ResUnet. Scientific Reports 11, 1 (Nov. 2021), 2292...
2021 doi
-
[24]
Sapun Parekh and Schwendy Mischa. 2022. EVICAN Dataset. https://doi.org/ 10.17617/3.AJBV1S Type: dataset
2022 doi
-
[25]
eLife 9 (Nov
High-resolution transcriptional and morphogenetic profiling of cells from micropatterned human ESC gastruloid cultures. eLife 9 (Nov. 2020), e59445. https://doi.org/10.7554/eLife.59445
2020 doi
-
[26]
Glastonbury, Henry Z
Joseph Paul Cohen, Genevieve Boucher, Craig A. Glastonbury, Henry Z. Lo, and Yoshua Bengio. 2017. Count-ception: Counting by Fully Convolutional Redundant Counting. 18–26. https://openaccess. thecvf.com/content_ICCV_2017_workshops/w1/html/Cohen_Count- ception_Counting_by_ICCV_...
2017
- [27]
- [28]
-
[29]
Ling Shao, Fan Zhu, and Xuelong Li. 2015. Transfer learning for visual catego- rization: a survey. IEEE transactions on neural networks and learning systems 26, 5 (May 2015), 1019–1034. https://doi.org/10.1109/TNNLS.2014.2330900
2015
- [30]
-
[31]
Schneider, Wayne S
Caroline A. Schneider, Wayne S. Rasband, and Kevin W. Eliceiri. 2012. NIH Image to ImageJ: 25 years of image analysis. Nature Methods 9, 7 (July 2012), 671–675. https://doi.org/10.1038/nmeth.2089 Number: 7 Publisher: Nature Publishing Group
2012 doi
-
[32]
Sakaguchi, and Surya K
Metin Uz, Maxsam Donta, Meryem Mededovic, Donald S. Sakaguchi, and Surya K. Mallapragada. 2019. Development of Gelatin and Graphene-Based Nerve Regeneration Conduits Using Three-Dimensional (3D) Printing Strate- gies for Electrical Transdifferentiation of Mesenchymal Stem Cell...
2019 doi
-
[33]
Hondred, Maxsam Donta, Juhyung Jung, Emily Kozik, Jonathan Green, Elizabeth J
Metin Uz, John A. Hondred, Maxsam Donta, Juhyung Jung, Emily Kozik, Jonathan Green, Elizabeth J. Sandquist, Donald S. Sakaguchi, Jonathan C. Claussen, and Surya Mallapragada. 2020. Determination of Electrical Stimuli Parameters To Transdifferentiate Genetically Engineered Mese...
2020 doi
-
[34]
Korsuk Sirinukunwattana, Shan E Ahmed Raza, Yee-Wah Tsang, David R. J. Snead, Ian A. Cree, and Nasir M. Rajpoot. 2016. Locality Sensitive Deep Learning for Detection and Classification of Nuclei in Routine Colon Cancer Histology Images. IEEE Transactions on Medical Imaging 35,...
2016
-
[35]
Alison Noble, and Andrew Zisserman
Weidi Xie, J. Alison Noble, and Andrew Zisserman. 2018. Microscopy cell count- ing and detection with fully convolutional regression networks. Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization 6, 3 (May 2018), 283–292. https://doi.org/10.1080...
2018
-
[36]
Yingying Zhang, Desen Zhou, Siqin Chen, Shenghua Gao, and Yi Ma. 2016. Single-Image Crowd Counting via Multi-Column Convolutional Neural Network. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 589–597. https://doi.org/10.1109/CVPR.2016.70 ISSN: 1063-6919
2016 doi
-
[37]
Ching-Wei Wang, Sheng-Chuan Huang, Yu-Ching Lee, Yu-Jie Shen, Shwu-Ing Meng, and Jeff L. Gaol. 2022. Deep learning for bone marrow cell detection and classification on whole-slide images. Medical Image Analysis 75 (2022), 102270. https://doi.org/10.1016/j.media.2021.102270
2022
- [2015]
-
[2020]
mSystems 5, 1 (2020), 10.1128/msystems.00840–19
Geometry-Aware Cell Detection with Deep Learning. mSystems 5, 1 (2020), 10.1128/msystems.00840–19. https://doi.org/10.1128/msystems.00840-19 arXiv:https://journals.asm.org/doi/pdf/10.1128/msystems.00840-19
2020 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.