REVIEW 6 major objections 7 minor 70 references
Multi-scale and Multi-path Cascaded Convolutional Network for Semantic Segmentation of Colorectal Polyps
T0 review · 6 major / 7 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A compact CNN, MMCC-Net, is claimed to outperform eight state-of-the-art models on pixel-level colorectal polyp segmentation across six public datasets using about 1.43 million parameters.
desk verdict A sensible lightweight CNN with thorough experiments, but the claimed statistical superiority over strong baselines is not supported by the reported confidence intervals and split protocol. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the dense multi-scale feature aggregation written as $DFA = Fa \otimes F1 \otimes F2 \otimes F4$, $DFB = AFi \otimes E1 \otimes E2 \otimes E4 \otimes AFB$, and $DFC = DFB \otimes AFA$, where $\otimes$ is depth-wise concatenation. Three parallel routes produce features at different dilation and stride factors, a mid-block adds an attention-filtered path, and a feature enhancer preserves low-level spatial cues. This cascade lets the decoder combine local edges, small-polyp details, and broad context, while the joint loss $L_{seg} = L_{Dice} + L_{BCE}$ with L2 smoothing on the Dice term addresses class imbalance.
What would settle it
Run a paired per-image significance test, such as a Wilcoxon signed-rank or bootstrap test on per-image Dice, between MMCC-Net and FCB-SwinV2 and PVT-CASCADE on a CVC-ClinicDB split that keeps frames from the same colonoscopy video entirely in either training or testing; if the difference is not significant at p < 0.05, the paper's central outperformance claim fails.
Extended reading notes
Core claim
The paper's central claim is that a carefully balanced convolutional network can outperform transformer-based and hybrid competitors on polyp segmentation without large parameter counts. MMCC-Net couples multi-scale and multi-path cascaded convolutions with three routes for feature fusion, two attention modules, and a feature enhancer, and it is trained with a joint Dice-plus-cross-entropy loss with L2 smoothing on the Dice term. Across Kvasir, CVC-ClinicDB, CVC-300, ETIS, CVC-ColonDB, and EndoCV2020, the authors report that MMCC-Net achieves the highest point estimates for Dice, MIoU, precision, recall, accuracy, and the lowest Hausdorff distance among the eight compared methods, with narrow confidence intervals. They interpret this as evidence that strong global context can come from cascaded convolutions and dense feature aggregation rather than from transformers.
Load-bearing premise
The paper's results stand on the assumption that the differences between MMCC-Net and the strongest baselines are real and not artifacts of random training variation, split choice, or data leakage between video frames; the reported confidence intervals overlap for some key metrics and no paired significance tests are shown.
Editorial extensions
If this is right
- At roughly 1.43 million parameters and 20.85 G FLOPs, the model could run on modest hardware or edge devices in colonoscopy suites.
- A joint Dice-plus-BCE loss with L2 smoothing can handle polyp/background imbalance without explicit class weighting, a recipe transferable to other lesion segmentation tasks.
- Multi-scale dense fusion with few filters per layer may generalize to small or irregular polyps because low-level edge cues are preserved through the feature enhancer.
- The same architecture could be adapted to video polyp segmentation or other lumen and tissue segmentation tasks that require boundary precision.
- The low parameter count and short training time make repeated retraining on hospital-specific data practical.
Reading between the lines
- Because the reported margins over the strongest baselines are often a few tenths of a percentage point and some confidence intervals overlap, a reader should treat the superiority claim as provisional until paired significance tests are reported.
- A natural next check is to evaluate MMCC-Net on a CVC-ClinicDB split that keeps frames from the same colonoscopy video entirely in either training or testing, following the data-leakage caveat the paper itself cites for FCB-SwinV2.
- The ablation pattern suggests the feature enhancer, rather than skip connections alone, drives much of the improvement; isolating that component on harder small-polyp datasets would be a direct stress test.
- The combination of Grad-CAM heatmaps and a compact architecture points toward a screening assistant role, though real clinical deployment would require testing on diverse imaging equipment and lighting conditions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes MMCC-Net, a lightweight CNN for colorectal polyp segmentation built from multi-scale multi-path cascaded convolutions, dense skip connections, two attention modules, and a feature enhancer, trained with a joint Dice and binary cross-entropy loss. Experiments are conducted on six public datasets (Kvasir, CVC-ClinicDB, CVC-300, ETIS, CVC-ColonDB, EndoCV2020) and compared with eight published models. The paper reports Dice scores from 77.43 to 94.45 and MIoU from 72.71 to 90.16 across datasets, approximately 1.43M parameters, repeated 10-run statistics, 5-fold cross-validation, ablations, loss/optimizer/LR sensitivity, HDD/AUC, and efficiency comparisons. The central claim is that MMCC-Net outperforms all eight SOTA models while being substantially more parameter-efficient.
Significance. The intended contribution is a parameter-efficient CNN that matches or beats transformer-based segmenters; such a model would have practical value for colonoscopy workflows. The paper's strengths are its breadth (six datasets, eight baselines), repeated-run design, ablation of architectural modules, and explicit discussion of failure cases and deployment issues. However, the statistical basis for the headline claim is not established: the largest advantages over the strongest baselines are tiny (0.02-0.13 Dice points), confidence intervals overlap, no paired significance tests are provided, and several reported intervals are arithmetically impossible. The additional risk of video-level leakage in the CVC-ClinicDB split, acknowledged in the paper's own discussion of FCB-SwinV2, could materially inflate the reported results. With corrected statistics and a leakage-free evaluation, the paper could be a solid empirical study, but the current evidence does not support the claimed superiority.
major comments (6)
- [Table 2] The claim that MMCC-Net "consistently outperforms" eight SOTA models is not supported by Table 2. On CVC-ClinicDB, MMCC-Net Dice is 94.45 ± 0.12 with 95% CI (94.19, 94.71) versus FCB-SwinV2 94.43 ± 0.13 (94.17, 94.69); the difference is 0.02 points and the intervals overlap almost entirely. On Kvasir, the corresponding difference is 0.13 Dice points with overlapping intervals. No paired significance test (Wilcoxon signed-rank, paired bootstrap, or corrected resampled t-test) is reported for any comparison, so the reported "superior performance" has no demonstrated statistical support. Please add paired tests over the 10 runs and report exact p-values or bootstrap CIs for each SOTA comparison.
- [Tables 2 and 4] The confidence intervals appear internally inconsistent. With n=10 runs and t_{0.025,9}=2.262, the Kvasir proposed Dice mean 92.65 and SD 0.13 imply a 95% CI of approximately (92.56, 92.74), not the reported (92.65, 93.25); the lower bound cannot equal the mean. Similar discrepancies appear in many rows (e.g., several CIs have both endpoints above the mean). Please recompute all CIs from the actual per-run results and state the critical value and formula used.
- [Section 4.1 / Table 1] The CVC-ClinicDB split is at risk of video-level data leakage. The dataset consists of 612 images from 29 colonoscopy videos, and Table 1 shows a 490/61/61 train/validation/test split, but the manuscript nowhere states that the split is video-aware. Section 2.2 itself notes that FCB-SwinV2 highlights video-sequence data leakage in CVC-ClinicDB. If frames from the same video appear in both training and test partitions, the reported ClinicDB numbers are inflated. Please either confirm that the split was performed at the video level and describe the procedure, or rerun Experiment 1 with a video-aware split.
- [Table 3] Table 3 (5-fold cross-validation) contradicts the text's claim that MMCC-Net "outperforms all other models regarding mDice, MIoU, precision, and recall across both datasets." On Kvasir, MMCC-Net's Dice is 92.29 ± 0.22, lower than PVT-CASCADE's 92.49 ± 0.31. This is a direct internal inconsistency in the central comparative claim; please correct the table or the text and discuss the discrepancy.
- [Section 3.2 / Eqs. (4)-(11)] The loss equations are not usable as written. Eq. (4) defines an L2 Dice loss, Eq. (8) defines Lseg = LDice + LBce, and Eq. (11) introduces α and γ in a placement that is dimensionally inconsistent with Eqs. (8)-(10); the grad-CAM text and Eq. (7) are also garbled. Since the joint loss and its hyperparameters (α=0.22, γ=1.9) are part of the method, please rewrite the loss derivation with consistent notation and verify every equation.
- [Section 4.3.1 / Table 8] The manuscript does not state whether α, γ, the learning rate, and other architectural choices were selected using the same test folds on which the final numbers are reported. If these hyperparameters were tuned on the CVC-ClinicDB and Kvasir test partitions, the reported confidence intervals are selection-conditional and would be optimistic. Please describe the model selection protocol used for each experiment.
minor comments (7)
- [Section 4.2] The text states that the study uses eight evaluation measures but then enumerates nine: Dice, accuracy, sensitivity, precision, specificity, IOU, AUC, CI, and HDD.
- [Figure 3] The architecture diagram is difficult to read at the resolution provided; please supply a higher-resolution version with all modules and pathways clearly labeled.
- [Eq. (17)] The Hausdorff distance formula is incorrect as printed: the second term should be max over b in B of min over a in A of ||b - a||, not another maximization over a in A.
- [Section 4.3.2] The citations "FCB-Former [27]" and "FCB-SwinV2 [32]" appear to be wrong; these should likely refer to references [47] and [52], respectively.
- [Tables 2, 5, and 6] Tables 5 and 6 report CVC-ClinicDB Dice of 94.41 and Kvasir Dice of 92.40, while Table 2 reports 94.45 and 92.65 for the same model; please clarify whether these are the same 10-run averages or different runs.
- [Section 4.3] The training-time description is inconsistent: the text says 80 epochs took about 8 hours with no improvement beyond 60 epochs (about 6 hours), while Table 10 reports a training time of 6.0 hours; please align these statements.
- [Equations (12)-(18)] Notation is inconsistent between "IOU" in Eq. (16) and "MIoU" in the tables; please unify the terminology and define HDD before first use in Section 4.2.
Circularity Check
No significant circularity: MMCC-Net's performance claims are empirical comparisons against independently published baselines, with no fitted parameter renamed as a prediction.
full rationale
This paper is an empirical architecture study, not a derivation. The central claim that MMCC-Net outperforms eight SOTA models is supported by direct experimental measurements (Dice, MIoU, etc.) on public datasets, with all baselines run under the same protocol using open-source code. The loss function in Equations (8)-(11) defines an optimization objective rather than deriving the reported segmentation outputs from the definition of the metric; no predicted quantity is equal to an input by construction. Hyperparameters such as alpha = 0.22, gamma = 1.9, and learning rate 1e-4 were selected empirically, but this is standard practice and does not amount to fitting a parameter and then calling a closely related quantity a prediction. The paper's self-citations (e.g., references 23-27) appear in the literature review and methodology context but are not load-bearing for the superiority claim, which rests on the reported comparisons and ablations. The critique that the reported confidence intervals are internally inconsistent and that no paired significance tests are provided is a statistical-validity concern, not a circularity concern. Likewise, the possible video-level data leakage in CVC-ClinicDB is a data-protocol issue, not a case where the paper's conclusion is equivalent to its inputs. Therefore, no circular step can be identified under the required standard of quoting a specific reduction.
Assumptions & free parameters
free parameters (4)
- alpha (loss weighting) =
0.22
- gamma (loss exponent) =
1.9
- learning rate =
1e-4
- dilation and stride factors for multi-scale branches =
1, 2, 4
assumptions (4)
- domain assumption Ground-truth annotations in the six datasets are correct and consistent.
- domain assumption Training on Kvasir and CVC-ClinicDB and testing on four other datasets measures cross-dataset generalization.
- standard math The 95% confidence intervals computed from 10 runs with a t-distribution are valid.
- domain assumption The public benchmark datasets are representative of clinical polyp imaging.
invented entities (1)
-
Feature Enhancer (FE) module
Cite this review
Pith. "Pith review of Multi-scale and Multi-path Cascaded Convolutional Network for Semantic Segmentation of Colorectal Polyps." pith.science (2026). https://pith.science/paper/Z4PXHIBV
@misc{pith2026241202443,
author = {Pith},
title = {Pith review of: Multi-scale and Multi-path Cascaded Convolutional Network for Semantic Segmentation of Colorectal Polyps},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z4PXHIBV}},
note = {Machine review of arXiv:2412.02443}
}
read the original abstract
Colorectal polyps are structural abnormalities of the gastrointestinal tract that can potentially become cancerous in some cases. The study introduces a novel framework for colorectal polyp segmentation named the Multi-Scale and Multi-Path Cascaded Convolution Network (MMCC-Net), aimed at addressing the limitations of existing models, such as inadequate spatial dependence representation and the absence of multi-level feature integration during the decoding stage by integrating multi-scale and multi-path cascaded convolutional techniques and enhances feature aggregation through dual attention modules, skip connections, and a feature enhancer. MMCC-Net achieves superior performance in identifying polyp areas at the pixel level. The Proposed MMCC-Net was tested across six public datasets and compared against eight SOTA models to demonstrate its efficiency in polyp segmentation. The MMCC-Net's performance shows Dice scores with confidence intervals ranging between (77.08, 77.56) and (94.19, 94.71) and Mean Intersection over Union (MIoU) scores with confidence intervals ranging from (72.20, 73.00) to (89.69, 90.53) on the six databases. These results highlight the model's potential as a powerful tool for accurate and efficient polyp segmentation, contributing to early detection and prevention strategies in colorectal cancer.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Selvaraj, J. and A. Jayanthy, Automatic polyp semantic segmentation using wireless capsule endoscopy images with various convolutional neural network and optimization techniques: a comparison and performance evaluation. Biomedical Engineering: Applications, Basis and Communications, 2023. 35(06): p. 2350026
work page 2023
-
[2]
Selvaraj, J. and A. Jayanthy, Design and development of artificial intelligence‐based application programming interface for early detection and diagnosis of colorectal cancer from wireless capsule endoscopy images. International Journal of Imaging Systems and Technology, 2024. 34(2): p. e23034
work page 2024
-
[3]
Liu, F., Z. Hua, J. Li, and L. Fan, MFBGR: Multi-scale feature boundary graph reasoning network for polyp segmentation. Engineering Applications of Artificial Intelligence, 2023. 123: p. 106213
work page 2023
-
[4]
Experimental Setup and Results 4.1 Databases We use six public datasets for comparison: EndoCV2020 [57], CVC-300 [58], ETIS [59], CVC-ColonDB [60], CVC-ClinicDB [61], and Kvasir [62]. These datasets are broadly used in evaluating the performance and efficacy of the various polyp segmentation methods under consideration. Extensive tests are carried out on ...
-
[5]
Discussion Automated polyp segmentation in colonoscopy images has become crucial in ensuring precise polypectomy. Consequently, researchers and medical professionals have recently paid increasing attention to developing such methods. In our research, through an in-depth analysis of integrating attention mechanisms, skip connections, and feature enhancemen...
-
[6]
The FE combines with the attention modules and dense skip path segment smaller polyps in the images
Conclusion MMCC-Net implements multi-scale features, two attention modules, and a feature enhancer (FE) to produce an effective multi-scale hierarchical feature representation. The FE combines with the attention modules and dense skip path segment smaller polyps in the images. The segmentation performance of MMCC-Net, which uses approximately 1.43 million...
- [7]
-
[8]
Yue, G., S. Li, R. Cong, T. Zhou, B. Lei, and T. Wang, Attention-Guided Pyramid Context Network for Polyp Segmentation in Colonoscopy Images. IEEE Transactions on Instrumentation and Measurement,
Show all 70 references
-
[9]
Zhu, J., M. Ge, Z. Chang, and W. Dong, CRCNet: Global-local context and multi-modality cross attention for polyp segmentation. Biomedical Signal Processing and Control, 2023. 83: p. 104593
2023
-
[10]
Staller, V
Lee, T.-C., K. Staller, V. Botoman, M.P. Pathipati, S. Varma, and B. Kuo, ChatGPT Answers Common Patient Questions About Colonoscopy. Gastroenterology, 2023
2023
-
[11]
Naqvi, and E
Khan, T.M., S.S. Naqvi, and E. Meijering, ESDMR-Net: A lightweight network with expand-squeeze and dual multiscale residual connections for medical image segmentation. Engineering Applications of Artificial Intelligence, 2024. 133: p. 107995
2024
-
[12]
Iqbal, K
Matloob Abbasi, M., S. Iqbal, K. Aurangzeb, M. Alhussein, and T.M. Khan, LMBiS-Net: A lightweight bidirectional skip connection based multipath CNN for retinal blood vessel segmentation. Scientific Reports, 2024. 14(1): p. 15219
2024
-
[13]
Zafar, A
Farooq, H., Z. Zafar, A. Saadat, T.M. Khan, S. Iqbal, and I. Razzak, LSSF-Net: Lightweight segmentation with self-awareness, spatial attention, and focal modulation. Artificial Intelligence in Medicine, 2024: p. 103012
2024
-
[14]
Naqvi, S
Naveed, A., S.S. Naqvi, S. Iqbal, I. Razzak, H.A. Khan, and T.M. Khan, RA-Net: Region-Aware Attention Network for Skin Lesion Segmentation. Cognitive Computation, 2024: p. 1-18
2024
-
[15]
Naqvi, T.M
Naveed, A., S.S. Naqvi, T.M. Khan, and I. Razzak, PCA: progressive class-wise attention for skin lesions diagnosis. Engineering Applications of Artificial Intelligence, 2024. 127: p. 107417
2024
-
[16]
Razzak, A
Mazher, M., I. Razzak, A. Qayyum, M. Tanveer, S. Beier, T. Khan, and S.A. Niederer, Self-supervised spatial–temporal transformer fusion based federated framework for 4D cardiovascular image segmentation. Information Fusion, 2024. 106: p. 102256
2024
-
[17]
Naqvi, T.M
Naveed, A., S.S. Naqvi, T.M. Khan, S. Iqbal, M.Y. Wani, and H.A. Khan, AD-Net: Attention-based dilated convolutional residual network with guided decoder for robust skin lesion segmentation. Neural Computing and Applications, 2024: p. 1-23
2024
-
[18]
Khan, S.S
Iqbal, S., T.M. Khan, S.S. Naqvi, A. Naveed, M. Usman, H.A. Khan, and I. Razzak, Ldmres-Net: a lightweight neural network for efficient medical image segmentation on iot and edge devices. IEEE Journal of Biomedical and Health Informatics, 2023
2023
-
[19]
Langah, H.A
Naqvi, S.S., Z.A. Langah, H.A. Khan, M.I. Khan, T. Bashir, M.I. Razzak, and T.M. Khan, Glan: Gan assisted lightweight attention network for biomedical imaging based diagnostics. Cognitive Computation,
-
[20]
Khan, S.S
Iqbal, S., T.M. Khan, S.S. Naqvi, and G. Holmes, MLR-Net: A multi-layer residual convolutional neural network for leather defect segmentation. Engineering Applications of Artificial Intelligence, 2023. 126: p. 107007
2023
-
[21]
Naqvi, A
Khan, T.M., S.S. Naqvi, A. Robles-Kelly, and I. Razzak, Retinal vessel segmentation via a Multi-resolution Contextual Network and adversarial learning. Neural Networks, 2023. 165: p. 310-320
2023
-
[22]
Naveed, S.S
Iqbal, S., K. Naveed, S.S. Naqvi, A. Naveed, and T.M. Khan, Robust retinal blood vessel segmentation using a patch-based statistical adaptive multi-scale line detector. Digital Signal Processing, 2023. 139: p. 104075
2023
-
[23]
Jinchao, S
Yaqub, M., F. Jinchao, S. Ahmed, A. Mehmood, I.S. Chuhan, M.A. Manan, and M.S. Pathan, DeepLabV3, IBCO-based ALCResNet: A fully automated classification, and grading system for brain tumor. Alexandria Engineering Journal, 2023. 76: p. 609-627
2023
-
[24]
Mazher, T
Qayyum, A., M. Mazher, T. Khan, and I. Razzak, Semi-supervised 3D-InceptionNet for segmentation and survival prediction of head and neck primary cancers. Engineering Applications of Artificial Intelligence,
-
[25]
Iqbal, S., T.M. Khan, K. Naveed, S.S. Naqvi, and S.J. Nawaz, Recent trends and advances in fundus image analysis: A review. Computers in Biology and Medicine, 2022. 151: p. 106277
2022
-
[26]
Khan, S.S
Arsalan, M., T.M. Khan, S.S. Naqvi, M. Nawaz, and I. Razzak, Prompt deep light-weight vessel segmentation network (PLVS-Net). IEEE/ACM Transactions on Computational Biology and Bioinformatics, 2022. 20(2): p. 1363-1371
2022
-
[27]
Jinchao, T.M
Manan, M.A., F. Jinchao, T.M. Khan, M. Yaqub, S. Ahmed, and I.S. Chuhan, Semantic segmentation of retinal exudates using a residual encoder–decoder architecture in diabetic retinopathy. Microscopy Research and Technique, 2023. 86(11): p. 1443-1460
2023
-
[28]
Jinchao, S
Manan, M.A., F. Jinchao, S. Ahmed, and A. Raheem. DPE-Net: Dual-Parallel Encoder Based Network for Semantic Segmentation of Polyps. in 2024 9th International Conference on Signal and Image Processing (ICSIP). 2024. IEEE
2024
-
[29]
Khan, A.Q., G. Sun, Y. Li, A. Bilal, and M.A. Manan, Optimizing Fully Convolutional Encoder-Decoder Network for Segmentation of Diabetic Eye Disease. Computers, Materials & Continua, 2023. 77(2)
2023
-
[30]
Raheem, A., Z. Yang, H. Yu, M. Yaqub, F. Sabah, S. Ahmed, . . . I.S. Chuhan, IPC-CNN: A Robust Solution for Precise Brain Tumor Segmentation Using Improved Privacy-Preserving Collaborative Convolutional Neural Network. KSII Transactions on Internet and Information Systems (TII...
2024
-
[31]
Iqbal, A
Abbasi, M.M., S. Iqbal, A. Naveed, T.M. Khan, S.S. Naqvi, and W. Khalid, LMBiS-Net: A Lightweight Multipath Bidirectional Skip Connection based CNN for Retinal Blood Vessel Segmentation. arXiv preprint arXiv:2309.04968, 2023
2023 arXiv
-
[32]
Lin, Y., X. Han, K. Chen, W. Zhang, and Q. Liu, CSwinDoubleU-Net: A double U-shaped network combined with convolution and Swin Transformer for colorectal polyp segmentation. Biomedical Signal Processing and Control, 2024. 89: p. 105749
2024
-
[33]
Wu, and Y.-L
Huang, C.-H., H.-Y. Wu, and Y.-L. Lin, Hardnet-mseg: A simple encoder-decoder polyp segmentation neural network that achieves over 0.9 mean dice and 86 fps. arXiv preprint arXiv:2101.07172, 2021
2021 arXiv
-
[34]
Robles-Kelly, and S.S
Khan, T.M., A. Robles-Kelly, and S.S. Naqvi. T-Net: A resource-constrained tiny convolutional neural network for medical image segmentation. in Proceedings of the IEEE/CVF winter conference on applications of computer vision. 2022
2022
-
[35]
Sala, C.-B
Yeung, M., E. Sala, C.-B. Schönlieb, and L. Rundo, Focus U-Net: A novel dual attention-gated CNN for polyp segmentation during colonoscopy. Computers in biology and medicine, 2021. 137: p. 104815
2021
-
[36]
Yang, Q., C. Geng, R. Chen, C. Pang, R. Han, L. Lyu, and Y. Zhang, DMU-Net: Dual-route mirroring U- Net with mutual learning for malignant thyroid nodule segmentation. Biomedical Signal Processing and Control, 2022. 77: p. 103805
2022
-
[37]
Guo, X., Z. Chen, J. Liu, and Y. Yuan, Non-equivalent images and pixels: Confidence-aware resampling with meta-learning mixup for polyp segmentation. Medical Image Analysis, 2022. 78: p. 102394
2022
-
[38]
Zhao, and Z
Wu, H., Z. Zhao, and Z. Wang, META-Unet: Multi-scale efficient transformer attention Unet for fast and high-accuracy polyp segmentation. IEEE Transactions on Automation Science and Engineering, 2023
2023
-
[39]
Fan, D.-P., G.-P. Ji, T. Zhou, G. Chen, H. Fu, J. Shen, and L. Shao. Pranet: Parallel reverse attention network for polyp segmentation. in Medical Image Computing and Computer Assisted Intervention– MICCAI 2020: 23rd International Conference, Lima, Peru, October 4–8, 2020, Pro...
2020
-
[40]
Bhattacharya, D., C. Betz, D. Eggert, and A. Schlaefer, Dual Parallel Reverse Attention Edge Network: DPRA-EdgeNet. Nordic Machine Intelligence, 2021. 1(1): p. 8-10
2021
-
[41]
Ta, N., H. Chen, Y. Lyu, and T. Wu, BLE-Net: boundary learning and enhancement network for polyp segmentation. Multimedia Systems, 2022: p. 1-14
2022
-
[42]
Lai, H., Y. Luo, G. Zhang, X. Shen, B. Li, and J. Lu, Toward accurate polyp segmentation with cascade boundary-guided attention. The Visual Computer, 2022: p. 1-17
2022
-
[43]
Lee, and D
Kim, T., H. Lee, and D. Kim. Uacanet: Uncertainty augmented context attention for polyp segmentation. in Proceedings of the 29th ACM International Conference on Multimedia. 2021
2021
-
[44]
Chou, D.-P
Ji, G.-P., Y.-C. Chou, D.-P. Fan, G. Chen, H. Fu, D. Jha, and L. Shao. Progressively normalized self- attention network for video polyp segmentation. in Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, S...
2021
-
[45]
Siddiquee, N
Zhou, Z., M.M.R. Siddiquee, N. Tajbakhsh, and J. Liang, Unet++: Redesigning skip connections to exploit multiscale features in image segmentation. IEEE transactions on medical imaging, 2019. 39(6): p. 1856- 1867
2019
-
[46]
Fan, D.-P., G.-P. Ji, T. Zhou, G. Chen, H. Fu, J. Shen, and L. Shao. Pranet: Parallel reverse attention network for polyp segmentation. in International conference on medical image computing and computer- assisted intervention. 2020. Springer
2020
-
[47]
Sanderson, E. and B.J. Matuszewski. FCN-transformer feature fusion for polyp segmentation. in Annual conference on medical image understanding and analysis. 2022. Springer
2022
-
[48]
Zhang, R., G. Li, Z. Li, S. Cui, D. Qian, and Y. Yu. Adaptive context selection for polyp segmentation. in Medical Image Computing and Computer Assisted Intervention–MICCAI 2020: 23rd International Conference, Lima, Peru, October 4–8, 2020, Proceedings, Part VI 23. 2020. Springer
2020
-
[49]
Tomar, N.K., D. Jha, U. Bagci, and S. Ali. TGANet: text-guided attention for improved polyp segmentation. in Medical Image Computing and Computer Assisted Intervention–MICCAI 2022: 25th International Conference, Singapore, September 18–22, 2022, Proceedings, Part III. 2022. Springer
2022
-
[50]
Chen, J., Y. Lu, Q. Yu, X. Luo, E. Adeli, Y. Wang, . . . Y. Zhou, Transunet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:2102.04306, 2021
2021 arXiv
-
[51]
Nguyen, M., T.T. Bui, Q. Van Nguyen, T.T. Nguyen, and T. Van Pham, LAPFormer: A Light and Accurate Polyp Segmentation Transformer. arXiv preprint arXiv:2210.04393, 2022
2022 arXiv
-
[52]
Oanh, N.T
Duc, N.T., N.T. Oanh, N.T. Thuy, T.M. Triet, and V.S. Dinh, Colonformer: An efficient transformer based method for colon polyp segmentation. IEEE Access, 2022. 10: p. 80575-80586
2022
-
[53]
Wang, D.-P
Dong, B., W. Wang, D.-P. Fan, J. Li, H. Fu, and L. Shao, Polyp-pvt: Polyp segmentation with pyramid vision transformers. arXiv preprint arXiv:2108.06932, 2021
2021 arXiv
-
[54]
Liu, X. and S. Song, Attention combined pyramid vision transformer for polyp segmentation. Biomedical Signal Processing and Control, 2024. 89: p. 105792
2024
-
[55]
Tomar, V
Jha, D., N.K. Tomar, V. Sharma, and U. Bagci. TransNetR: transformer-based residual network for polyp segmentation with multi-center out-of-distribution testing. in Medical Imaging with Deep Learning. 2024. PMLR
2024
-
[56]
Fitzgerald, K. and B. Matuszewski, FCB-SwinV2 Transformer for Polyp Segmentation. arXiv preprint arXiv:2302.01027, 2023
2023 arXiv
-
[57]
Shibao, H
Chaoyang, Z., S. Shibao, H. Wenmao, and Z. Pengcheng, FDR-TransUNet: A novel encoder-decoder architecture with vision transformer for improved medical image segmentation. Computers in Biology and Medicine, 2024. 169: p. 107858
2024
-
[58]
Wang, Y., Z. Deng, Q. Lou, S. Hu, K.-s. Choi, and S. Wang, Cooperation Learning Enhanced Colonic Polyp Segmentation Based on Transformer-CNN Fusion. arXiv preprint arXiv:2301.06892, 2023
2023 arXiv
-
[59]
Histace, O
Silva, J., A. Histace, O. Romain, X. Dray, and B. Granado, Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer. International journal of computer assisted radiology and surgery, 2014. 9: p. 283-293
2014
-
[60]
Despite UNet's implementation of skip connections to enhance feature integration, it often misinterprets surrounding tissues
and EndoCV2020 [57] datasets, characterized by their complex backgrounds, blurriness, and low contrast. Despite UNet's implementation of skip connections to enhance feature integration, it often misinterprets surrounding tissues. HarDNet-MSEG [29] captures multi-scale features...
-
[61]
Rahman, M.M. and R. Marculescu. Medical image segmentation via cascaded attention decoding. in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2023
2023
-
[62]
Ali, S., C. Daul, J. Rittscher, D. Stoyanov, and E. Grisan. Preface to: EndoCV2020Computer Vision in Endoscopy. in CEUR Workshop Proceedings. 2020. CEUR Workshop Proceedings
2020
-
[63]
Bernal, F.J
Vázquez, D., J. Bernal, F.J. Sánchez, G. Fernández-Esparrach, A.M. López, A. Romero, . . . A. Courville, A benchmark for endoluminal scene segmentation of colonoscopy images. Journal of healthcare engineering, 2017. 2017
2017
-
[64]
Gurudu, and J
Tajbakhsh, N., S.R. Gurudu, and J. Liang, Automated polyp detection in colonoscopy videos using shape and context information. IEEE transactions on medical imaging, 2015. 35(2): p. 630-644
2015
-
[65]
Sánchez, G
Bernal, J., F.J. Sánchez, G. Fernández-Esparrach, D. Gil, C. Rodríguez, and F. Vilariño, WM-DOVA maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians. Computerized medical imaging and graphics, 2015. 43: p. 99-111
2015
-
[66]
Smedsrud, M.A
Jha, D., P.H. Smedsrud, M.A. Riegler, P. Halvorsen, T. de Lange, D. Johansen, and H.D. Johansen. Kvasir- seg: A segmented polyp dataset. in MultiMedia Modeling: 26th International Conference, MMM 2020, Daejeon, South Korea, January 5–8, 2020, Proceedings, Part II 26. 2020. Springer
2020
-
[67]
Zhang, Y
Xia, H., M. Zhang, Y. Tan, and C. Xia, MCGNet: Multi-level consistency guided polyp segmentation. Biomedical Signal Processing and Control, 2023. 86: p. 105343
2023
-
[68]
Iqbal, A. and M. Sharif, BTS-ST: Swin transformer network for segmentation and classification of multimodality breast cancer images. Knowledge-Based Systems, 2023. 267: p. 110393
2023
-
[69]
Fischer, and T
Ronneberger, O., P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. in Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18. 20...
2015
-
[70]
Selvaraj, J. and S. Umapathy, CRPU-NET: a deep learning model based semantic segmentation for the detection of colorectal polyp in lower gastrointestinal tract. Biomedical Physics & Engineering Express,
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.