Pith. sign in

REVIEW 5 major objections 5 minor 64 references

TSUBF-Net: Trans-Spatial UNet-like Network with Bi-direction Fusion for Segmentation of Adenoid Hypertrophy in CT

T0 review · 5 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read TSUBF-Net, a 3D U-shaped segmentation network with trans-spatial attention, bidirectional fusion, and a Sobel-gradient smoothness loss, outperforms prior methods for segmenting adenoid hypertrophy in CT scans, reporting DSC 92.26, IoU…

desk verdict Plausible first 3D adenoid segmentation paper whose headline AHSD SOTA claim rests on one private split, while the public-dataset results are the more credible evidence. read the letter →

arxiv 2412.00787 v1 pith:QJBNWIOY submitted 2024-12-01 eess.IV cs.AIcs.CV

classification eess.IVcs.AIcs.CV
keywords adenoidhypertrophy3DmedicalimagesegmentationCTtrans-spatialattentionfeaturefusionSobellossU-shapednetworkpediatricairway
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

TSUBF-Net is a 3D U-shaped segmentation network built for a task where previous work has been sparse: delineating adenoid hypertrophy in pediatric head CT. The paper argues that adenoid boundaries are unusually blurred, so the network adds a Trans-Spatial Perception (TSP) module that attends across CT slices, a Bi-directional Sampling Collaborated Fusion (BSCF) module that merges down-sampled and up-sampled features, and a Sobel-based loss that penalizes rough surfaces. On the authors' private 240-patient AHSD dataset, TSUBF-Net reports the best scores among the compared methods, with HD95 7.03, IoU 85.63, and DSC 92.26, beating nnUNet and UNETR++. If the result holds, it would give surgeons a quantitative, volumetric view of the adenoid before ablation surgery, where today they rely largely on endoscopic views and experience.

What carries the argument

The load-bearing pieces are three. The Trans-Spatial Perception (TSP) module is a four-head attention mechanism: one channel head plus three inter-layer spatial heads that attend along height, width, and depth, with shared query and key weights, so a voxel can gather context from neighboring slices. The Bi-directional Sampling Collaborated Fusion (BSCF) module replaces plain skip connections: up-sampled and down-sampled features are first aligned by a shared 3x3x3 convolution, then up-sampled features act as queries against down-sampled keys and values so contour information corrects segmentation features. The Sobel loss convolves the prediction with 3D Sobel kernels along x, y, and z and adds the gradient magnitude to the soft Dice and cross-entropy loss, explicitly favoring smooth surfaces.

What would settle it

Take the 38 held-out AHSD CT volumes, have a second clinician independently redraw the adenoid boundaries, and recompute HD95 and DSC for TSUBF-Net and UNETR++ against that new ground truth; if the margin over UNETR++ shrinks below the reported 3.99 DSC points or flips on HD95, the boundary-superiority claim is not robust. Alternatively, retrain on several random 189/38 splits and check whether TSUBF-Net wins every time.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that the specific combination of TSP, BSCF, and Sobel loss, rather than any single component, is what solves fuzzy-boundary adenoid segmentation. The ablation table shows that TSP alone slightly lowers DSC relative to the UNETR++ baseline (87.78 versus 88.22) while cutting FLOPs; adding BSCF raises DSC to 90.91; adding both reaches 91.49; and the Sobel loss with lambda 0.1 brings the model to 92.26, a 4.57-point improvement over baseline. The paper reports TSUBF-Net as superior to all compared methods on AHSD, with HD95 of 7.03 versus 10.13 for UNETR++ and 8.16 for nnUNet, and also reports the best DSC on MSD-Lung (83.69) while being close to the best on ACDC (92.68, versus UNETR++ at 92.83). All reported AHSD numbers are single-model, without pre-training or ensembling.

Load-bearing premise

The load-bearing premise is that the manual ground-truth boundaries in the AHSD dataset are correct and that one random split into 189 training and 38 test patients represents the full variety of pediatric adenoid CT anatomy.

Editorial extensions

If this is right

  • On AHSD, the full model's DSC advantage over UNETR++ is 3.99 points and over nnUNet is 1.13 points, with the largest reported gap in HD95 (7.03 versus 10.13 and 8.16), so the claim is specifically a boundary-accuracy improvement.
  • The same architecture transfers to a different small-target CT task: on MSD-Lung, TSUBF-Net reports DSC 83.69 versus 80.68 for UNETR++ and 80.14 for MedNeXt.
  • On ACDC, an MRI dataset with smooth organ boundaries, TSUBF-Net's mean DSC of 92.68 is essentially tied with UNETR++ at 92.83, suggesting the smoothness prior is neutral or slightly detrimental when boundaries are already well defined.
  • The ablation values imply that the BSCF fusion module, rather than the TSP attention module, accounts for most of the accuracy gain, while TSP chiefly buys lower computational cost.
  • The paper's stated future direction is to extend segmentation to tonsils, turbinates, and epiglottis to build a complete sleep-apnea database, which the architecture is positioned to support.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The superiority claim rests on one private dataset with a single random 189/38 split and no inter-observer variability estimate; re-splitting or independent re-annotation could narrow or reverse the reported margins, so the headline numbers should be read as provisional until the dataset is released or validated externally.
  • The pattern across datasets suggests a trade-off: explicit boundary-smoothness regularization helps when the target has genuinely indistinct borders, such as adenoids and lung tumors, but can slightly hurt on clearly bounded organs like cardiac MRI; a learned or per-dataset lambda, which the paper itself floats, would be the natural extension.
  • The BSCF mechanism of letting up-sampled features query down-sampled features is a general recipe for any U-shaped segmenter where contour and semantics live at different scales; it could be grafted onto other backbones and tested on other ill-defined structures such as tumor margins or ground-glass opacities.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes TSUBF-Net, a 3D U-shaped segmentation network for adenoid hypertrophy in CT, with three contributions: the Trans-Spatial Perception (TSP) module, the Bi-directional Sampling Collaborated Fusion (BSCF) module, and a Sobel-gradient smoothness term added to the loss. The method is evaluated on a private adenoid dataset (AHSD), where it is claimed to outperform prior state-of-the-art methods (HD95 7.03, IoU 85.63, DSC 92.26), and on two public benchmarks, ACDC and MSD-Lung, where it is competitive or favorable. The paper also reports ablation studies showing the contribution of each module and the sensitivity of the Sobel loss weight.

Significance. If the empirical claims are accepted, the paper offers a useful architectural recipe for 3D segmentation of structures with weak or ambiguous boundaries, and it addresses a genuinely under-studied clinical task. The inclusion of public-dataset results (ACDC, MSD-Lung) is a real strength and provides some evidence that the proposed modules generalize beyond the private AHSD data. The ablations, despite the issues noted below, also give a clear decomposition of the contributions of TSP and BSCF. However, the central SOTA claim is an empirical benchmark result on a private dataset, and the current experimental design does not yet support that claim as stated: there are no error bars, the test set is used for hyperparameter selection, the ground-truth annotation protocol is unspecified, and one key metric definition appears mis-specified.

major comments (5)
  1. [Section 4.1 and Table 2] The headline superiority claim rests on a single random split of a private dataset (189 training / 38 test volumes) with no error bars, no multiple seeds, and no cross-validation. The reported advantage over nnUNet is small in absolute terms (DSC 92.26 vs 91.13, HD95 7.03 vs 8.16), and HD95 is an outlier-sensitive boundary metric on a test set of only 38 volumes. The authors should report per-volume distributions, confidence intervals, and ideally results across multiple splits or seeds, so the reader can assess whether the advantage is stable.
  2. [Section 4.3.4 and Table 5] The Sobel loss weight lambda is selected by evaluating lambda in {1.0, 0.5, 0.1} on the AHSD test set, since no validation split is described. This means the reported 4.57% enhancement over the baseline includes selection bias, and the comparison with fixed-hyperparameter baselines is not fully fair. The authors should either introduce a separate validation split for hyperparameter selection or use nested cross-validation, and should describe the selection procedure explicitly.
  3. [Section 4.4 vs Table 5] There is a direct internal contradiction about which lambda value is best. Section 4.4 states that 'when the lambda is set to 1.0, the three evaluation indicators all get the best results,' but Table 5 shows lambda = 0.1 gives the best HD95 (7.03), IoU (85.63), and DSC (92.26), while lambda = 1.0 gives HD95 7.78, IoU 84.63, and DSC 91.68. This contradiction must be resolved; as written, it undermines the reliability of the ablation conclusions.
  4. [Equation (10)] The HD95 definition appears mis-specified. The standard definition is HD95(Y,P) = max( h95(Y,P), h95(P,Y) ), where h95 is the 95th percentile of the directed distances. Equation (10) writes HD95(Y,P) = max( dYP + dPY ), and the surrounding text describes 'maximum 95th percentile distance' in a way that conflates the two directed distances. The authors should correct the equation and state precisely which implementation was used, since this metric is central to the claimed improvement.
  5. [Section 4.1] No annotation protocol is described for AHSD: the paper does not state the number of annotators, whether there was any inter-observer agreement assessment, or how the clinically ambiguous posterior boundary was defined and resolved. Given that the paper repeatedly emphasizes that this boundary is indistinct and clinically important, the quality and consistency of the ground-truth labels are load-bearing for the reported benchmark. This information should be added, or the corresponding limitation should be explicitly acknowledged.
minor comments (5)
  1. [Section 5] The conclusion mislabels the datasets: it refers to 'ACDC(CT)' and 'MSD-Lung(MRI)', whereas Section 4 correctly identifies ACDC as MRI and MSD-Lung as CT.
  2. [Section 2.4] There is a typo: 'Swim-Transformer' should be 'Swin-Transformer'.
  3. [Equation (5)] The loss expression in Equation (5) lacks clear parentheses; as typeset, it is not clear whether the cross-entropy term is inside or outside the sum over classes. Please reformat.
  4. [Section 4.3.4] The text defines FLOPs as 'Floating Point Operations Per Second', but the table reports FLOPs in G (giga floating-point operations), which is a count, not a rate. Please correct the terminology.
  5. [Equation (2)] The attention formula appears to have a typo: it should be Softmax( Q K^T / sqrt(d) ) * V rather than Softmax( Qs, K^T / sqrt(d) ) * V.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims are empirical benchmark comparisons, not derivations that reduce to their own inputs.

full rationale

The paper's central claim is that TSUBF-Net outperforms state-of-the-art methods on the private AHSD dataset and is competitive on ACDC and MSD-Lung. This is an empirical evaluation claim, not a derivation from first principles, so the core circularity patterns do not apply. The TSP module, BSCF module, and Sobel loss are architectural and loss-function modifications; their reported contributions come from ablation experiments (Table 5), and no equation or result is defined in terms of the target metric. The Sobel loss weight lambda=0.1 is selected by evaluating on the AHSD test set, which introduces selection bias and may inflate the reported improvement, but this is a methodological weakness in benchmarking rather than a circular reduction: the paper does not fit a parameter and then present a closely related quantity as an independent prediction. There is no load-bearing self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation; the coauthor-related references [47], [48], and [62] are not used to justify the central superiority claim. The internal contradiction between Section 4.4 ("lambda = 1.0 ... best results") and Table 5 (lambda = 0.1 best) is a correctness or reporting risk, not circularity. Because the headline result is tested against external public benchmarks and its components are not defined in terms of the outcome they purport to predict, the derivation chain is self-contained with respect to circularity. A score of 0 is therefore appropriate, with the caveat that private-data evaluation and test-set hyperparameter selection are separate validity concerns.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on a private dataset with assumed label accuracy, a single split, and fair baseline comparisons. The Sobel loss weight is the only numerically fitted parameter that directly affects the reported gain. No new physical entities are introduced.

free parameters (3)
  • Sobel loss weight lambda = 0.1
    Weight of the Sobel loss; selected by testing 1.0, 0.5, and 0.1 on the test set (Table 5). The reported DSC 92.26 depends on this value.
  • TSP output channel compression ratio = 1/4
    The spatial and channel attention outputs are compressed to one quarter of their original size by linear layers (Section 3.2), a hand-chosen design not justified by experiments.
  • Number of attention heads in TSP = 4 (1 channel + 3 spatial)
    Chosen by architecture design; no ablation on head count or sharing of Q/K matrices is reported.
assumptions (4)
  • domain assumption The AHSD manual segmentation labels are accurate and clinically consistent.
    Section 4.1 describes the dataset but reports no inter-observer agreement or label validation; the entire supervised training and evaluation rely on these labels.
  • domain assumption The single random 189/38 train/test split of AHSD is representative and no data leakage occurred.
    Section 4.1 states the split was random; no cross-validation, no multiple seeds, and the 13 normal cases are not clearly used in training or testing.
  • domain assumption Baseline methods in Tables 2-4 were trained under equivalent settings and are reported without disadvantage.
    Section 4.2 specifies settings for TSUBF-Net but not per-method tuning; comparison fairness is assumed.
  • domain assumption The 3D Sobel kernels are a valid regularizer for boundary smoothness in this task.
    Section 3.4 introduces Eqs. 6-8 without evidence that minimizing these gradients improves clinical accuracy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TSUBF-Net: Trans-Spatial UNet-like Network with Bi-direction Fusion for Segmentation of Adenoid Hypertrophy in CT." pith.science (2026). https://pith.science/paper/QJBNWIOY

@misc{pith2026241200787,
  author       = {Pith},
  title        = {Pith review of: TSUBF-Net: Trans-Spatial UNet-like Network with Bi-direction Fusion for Segmentation of Adenoid Hypertrophy in CT},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QJBNWIOY}},
  note         = {Machine review of arXiv:2412.00787}
}
read the original abstract

Adenoid hypertrophy stands as a common cause of obstructive sleep apnea-hypopnea syndrome in children. It is characterized by snoring, nasal congestion, and growth disorders. Computed Tomography (CT) emerges as a pivotal medical imaging modality, utilizing X-rays and advanced computational techniques to generate detailed cross-sectional images. Within the realm of pediatric airway assessments, CT imaging provides an insightful perspective on the shape and volume of enlarged adenoids. Despite the advances of deep learning methods for medical imaging analysis, there remains an emptiness in the segmentation of adenoid hypertrophy in CT scans. To address this research gap, we introduce TSUBF-Nett (Trans-Spatial UNet-like Network based on Bi-direction Fusion), a 3D medical image segmentation framework. TSUBF-Net is engineered to effectively discern intricate 3D spatial interlayer features in CT scans and enhance the extraction of boundary-blurring features. Notably, we propose two innovative modules within the U-shaped network architecture:the Trans-Spatial Perception module (TSP) and the Bi-directional Sampling Collaborated Fusion module (BSCF).These two modules are in charge of operating during the sampling process and strategically fusing down-sampled and up-sampled features, respectively. Furthermore, we introduce the Sobel loss term, which optimizes the smoothness of the segmentation results and enhances model accuracy. Extensive 3D segmentation experiments are conducted on several datasets. TSUBF-Net is superior to the state-of-the-art methods with the lowest HD95: 7.03, IoU:85.63, and DSC: 92.26 on our own AHSD dataset. The results in the other two public datasets also demonstrate that our methods can robustly and effectively address the challenges of 3D segmentation in CT scans.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 52 canonical work pages

  1. [1]

    Prevalence of adenoid hypertrophy: A systematic 23 review and meta-analysis,

    L. Pereira, J. Monyror, F. T. Almeida, F. R. Almeida, E. Guerra, C. Flores-Mir, and C. Pachˆ eco-Pereira, “Prevalence of adenoid hypertrophy: A systematic 23 review and meta-analysis,” Sleep Medicine Reviews , vol. 38, pp. 101–112,

  2. [2]

    The effect of mode of breathing on craniofacial growth—revisited,

    T. Peltom¨ aki, “The effect of mode of breathing on craniofacial growth—revisited,” The European Journal of Orthodontics , vol. 29, no. 5, pp. 426–429, 2007

  3. [3]

    Pharyngitis and adenotonsillar disease,

    B. J. Wiatrak and A. L. Woolley, “Pharyngitis and adenotonsillar disease,” Otolaryngology Head and Neck Surgery, ed , vol. 3, pp. 188–215, 2005

  4. [4]

    Polysomnography findings in preschool children with obstructive sleep apnea are affected by growth and developmen- tal level,

    C. Lu, C. Sun, Y. Xu, C. Chen, and Q. Li, “Polysomnography findings in preschool children with obstructive sleep apnea are affected by growth and developmen- tal level,” International Journal of Pediatric Otorhinolaryngology , vol. 162, p. 111310, 2022

  5. [5]

    Non-surgical treatment of adenoidal hypertrophy: the role of treat- ing ige-mediated inflammation,

    G. Scadding, “Non-surgical treatment of adenoidal hypertrophy: the role of treat- ing ige-mediated inflammation,” Pediatric allergy and immunology, vol. 21, no. 8, pp. 1095–1106, 2010

  6. [6]

    Agree- ment between cone-beam computed tomography and nasoendoscopy evaluations of adenoid hypertrophy,

    M. P. Major, M. Witmans, H. El-Hakim, P. W. Major, and C. Flores-Mir, “Agree- ment between cone-beam computed tomography and nasoendoscopy evaluations of adenoid hypertrophy,” American Journal of Orthodontics and Dentofacial Orthopedics, vol. 146, no. 4, pp. 451–459, 2014

  7. [7]

    Pediatric endoscopic transnasal adenoid ablation,

    J. J. Shin and C. J. Hartnick, “Pediatric endoscopic transnasal adenoid ablation,” Annals of Otology, Rhinology & Laryngology , vol. 112, no. 6, pp. 511–514, 2003

  8. [8]

    Endoscopic adenoidectomy with the microde- brider,

    E. Yanagisawa and E. M. Weaver, “Endoscopic adenoidectomy with the microde- brider,” Ear, Nose & Throat Journal , vol. 76, no. 2, pp. 72–74, 1997

Show all 64 references
  1. [9]

    Radiographic evaluation of ade- noidal size in children: adenoidal-nasopharyngeal ratio,

    M. Fujioka, L. W. Young, and B. Girdany, “Radiographic evaluation of ade- noidal size in children: adenoidal-nasopharyngeal ratio,” American journal of roentgenology, vol. 133, no. 3, pp. 401–404, 1979

  2. [10]

    Size assessment of adenoid and nasopharyngeal airway by acoustic rhinometry in chil- dren,

    J.-H. Cho, D.-H. Lee, N.-S. Lee, Y.-S. Won, H.-R. Yoon, and B.-D. Suh, “Size assessment of adenoid and nasopharyngeal airway by acoustic rhinometry in chil- dren,” The Journal of Laryngology & Otology , vol. 113, no. 10, pp. 899–905, 1999

  3. [11]

    The adenoidal-nasopharyngeal ratio (an ratio): its validity in selecting children for adenoidectomy,

    S. Elwany, “The adenoidal-nasopharyngeal ratio (an ratio): its validity in selecting children for adenoidectomy,” The Journal of Laryngology & Otology , vol. 101, no. 6, pp. 569–573, 1987

  4. [12]

    Comparative analysis of conventional cold curettage versus endoscopic assisted coblation adenoidectomy,

    R. Bidaye, N. Vaid, and K. Desarda, “Comparative analysis of conventional cold curettage versus endoscopic assisted coblation adenoidectomy,” The Journal of Laryngology & Otology, vol. 133, no. 4, pp. 294–299, 2019. 24

  5. [13]

    Obstructive adenoid tissue: an indication for powered-shaver adenoidectomy,

    T. Havas and D. Lowinger, “Obstructive adenoid tissue: an indication for powered-shaver adenoidectomy,” Archives of Otolaryngology–Head & Neck Surgery, vol. 128, no. 7, pp. 789–791, 2002

  6. [14]

    Evaluation of adenoid hyper- trophy with ultrasonography,

    Y. Wang, H. Jiao, C. Mi, G. Yang, and T. Han, “Evaluation of adenoid hyper- trophy with ultrasonography,” The Indian Journal of Pediatrics , vol. 87, pp. 910–915, 2020

  7. [15]

    Automated adenoid hypertrophy assessment with lateral cephalometry in children based on artificial intelligence,

    T. Zhao, J. Zhou, J. Yan, L. Cao, Y. Cao, F. Hua, and H. He, “Automated adenoid hypertrophy assessment with lateral cephalometry in children based on artificial intelligence,” Diagnostics, vol. 11, no. 8, p. 1386, 2021

  8. [16]

    A deep-learning-based approach for adenoid hypertrophy diagnosis,

    Y. Shen, X. Li, X. Liang, H. Xu, C. Li, Y. Yu, and B. Qiu, “A deep-learning-based approach for adenoid hypertrophy diagnosis,” Medical Physics, vol. 47, no. 5, pp. 2171–2181, 2020

  9. [17]

    An efficient deep model for children sleep apnea detection using snoring signals,

    Z. Liang, Y. Zhou, L. Ding, and X. Chen, “An efficient deep model for children sleep apnea detection using snoring signals,” in Proceedings of the 3rd Inter- national Symposium on Artificial Intelligence for Medicine Sciences , 2022, pp. 551–558

  10. [18]

    Automated radiographic evaluation of adenoid hypertrophy based on vgg-lite,

    J. Liu, S. Li, Y. Cai, D. Lan, Y. Lu, W. Liao, S. Ying, and Z. Zhao, “Automated radiographic evaluation of adenoid hypertrophy based on vgg-lite,” Journal of Dental Research, vol. 100, no. 12, pp. 1337–1343, 2021

  11. [19]

    Con- trastive learning-based adenoid hypertrophy grading network using nasoendo- scopic image,

    S. Zheng, X. Li, M. Bi, Y. Wang, H. Liu, X. Feng, Y. Fan, and L. Shen, “Con- trastive learning-based adenoid hypertrophy grading network using nasoendo- scopic image,” in 2022 IEEE 35th International Symposium on Computer-Based Medical Systems (CBMS) . IEEE, 2022, pp. 377–382

  12. [20]

    Adenoid segmentation in x-ray images using u-net,

    A. A. Alshbishiri, M. A. Marghalani, H. A. Khan, R. G. Ahmad, M. A. Alqarni, and M. M. Khan, “Adenoid segmentation in x-ray images using u-net,” in 2021 National Computing Colleges Conference (NCCC) . IEEE, 2021, pp. 1–6

  13. [21]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer- Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18 ....

  14. [22]

    Dense-unet: a novel multiphoton in vivo cellular image segmentation model based on a convolutional neural network,

    S. Cai, Y. Tian, H. Lui, H. Zeng, Y. Wu, and G. Chen, “Dense-unet: a novel multiphoton in vivo cellular image segmentation model based on a convolutional neural network,” Quantitative imaging in medicine and surgery , vol. 10, no. 6, p. 1275, 2020. 25

  15. [23]

    Unet 3+: A full-scale connected unet for medical image segmenta- tion,

    H. Huang, L. Lin, R. Tong, H. Hu, Q. Zhang, Y. Iwamoto, X. Han, Y.-W. Chen, and J. Wu, “Unet 3+: A full-scale connected unet for medical image segmenta- tion,” in ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2020, p...

  16. [24]

    Unet++: A nested u-net architecture for medical image segmentation,

    Z. Zhou, M. M. Rahman Siddiquee, N. Tajbakhsh, and J. Liang, “Unet++: A nested u-net architecture for medical image segmentation,” in Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support: 4th International Workshop, DLMIA 2018, and 8th ...

  17. [25]

    3d u-net: learning dense volumetric segmentation from sparse annotation,

    ¨O. C ¸ i¸ cek, A. Abdulkadir, S. S. Lienkamp, T. Brox, and O. Ronneberger, “3d u-net: learning dense volumetric segmentation from sparse annotation,” in Med- ical Image Computing and Computer-Assisted Intervention–MICCAI 2016: 19th International Conference, Athens, Greece, Oc...

  18. [26]

    nnu-net: a self-configuring method for deep learning-based biomedical image segmentation,

    F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, and K. H. Maier-Hein, “nnu-net: a self-configuring method for deep learning-based biomedical image segmentation,” Nature methods, vol. 18, no. 2, pp. 203–211, 2021

  19. [27]

    Hierarchical 3d fully convolutional networks for multi-organ segmentation,

    H. R. Roth, H. Oda, Y. Hayashi, M. Oda, N. Shimizu, M. Fujiwara, K. Mis- awa, and K. Mori, “Hierarchical 3d fully convolutional networks for multi-organ segmentation,” arXiv preprint arXiv:1704.06382 , 2017

  20. [28]

    Attention gated networks: Learning to leverage salient regions in medical images,

    J. Schlemper, O. Oktay, M. Schaap, M. Heinrich, B. Kainz, B. Glocker, and D. Rueckert, “Attention gated networks: Learning to leverage salient regions in medical images,” Medical image analysis , vol. 53, pp. 197–207, 2019

  21. [29]

    Modified u-net (mu-net) with incorporation of object-dependent high level features for improved liver and liver-tumor segmentation in ct images,

    H. Seo, C. Huang, M. Bassenne, R. Xiao, and L. Xing, “Modified u-net (mu-net) with incorporation of object-dependent high level features for improved liver and liver-tumor segmentation in ct images,” IEEE transactions on medical imaging , vol. 39, no. 5, pp. 1316–1325, 2019

  22. [30]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020

  23. [31]

    Visual transformers: Token-based image represen- tation and processing for computer vision,

    B. Wu, C. Xu, X. Dai, A. Wan, P. Zhang, Z. Yan, M. Tomizuka, J. Gonzalez, K. Keutzer, and P. Vajda, “Visual transformers: Token-based image represen- tation and processing for computer vision,” arXiv preprint arXiv:2006.03677 , 2020

  24. [32]

    Generating long sequences with sparse transformers,

    R. Child, S. Gray, A. Radford, and I. Sutskever, “Generating long sequences with sparse transformers,” arXiv preprint arXiv:1904.10509 , 2019. 26

  25. [33]

    Reformer: The efficient transformer,

    N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” arXiv preprint arXiv:2001.04451 , 2020

  26. [34]

    Swin- unet: Unet-like pure transformer for medical image segmentation,

    H. Cao, Y. Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang, “Swin- unet: Unet-like pure transformer for medical image segmentation,” in European conference on computer vision . Springer, 2022, pp. 205–218

  27. [35]

    Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images,

    A. Hatamizadeh, V. Nath, Y. Tang, D. Yang, H. R. Roth, and D. Xu, “Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images,” in International MICCAI Brainlesion Workshop . Springer, 2021, pp. 272–284

  28. [36]

    Unetr: Transformers for 3d medical image segmentation,

    A. Hatamizadeh, Y. Tang, V. Nath, D. Yang, A. Myronenko, B. Landman, H. R. Roth, and D. Xu, “Unetr: Transformers for 3d medical image segmentation,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2022, pp. 574–584

  29. [37]

    Unetr++: Delving into efficient and accurate 3d medical image segmentation,

    A. M. Shaker, M. Maaz, H. Rasheed, S. Khan, M.-H. Yang, and F. S. Khan, “Unetr++: Delving into efficient and accurate 3d medical image segmentation,” IEEE Transactions on Medical Imaging , 2024

  30. [38]

    Convolution-free medical image segmentation using transformers,

    D. Karimi, S. D. Vasylechko, and A. Gholipour, “Convolution-free medical image segmentation using transformers,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedi...

  31. [39]

    Transbts: Multimodal brain tumor segmentation using transformer,

    W. Wang, C. Chen, M. Ding, H. Yu, S. Zha, and J. Li, “Transbts: Multimodal brain tumor segmentation using transformer,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Confer- ence, Strasbourg, France, September 27–October 1, 2021,...

  32. [40]

    Cotr: Efficiently bridging cnn and transformer for 3d medical image segmentation,

    Y. Xie, J. Zhang, C. Shen, and Y. Xia, “Cotr: Efficiently bridging cnn and transformer for 3d medical image segmentation,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Confer- ence, Strasbourg, France, September 27–October 1, 20...

  33. [41]

    nnformer: Interleaved transformer for volumetric segmentation,

    H.-Y. Zhou, J. Guo, Y. Zhang, L. Yu, L. Wang, and Y. Yu, “nnformer: Interleaved transformer for volumetric segmentation,” arXiv preprint arXiv:2109.03201 , 2021

  34. [42]

    Springer, 2021, pp. 171–180

  35. [43]

    Sobel edge detection algorithm,

    S. Gupta and S. G. Mazumdar, “Sobel edge detection algorithm,” International journal of computer science and management Research , vol. 2, no. 2, pp. 1578– 1583, 2013

  36. [44]

    V-net: Fully convolutional neural net- works for volumetric medical image segmentation,

    F. Milletari, N. Navab, and S.-A. Ahmadi, “V-net: Fully convolutional neural net- works for volumetric medical image segmentation,” in 2016 fourth international conference on 3D vision (3DV) . Ieee, 2016, pp. 565–571. 27

  37. [45]

    Merged u-net for bone tumors x-ray images segmentation,

    Z. Xie, K. Zhao, X. Yan, S. Wu, J. Mei, and H. Lu, “Merged u-net for bone tumors x-ray images segmentation,” in 2022 IEEE International Conference on Image Processing (ICIP). IEEE, 2022, pp. 1276–1280

  38. [46]

    Deep learning techniques for automatic mri cardiac multi-structures segmentation and diagnosis: is the problem solved?

    O. Bernard, A. Lalande, C. Zotti, F. Cervenansky, X. Yang, P.-A. Heng, I. Cetin, K. Lekadir, O. Camara, M. A. G. Ballester et al. , “Deep learning techniques for automatic mri cardiac multi-structures segmentation and diagnosis: is the problem solved?” IEEE transactions on med...

  39. [47]

    Rethinking exemplars for continual semantic segmentation in endoscopy scenes: Entropy-based mini-batch pseudo-replay,

    G. Wang, L. Bai, Y. Wu, T. Chen, and H. Ren, “Rethinking exemplars for continual semantic segmentation in endoscopy scenes: Entropy-based mini-batch pseudo-replay,” Computers in Biology and Medicine , vol. 165, p. 107412, 2023

  40. [48]

    Monai: An open-source framework for deep learning in healthcare,

    M. J. Cardoso, W. Li, R. Brown, N. Ma, E. Kerfoot, Y. Wang, B. Murrey, A. Myronenko, C. Zhao, D. Yang, V. Nath, Y. He, Z. Xu, A. Hatamizadeh, A. Myronenko, W. Zhu, Y. Liu, M. Zheng, Y. Tang, I. Yang, M. Zephyr, B. Hashemian, S. Alle, M. Z. Darestani, C. Budd, M. Modat, T. Verc...

  41. [49]

    Transunet: Transformers make strong encoders for medical image segmentation,

    J. Chen, Y. Lu, Q. Yu, X. Luo, E. Adeli, Y. Wang, L. Lu, A. L. Yuille, and Y. Zhou, “Transunet: Transformers make strong encoders for medical image segmentation,” 2021

  42. [50]

    Domain adaptive sim-to-real segmentation of oropharyngeal organs,

    G. Wang, T.-A. Ren, J. Lai, L. Bai, and H. Ren, “Domain adaptive sim-to-real segmentation of oropharyngeal organs,” arXiv preprint arXiv:2305.10883 , 2023

  43. [51]

    Deep network-based comprehensive parotid gland tumor detection,

    K. M. Sunnetci, E. Kaba, F. B. Celiker, and A. Alkan, “Deep network-based comprehensive parotid gland tumor detection,” Academic Radiology, vol. 31, no. 1, pp. 157–167, 2024. [Online]. Available: https://www.sciencedirect.com/ science/article/pii/S107663322300226X 28

  44. [52]

    Encoder-decoder with atrous separable convolution for semantic image segmentation,

    L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in Pro- ceedings of the European conference on computer vision (ECCV) , 2018, pp. 801–818

  45. [53]

    Mednext: Transformer-driven scaling of convnets for medical image segmentation,

    S. Roy, G. Koehler, C. Ulrich, M. Baumgartner, J. Petersen, F. Isensee, P. F. Jaeger, and K. Maier-Hein, “Mednext: Transformer-driven scaling of convnets for medical image segmentation,” 2024. [Online]. Available: https://arxiv.org/abs/2303.09975

  46. [54]

    Agileformer: Spatially agile transformer unet for medical image segmentation,

    P. Qiu, J. Yang, S. Kumar, S. S. Ghosh, and A. Sotiras, “Agileformer: Spatially agile transformer unet for medical image segmentation,” 2024. [Online]. Available: https://arxiv.org/abs/2404.00122

  47. [55]

    The medical segmentation decathlon,

    M. Antonelli, A. Reinke, S. Bakas, K. Farahani, A. Kopp-Schneider, B. A. Landman, G. Litjens, B. Menze, O. Ronneberger, R. M. Summers, B. van Ginneken, M. Bilello, P. Bilic, P. F. Christ, R. K. G. Do, M. J. Gollub, S. H. Heckers, H. Huisman, W. R. Jarnagin, M. K. McHugo, S. Na...

  48. [56]

    Missformer: An effective transformer for 2d medical image segmentation,

    X. Huang, Z. Deng, D. Li, X. Yuan, and Y. Fu, “Missformer: An effective transformer for 2d medical image segmentation,” IEEE Transactions on Medical Imaging, vol. 42, no. 5, pp. 1484–1494, 2022

  49. [57]

    Statistical validation of image segmentation quality based on a spatial overlap index: Scientific reports,

    K. H. Zou, S. K. Warfield, A. Bharatha, C. M. C. Tempany, M. R. Kaus, S. J. Haker, W. M. Wells, F. A. Jolesz, and R. Kikinis, “Statistical validation of image segmentation quality based on a spatial overlap index: Scientific reports,” Academic Radiology, vol. 11, no. 2, Feb. 2...

  50. [58]

    Swin- unet: Unet-like pure transformer for medical image segmentation,

    H. Cao, Y. Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang, “Swin- unet: Unet-like pure transformer for medical image segmentation,” 2021

  51. [59]

    Neural machine translation by jointly learning to align and translate,

    D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” arXiv preprint arXiv:1409.0473 , 2014

  52. [60]

    Metrics for evaluating 3d medical image segmentation: analysis, selection, and tool,

    A. A. Taha and A. Hanbury, “Metrics for evaluating 3d medical image segmentation: analysis, selection, and tool,” BMC Medical Imaging, vol. 15, no. 1, p. 29, 2015. [Online]. Available: https://doi.org/10.1186/s12880-015-0070-3

  53. [61]

    A computational approach to edge detection,

    J. Canny, “A computational approach to edge detection,” in Readings in Computer Vision , M. A. Fischler and O. Firschein, Eds., 1987, pp. 184–203

  54. [62]

    The pascal visual object classes (voc) challenge,

    M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” International journal of computer vision, vol. 88, no. 2, pp. 303–338, 2010. 29

  55. [64]

    Perception rein- forcement using auxiliary learning feature fusion: A modified yolov8 for head detection,

    J. Chen, G. Wang, W. Liu, X. Zhong, Y. Tian, and Z. Wu, “Perception rein- forcement using auxiliary learning feature fusion: A modified yolov8 for head detection,” in 2023 China Automation Congress (CAC) . IEEE, 2023, pp. 4709–4714. 30

  56. [2018]

    Available: https://www.sciencedirect.com/science/article/pii/ S108707921630137X

    [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S108707921630137X

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.