Pith. sign in

REVIEW 3 major objections 6 minor 36 references

A boundary-focused 3D network predicts perineural invasion from MRI more accurately than standard convolutional or transformer models.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 07:47 UTC pith:L67O6XDL

load-bearing objection Competent task-aligned hybrid for a hard clinical label: real within-cohort AUC lift, honest limits, but no uncertainty and single-center data keep it engineering-scale. the 3 major comments →

arxiv 2607.10992 v1 pith:L67O6XDL submitted 2026-07-13 cs.CV cs.AI

LoSA-Net: A Localized and Scale-Adaptive Network for Boundary-Sensitive Prediction of Perineural Invasion in 3D MRI

classification cs.CV cs.AI
keywords perineural invasionboundary-aware learningmulti-scale feature fusion3D MRIcholangiocarcinomaneighborhood attentionmedical image analysis
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Perineural invasion (PNI) signals aggressive tumor behavior and can change surgical planning, yet its MRI signs are faint, elongated, and easily confused with vessels or ducts. Standard volumetric networks lose those thin cues through downsampling or global attention. This paper introduces LoSA-Net, a four-stage encoder that keeps nerve-aligned detail with localized talking-head attention, adapts its receptive field with multi-scale depthwise mixing, and keeps coarse semantics aligned with fine edges across stages. On contrast-enhanced MRI from 168 cholangiocarcinoma patients it reaches a mean AUC of 0.7567 and beats matched CNN and transformer baselines. If the result holds, preoperative imaging could supply a usable PNI risk estimate before surgery rather than waiting for pathology.

Core claim

On a single-center cohort of 168 patients with pathologically confirmed cholangiocarcinoma, LoSA-Net’s combination of Talking Neighborhood Attention, Scale-Adaptive Feature Mixing, and Cross-Scale Refinement and Alignment yields an AUC of 0.7567 for binary PNI prediction and outperforms representative 3D ResNet, DenseNet, EfficientNet, Swin, and Neighborhood Attention models trained under identical preprocessing and optimization.

What carries the argument

Talking Neighborhood Attention (localized self-attention plus head-wise mixing) that preserves continuity along thin nerve-caliber structures while still mixing directional cues; together with Scale-Adaptive Feature Mixing and Cross-Scale Refinement and Alignment it keeps weak boundary signals from being erased by downsampling.

Load-bearing premise

Tumor-centered MRI crops actually contain enough visible, learnable signs of pathologically defined PNI, even though many invasions are microscopic and below the scanner’s resolution.

What would settle it

An independent multi-center test set of cholangiocarcinoma MRI in which LoSA-Net’s AUC falls to the level of the matched CNN or transformer baselines (around 0.68–0.70) while the same architecture still works on other boundary-sensitive tasks.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Preoperative PNI risk scores could be computed from routine hepatobiliary-phase MRI and used to flag cases that may need wider surgical margins.
  • The same three modules can be dropped into other 3D medical networks that must detect thin, low-contrast tubular structures.
  • Ablation results imply that zero-initialized cross-scale gates and multi-kernel depthwise branches are transferable design choices for any hierarchical volumetric encoder.
  • Grad-CAM maps that light up tumor borders and adjacent tubular anatomy give radiologists a visual check that the network is attending to clinically plausible regions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If microscopic PNI is truly invisible on MRI, the measured AUC may largely reflect peri-tumoral texture or vessel proximity rather than true neural invasion; multi-modal fusion with diffusion or high-resolution nerve imaging would test that.
  • The same locality-plus-scale recipe should transfer to other cancers where PNI is prognostic (pancreas, head-and-neck) once larger multi-center cohorts exist.
  • Tumor-centered cropping discards distant nerve pathways; full-volume or multi-crop inference could raise ceiling performance without changing the architecture.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript proposes LoSA-Net, a four-stage 3D volumetric encoder for binary preoperative prediction of perineural invasion (PNI) from contrast-enhanced hepatobiliary-phase T1 MRI in cholangiocarcinoma. The architecture couples Talking Neighborhood Attention (localized self-attention with talking-head mixing), Scale-Adaptive Feature Mixing (multi-branch depthwise convolutions with a dynamic scale selector), and Cross-Scale Refinement and Alignment (bidirectional coarse–fine residual paths with Laplacian edge prior and zero-initialized gates). On a single-center cohort of 168 patients (67 PNI-positive) after exclusions from 306 candidates, using patient-level 5-fold CV on tumor-centered 96×96×48 crops, LoSA-Net reports mean AUC 0.7567 and outperforms matched 3D CNN and transformer baselines (best ~0.695). Ablations attribute gains to each module; Grad-CAM maps emphasize peri-tumoral boundaries.

Significance. If the within-cohort ranking is robust, the work is a useful, well-motivated contribution to boundary-aware 3D medical imaging: it targets a clinically relevant but imaging-subtle endpoint, supplies inductive biases (locality, multi-scale depthwise mixing, cross-scale boundary feedback) that match the stated failure modes of strided CNNs and global transformers, and reports matched baselines plus module ablations rather than only a single end-to-end number. The explicit discussion of microscopic PNI below MRI resolution is appropriately cautious. Significance remains primarily methodological and cohort-internal; clinical or multi-center impact is not yet established.

major comments (3)
  1. Tables 1–2 and §3.2–3.3 report only mean 5-fold AUCs (LoSA-Net 0.7567 vs best baseline 0.6951; ablations down to ~0.70) with no fold-wise standard deviations, confidence intervals, or paired statistical tests. With N=168 (~33–34 patients/fold) and class imbalance (67/101), fold variance can plausibly exceed the ~0.06 absolute gap. The central claim that TNA/SAFM/CSRA yield superior discrimination under matched settings is therefore not yet shown to be statistically reliable; please report per-fold AUCs (or mean±SD), bootstrap CIs, and a paired test (e.g., DeLong or permutation) against the strongest baseline and key ablations.
  2. §3.1 and §4: the cohort is single-center after excluding 138/306 candidates (artifacts, missing sequences, incorrect labels, tumor burden >70%), with no external or multi-protocol test set. The headline comparison and module attributions are therefore conditional on this protocol and exclusion policy. At minimum, quantify sensitivity of the ranking to the exclusion criteria or hold out a protocol/time-based split; ideally add an external site or leave-one-scanner analysis so the claimed architectural advantage is not confounded with site-specific contrast and crop statistics.
  3. §4 and §3.1 acknowledge that pathologically defined PNI may be microscopic and below MRI resolution, while labels come from postoperative pathology rather than imaging-visible ground truth. The paper’s narrative that TNA/SAFM/CSRA recover “nerve-aligned” / “boundary-sensitive” perineural cues (Abstract, §2.1–2.3, Grad-CAM §3.4) therefore rests on an untested assumption that visible peri-tumoral edges are reliable correlates of the label. Please either (i) stratify performance by imaging-visible vs occult PNI if such annotation exists, or (ii) substantially temper causal language about recovering perineural trajectories and treat the task more carefully as image-based risk prediction under label noise.
minor comments (6)
  1. Fig. 1 stage-3 block is labeled “SFAM” instead of “SAFM”; correct the typo for consistency with the text and other stages.
  2. Eqs. (1)–(6) (TNA) and (7)–(11) (SAFM) would benefit from explicit neighborhood size, head count, and channel dimensions used in experiments so the architecture is fully reproducible from the paper alone.
  3. Table 1 lists only AUC; reporting sensitivity/specificity or balanced accuracy at a fixed operating point would help assess clinical utility under class imbalance.
  4. §3.1: state the exact class-balance α used in focal loss and the TNA neighborhood kernel size (NAT-TEN) for completeness.
  5. Fig. 5 Grad-CAM: clarify whether maps are from correctly classified cases only and whether any quantitative localization metric (e.g., overlap with annotated peri-tumoral vessels/ducts) was computed.
  6. Title/abstract use “Boundary-Sensitive Prediction”; ensure the same hyphenation and capitalization of LoSA-Net / TNA / SAFM / CSRA are consistent throughout (including arXiv header “INV ASION”).

Circularity Check

0 steps flagged

No significant circularity: empirical architecture paper with held-out evaluation and independent baselines.

full rationale

LoSA-Net is a standard empirical deep-learning architecture paper. Its central claims are measured AUCs (0.7567 mean 5-fold) obtained by training a volumetric encoder on patient-level cross-validation splits of 168 MRI cases and comparing against matched 3D CNN and transformer baselines under identical preprocessing and optimization (Tables 1–2, §3.2–3.3). Ablations re-train after removing or varying TNA/SAFM/CSRA and re-measure the same metric; none of these numbers is forced by construction from a fitted parameter that is then re-labeled a “prediction.” Module definitions (localized neighborhood attention with talking-head mixing, multi-scale depthwise branches with a dynamic selector, zero-initialized cross-scale scalars) are ordinary architectural choices, not self-definitional identities that equate inputs to the reported AUC. Citations (NAT [21], talking-heads [28], focal loss [30], Grad-CAM [31], etc.) are to external prior work and do not supply load-bearing uniqueness theorems authored by the present team. Design hyperparameters (kernel set {3,5,7}, zero-init of λ/μ/ν) are not circular derivations. The evaluation chain is therefore self-contained against external benchmarks; no step reduces by the paper’s own equations or self-citation to its own inputs.

Axiom & Free-Parameter Ledger

6 free parameters · 4 axioms · 3 invented entities

The central claim rests on empirical training choices and domain assumptions about MRI-visible PNI, not on free-form physical constants. Free parameters are standard ML hyperparameters and architectural knobs. Axioms are domain assumptions (PNI has learnable MRI correlates; pathology labels are usable targets) plus standard math operators. Invented entities are the three named modules, which are architectural constructs rather than physical objects; they have no independent evidence outside this paper’s ablations.

free parameters (6)
  • SAFM kernel set K
    Hand-chosen multi-scale depthwise kernels {3,5,7}; ablations show sensitivity to single-kernel choices.
  • AdamW learning rate and schedule
    Initial LR 1e-4 with cosine decay for 200 epochs; not derived, selected for training.
  • Focal loss gamma and class-balance alpha
    γ=2 fixed; α dataset-derived. Directly shapes optimization under imbalance.
  • CSRA scalars λ, μ, ν initialization
    Zero-init chosen for stability; nonzero Gaussian init degrades AUC in ablation.
  • Input crop / voxel grid 96×96×48
    Resampling and tumor-centered crop size are design choices that define the model’s field of view.
  • TNA neighborhood size / head count
    Local neighborhood and talking-head mixing matrices are architectural free choices (exact N(i) size not fully specified).
axioms (4)
  • domain assumption Pathologically confirmed PNI labels are valid supervision targets for preoperative MRI prediction even when invasion may be microscopic.
    Stated in Introduction and Discussion; load-bearing for treating AUC as clinically meaningful.
  • domain assumption Tumor-centered peri-tumoral crops contain the imaging correlates of perineural spread needed for discrimination.
    Dataset §3.1; ablation shows uncropped input drops AUC to 0.6493.
  • standard math Localized self-attention, multi-scale depthwise mixing, and residual cross-scale alignment are valid operators for 3D feature learning (standard ML math).
    Equations (1)–(15); standard softmax attention, convolutions, trilinear resampling, Laplacian edge prior.
  • ad hoc to paper Single-center hepatobiliary-phase T1 protocol after artifact/label exclusions is representative enough for the reported comparison.
    §3.1 exclusion from 306 to 168; no multi-center assumption is tested.
invented entities (3)
  • Talking Neighborhood Attention (TNA) no independent evidence
    purpose: Preserve nerve-aligned local detail via neighborhood attention plus pre/post-softmax head mixing.
    Named module combining NAT-style locality with talking-heads; evidence is internal ablation only.
  • Scale-Adaptive Feature Mixing (SAFM) no independent evidence
    purpose: Adapt effective receptive field with multi-branch depthwise kernels and a dynamic selector.
    Architectural construct; gains shown only within this paper’s ablations.
  • Cross-Scale Refinement and Alignment (CSRA) no independent evidence
    purpose: Align coarse semantics and fine boundary cues across stages with zero-init gated residuals and Laplacian edge prior.
    Paper-specific stage coupling; no external independent validation.

pith-pipeline@v1.1.0-grok45 · 13283 in / 3513 out tokens · 29568 ms · 2026-07-14T07:47:53.114853+00:00 · methodology

0 comments
read the original abstract

Perineural invasion (PNI) is a clinically relevant indicator of tumor aggressiveness and can influence surgical decision-making, motivating interest in reliable preoperative assessment. The subtle MRI features of PNI, however, often resemble nearby anatomy, complicating noninvasive prediction. These fine perineural cues are easily attenuated by routine downsampling or overly global feature aggregation, reducing the effectiveness of conventional volumetric models. We present LoSA-Net, a localized and scale-adaptive architecture for boundary-sensitive PNI prediction in 3D MRI. Talking Neighborhood Attention (TNA) preserves nerve-aligned detail through localized self-attention with head-wise mixing, and Scale-Adaptive Feature Mixing (SAFM) modulates the receptive field using multi-scale depthwise processing. Cross-Scale Refinement and Alignment (CSRA) maintains consistency between semantic context and high-resolution boundaries across stages. In contrast-enhanced MRI scans from 168 patients with cholangiocarcinoma, LoSA-Net achieves an AUC of 0.7567 and outperforms representative convolutional and transformer baselines under matched preprocessing and optimization settings.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

36 extracted references · 3 linked inside Pith

  1. [1]

    INTRODUCTION Perineural invasion (PNI) is the spread of tumor cells along or within the nerve sheath and is associated with pain, neurologic deficits, and poor survival across multiple malignancies [1, 2]. PNI may war- rant more aggressive margin considerations during surgery [3, 4], suggesting that reliable noninvasive preoperative identification could h...

  2. [2]

    METHODOLOGY LoSA-Net is a four-stage volumetric encoder (Fig. 1). An overlap- ping convolutional tokenizer maps the input into a dense grid of tokens. Each stage stacks blocks that couple Talking Neighborhood Attention (TNA) with Scale-Adaptive Feature Mixing (SAFM). Cross-Scale Refinement and Alignment (CSRA) connects adjacent arXiv:2607.10992v1 [cs.CV] ...

  3. [3]

    Dataset and Implementation We utilized contrast-enhanced, hepatobiliary-phase T1-weighted MRI volumes from Samsung Medical Center, acquired over approx- imately 10 years

    EXPERIMENTS AND RESULTS 3.1. Dataset and Implementation We utilized contrast-enhanced, hepatobiliary-phase T1-weighted MRI volumes from Samsung Medical Center, acquired over approx- imately 10 years. From 306 initial candidates, scans with severe motion artifacts, missing sequences, incorrect labels, or tumor bur- den exceeding 70% of liver volume were ex...

  4. [4]

    DISCUSSION AND CONCLUSION This study addressed the problem of predicting PNI in 3D MRI, where the relevant imaging patterns are subtle, elongated, and often obscured by nearby vascular or ductal structures. Across a cohort of patients with cholangiocarcinoma, LoSA-Net achieved higher dis- crimination than representative convolutional and transformer mod- ...

  5. [5]

    2020-0-01305)

    ACKNOWLEDGMENTS This work was supported by the Institute of Information & Commu- nications Technology Planning & Evaluation (IITP), funded by the Korea government (MSIT), under the Artificial Intelligence Semi- conductor Support Program to nurture the best talents (IITP-2023- RS-2023-00256081) and the grant for the Development of an AI Deep Learning Proce...

  6. [6]

    Perineural invasion in cancer: a review of the literature,

    C. Liebig, G. Ayala, J. A. Wilks, D. H. Berger, and D. Albo, “Perineural invasion in cancer: a review of the literature,”Can- cer: Interdisciplinary International Journal of the American Cancer Society, vol. 115, no. 15, pp. 3379–3391, 2009

  7. [7]

    Perineural invasion and perineural tumor spread in head and neck cancer,

    R. L. Bakst, C. M. Glastonbury, U. Parvathaneni, N. Katabi, K. S. Hu, and S. S. Yom, “Perineural invasion and perineural tumor spread in head and neck cancer,”International Journal of Radiation Oncology* Biology* Physics, vol. 103, no. 5, pp. 1109–1124, 2019

  8. [8]

    Z. Liu, C. Luo, X. Chen, Y . Feng, J. Feng, R. Zhang, F. Ouyang, X. Li, Z. Tan, L. Denget al., “Noninvasive predic- tion of perineural invasion in intrahepatic cholangiocarcinoma by clinicoradiological features and computed tomography ra- diomics based on interpretable machine learning: a multicenter cohort study,”International Journal of Surgery, vol. 11...

  9. [9]

    Influence of surgical margins on overall survival after resection of intrahep- atic cholangiocarcinoma: A meta-analysis,

    H. Tang, W. Lu, B. Li, X. Meng, and J. Dong, “Influence of surgical margins on overall survival after resection of intrahep- atic cholangiocarcinoma: A meta-analysis,”Medicine, vol. 95, no. 35, p. e4621, 2016

  10. [10]

    Perineural invasion and spread in common abdominopelvic diseases: imaging diagnosis and clinical significance,

    W. Tu, R. V . Gottumukkala, N. Schieda, L. Lavall ´ee, B. A. Adam, and S. G. Silverman, “Perineural invasion and spread in common abdominopelvic diseases: imaging diagnosis and clinical significance,”Radiographics, vol. 43, no. 7, p. e220148, 2023

  11. [11]

    Diagnostic accu- racy of mri in detecting the perineural spread of head and neck tumors: a systematic review and meta-analysis,

    U. Abdullaeva, B. Pape, and J. Hirvonen, “Diagnostic accu- racy of mri in detecting the perineural spread of head and neck tumors: a systematic review and meta-analysis,”Diagnostics, vol. 14, no. 1, p. 113, 2024

  12. [12]

    Zhang, W

    W. Zhang, W. Zhang, X. Li, X. Cao, G. Yang, and H. Zhang, “Predicting tumor perineural invasion status in high-grade prostate cancer based on a clinical–radiomics model incorpo- rating t2-weighted and diffusion-weighted magnetic resonance images,”Cancers, vol. 15, no. 1, p. 86, 2022

  13. [13]

    Ct-based ra- diomics analysis for noninvasive prediction of perineural inva- sion of perihilar cholangiocarcinoma,

    P.-C. Zhan, P.-j. Lyu, Z. Li, X. Liu, H.-X. Wang, N.-N. Liu, Y . Zhang, W. Huang, Y . Chen, and J.-b. Gao, “Ct-based ra- diomics analysis for noninvasive prediction of perineural inva- sion of perihilar cholangiocarcinoma,”Frontiers in Oncology, vol. 12, p. 900478, 2022

  14. [14]

    Making radiomics more reproducible across scan- ner and imaging protocol variations: a review of harmonization methods,

    S. A. Mali, A. Ibrahim, H. C. Woodruff, V . Andrearczyk, H. M ¨uller, S. Primakov, Z. Salahuddin, A. Chatterjee, and P. Lambin, “Making radiomics more reproducible across scan- ner and imaging protocol variations: a review of harmonization methods,”Journal of personalized medicine, vol. 11, no. 9, p. 842, 2021

  15. [15]

    Reproducibil- ity and generalizability in radiomics modeling: possible strate- gies in radiologic and statistical perspectives,

    J. E. Park, S. Y . Park, H. J. Kim, and H. S. Kim, “Reproducibil- ity and generalizability in radiomics modeling: possible strate- gies in radiologic and statistical perspectives,”Korean journal of radiology, vol. 20, no. 7, pp. 1124–1137, 2019

  16. [16]

    Enhancing radiomics reproducibility: Deep learning-based harmonization of abdominal computed tomography (ct) images,

    S. B. Lee, Y . Hong, Y . J. Cho, D. Jeong, J. Lee, J. W. Choi, J. Y . Hwang, S. Lee, Y . H. Choi, and J.-E. Cheon, “Enhancing radiomics reproducibility: Deep learning-based harmonization of abdominal computed tomography (ct) images,”Bioengi- neering, vol. 11, no. 12, p. 1212, 2024

  17. [17]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770– 778

  18. [18]

    Densely connected convolutional networks,

    G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” inProceedings of the IEEE conference on computer vision and pattern recog- nition, 2017, pp. 4700–4708

  19. [19]

    Efficientnet: Rethinking model scaling for convolutional neural networks,

    M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” inInternational conference on machine learning. PMLR, 2019, pp. 6105–6114

  20. [20]

    A comprehen- sive review of u-net and its variants: Advances and applica- tions in medical image segmentation,

    W. Jiangtao, N. I. R. Ruhaiyem, and F. Panpan, “A comprehen- sive review of u-net and its variants: Advances and applica- tions in medical image segmentation,”IET Image Processing, vol. 19, no. 1, p. e70019, 2025

  21. [21]

    Dra-net: Medical image seg- mentation based on adaptive feature extraction and region-level information fusion,

    Z. Huang, L. Wang, and L. Xu, “Dra-net: Medical image seg- mentation based on adaptive feature extraction and region-level information fusion,”Scientific Reports, vol. 14, no. 1, p. 9714, 2024

  22. [22]

    Transformers in vision: A survey,

    S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,”ACM computing surveys (CSUR), vol. 54, no. 10s, pp. 1–41, 2022

  23. [23]

    Transformers in medical imaging: A survey,

    F. Shamshad, S. Khan, S. W. Zamir, M. H. Khan, M. Hayat, F. S. Khan, and H. Fu, “Transformers in medical imaging: A survey,”Medical image analysis, vol. 88, p. 102802, 2023

  24. [24]

    Transforming medical imaging with transform- ers? a comparative review of key properties, current pro- gresses, and future perspectives,

    J. Li, J. Chen, Y . Tang, C. Wang, B. A. Landman, and S. K. Zhou, “Transforming medical imaging with transform- ers? a comparative review of key properties, current pro- gresses, and future perspectives,”Medical image analysis, vol. 85, p. 102762, 2023

  25. [25]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inProceedings of the IEEE/CVF in- ternational conference on computer vision, 2021, pp. 10 012– 10 022

  26. [26]

    Neighborhood attention transformer,

    A. Hassani, S. Walton, J. Li, S. Li, and H. Shi, “Neighborhood attention transformer,” inProceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, 2023, pp. 6185–6194

  27. [27]

    Going deeper with convolutions,

    C. Szegedy, W. Liu, Y . Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V . Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 1–9

  28. [28]

    Aggre- gated residual transformations for deep neural networks,

    S. Xie, R. Girshick, P. Doll ´ar, Z. Tu, and K. He, “Aggre- gated residual transformations for deep neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1492–1500

  29. [29]

    Squeeze-and-excitation net- works,

    J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation net- works,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132–7141

  30. [30]

    Selective kernel net- works,

    X. Li, W. Wang, X. Hu, and J. Yang, “Selective kernel net- works,” inProceedings of the IEEE/CVF conference on com- puter vision and pattern recognition, 2019, pp. 510–519

  31. [31]

    Xception: Deep learning with depthwise separable convolutions,

    F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” inProceedings of the IEEE conference on com- puter vision and pattern recognition, 2017, pp. 1251–1258

  32. [32]

    Mobilenetv2: Inverted residuals and linear bottle- necks,

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.- C. Chen, “Mobilenetv2: Inverted residuals and linear bottle- necks,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4510–4520

  33. [33]

    Talking- heads attention,

    N. Shazeer, Z. Lan, Y . Cheng, N. Ding, and L. Hou, “Talking- heads attention,”arXiv preprint arXiv:2003.02436, 2020

  34. [34]

    Decoupled weight decay regular- ization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regular- ization,”arXiv preprint arXiv:1711.05101, 2017

  35. [35]

    Focal loss for dense object detection,

    T.-Y . Lin, P. Goyal, R. Girshick, K. He, and P. Doll ´ar, “Focal loss for dense object detection,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 2980– 2988

  36. [36]

    Grad-cam: Visual explanations from deep net- works via gradient-based localization,

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep net- works via gradient-based localization,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 618–626