Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Axis-wise Gaussian priors from vision-language features, injected into axial attention, raise cervical cytology classification to 99.48% and 96.08% accuracy on two public Pap-smear datasets.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 12:57 UTC pith:TXXHDIZI

load-bearing objection Clean engineering synthesis of CLIP + Gaussian MoE axis priors + axial attention that posts strong numbers on two small cytology sets; ablations help, but single-split uncertainty leaves the geometry claim only partially secured. the 3 major comments →

arxiv 2607.10278 v1 pith:TXXHDIZI submitted 2026-07-11 cs.CV

Geometry-aware Gaussian Prior and Axial Attention for Cervical Cytology Image Classification

classification cs.CV
keywords Cervical cancer screeningCervical cytology image classificationPap smearGaussian priorsAxial self-attentionCLIPGeometry-aware attention
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Automated reading of Pap-smear images is still limited by the fact that normal, precancerous and cancerous cells look very similar and sit in complex spatial patterns. Convolution networks see local texture but miss long-range structure; plain attention sees global context but has no built-in notion of cellular geometry. This paper claims that the missing ingredient is a set of direction-sensitive Gaussian priors, generated by routing a CLIP global semantic token through a mixture of Gaussian experts, then multiplied into horizontal and vertical axial attention. The resulting geometry-aware model reaches 99.48% accuracy on the Mendeley liquid-based cytology set and 96.08% on SIPaKMeD, while the learned priors light up the same nuclear and cellular regions a cytologist would examine. If the claim holds, screening systems can become both more accurate and more interpretable without hand-crafted geometric rules.

Core claim

The authors show that structural regularities of cervical cells—nuclear alignment and spatial organisation—can be captured by Top-K Gaussian experts conditioned only on a pretrained CLIP global token, and that multiplying those axis-wise priors into axial self-attention produces a classifier that outperforms recent CNN and hybrid baselines on two standard cytology benchmarks while highlighting diagnostically relevant regions.

What carries the argument

Geometry-aware Gaussian Prior Module: a router selects Top-K Gaussian experts from a CLIP-derived global descriptor, forms directional weighting maps, converts them into pairwise cosine-similarity prior matrices, and multiplies those matrices into horizontal and vertical axial attention scores.

Load-bearing premise

That a handful of Gaussians steered by a single CLIP global token actually encode the clinically important spatial patterns of cells, rather than merely memorising the spatial statistics of the two training sets.

What would settle it

Train the identical architecture on a third, multi-centre cytology collection whose nuclear spacing and clustering statistics differ from Mendeley and SIPaKMeD; if accuracy collapses or the visualised priors stop highlighting nuclei, the structural-prior claim is falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Screening pipelines can replace pure CNN or plain Transformer backbones with a CLIP-plus-Gaussian-axial block and expect higher accuracy and more balanced precision-recall on liquid-based and single-cell Pap images.
  • Visualisation of the learned Gaussian maps can serve as a lightweight, built-in explanation layer that points a cytologist to the same regions the model used for its decision.
  • The same axis-wise Gaussian-expert construction can be dropped into other high-resolution medical tasks that need directional structural bias without quadratic full self-attention cost.
  • Ablation results imply that three Transformer stages and a moderate Top-K (6–8) already give the best accuracy–stability trade-off, guiding practical hyper-parameter choices.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the Gaussians truly encode nuclear geometry, the same module should transfer to other cytology domains (urine, breast fine-needle aspirates) with only light fine-tuning of the router.
  • The multiplicative prior modulation may be replaceable by an additive relative bias; a controlled swap would test whether the Gaussian shape itself, rather than any soft spatial prior, is necessary.
  • Because the method relies on CLIP’s semantic abstraction, any future vision-language backbone that better separates cellular morphology should raise the ceiling further without redesigning the attention block.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The manuscript proposes a geometry-aware framework for cervical cytology image classification that extracts features with a pretrained CLIP-ViT backbone, generates axis-wise Gaussian priors via a Top-K mixture of Gaussian experts conditioned on the global semantic token (Eqs. 3–11, §3.2), and multiplicatively injects those priors into axial self-attention (Eqs. 18–19, §3.3) before a lightweight classification head. On the Mendeley LBC and SIPaKMeD public datasets the method reports 99.48 % accuracy / 99.27 % F1 and 96.08 % accuracy / 95.40 % F1, respectively, outperforming a range of recent CNN and hybrid baselines (Table 2). Ablations isolate Transformer depth, the geometric-prior module, expert count, and the two axial directions (Tables 3–6), and qualitative visualizations of priors and confidence scores are supplied.

Significance. If the claimed gains are robust, the work supplies a concrete, domain-motivated inductive bias—direction-sensitive Gaussian structural priors—for a clinically important screening task that still relies heavily on manual microscopy. The architecture is clearly specified, the ablations systematically remove each proposed component, and the interpretability narrative (priors highlighting cellular regions) is a useful addition for decision-support systems. The contribution is incremental rather than foundational, but it is a legitimate engineering advance for automated Pap-smear analysis provided the numerical superiority can be shown to be stable.

major comments (3)
  1. Tables 2–6 and §4–5 report all metrics on a single fixed 80/20 split of two small collections (Mendeley LBC ≈ 963 images, SIPaKMeD ≈ 4 049 images) with no multi-seed averages, k-fold cross-validation, or confidence intervals. Absolute margins over already-strong baselines are 0.1–1.2 %. Under these conditions the headline numbers (99.48 % / 96.08 %) cannot be distinguished from split luck or from the CLIP + DFormer backbone alone; the causal link between the geometry-aware construction and the claimed gains remains unsecured.
  2. The ablations never isolate the Gaussian modulation itself. Table 4 removes the entire geometric-prior module and Table 6 removes axial axes, but there is no control that retains axial attention while replacing the multiplicative Gaussian terms (Eqs. 18–19) by identity or by a non-Gaussian positional bias. Consequently it is impossible to attribute the observed lifts specifically to the Top-K Gaussian experts (Eqs. 3–11) rather than to axial attention or the pretrained backbone.
  3. §3.2 and the interpretability claim rest on the premise that a single CLIP global token supplies sufficient semantic guidance to route and parameterize Gaussians that capture clinically relevant nuclear alignment and cellular organization. No quantitative localization metric (e.g., overlap with nuclei or expert-annotated ROIs) is reported; the visual analysis is purely qualitative. Without such evidence the clinical-utility narrative remains an untested assumption.
minor comments (4)
  1. Related-work §2.2 is dominated by self-citations on captioning, video understanding and multimodal generation that are only loosely connected to cytology classification; a tighter focus on medical image classification and geometric priors would improve relevance.
  2. Notation for the prior matrices is overloaded: G_W / G_H first denote 1-D Gaussian distributions (Eq. 11) and later pairwise cosine-similarity matrices (Eqs. 13–14). Distinct symbols would avoid confusion.
  3. Implementation details (§4.3) list many free hyperparameters (Top-K, E, γ, scale factor 9, warm-up lengths, peak LRs) without sensitivity analysis beyond the Top-K table; a short discussion of robustness would strengthen reproducibility.
  4. Figure 1 and the graphical abstract are described but their visual content is not fully self-explanatory in the text; captions should explicitly state what structural cues the priors are intended to highlight.

Circularity Check

1 steps flagged

No derivation circularity: accuracy claims are empirical measurements on held-out splits, not quantities defined by the Gaussian priors or axial modulation.

specific steps
  1. self citation load bearing [§2.2 Multimodal Vision and Structure-aware Representation (refs [6]–[58])]
    "These studies motivate our use of pretrained semantic features as a high-level guide for learning structure-sensitive cytology representations. ... Our method adapts this insight to cervical cytology by using Gaussian experts to inject direction-sensitive geometric priors into axial attention."

    A large fraction of the multimodal/structure-aware citations are by overlapping authors (Chen, Mao, Ye, etc.). They supply motivational framing rather than a uniqueness theorem or a numerical result that the present accuracy claims depend on. The cytology numbers themselves are measured on Mendeley LBC and SIPaKMeD, so the self-citations are not load-bearing for the strongest claim; the step is therefore only minor.

full rationale

The paper proposes an architecture (CLIP backbone + Gaussian expert priors + axial attention) and reports classification metrics on public cytology datasets. The Gaussian parameters (means, widths, Top-K routing) are learned end-to-end from training data and then evaluated on a fixed 80/20 test split; nothing in Eqs. 3–24 defines accuracy, F1, or the reported margins in terms of a fitted quantity that is later re-presented as a prediction. Ablations (Tables 3–6) remove modules and re-measure the same external metrics, which is ordinary empirical validation rather than self-definitional reduction. Related-work citations include many papers from the same group, but those works address captioning, report generation, and multimodal tasks; none of them supply the cytology accuracy numbers or a uniqueness theorem that forces the present design. The design choices (Gaussian experts, axial decomposition) are inductive biases, not tautologies. Consequently the central claim does not reduce by construction to its inputs. Minor self-citation density exists but is not load-bearing for the reported results, yielding a score of 1.

Axiom & Free-Parameter Ledger

5 free parameters · 3 axioms · 2 invented entities

The central performance claim rests on a modest set of standard mathematical objects, a domain modeling assumption about cellular geometry, several hand-chosen hyper-parameters, and two newly introduced modules whose only evidence is the paper’s own ablations.

free parameters (5)
  • Top-K experts = 8 / 6
    Chosen separately per dataset (K=8 for LBC, K=6 for SIPaKMeD) after ablation; directly controls which Gaussians contribute to the prior.
  • Number of Gaussian experts E = 10
    Fixed at 10; ablation shows sensitivity of final accuracy to this choice.
  • Gaussian decay coefficient gamma
    Multiplies the prior inside the attention softmax (Eqs. 18–19); value not reported but required for the modulation to be non-trivial.
  • Peak learning rates and warm-up epochs = 5e-6 / 1e-5
    Dataset-specific (5e-6 / 20 epochs for LBC; 1e-5 / 40 epochs for SIPaKMeD); affect convergence to the reported numbers.
  • Gaussian scale normalization factor = 9
    Set to 9; rescales expert widths before aggregation.
axioms (3)
  • standard math Standard scaled-dot-product attention and Gaussian density formulas are valid feature aggregators.
    Used throughout §§3.2–3.3 without further justification.
  • domain assumption Cervical cells exhibit diagnostically useful structural regularities (nuclear alignment, spatial organization) that can be captured by axis-aligned Gaussian maps.
    Stated in abstract and §3.2; underpins the entire prior-generation design.
  • ad hoc to paper A single CLIP global token supplies sufficient semantic guidance to route and parameterize the Gaussian experts for cytology images.
    Introduced in Eqs. 3–7; no external evidence that CLIP’s natural-image pretraining transfers this way.
invented entities (2)
  • Geometry-aware Gaussian Prior Module (Gaussian experts + router) no independent evidence
    purpose: Generate direction-sensitive spatial priors from global semantics to re-weight features and build geometry-aware attention matrices.
    Core novel module of the paper; independent evidence limited to internal ablations and qualitative maps.
  • Gaussian-enhanced Axial Self-Attention no independent evidence
    purpose: Modulate standard axial attention scores by the learned Gaussian priors so that long-range dependencies respect structural constraints.
    Second core module; again supported only by the paper’s own experiments.

pith-pipeline@v1.1.0-grok45 · 24152 in / 3005 out tokens · 31329 ms · 2026-07-14T12:57:18.092396+00:00 · methodology

0 comments
read the original abstract

Accurate cervical cytology image classification is a key component of automated cervical cancer screening, where reliable recognition of normal, precancerous, and cancer-associated cellular patterns from Pap smear images can improve screening efficiency and diagnostic consistency. However, this task remains challenging because cervical cells exhibit complex morphology, subtle intra-class variations, and strong inter-class similarities. Existing convolution-based models capture local texture well but have limited ability to model long-range relationships, whereas attention-based models provide broader context but often lack explicit structural guidance. To address these limitations, we propose a geometry-aware classification framework for cervical cancer screening-oriented cytology image analysis, incorporating semantic abstraction and structural priors learned from pre-trained vision-language features. The method uses Gaussian expert modules to generate axis-wise priors from global semantic information, capturing structural regularities such as nuclear alignment and cellular spatial organization. These priors are embedded into an axial self-attention module to modulate similarity computation along horizontal and vertical directions, improving long-range dependency modeling and structure-sensitive feature interaction. Experiments on the Mendeley liquid-based cytology and SIPaKMeD datasets show that the proposed method achieves 99.48% accuracy on the former and 96.08% on the latter, with balanced gains in recall, precision, and overall classification performance. Visual analysis further shows that the learned priors highlight diagnostically relevant cellular regions, demonstrating the potential of the proposed framework as a screening-oriented decision-support tool for cervical cytology.

Figures

Figures reproduced from arXiv: 2607.10278 by Cheng Ye, Nenan Lyu, Weidong Chen, Yating Li, Zhendong Mao.

Figure 1
Figure 1. Figure 1: Comparison between previous cervical cytology classification pipelines and [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overall architecture of the proposed geometry-aware cervical cancer [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Comparison with baseline methods on the Mendeley LBC dataset. [PITH_FULL_IMAGE:figures/full_fig_p020_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Comparison with baseline methods on the SIPaKMeD dataset. [PITH_FULL_IMAGE:figures/full_fig_p021_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Confusion matrices of the proposed method on the Mendeley LBC and [PITH_FULL_IMAGE:figures/full_fig_p022_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Micro-average ROC curves of the proposed method on two datasets: (a) [PITH_FULL_IMAGE:figures/full_fig_p023_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Visualization of class-wise prediction confidence (%) on two cervical cy [PITH_FULL_IMAGE:figures/full_fig_p024_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HTT-Net: Hierarchical Text-guided Transition Modeling for Surgical Video Phase Recognition

    cs.CV 2026-07 conditional novelty 6.0

    A hierarchical text-guided network that constructs and calibrates phase segments improves surgical phase recognition, setting a high Jaccard on Cholec80 and reporting large gains on a private LCRS-100 benchmark.

Reference graph

Works this paper leans on

88 extracted references · 14 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Global strategy to accelerate the elimination of cervical cancer as a public health problem

    World Health Organization. Global strategy to accelerate the elimination of cervical cancer as a public health problem . World Health Organization, 2020

  2. [2]

    Summers, Shaoxiong Liu, and Jianhua Y ao

    Ling Zhang, Le Lu, Isabella Nogues, Ronald M. Summers, Shaoxiong Liu, and Jianhua Y ao. DeepPap: Deep Convolutional Networks for Cervical Cell Clas- sification. 2017, IEEE Journal of Biomedical and Health Informatics , 21 (6): 1633–1643

  3. [3]

    Cervical cytology classification using PCA and GWO enhanced deep features selection

    Hritam Basak, Rohit Kundu, Sukanta Chakraborty, and Nibaran Das. Cervical cytology classification using PCA and GWO enhanced deep features selection. 2021, SN Computer Science, 2 (5): 369

  4. [4]

    An image is worth 16x16 words: Transformers for image recognition at scale, 2020

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale, 2020. arXiv preprint arXiv:2010.11929

  5. [5]

    TransMIL: Transformer based correlated multiple instance learning for whole slide image classification

    Zhuchen Shao, Hao Bian, Y ang Chen, Yifeng Wang, Jian Zhang, and Xiangyang Ji. TransMIL: Transformer based correlated multiple instance learning for whole slide image classification. 2021, Advances in Neural Information Processing Systems, 34: 2136–2147

  6. [6]

    Bootstrapping Large Language Models for Radiology Report Generation

    Chang Liu, Yuanhe Tian, Weidong Chen, Y an Song, and Y ongdong Zhang. Bootstrapping Large Language Models for Radiology Report Generation. 2024, Proceedings of the AAAI Conference on Artificial Intelligence , 38 (17): 18635– 18643

  7. [7]

    Im- proving radiology report generation with multi-grained abnormality prediction

    Yuda Jin, Weidong Chen, Yuanhe Tian, Y an Song, and Chenggang Y an. Im- proving radiology report generation with multi-grained abnormality prediction. 2024, Neurocomputing, 600: 128122

  8. [8]

    dong Mao

    Yuda Jin, Weidong Chen, Yuanhe Tian, Y an Song, Chenggang Y an, and Zhen- 28 JUSTC Geometry-aware Cytology Classification Yating Li et al . dong Mao. Improving Radiology Report Generation with D 2-Net: When Diffu- sion Meets Discriminator. In: ICASSP 2024 - 2024 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP) , 2215–221...

  9. [9]

    Im- proving Image Captioning via Predicting Structured Concepts

    Ting Wang, Weidong Chen, Yuanhe Tian, Y an Song, and Zhendong Mao. Im- proving Image Captioning via Predicting Structured Concepts. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 360–370. Association for Computational Linguistics, 2023

  10. [10]

    Ex- ploring Visual Relationships via Transformer-based Graphs for Enhanced Image Captioning

    Jingyu Li, Zhendong Mao, Hao Li, Weidong Chen, and Y ongdong Zhang. Ex- ploring Visual Relationships via Transformer-based Graphs for Enhanced Image Captioning. 2024, ACM Transactions on Multimedia Computing, Communica- tions, and Applications, 20 (5): 1–23

  11. [11]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classifica- tion with deep convolutional neural networks. 2012, Advances in Neural Infor- mation Processing Systems, 25

  12. [12]

    Network In Network, 2013

    Min Lin, Qiang Chen, and Shuicheng Y an. Network In Network, 2013. arXiv preprint arXiv:1312.4400

  13. [13]

    Going Deeper With Convolutions

    Christian Szegedy, Wei Liu, Y angqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabi- novich. Going Deeper With Convolutions. In: Proceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition , 1–9, 2015

  14. [14]

    Very deep convolutional networks for large-scale image recognition, 2014

    Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition, 2014. arXiv preprint arXiv:1409.1556

  15. [15]

    Deep Residual Learning for Image Recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–778, 2016

  16. [16]

    Segmenting Retinal Blood Vessels With Deep Neural Networks

    Paweł Liskowski and Krzysztof Krawiec. Segmenting Retinal Blood Vessels With Deep Neural Networks. 2016, IEEE Transactions on Medical Imaging, 35 (11): 2369–2380. 29 JUSTC Geometry-aware Cytology Classification Yating Li et al

  17. [17]

    Novoa, Justin Ko, Susan M

    Andre Esteva, Brett Kuprel, Roberto A. Novoa, Justin Ko, Susan M. Swetter, Helen M. Blau, and Sebastian Thrun. Dermatologist-level classification of skin cancer with deep neural networks. 2017, Nature, 542 (7639): 115–118

  18. [18]

    Lungren, and Andrew Y

    Pranav Rajpurkar, Jeremy Irvin, Kaylie Zhu, Brandon Y ang, Hershel Mehta, Tony Duan, Daisy Ding, Aarti Bagul, Curtis Langlotz, Katie Shpanskaya, Matthew P . Lungren, and Andrew Y . Ng. CheXNet: Radiologist-Level Pneu- monia Detection on Chest X-Rays with Deep Learning, 2017. arXiv preprint arXiv:1711.05225

  19. [19]

    Weinberger

    Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q. Weinberger. Densely connected convolutional networks. In: Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition , 4700–4708, 2017

  20. [20]

    EfficientNet: Rethinking model scaling for convo- lutional neural networks

    Mingxing Tan and Quoc Le. EfficientNet: Rethinking model scaling for convo- lutional neural networks. In: International Conference on Machine Learning , 6105–6114. PMLR, 2019

  21. [21]

    Sahu and Ramgopal Kashyap

    Hemlata P . Sahu and Ramgopal Kashyap. Fine_Denseiganet: Automatic Medi- cal Image Classification in Chest CT Scan Using Hybrid Deep Learning Frame- work. 2025, International Journal of Image and Graphics , 25 (01): 2550004

  22. [22]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention Is All Y ou Need. In: Advances in Neural Information Processing Systems, volume 30, 2017

  23. [23]

    Training data-efficient image transformers and distillation through attention

    Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers and distillation through attention. In: International Conference on Machine Learn- ing, 10347–10357. PMLR, 2021

  24. [24]

    DynamicViT: Efficient vision transformers with dynamic token sparsifi- cation

    Y ongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. DynamicViT: Efficient vision transformers with dynamic token sparsifi- cation. 2021, Advances in Neural Information Processing Systems , 34: 13937– 13949

  25. [25]

    TransMed: Transformers advance multi- 30 JUSTC Geometry-aware Cytology Classification Yating Li et al

    Yin Dai, Yifan Gao, and Fayu Liu. TransMed: Transformers advance multi- 30 JUSTC Geometry-aware Cytology Classification Yating Li et al . modal medical image classification. 2021, Diagnostics, 11 (8): 1384

  26. [26]

    Lesion-aware transformers for diabetic retinopathy grading

    Rui Sun, Yihao Li, Tianzhu Zhang, Zhendong Mao, Feng Wu, and Y ongdong Zhang. Lesion-aware transformers for diabetic retinopathy grading. In:Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10933–10942, 2021

  27. [27]

    Swin Transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 9992–10002, 2021

  28. [28]

    Roth, and Daguang Xu

    Ali Hatamizadeh, Vishwesh Nath, Yucheng Tang, Dong Y ang, Holger R. Roth, and Daguang Xu. Swin UNETR: Swin transformers for semantic segmentation of brain tumors in MRI images. In: International MICCAI Brainlesion Work- shop, 272–284. Springer, 2022

  29. [29]

    Contour-Augmented Concept Prediction Network for Image Captioning , 180–

    Ting Wang, Weidong Chen, Jingyu Li, Yixing Peng, and Zhendong Mao. Contour-Augmented Concept Prediction Network for Image Captioning , 180–

  30. [30]

    Springer Nature Switzerland, 2023

  31. [31]

    Prompting Few- shot Multi-hop Question Generation via Comprehending T ype-aware Semantics

    Zefeng Lin, Weidong Chen, Y an Song, and Y ongdong Zhang. Prompting Few- shot Multi-hop Question Generation via Comprehending T ype-aware Semantics. In: Findings of the Association for Computational Linguistics: NAACL 2024 , 3730–3740. Association for Computational Linguistics, 2024

  32. [32]

    Text Style Transfer with Contrastive Transfer Pattern Mining

    Jingxuan Han, Quan Wang, Licheng Zhang, Weidong Chen, Y an Song, and Zhendong Mao. Text Style Transfer with Contrastive Transfer Pattern Mining. In: Proceedings of the 61st Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers) , 7914–7927. Association for Com- putational Linguistics, 2023

  33. [33]

    End-to-end Aspect- based Sentiment Analysis with Combinatory Categorial Grammar

    Yuanhe Tian, Weidong Chen, Bo Hu, Y an Song, and Fei Xia. End-to-end Aspect- based Sentiment Analysis with Combinatory Categorial Grammar. In: Findings of the Association for Computational Linguistics: ACL 2023, 13597–13609. As- sociation for Computational Linguistics, 2023. 31 JUSTC Geometry-aware Cytology Classification Yating Li et al

  34. [34]

    Cascade Cross-modal Attention Network for Video Actor and Action Segmentation from a Sentence

    Weidong Chen, Guorong Li, Xinfeng Zhang, Hongyang Yu, Shuhui Wang, and Qingming Huang. Cascade Cross-modal Attention Network for Video Actor and Action Segmentation from a Sentence. In: Proceedings of the 29th ACM International Conference on Multimedia, MM ’21, 4053–4062. ACM, 2021

  35. [35]

    Multi-Attention Network for Com- pressed Video Referring Object Segmentation

    Weidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han, Shuhui Wang, Laiyun Qing, Qingming Huang, and Guorong Li. Multi-Attention Network for Com- pressed Video Referring Object Segmentation. In: Proceedings of the 30th ACM International Conference on Multimedia, MM ’22, 4416–4425. ACM, 2022

  36. [36]

    Weakly Supervised Text-based Actor-Action Video Segmentation by Clip-level Multi-instance Learning

    Weidong Chen, Guorong Li, Xinfeng Zhang, Shuhui Wang, Liang Li, and Qing- ming Huang. Weakly Supervised Text-based Actor-Action Video Segmentation by Clip-level Multi-instance Learning. 2023, ACM Transactions on Multimedia Computing, Communications, and Applications , 19 (1): 1–22

  37. [37]

    Towards Efficient Partially Relevant Video Retrieval With Active Moment Discovering

    Peipei Song, Long Zhang, Long Lan, Weidong Chen, Dan Guo, Xun Y ang, and Meng Wang. Towards Efficient Partially Relevant Video Retrieval With Active Moment Discovering. 2025, IEEE Transactions on Multimedia, 27: 6740–6751

  38. [38]

    Dual- path Collaborative Generation Network for Emotional Video Captioning

    Cheng Y e, Weidong Chen, Jingyu Li, Lei Zhang, and Zhendong Mao. Dual- path Collaborative Generation Network for Emotional Video Captioning. In: Proceedings of the 32nd ACM International Conference on Multimedia , MM ’24, 496–505. ACM, 2024

  39. [39]

    Improving Video Summarization by Exploring the Coherence Between Corresponding Captions

    Cheng Y e, Weidong Chen, Bo Hu, Lei Zhang, Y ongdong Zhang, and Zhendong Mao. Improving Video Summarization by Exploring the Coherence Between Corresponding Captions. 2025, IEEE Transactions on Image Processing , 34: 5369–5384

  40. [40]

    Multi-round Mutual Emotion-Cause Pair Extraction for Emotion- Attributed Video Captioning

    Cheng Y e, Weidong Chen, Peipei Song, Xinyan Liu, Lei Zhang, and Zhen- dong Mao. Multi-round Mutual Emotion-Cause Pair Extraction for Emotion- Attributed Video Captioning. In: Proceedings of the 33rd ACM International Conference on Multimedia, MM ’25, 3320–3329. ACM, 2025

  41. [41]

    Subjective-Objective Emotion-Correlated Generation Network for 32 JUSTC Geometry-aware Cytology Classification Yating Li et al

    Weidong Chen, Cheng Y e, Peipei Song, Lei Zhang, Y ongdong Zhang, and Zhen- dong Mao. Subjective-Objective Emotion-Correlated Generation Network for 32 JUSTC Geometry-aware Cytology Classification Yating Li et al . Subjective Video Captioning. 2025, IEEE Transactions on Image Processing , 35: 540–555

  42. [42]

    Stimuli-Aware Emotion Adaptor for Enhancing LLM in Affective Explanation Captioning

    Zhiyan Zhang, Peipei Song, Jinpeng Hu, Weidong Chen, Lin Ni, and Xun Y ang. Stimuli-Aware Emotion Adaptor for Enhancing LLM in Affective Explanation Captioning. In: ICASSP 2026 - 2026 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP) , 10662–10666. IEEE, 2026

  43. [43]

    Bridging Subjectivity in Affective Explanation Captioning via Consensus-Prompted Emotion Reasoning

    Peipei Song, Zhiyan Zhang, Weidong Chen, Jinpeng Hu, Xun Y ang, and Xi- aojun Chang. Bridging Subjectivity in Affective Explanation Captioning via Consensus-Prompted Emotion Reasoning. 2026, IEEE Transactions on Image Processing, 35: 6760–6774

  44. [44]

    A Multi-Agent Framework with Structured Reasoning and Reflective Re- finement for Multimodal Empathetic Response Generation,2026

    Liping Wang, Cheng Y e, Weidong Chen, Peipei Song, Bo Hu, and Zhendong Mao. A Multi-Agent Framework with Structured Reasoning and Reflective Re- finement for Multimodal Empathetic Response Generation,2026. arXiv preprint arXiv:2604.18988

  45. [45]

    FACE-net: Factual Calibration and Emo- tion Augmentation for Retrieval-enhanced Emotional Video Captioning, 2026

    Weidong Chen, Cheng Y e, Zhendong Mao, Peipei Song, Xinyan Liu, Lei Zhang, Xiaojun Chang, and Y ongdong Zhang. FACE-net: Factual Calibration and Emo- tion Augmentation for Retrieval-enhanced Emotional Video Captioning, 2026. arXiv preprint arXiv:2603.17455

  46. [46]

    Towards Accurate Emotion-Attributed Video Caption- ing via Fine-grained Emotion-Cause Pair Extraction, 2026

    Weidong Chen, Cheng Y e, Zhendong Mao, Liping Wang, Xinyan Liu, and Y ongdong Zhang. Towards Accurate Emotion-Attributed Video Caption- ing via Fine-grained Emotion-Cause Pair Extraction, 2026. arXiv preprint arXiv:2606.08566

  47. [47]

    Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning, 2026

    Zihan Meng, Dexiang Hong, Weidong Chen, Ziyu Zhou, Bo Hu, and Zhendong Mao. Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning, 2026. arXiv preprint arXiv:2606.10533

  48. [48]

    Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly Detection

    Xiaoyu Huang, Weidong Chen, Bo Hu, and Zhendong Mao. Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly Detection. 2025, Proceedings of the AAAI Conference on Artificial Intelligence, 39 (16): 17476–17484. 33 JUSTC Geometry-aware Cytology Classification Yating Li et al

  49. [49]

    CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design

    Hui Zhang, Dexiang Hong, Maoke Y ang, Yutao Cheng, Zhao Zhang, Weidong Chen, Jie Shao, Xinglong Wu, Zuxuan Wu, and Yu-Gang Jiang. CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design. In: International Conference on Learning Representations , 2026

  50. [50]

    CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation, 2025

    Zhao Zhang, Yutao Cheng, Dexiang Hong, Maoke Y ang, Gonglei Shi, Lei Ma, Hui Zhang, Jie Shao, and Xinglong Wu. CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation, 2025. arXiv preprint arXiv:2506.10890

  51. [51]

    CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers, 2026

    Weidong Chen, Dexiang Hong, Zhendong Mao, Yutao Cheng, Xinyan Liu, Lei Zhang, and Y ongdong Zhang. CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers, 2026. arXiv preprint arXiv:2604.19632

  52. [52]

    Rethinking Pseudo Word Learning in Zero-Shot Composed Image Retrieval: From an Object-Aware Perspective

    Zhe Li, Lei Zhang, Kun Zhang, Weidong Chen, Y ongdong Zhang, and Zhen- dong Mao. Rethinking Pseudo Word Learning in Zero-Shot Composed Image Retrieval: From an Object-Aware Perspective. In: Proceedings of the 48th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’25, 833–843. ACM, 2025

  53. [53]

    Combatting Data Imbalance and Noise in Micro-Action Recognition

    Chuang Wang, Weidong Chen, Xu Cui, Yiming Zhao, Zhaobo Qi, Pengqi Huang, Xinyan Liu, and Weigang Zhang. Combatting Data Imbalance and Noise in Micro-Action Recognition. In: Proceedings of the 33rd ACM International Conference on Multimedia, MM ’25, 14229–14235. ACM, 2025

  54. [54]

    Query-based Collaborative Multimodal Token Pruning for Audio-Visual Question Answering

    Xilin Qin, Dexiang Hong, Weidong Chen, Cheng Y e, Xinyan Liu, Peipei Song, and Lei Zhang. Query-based Collaborative Multimodal Token Pruning for Audio-Visual Question Answering. In: 2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics (AIHCIR), 1–

  55. [55]

    Hierarchical Knowledge Distillation for Cross-Lingual Stance Detec- tion

    Qiuli Zhou, Jingyuan Y ao, Shengeng Tang, Weidong Chen, Lechao Cheng, and Jun Tang. Hierarchical Knowledge Distillation for Cross-Lingual Stance Detec- tion. In: 2025 4th International Conference on Artificial Intelligence, Human- Computer Interaction and Robotics (AIHCIR) , 1–5. IEEE, 2025. 34 JUSTC Geometry-aware Cytology Classification Yating Li et al

  56. [56]

    EmoVerse: A MLLMs-Driven Emotion Represen- tation Dataset for Interpretable Visual Emotion Analysis, 2025

    Yijie Guo, Dexiang Hong, Weidong Chen, Zihan She, Cheng Y e, Xiaojun Chang, and Zhendong Mao. EmoVerse: A MLLMs-Driven Emotion Represen- tation Dataset for Interpretable Visual Emotion Analysis, 2025. arXiv preprint arXiv:2511.12554

  57. [57]

    Matching Street View and Satellite Images via Drone Imagery and Semantic Descriptions

    Xinyan Liu, Weidong Chen, Zhaobo Qi, Beichen Zhang, and Weigang Zhang. Matching Street View and Satellite Images via Drone Imagery and Semantic Descriptions. In: Proceedings of the 3rd International Workshop on UAV s in Multimedia: Capturing the World from a New Perspective , 4–9. ACM, 2025

  58. [58]

    Difference-Aware Iterative Reasoning Network for Key Relation Detection

    Bowen Zhao, Weidong Chen, Bo Hu, Hongtao Xie, and Zhendong Mao. Difference-Aware Iterative Reasoning Network for Key Relation Detection. In: 2023 IEEE International Conference on Multimedia and Expo (ICME) , 276–

  59. [59]

    Sentiment- Oriented Transformer-Based Variational Autoencoder Network for Live Video Commenting

    Fengyi Fu, Shancheng Fang, Weidong Chen, and Zhendong Mao. Sentiment- Oriented Transformer-Based Variational Autoencoder Network for Live Video Commenting. 2024, ACM Transactions on Multimedia Computing, Communi- cations, and Applications , 20 (4): 1–24

  60. [60]

    CMT: Convolutional neural networks meet vision transformers

    Jianyuan Guo, Kai Han, Han Wu, Y ehui Tang, Xinghao Chen, Yunhe Wang, and Chang Xu. CMT: Convolutional neural networks meet vision transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12165–12175, 2022

  61. [61]

    PVT v2: Improved baselines with pyramid vision transformer

    Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. PVT v2: Improved baselines with pyramid vision transformer. 2022, Computational Visual Media, 8 (3): 415–424

  62. [62]

    A ConvNet for the 2020s

    Zhuang Liu, Hanzi Mao, Chao- Yuan Wu, Christoph Feichtenhofer, Trevor Dar- rell, and Saining Xie. A ConvNet for the 2020s. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 11966– 11976, 2022

  63. [63]

    Smith, and Mike Lewis

    Ofir Press, Noah A. Smith, and Mike Lewis. Train short, test long: Atten- tion with linear biases enables input length extrapolation, 2021. arXiv preprint 35 JUSTC Geometry-aware Cytology Classification Yating Li et al . arXiv:2108.12409

  64. [64]

    Self-attention with relative position representations, 2018

    Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations, 2018. arXiv preprint arXiv:1803.02155

  65. [65]

    Ultra-high resolution segmentation via boundary-enhanced patch-merging trans- former

    Haopeng Sun, Yingwei Zhang, Lumin Xu, Sheng Jin, and Yiqiang Chen. Ultra-high resolution segmentation via boundary-enhanced patch-merging trans- former. In: Proceedings of the AAAI Conference on Artificial Intelligence , vol- ume 39, 7087–7095, 2025

  66. [66]

    Yuille, and Yuyin Zhou

    Jieneng Chen, Y ongyi Lu, Qihang Yu, Xiangde Luo, Ehsan Adeli, Y an Wang, Le Lu, Alan L. Yuille, and Yuyin Zhou. TransUNet: Transformers make strong encoders for medical image segmentation, 2021. arXiv preprint arXiv:2102.04306

  67. [67]

    CoTr: Efficiently bridging CNN and Transformer for 3D medical image segmentation

    Yutong Xie, Jianpeng Zhang, Chunhua Shen, and Y ong Xia. CoTr: Efficiently bridging CNN and Transformer for 3D medical image segmentation. In: In- ternational Conference on Medical Image Computing and Computer-Assisted Intervention, 171–180. Springer, 2021

  68. [68]

    Mahanta, Himakshi Borah, and Chandana Ray Das

    Elima Hussain, Lipi B. Mahanta, Himakshi Borah, and Chandana Ray Das. Liq- uid based-cytology Pap smear dataset for automated multi-class diagnosis of pre-cancerous and cervical cancer lesions. 2020, Data in Brief , 30: 105589

  69. [69]

    Plissiti, P

    Marina E. Plissiti, P . Dimitrakopoulos, G. Sfikas, Christophoros Nikou, O. Krikoni, and A. Charchanti. SIPaKMeD: A New Dataset for Feature and Image Based Classification of Normal and Pathological Cervical Cells in Pap Smear Images. In: 2018 25th IEEE International Conference on Image Pro- cessing, 3144–3148, 2018

  70. [70]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, and Jack Clark. Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning , 8748–8763. PMLR, 2021

  71. [71]

    Question-Aware Gaussian Experts for Audio-Visual Question 36 JUSTC Geometry-aware Cytology Classification Yating Li et al

    Hongyeob Kim, Inyoung Jung, Dayoon Suh, Y oujia Zhang, Sangmin Lee, and Sungeun Hong. Question-Aware Gaussian Experts for Audio-Visual Question 36 JUSTC Geometry-aware Cytology Classification Yating Li et al . Answering. In: Proceedings of the Computer Vision and Pattern Recognition Conference, 13681–13690, 2025

  72. [72]

    DFormerV2: Geometry self-attention for RGBD semantic segmentation

    Bo-Wen Yin, Jiao-Long Cao, Ming-Ming Cheng, and Qibin Hou. DFormerV2: Geometry self-attention for RGBD semantic segmentation. In: Proceedings of the Computer Vision and Pattern Recognition Conference, 19345–19355, 2025

  73. [73]

    Automated cervical cancer cell diagnosis via grid search-optimized multi-CNN ensemble networks

    Omair Bilal, Arash Hekmat, and Saif Ur Rehman Khan. Automated cervical cancer cell diagnosis via grid search-optimized multi-CNN ensemble networks. 2025, Network Modeling Analysis in Health Informatics and Bioinformatics , 14 (1): 67

  74. [74]

    The utilization of padding scheme on convolutional neural network for cervical cell images classification

    Toto Haryanto, Imas Sukaesih Sitanggang, Muhammad Ashyar Agmalaro, and Riries Rulaningtyas. The utilization of padding scheme on convolutional neural network for cervical cell images classification. In: 2020 International Confer- ence on Computer Engineering, Network, and Intelligent Multimedia , 34–38. IEEE, 2020

  75. [75]

    PathoCoder: Rethinking the Flaws of Patch-Based Learning for Multi-Class Classification in Computational Pathology

    Ferdaous Idlahcen, Pierjos Francis Colere Mboukou, Ali Idri, and Hicham El At- tar. PathoCoder: Rethinking the Flaws of Patch-Based Learning for Multi-Class Classification in Computational Pathology. 2025, Microscopy Research and Technique, 88 (6): 1712–1726

  76. [76]

    Enhancing cervical cancer diag- nosis: Integrated attention-transformer system with weakly supervised learning

    Ashfaque Khowaja, Beiji Zou, and Xiaoyan Kui. Enhancing cervical cancer diag- nosis: Integrated attention-transformer system with weakly supervised learning. 2024, Image and Vision Computing, 149: 105193

  77. [77]

    Perbandingan Arsitektur ResNet50 dan ResNet101 dalam Klasifikasi Kanker Serviks pada Citra Pap Smear

    Za’imatun Niswati, Rahayuning Hardatin, Meia Noer Muslimah, and Siti Nur Hasanah. Perbandingan Arsitektur ResNet50 dan ResNet101 dalam Klasifikasi Kanker Serviks pada Citra Pap Smear. 2021, Faktor Exacta, 14 (3): 160–167

  78. [78]

    Privacy preserved cervical can- cer detection using convolutional neural networks applied to pap smear im- ages

    Shtwai Alsubai, Abdullah Alqahtani, Mohemmed Sha, Ahmad Almadhor, Sidra Abbas, Huma Mughal, and Michal Gregus. Privacy preserved cervical can- cer detection using convolutional neural networks applied to pap smear im- ages. 2023, Computational and Mathematical Methods in Medicine , 2023 (1): 9676206. 37 JUSTC Geometry-aware Cytology Classification Yating Li et al

  79. [79]

    Differential evolution optimization based ensemble framework for accurate cervical cancer diagnosis

    Omair Bilal, Sohaib Asif, Ming Zhao, Y angfan Li, Fengxiao Tang, and Yusen Zhu. Differential evolution optimization based ensemble framework for accurate cervical cancer diagnosis. 2024, Applied Soft Computing , 167: 112366

  80. [80]

    Deep Learning Enabled Segmentation, Classification and Risk Assess- ment of Cervical Cancer, 2025

    Abdul Samad Shaik, Shashaank Mattur Aswatha, and Rahul Jashvantbhai Pandya. Deep Learning Enabled Segmentation, Classification and Risk Assess- ment of Cervical Cancer, 2025. arXiv preprint arXiv:2505.15505

Showing first 80 references.