REVIEW 3 major objections 4 minor 1 cited by
Axis-wise Gaussian priors from vision-language features, injected into axial attention, raise cervical cytology classification to 99.48% and 96.08% accuracy on two public Pap-smear datasets.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 12:57 UTC pith:TXXHDIZI
load-bearing objection Clean engineering synthesis of CLIP + Gaussian MoE axis priors + axial attention that posts strong numbers on two small cytology sets; ablations help, but single-split uncertainty leaves the geometry claim only partially secured. the 3 major comments →
Geometry-aware Gaussian Prior and Axial Attention for Cervical Cytology Image Classification
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The authors show that structural regularities of cervical cells—nuclear alignment and spatial organisation—can be captured by Top-K Gaussian experts conditioned only on a pretrained CLIP global token, and that multiplying those axis-wise priors into axial self-attention produces a classifier that outperforms recent CNN and hybrid baselines on two standard cytology benchmarks while highlighting diagnostically relevant regions.
What carries the argument
Geometry-aware Gaussian Prior Module: a router selects Top-K Gaussian experts from a CLIP-derived global descriptor, forms directional weighting maps, converts them into pairwise cosine-similarity prior matrices, and multiplies those matrices into horizontal and vertical axial attention scores.
Load-bearing premise
That a handful of Gaussians steered by a single CLIP global token actually encode the clinically important spatial patterns of cells, rather than merely memorising the spatial statistics of the two training sets.
What would settle it
Train the identical architecture on a third, multi-centre cytology collection whose nuclear spacing and clustering statistics differ from Mendeley and SIPaKMeD; if accuracy collapses or the visualised priors stop highlighting nuclei, the structural-prior claim is falsified.
If this is right
- Screening pipelines can replace pure CNN or plain Transformer backbones with a CLIP-plus-Gaussian-axial block and expect higher accuracy and more balanced precision-recall on liquid-based and single-cell Pap images.
- Visualisation of the learned Gaussian maps can serve as a lightweight, built-in explanation layer that points a cytologist to the same regions the model used for its decision.
- The same axis-wise Gaussian-expert construction can be dropped into other high-resolution medical tasks that need directional structural bias without quadratic full self-attention cost.
- Ablation results imply that three Transformer stages and a moderate Top-K (6–8) already give the best accuracy–stability trade-off, guiding practical hyper-parameter choices.
Where Pith is reading between the lines
- If the Gaussians truly encode nuclear geometry, the same module should transfer to other cytology domains (urine, breast fine-needle aspirates) with only light fine-tuning of the router.
- The multiplicative prior modulation may be replaceable by an additive relative bias; a controlled swap would test whether the Gaussian shape itself, rather than any soft spatial prior, is necessary.
- Because the method relies on CLIP’s semantic abstraction, any future vision-language backbone that better separates cellular morphology should raise the ceiling further without redesigning the attention block.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a geometry-aware framework for cervical cytology image classification that extracts features with a pretrained CLIP-ViT backbone, generates axis-wise Gaussian priors via a Top-K mixture of Gaussian experts conditioned on the global semantic token (Eqs. 3–11, §3.2), and multiplicatively injects those priors into axial self-attention (Eqs. 18–19, §3.3) before a lightweight classification head. On the Mendeley LBC and SIPaKMeD public datasets the method reports 99.48 % accuracy / 99.27 % F1 and 96.08 % accuracy / 95.40 % F1, respectively, outperforming a range of recent CNN and hybrid baselines (Table 2). Ablations isolate Transformer depth, the geometric-prior module, expert count, and the two axial directions (Tables 3–6), and qualitative visualizations of priors and confidence scores are supplied.
Significance. If the claimed gains are robust, the work supplies a concrete, domain-motivated inductive bias—direction-sensitive Gaussian structural priors—for a clinically important screening task that still relies heavily on manual microscopy. The architecture is clearly specified, the ablations systematically remove each proposed component, and the interpretability narrative (priors highlighting cellular regions) is a useful addition for decision-support systems. The contribution is incremental rather than foundational, but it is a legitimate engineering advance for automated Pap-smear analysis provided the numerical superiority can be shown to be stable.
major comments (3)
- Tables 2–6 and §4–5 report all metrics on a single fixed 80/20 split of two small collections (Mendeley LBC ≈ 963 images, SIPaKMeD ≈ 4 049 images) with no multi-seed averages, k-fold cross-validation, or confidence intervals. Absolute margins over already-strong baselines are 0.1–1.2 %. Under these conditions the headline numbers (99.48 % / 96.08 %) cannot be distinguished from split luck or from the CLIP + DFormer backbone alone; the causal link between the geometry-aware construction and the claimed gains remains unsecured.
- The ablations never isolate the Gaussian modulation itself. Table 4 removes the entire geometric-prior module and Table 6 removes axial axes, but there is no control that retains axial attention while replacing the multiplicative Gaussian terms (Eqs. 18–19) by identity or by a non-Gaussian positional bias. Consequently it is impossible to attribute the observed lifts specifically to the Top-K Gaussian experts (Eqs. 3–11) rather than to axial attention or the pretrained backbone.
- §3.2 and the interpretability claim rest on the premise that a single CLIP global token supplies sufficient semantic guidance to route and parameterize Gaussians that capture clinically relevant nuclear alignment and cellular organization. No quantitative localization metric (e.g., overlap with nuclei or expert-annotated ROIs) is reported; the visual analysis is purely qualitative. Without such evidence the clinical-utility narrative remains an untested assumption.
minor comments (4)
- Related-work §2.2 is dominated by self-citations on captioning, video understanding and multimodal generation that are only loosely connected to cytology classification; a tighter focus on medical image classification and geometric priors would improve relevance.
- Notation for the prior matrices is overloaded: G_W / G_H first denote 1-D Gaussian distributions (Eq. 11) and later pairwise cosine-similarity matrices (Eqs. 13–14). Distinct symbols would avoid confusion.
- Implementation details (§4.3) list many free hyperparameters (Top-K, E, γ, scale factor 9, warm-up lengths, peak LRs) without sensitivity analysis beyond the Top-K table; a short discussion of robustness would strengthen reproducibility.
- Figure 1 and the graphical abstract are described but their visual content is not fully self-explanatory in the text; captions should explicitly state what structural cues the priors are intended to highlight.
Circularity Check
No derivation circularity: accuracy claims are empirical measurements on held-out splits, not quantities defined by the Gaussian priors or axial modulation.
specific steps
-
self citation load bearing
[§2.2 Multimodal Vision and Structure-aware Representation (refs [6]–[58])]
"These studies motivate our use of pretrained semantic features as a high-level guide for learning structure-sensitive cytology representations. ... Our method adapts this insight to cervical cytology by using Gaussian experts to inject direction-sensitive geometric priors into axial attention."
A large fraction of the multimodal/structure-aware citations are by overlapping authors (Chen, Mao, Ye, etc.). They supply motivational framing rather than a uniqueness theorem or a numerical result that the present accuracy claims depend on. The cytology numbers themselves are measured on Mendeley LBC and SIPaKMeD, so the self-citations are not load-bearing for the strongest claim; the step is therefore only minor.
full rationale
The paper proposes an architecture (CLIP backbone + Gaussian expert priors + axial attention) and reports classification metrics on public cytology datasets. The Gaussian parameters (means, widths, Top-K routing) are learned end-to-end from training data and then evaluated on a fixed 80/20 test split; nothing in Eqs. 3–24 defines accuracy, F1, or the reported margins in terms of a fitted quantity that is later re-presented as a prediction. Ablations (Tables 3–6) remove modules and re-measure the same external metrics, which is ordinary empirical validation rather than self-definitional reduction. Related-work citations include many papers from the same group, but those works address captioning, report generation, and multimodal tasks; none of them supply the cytology accuracy numbers or a uniqueness theorem that forces the present design. The design choices (Gaussian experts, axial decomposition) are inductive biases, not tautologies. Consequently the central claim does not reduce by construction to its inputs. Minor self-citation density exists but is not load-bearing for the reported results, yielding a score of 1.
Axiom & Free-Parameter Ledger
free parameters (5)
- Top-K experts =
8 / 6
- Number of Gaussian experts E =
10
- Gaussian decay coefficient gamma
- Peak learning rates and warm-up epochs =
5e-6 / 1e-5
- Gaussian scale normalization factor =
9
axioms (3)
- standard math Standard scaled-dot-product attention and Gaussian density formulas are valid feature aggregators.
- domain assumption Cervical cells exhibit diagnostically useful structural regularities (nuclear alignment, spatial organization) that can be captured by axis-aligned Gaussian maps.
- ad hoc to paper A single CLIP global token supplies sufficient semantic guidance to route and parameterize the Gaussian experts for cytology images.
invented entities (2)
-
Geometry-aware Gaussian Prior Module (Gaussian experts + router)
no independent evidence
-
Gaussian-enhanced Axial Self-Attention
no independent evidence
read the original abstract
Accurate cervical cytology image classification is a key component of automated cervical cancer screening, where reliable recognition of normal, precancerous, and cancer-associated cellular patterns from Pap smear images can improve screening efficiency and diagnostic consistency. However, this task remains challenging because cervical cells exhibit complex morphology, subtle intra-class variations, and strong inter-class similarities. Existing convolution-based models capture local texture well but have limited ability to model long-range relationships, whereas attention-based models provide broader context but often lack explicit structural guidance. To address these limitations, we propose a geometry-aware classification framework for cervical cancer screening-oriented cytology image analysis, incorporating semantic abstraction and structural priors learned from pre-trained vision-language features. The method uses Gaussian expert modules to generate axis-wise priors from global semantic information, capturing structural regularities such as nuclear alignment and cellular spatial organization. These priors are embedded into an axial self-attention module to modulate similarity computation along horizontal and vertical directions, improving long-range dependency modeling and structure-sensitive feature interaction. Experiments on the Mendeley liquid-based cytology and SIPaKMeD datasets show that the proposed method achieves 99.48% accuracy on the former and 96.08% on the latter, with balanced gains in recall, precision, and overall classification performance. Visual analysis further shows that the learned priors highlight diagnostically relevant cellular regions, demonstrating the potential of the proposed framework as a screening-oriented decision-support tool for cervical cytology.
Figures
Forward citations
Cited by 1 Pith paper
-
HTT-Net: Hierarchical Text-guided Transition Modeling for Surgical Video Phase Recognition
A hierarchical text-guided network that constructs and calibrates phase segments improves surgical phase recognition, setting a high Jaccard on Cholec80 and reporting large gains on a private LCRS-100 benchmark.
Reference graph
Works this paper leans on
-
[1]
Global strategy to accelerate the elimination of cervical cancer as a public health problem
World Health Organization. Global strategy to accelerate the elimination of cervical cancer as a public health problem . World Health Organization, 2020
2020
-
[2]
Summers, Shaoxiong Liu, and Jianhua Y ao
Ling Zhang, Le Lu, Isabella Nogues, Ronald M. Summers, Shaoxiong Liu, and Jianhua Y ao. DeepPap: Deep Convolutional Networks for Cervical Cell Clas- sification. 2017, IEEE Journal of Biomedical and Health Informatics , 21 (6): 1633–1643
2017
-
[3]
Cervical cytology classification using PCA and GWO enhanced deep features selection
Hritam Basak, Rohit Kundu, Sukanta Chakraborty, and Nibaran Das. Cervical cytology classification using PCA and GWO enhanced deep features selection. 2021, SN Computer Science, 2 (5): 369
2021
-
[4]
An image is worth 16x16 words: Transformers for image recognition at scale, 2020
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale, 2020. arXiv preprint arXiv:2010.11929
Pith/arXiv arXiv 2020
-
[5]
TransMIL: Transformer based correlated multiple instance learning for whole slide image classification
Zhuchen Shao, Hao Bian, Y ang Chen, Yifeng Wang, Jian Zhang, and Xiangyang Ji. TransMIL: Transformer based correlated multiple instance learning for whole slide image classification. 2021, Advances in Neural Information Processing Systems, 34: 2136–2147
2021
-
[6]
Bootstrapping Large Language Models for Radiology Report Generation
Chang Liu, Yuanhe Tian, Weidong Chen, Y an Song, and Y ongdong Zhang. Bootstrapping Large Language Models for Radiology Report Generation. 2024, Proceedings of the AAAI Conference on Artificial Intelligence , 38 (17): 18635– 18643
2024
-
[7]
Im- proving radiology report generation with multi-grained abnormality prediction
Yuda Jin, Weidong Chen, Yuanhe Tian, Y an Song, and Chenggang Y an. Im- proving radiology report generation with multi-grained abnormality prediction. 2024, Neurocomputing, 600: 128122
2024
-
[8]
dong Mao
Yuda Jin, Weidong Chen, Yuanhe Tian, Y an Song, Chenggang Y an, and Zhen- 28 JUSTC Geometry-aware Cytology Classification Yating Li et al . dong Mao. Improving Radiology Report Generation with D 2-Net: When Diffu- sion Meets Discriminator. In: ICASSP 2024 - 2024 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP) , 2215–221...
2024
-
[9]
Im- proving Image Captioning via Predicting Structured Concepts
Ting Wang, Weidong Chen, Yuanhe Tian, Y an Song, and Zhendong Mao. Im- proving Image Captioning via Predicting Structured Concepts. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 360–370. Association for Computational Linguistics, 2023
2023
-
[10]
Ex- ploring Visual Relationships via Transformer-based Graphs for Enhanced Image Captioning
Jingyu Li, Zhendong Mao, Hao Li, Weidong Chen, and Y ongdong Zhang. Ex- ploring Visual Relationships via Transformer-based Graphs for Enhanced Image Captioning. 2024, ACM Transactions on Multimedia Computing, Communica- tions, and Applications, 20 (5): 1–23
2024
-
[11]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classifica- tion with deep convolutional neural networks. 2012, Advances in Neural Infor- mation Processing Systems, 25
2012
-
[12]
Min Lin, Qiang Chen, and Shuicheng Y an. Network In Network, 2013. arXiv preprint arXiv:1312.4400
Pith/arXiv arXiv 2013
-
[13]
Going Deeper With Convolutions
Christian Szegedy, Wei Liu, Y angqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabi- novich. Going Deeper With Convolutions. In: Proceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition , 1–9, 2015
2015
-
[14]
Very deep convolutional networks for large-scale image recognition, 2014
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition, 2014. arXiv preprint arXiv:1409.1556
Pith/arXiv arXiv 2014
-
[15]
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–778, 2016
2016
-
[16]
Segmenting Retinal Blood Vessels With Deep Neural Networks
Paweł Liskowski and Krzysztof Krawiec. Segmenting Retinal Blood Vessels With Deep Neural Networks. 2016, IEEE Transactions on Medical Imaging, 35 (11): 2369–2380. 29 JUSTC Geometry-aware Cytology Classification Yating Li et al
2016
-
[17]
Novoa, Justin Ko, Susan M
Andre Esteva, Brett Kuprel, Roberto A. Novoa, Justin Ko, Susan M. Swetter, Helen M. Blau, and Sebastian Thrun. Dermatologist-level classification of skin cancer with deep neural networks. 2017, Nature, 542 (7639): 115–118
2017
-
[18]
Pranav Rajpurkar, Jeremy Irvin, Kaylie Zhu, Brandon Y ang, Hershel Mehta, Tony Duan, Daisy Ding, Aarti Bagul, Curtis Langlotz, Katie Shpanskaya, Matthew P . Lungren, and Andrew Y . Ng. CheXNet: Radiologist-Level Pneu- monia Detection on Chest X-Rays with Deep Learning, 2017. arXiv preprint arXiv:1711.05225
Pith/arXiv arXiv 2017
-
[19]
Weinberger
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q. Weinberger. Densely connected convolutional networks. In: Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition , 4700–4708, 2017
2017
-
[20]
EfficientNet: Rethinking model scaling for convo- lutional neural networks
Mingxing Tan and Quoc Le. EfficientNet: Rethinking model scaling for convo- lutional neural networks. In: International Conference on Machine Learning , 6105–6114. PMLR, 2019
2019
-
[21]
Sahu and Ramgopal Kashyap
Hemlata P . Sahu and Ramgopal Kashyap. Fine_Denseiganet: Automatic Medi- cal Image Classification in Chest CT Scan Using Hybrid Deep Learning Frame- work. 2025, International Journal of Image and Graphics , 25 (01): 2550004
2025
-
[22]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention Is All Y ou Need. In: Advances in Neural Information Processing Systems, volume 30, 2017
2017
-
[23]
Training data-efficient image transformers and distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers and distillation through attention. In: International Conference on Machine Learn- ing, 10347–10357. PMLR, 2021
2021
-
[24]
DynamicViT: Efficient vision transformers with dynamic token sparsifi- cation
Y ongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. DynamicViT: Efficient vision transformers with dynamic token sparsifi- cation. 2021, Advances in Neural Information Processing Systems , 34: 13937– 13949
2021
-
[25]
TransMed: Transformers advance multi- 30 JUSTC Geometry-aware Cytology Classification Yating Li et al
Yin Dai, Yifan Gao, and Fayu Liu. TransMed: Transformers advance multi- 30 JUSTC Geometry-aware Cytology Classification Yating Li et al . modal medical image classification. 2021, Diagnostics, 11 (8): 1384
2021
-
[26]
Lesion-aware transformers for diabetic retinopathy grading
Rui Sun, Yihao Li, Tianzhu Zhang, Zhendong Mao, Feng Wu, and Y ongdong Zhang. Lesion-aware transformers for diabetic retinopathy grading. In:Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10933–10942, 2021
2021
-
[27]
Swin Transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 9992–10002, 2021
2021
-
[28]
Roth, and Daguang Xu
Ali Hatamizadeh, Vishwesh Nath, Yucheng Tang, Dong Y ang, Holger R. Roth, and Daguang Xu. Swin UNETR: Swin transformers for semantic segmentation of brain tumors in MRI images. In: International MICCAI Brainlesion Work- shop, 272–284. Springer, 2022
2022
-
[29]
Contour-Augmented Concept Prediction Network for Image Captioning , 180–
Ting Wang, Weidong Chen, Jingyu Li, Yixing Peng, and Zhendong Mao. Contour-Augmented Concept Prediction Network for Image Captioning , 180–
-
[30]
Springer Nature Switzerland, 2023
2023
-
[31]
Prompting Few- shot Multi-hop Question Generation via Comprehending T ype-aware Semantics
Zefeng Lin, Weidong Chen, Y an Song, and Y ongdong Zhang. Prompting Few- shot Multi-hop Question Generation via Comprehending T ype-aware Semantics. In: Findings of the Association for Computational Linguistics: NAACL 2024 , 3730–3740. Association for Computational Linguistics, 2024
2024
-
[32]
Text Style Transfer with Contrastive Transfer Pattern Mining
Jingxuan Han, Quan Wang, Licheng Zhang, Weidong Chen, Y an Song, and Zhendong Mao. Text Style Transfer with Contrastive Transfer Pattern Mining. In: Proceedings of the 61st Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers) , 7914–7927. Association for Com- putational Linguistics, 2023
2023
-
[33]
End-to-end Aspect- based Sentiment Analysis with Combinatory Categorial Grammar
Yuanhe Tian, Weidong Chen, Bo Hu, Y an Song, and Fei Xia. End-to-end Aspect- based Sentiment Analysis with Combinatory Categorial Grammar. In: Findings of the Association for Computational Linguistics: ACL 2023, 13597–13609. As- sociation for Computational Linguistics, 2023. 31 JUSTC Geometry-aware Cytology Classification Yating Li et al
2023
-
[34]
Cascade Cross-modal Attention Network for Video Actor and Action Segmentation from a Sentence
Weidong Chen, Guorong Li, Xinfeng Zhang, Hongyang Yu, Shuhui Wang, and Qingming Huang. Cascade Cross-modal Attention Network for Video Actor and Action Segmentation from a Sentence. In: Proceedings of the 29th ACM International Conference on Multimedia, MM ’21, 4053–4062. ACM, 2021
2021
-
[35]
Multi-Attention Network for Com- pressed Video Referring Object Segmentation
Weidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han, Shuhui Wang, Laiyun Qing, Qingming Huang, and Guorong Li. Multi-Attention Network for Com- pressed Video Referring Object Segmentation. In: Proceedings of the 30th ACM International Conference on Multimedia, MM ’22, 4416–4425. ACM, 2022
2022
-
[36]
Weakly Supervised Text-based Actor-Action Video Segmentation by Clip-level Multi-instance Learning
Weidong Chen, Guorong Li, Xinfeng Zhang, Shuhui Wang, Liang Li, and Qing- ming Huang. Weakly Supervised Text-based Actor-Action Video Segmentation by Clip-level Multi-instance Learning. 2023, ACM Transactions on Multimedia Computing, Communications, and Applications , 19 (1): 1–22
2023
-
[37]
Towards Efficient Partially Relevant Video Retrieval With Active Moment Discovering
Peipei Song, Long Zhang, Long Lan, Weidong Chen, Dan Guo, Xun Y ang, and Meng Wang. Towards Efficient Partially Relevant Video Retrieval With Active Moment Discovering. 2025, IEEE Transactions on Multimedia, 27: 6740–6751
2025
-
[38]
Dual- path Collaborative Generation Network for Emotional Video Captioning
Cheng Y e, Weidong Chen, Jingyu Li, Lei Zhang, and Zhendong Mao. Dual- path Collaborative Generation Network for Emotional Video Captioning. In: Proceedings of the 32nd ACM International Conference on Multimedia , MM ’24, 496–505. ACM, 2024
2024
-
[39]
Improving Video Summarization by Exploring the Coherence Between Corresponding Captions
Cheng Y e, Weidong Chen, Bo Hu, Lei Zhang, Y ongdong Zhang, and Zhendong Mao. Improving Video Summarization by Exploring the Coherence Between Corresponding Captions. 2025, IEEE Transactions on Image Processing , 34: 5369–5384
2025
-
[40]
Multi-round Mutual Emotion-Cause Pair Extraction for Emotion- Attributed Video Captioning
Cheng Y e, Weidong Chen, Peipei Song, Xinyan Liu, Lei Zhang, and Zhen- dong Mao. Multi-round Mutual Emotion-Cause Pair Extraction for Emotion- Attributed Video Captioning. In: Proceedings of the 33rd ACM International Conference on Multimedia, MM ’25, 3320–3329. ACM, 2025
2025
-
[41]
Subjective-Objective Emotion-Correlated Generation Network for 32 JUSTC Geometry-aware Cytology Classification Yating Li et al
Weidong Chen, Cheng Y e, Peipei Song, Lei Zhang, Y ongdong Zhang, and Zhen- dong Mao. Subjective-Objective Emotion-Correlated Generation Network for 32 JUSTC Geometry-aware Cytology Classification Yating Li et al . Subjective Video Captioning. 2025, IEEE Transactions on Image Processing , 35: 540–555
2025
-
[42]
Stimuli-Aware Emotion Adaptor for Enhancing LLM in Affective Explanation Captioning
Zhiyan Zhang, Peipei Song, Jinpeng Hu, Weidong Chen, Lin Ni, and Xun Y ang. Stimuli-Aware Emotion Adaptor for Enhancing LLM in Affective Explanation Captioning. In: ICASSP 2026 - 2026 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP) , 10662–10666. IEEE, 2026
2026
-
[43]
Bridging Subjectivity in Affective Explanation Captioning via Consensus-Prompted Emotion Reasoning
Peipei Song, Zhiyan Zhang, Weidong Chen, Jinpeng Hu, Xun Y ang, and Xi- aojun Chang. Bridging Subjectivity in Affective Explanation Captioning via Consensus-Prompted Emotion Reasoning. 2026, IEEE Transactions on Image Processing, 35: 6760–6774
2026
-
[44]
Liping Wang, Cheng Y e, Weidong Chen, Peipei Song, Bo Hu, and Zhendong Mao. A Multi-Agent Framework with Structured Reasoning and Reflective Re- finement for Multimodal Empathetic Response Generation,2026. arXiv preprint arXiv:2604.18988
Pith/arXiv arXiv 2026
-
[45]
Weidong Chen, Cheng Y e, Zhendong Mao, Peipei Song, Xinyan Liu, Lei Zhang, Xiaojun Chang, and Y ongdong Zhang. FACE-net: Factual Calibration and Emo- tion Augmentation for Retrieval-enhanced Emotional Video Captioning, 2026. arXiv preprint arXiv:2603.17455
arXiv 2026
-
[46]
Weidong Chen, Cheng Y e, Zhendong Mao, Liping Wang, Xinyan Liu, and Y ongdong Zhang. Towards Accurate Emotion-Attributed Video Caption- ing via Fine-grained Emotion-Cause Pair Extraction, 2026. arXiv preprint arXiv:2606.08566
Pith/arXiv arXiv 2026
-
[47]
Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning, 2026
Zihan Meng, Dexiang Hong, Weidong Chen, Ziyu Zhou, Bo Hu, and Zhendong Mao. Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning, 2026. arXiv preprint arXiv:2606.10533
Pith/arXiv arXiv 2026
-
[48]
Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly Detection
Xiaoyu Huang, Weidong Chen, Bo Hu, and Zhendong Mao. Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly Detection. 2025, Proceedings of the AAAI Conference on Artificial Intelligence, 39 (16): 17476–17484. 33 JUSTC Geometry-aware Cytology Classification Yating Li et al
2025
-
[49]
CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design
Hui Zhang, Dexiang Hong, Maoke Y ang, Yutao Cheng, Zhao Zhang, Weidong Chen, Jie Shao, Xinglong Wu, Zuxuan Wu, and Yu-Gang Jiang. CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design. In: International Conference on Learning Representations , 2026
2026
-
[50]
CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation, 2025
Zhao Zhang, Yutao Cheng, Dexiang Hong, Maoke Y ang, Gonglei Shi, Lei Ma, Hui Zhang, Jie Shao, and Xinglong Wu. CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation, 2025. arXiv preprint arXiv:2506.10890
Pith/arXiv arXiv 2025
-
[51]
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers, 2026
Weidong Chen, Dexiang Hong, Zhendong Mao, Yutao Cheng, Xinyan Liu, Lei Zhang, and Y ongdong Zhang. CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers, 2026. arXiv preprint arXiv:2604.19632
Pith/arXiv arXiv 2026
-
[52]
Rethinking Pseudo Word Learning in Zero-Shot Composed Image Retrieval: From an Object-Aware Perspective
Zhe Li, Lei Zhang, Kun Zhang, Weidong Chen, Y ongdong Zhang, and Zhen- dong Mao. Rethinking Pseudo Word Learning in Zero-Shot Composed Image Retrieval: From an Object-Aware Perspective. In: Proceedings of the 48th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’25, 833–843. ACM, 2025
2025
-
[53]
Combatting Data Imbalance and Noise in Micro-Action Recognition
Chuang Wang, Weidong Chen, Xu Cui, Yiming Zhao, Zhaobo Qi, Pengqi Huang, Xinyan Liu, and Weigang Zhang. Combatting Data Imbalance and Noise in Micro-Action Recognition. In: Proceedings of the 33rd ACM International Conference on Multimedia, MM ’25, 14229–14235. ACM, 2025
2025
-
[54]
Query-based Collaborative Multimodal Token Pruning for Audio-Visual Question Answering
Xilin Qin, Dexiang Hong, Weidong Chen, Cheng Y e, Xinyan Liu, Peipei Song, and Lei Zhang. Query-based Collaborative Multimodal Token Pruning for Audio-Visual Question Answering. In: 2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics (AIHCIR), 1–
2025
-
[55]
Hierarchical Knowledge Distillation for Cross-Lingual Stance Detec- tion
Qiuli Zhou, Jingyuan Y ao, Shengeng Tang, Weidong Chen, Lechao Cheng, and Jun Tang. Hierarchical Knowledge Distillation for Cross-Lingual Stance Detec- tion. In: 2025 4th International Conference on Artificial Intelligence, Human- Computer Interaction and Robotics (AIHCIR) , 1–5. IEEE, 2025. 34 JUSTC Geometry-aware Cytology Classification Yating Li et al
2025
-
[56]
Yijie Guo, Dexiang Hong, Weidong Chen, Zihan She, Cheng Y e, Xiaojun Chang, and Zhendong Mao. EmoVerse: A MLLMs-Driven Emotion Represen- tation Dataset for Interpretable Visual Emotion Analysis, 2025. arXiv preprint arXiv:2511.12554
Pith/arXiv arXiv 2025
-
[57]
Matching Street View and Satellite Images via Drone Imagery and Semantic Descriptions
Xinyan Liu, Weidong Chen, Zhaobo Qi, Beichen Zhang, and Weigang Zhang. Matching Street View and Satellite Images via Drone Imagery and Semantic Descriptions. In: Proceedings of the 3rd International Workshop on UAV s in Multimedia: Capturing the World from a New Perspective , 4–9. ACM, 2025
2025
-
[58]
Difference-Aware Iterative Reasoning Network for Key Relation Detection
Bowen Zhao, Weidong Chen, Bo Hu, Hongtao Xie, and Zhendong Mao. Difference-Aware Iterative Reasoning Network for Key Relation Detection. In: 2023 IEEE International Conference on Multimedia and Expo (ICME) , 276–
2023
-
[59]
Sentiment- Oriented Transformer-Based Variational Autoencoder Network for Live Video Commenting
Fengyi Fu, Shancheng Fang, Weidong Chen, and Zhendong Mao. Sentiment- Oriented Transformer-Based Variational Autoencoder Network for Live Video Commenting. 2024, ACM Transactions on Multimedia Computing, Communi- cations, and Applications , 20 (4): 1–24
2024
-
[60]
CMT: Convolutional neural networks meet vision transformers
Jianyuan Guo, Kai Han, Han Wu, Y ehui Tang, Xinghao Chen, Yunhe Wang, and Chang Xu. CMT: Convolutional neural networks meet vision transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12165–12175, 2022
2022
-
[61]
PVT v2: Improved baselines with pyramid vision transformer
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. PVT v2: Improved baselines with pyramid vision transformer. 2022, Computational Visual Media, 8 (3): 415–424
2022
-
[62]
A ConvNet for the 2020s
Zhuang Liu, Hanzi Mao, Chao- Yuan Wu, Christoph Feichtenhofer, Trevor Dar- rell, and Saining Xie. A ConvNet for the 2020s. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 11966– 11976, 2022
2022
-
[63]
Ofir Press, Noah A. Smith, and Mike Lewis. Train short, test long: Atten- tion with linear biases enables input length extrapolation, 2021. arXiv preprint 35 JUSTC Geometry-aware Cytology Classification Yating Li et al . arXiv:2108.12409
Pith/arXiv arXiv 2021
-
[64]
Self-attention with relative position representations, 2018
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations, 2018. arXiv preprint arXiv:1803.02155
Pith/arXiv arXiv 2018
-
[65]
Ultra-high resolution segmentation via boundary-enhanced patch-merging trans- former
Haopeng Sun, Yingwei Zhang, Lumin Xu, Sheng Jin, and Yiqiang Chen. Ultra-high resolution segmentation via boundary-enhanced patch-merging trans- former. In: Proceedings of the AAAI Conference on Artificial Intelligence , vol- ume 39, 7087–7095, 2025
2025
-
[66]
Jieneng Chen, Y ongyi Lu, Qihang Yu, Xiangde Luo, Ehsan Adeli, Y an Wang, Le Lu, Alan L. Yuille, and Yuyin Zhou. TransUNet: Transformers make strong encoders for medical image segmentation, 2021. arXiv preprint arXiv:2102.04306
Pith/arXiv arXiv 2021
-
[67]
CoTr: Efficiently bridging CNN and Transformer for 3D medical image segmentation
Yutong Xie, Jianpeng Zhang, Chunhua Shen, and Y ong Xia. CoTr: Efficiently bridging CNN and Transformer for 3D medical image segmentation. In: In- ternational Conference on Medical Image Computing and Computer-Assisted Intervention, 171–180. Springer, 2021
2021
-
[68]
Mahanta, Himakshi Borah, and Chandana Ray Das
Elima Hussain, Lipi B. Mahanta, Himakshi Borah, and Chandana Ray Das. Liq- uid based-cytology Pap smear dataset for automated multi-class diagnosis of pre-cancerous and cervical cancer lesions. 2020, Data in Brief , 30: 105589
2020
-
[69]
Plissiti, P
Marina E. Plissiti, P . Dimitrakopoulos, G. Sfikas, Christophoros Nikou, O. Krikoni, and A. Charchanti. SIPaKMeD: A New Dataset for Feature and Image Based Classification of Normal and Pathological Cervical Cells in Pap Smear Images. In: 2018 25th IEEE International Conference on Image Pro- cessing, 3144–3148, 2018
2018
-
[70]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, and Jack Clark. Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning , 8748–8763. PMLR, 2021
2021
-
[71]
Question-Aware Gaussian Experts for Audio-Visual Question 36 JUSTC Geometry-aware Cytology Classification Yating Li et al
Hongyeob Kim, Inyoung Jung, Dayoon Suh, Y oujia Zhang, Sangmin Lee, and Sungeun Hong. Question-Aware Gaussian Experts for Audio-Visual Question 36 JUSTC Geometry-aware Cytology Classification Yating Li et al . Answering. In: Proceedings of the Computer Vision and Pattern Recognition Conference, 13681–13690, 2025
2025
-
[72]
DFormerV2: Geometry self-attention for RGBD semantic segmentation
Bo-Wen Yin, Jiao-Long Cao, Ming-Ming Cheng, and Qibin Hou. DFormerV2: Geometry self-attention for RGBD semantic segmentation. In: Proceedings of the Computer Vision and Pattern Recognition Conference, 19345–19355, 2025
2025
-
[73]
Automated cervical cancer cell diagnosis via grid search-optimized multi-CNN ensemble networks
Omair Bilal, Arash Hekmat, and Saif Ur Rehman Khan. Automated cervical cancer cell diagnosis via grid search-optimized multi-CNN ensemble networks. 2025, Network Modeling Analysis in Health Informatics and Bioinformatics , 14 (1): 67
2025
-
[74]
The utilization of padding scheme on convolutional neural network for cervical cell images classification
Toto Haryanto, Imas Sukaesih Sitanggang, Muhammad Ashyar Agmalaro, and Riries Rulaningtyas. The utilization of padding scheme on convolutional neural network for cervical cell images classification. In: 2020 International Confer- ence on Computer Engineering, Network, and Intelligent Multimedia , 34–38. IEEE, 2020
2020
-
[75]
PathoCoder: Rethinking the Flaws of Patch-Based Learning for Multi-Class Classification in Computational Pathology
Ferdaous Idlahcen, Pierjos Francis Colere Mboukou, Ali Idri, and Hicham El At- tar. PathoCoder: Rethinking the Flaws of Patch-Based Learning for Multi-Class Classification in Computational Pathology. 2025, Microscopy Research and Technique, 88 (6): 1712–1726
2025
-
[76]
Enhancing cervical cancer diag- nosis: Integrated attention-transformer system with weakly supervised learning
Ashfaque Khowaja, Beiji Zou, and Xiaoyan Kui. Enhancing cervical cancer diag- nosis: Integrated attention-transformer system with weakly supervised learning. 2024, Image and Vision Computing, 149: 105193
2024
-
[77]
Perbandingan Arsitektur ResNet50 dan ResNet101 dalam Klasifikasi Kanker Serviks pada Citra Pap Smear
Za’imatun Niswati, Rahayuning Hardatin, Meia Noer Muslimah, and Siti Nur Hasanah. Perbandingan Arsitektur ResNet50 dan ResNet101 dalam Klasifikasi Kanker Serviks pada Citra Pap Smear. 2021, Faktor Exacta, 14 (3): 160–167
2021
-
[78]
Privacy preserved cervical can- cer detection using convolutional neural networks applied to pap smear im- ages
Shtwai Alsubai, Abdullah Alqahtani, Mohemmed Sha, Ahmad Almadhor, Sidra Abbas, Huma Mughal, and Michal Gregus. Privacy preserved cervical can- cer detection using convolutional neural networks applied to pap smear im- ages. 2023, Computational and Mathematical Methods in Medicine , 2023 (1): 9676206. 37 JUSTC Geometry-aware Cytology Classification Yating Li et al
2023
-
[79]
Differential evolution optimization based ensemble framework for accurate cervical cancer diagnosis
Omair Bilal, Sohaib Asif, Ming Zhao, Y angfan Li, Fengxiao Tang, and Yusen Zhu. Differential evolution optimization based ensemble framework for accurate cervical cancer diagnosis. 2024, Applied Soft Computing , 167: 112366
2024
-
[80]
Deep Learning Enabled Segmentation, Classification and Risk Assess- ment of Cervical Cancer, 2025
Abdul Samad Shaik, Shashaank Mattur Aswatha, and Rahul Jashvantbhai Pandya. Deep Learning Enabled Segmentation, Classification and Risk Assess- ment of Cervical Cancer, 2025. arXiv preprint arXiv:2505.15505
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.