Pith. sign in

REVIEW 2 cited by

U-Net Transformer: Self and Cross Attention for Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.06104 v2 pith:454ZODQ5 submitted 2021-03-10 eess.IV cs.CV

classification eess.IVcs.CV
keywords segmentationu-transformerattentioncross-attentionfeaturesimageu-netbrought
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Medical image segmentation remains particularly challenging for complex and low-contrast anatomical structures. In this paper, we introduce the U-Transformer network, which combines a U-shaped architecture for image segmentation with self- and cross-attention from Transformers. U-Transformer overcomes the inability of U-Nets to model long-range contextual interactions and spatial dependencies, which are arguably crucial for accurate segmentation in challenging contexts. To this end, attention mechanisms are incorporated at two main levels: a self-attention module leverages global interactions between encoder features, while cross-attention in the skip connections allows a fine spatial recovery in the U-Net decoder by filtering out non-semantic features. Experiments on two abdominal CT-image datasets show the large performance gain brought out by U-Transformer compared to U-Net and local Attention U-Nets. We also highlight the importance of using both self- and cross-attention, and the nice interpretability features brought out by U-Transformer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Head Explainer: A General Framework to Improve Explainability in CNNs and Transformers

    cs.CV 2025-01 reject novelty 3.0 of 10

    MHEX inserts attention-gated deep-supervision heads into ResNet and BERT and derives saliency maps from the product of the head weights, claiming better accuracy and more detailed explanations.

  2. A Comparative Study of U-Net Architectures for Change Detection in Satellite Images

    cs.CV 2025-06 reject novelty 2.0 of 10

    A literature survey of U-Net variants for remote sensing change detection that compiles reported results from prior papers without running any new experiments.

Pith tools