Pith. sign in

REVIEW 1 cited by

Focal-UNet: UNet-like Focal Modulation for Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.09263 v1 pith:SES4M2DL submitted 2022-12-19 eess.IV cs.CV

classification eess.IVcs.CV
keywords architecturefocallocalproposedu-shapedbeendatasetdice
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, many attempts have been made to construct a transformer base U-shaped architecture, and new methods have been proposed that outperformed CNN-based rivals. However, serious problems such as blockiness and cropped edges in predicted masks remain because of transformers' patch partitioning operations. In this work, we propose a new U-shaped architecture for medical image segmentation with the help of the newly introduced focal modulation mechanism. The proposed architecture has asymmetric depths for the encoder and decoder. Due to the ability of the focal module to aggregate local and global features, our model could simultaneously benefit the wide receptive field of transformers and local viewing of CNNs. This helps the proposed method balance the local and global feature usage to outperform one of the most powerful transformer-based U-shaped models called Swin-UNet. We achieved a 1.68% higher DICE score and a 0.89 better HD metric on the Synapse dataset. Also, with extremely limited data, we had a 4.25% higher DICE score on the NeoPolyp dataset. Our implementations are available at: https://github.com/givkashi/Focal-UNet

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedFormer: Hierarchical Medical Vision Transformer with Content-Aware Dual Sparse Selection Attention

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A two-stage sparse attention mechanism, selecting top regions then top pixels per query, improves accuracy on multiple medical imaging benchmarks with lower compute than full attention.

Pith tools