Pith. sign in

REVIEW 2 cited by

MobileUtr: Revisiting the relationship between light-weight CNN and Transformer for efficient medical image segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.01740 v1 pith:6BCPUEOJ submitted 2023-12-04 eess.IV cs.CV

classification eess.IVcs.CV
keywords medicalimageefficientmobileutrsegmentationtransformercnnsinformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Due to the scarcity and specific imaging characteristics in medical images, light-weighting Vision Transformers (ViTs) for efficient medical image segmentation is a significant challenge, and current studies have not yet paid attention to this issue. This work revisits the relationship between CNNs and Transformers in lightweight universal networks for medical image segmentation, aiming to integrate the advantages of both worlds at the infrastructure design level. In order to leverage the inductive bias inherent in CNNs, we abstract a Transformer-like lightweight CNNs block (ConvUtr) as the patch embeddings of ViTs, feeding Transformer with denoised, non-redundant and highly condensed semantic information. Moreover, an adaptive Local-Global-Local (LGL) block is introduced to facilitate efficient local-to-global information flow exchange, maximizing Transformer's global context information extraction capabilities. Finally, we build an efficient medical image segmentation model (MobileUtr) based on CNN and Transformer. Extensive experiments on five public medical image datasets with three different modalities demonstrate the superiority of MobileUtr over the state-of-the-art methods, while boasting lighter weights and lower computational cost. Code is available at https://github.com/FengheTan9/MobileUtr.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A frozen pre-trained LLM layer inserted between a CNN encoder and decoder improves medical image segmentation across ultrasound, dermoscopy, polyp, and CT benchmarks with few added trainable parameters.

  2. U-RWKV: Lightweight medical image segmentation with direction-adaptive RWKV

    eess.IV 2025-07 conditional novelty 5.0 of 10

    U-RWKV is a lightweight U-shaped medical image segmenter that combines multi-directional RWKV scanning with stage-adaptive channel recalibration, reporting competitive Dice scores with about three million parameters.

Pith tools