Pith. sign in

REVIEW 2 cited by

Multi-scale Hierarchical Vision Transformer with Cascaded Attention Decoding for Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.16892 v1 pith:5CPBRG6U submitted 2023-03-29 cs.CV

classification cs.CV
keywords imagemedicalmeritsegmentationaggregationattentioncascadeddecoding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers have shown great success in medical image segmentation. However, transformers may exhibit a limited generalization ability due to the underlying single-scale self-attention (SA) mechanism. In this paper, we address this issue by introducing a Multi-scale hiERarchical vIsion Transformer (MERIT) backbone network, which improves the generalizability of the model by computing SA at multiple scales. We also incorporate an attention-based decoder, namely Cascaded Attention Decoding (CASCADE), for further refinement of multi-stage features generated by MERIT. Finally, we introduce an effective multi-stage feature mixing loss aggregation (MUTATION) method for better model training via implicit ensembling. Our experiments on two widely used medical image segmentation benchmarks (i.e., Synapse Multi-organ, ACDC) demonstrate the superior performance of MERIT over state-of-the-art methods. Our MERIT architecture and MUTATION loss aggregation can be used with downstream medical image and semantic segmentation tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A dual-encoder segmentation network fusing original and fuzzy-enhanced images with bidirectional cross-attention reports 93.25 Dice on ACDC and 85.49 Dice on Synapse.

  2. Pixel-wise Modulated Dice Loss for Medical Image Segmentation

    eess.IV 2025-06 conditional novelty 4.0 of 10

    A pixel-wise modulated Dice loss, weighting each pixel by its prediction error, reports improved segmentation accuracy on Kvasir, ACDC, and MSSEG benchmarks.

Pith tools