Pith. sign in

REVIEW 3 cited by

Hi-End-MAE: Hierarchical encoder-driven masked autoencoders are stronger vision learners for medical image segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.08347 v1 pith:NGC5PE4Y submitted 2025-02-12 cs.CV

classification cs.CV
keywords medicalhi-end-maeacrosshierarchicalimagedownstreamencoder-drivenlayers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Medical image segmentation remains a formidable challenge due to the label scarcity. Pre-training Vision Transformer (ViT) through masked image modeling (MIM) on large-scale unlabeled medical datasets presents a promising solution, providing both computational efficiency and model generalization for various downstream tasks. However, current ViT-based MIM pre-training frameworks predominantly emphasize local aggregation representations in output layers and fail to exploit the rich representations across different ViT layers that better capture fine-grained semantic information needed for more precise medical downstream tasks. To fill the above gap, we hereby present Hierarchical Encoder-driven MAE (Hi-End-MAE), a simple yet effective ViT-based pre-training solution, which centers on two key innovations: (1) Encoder-driven reconstruction, which encourages the encoder to learn more informative features to guide the reconstruction of masked patches; and (2) Hierarchical dense decoding, which implements a hierarchical decoding structure to capture rich representations across different layers. We pre-train Hi-End-MAE on a large-scale dataset of 10K CT scans and evaluated its performance across seven public medical image segmentation benchmarks. Extensive experiments demonstrate that Hi-End-MAE achieves superior transfer learning capabilities across various downstream tasks, revealing the potential of ViT in medical imaging applications. The code is available at: https://github.com/FengheTan9/Hi-End-MAE

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SimCroP: Radiograph Representation Learning with Similarity-driven Cross-granularity Pre-training

    cs.CV 2025-09 conditional novelty 6.0 of 10

    SimCroP learns chest-CT representations by aligning each report sentence to its most similar visual patches and fusing whole-scan and word-patch features, reporting higher classification and segmentation scores than s...

  2. Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A frozen pre-trained LLM layer inserted between a CNN encoder and decoder improves medical image segmentation across ultrasound, dermoscopy, polyp, and CT benchmarks with few added trainable parameters.

  3. U-RWKV: Lightweight medical image segmentation with direction-adaptive RWKV

    eess.IV 2025-07 conditional novelty 5.0 of 10

    U-RWKV is a lightweight U-shaped medical image segmenter that combines multi-directional RWKV scanning with stage-adaptive channel recalibration, reporting competitive Dice scores with about three million parameters.

Pith tools