Pith. sign in

REVIEW 4 cited by

EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.00863 v1 pith:K3LCTZDN submitted 2023-12-01 cs.CV

classification cs.CV
keywords imageanythingpretrainingsegmentapplicationsefficientsamsmaskedmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Segment Anything Model (SAM) has emerged as a powerful tool for numerous vision applications. A key component that drives the impressive performance for zero-shot transfer and high versatility is a super large Transformer model trained on the extensive high-quality SA-1B dataset. While beneficial, the huge computation cost of SAM model has limited its applications to wider real-world applications. To address this limitation, we propose EfficientSAMs, light-weight SAM models that exhibits decent performance with largely reduced complexity. Our idea is based on leveraging masked image pretraining, SAMI, which learns to reconstruct features from SAM image encoder for effective visual representation learning. Further, we take SAMI-pretrained light-weight image encoders and mask decoder to build EfficientSAMs, and finetune the models on SA-1B for segment anything task. We perform evaluations on multiple vision tasks including image classification, object detection, instance segmentation, and semantic object detection, and find that our proposed pretraining method, SAMI, consistently outperforms other masked image pretraining methods. On segment anything task such as zero-shot instance segmentation, our EfficientSAMs with SAMI-pretrained lightweight image encoders perform favorably with a significant gain (e.g., ~4 AP on COCO/LVIS) over other fast SAM models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Zero-Shot Polygon Matching with Pre-trained Models for Pose Estimation and Polygon Cloud from Challenging Stereo

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Polygon regions in stereo image pairs can be matched without training by combining SAM masks, pyramid-guided search, and Hungarian-based local matching.

  2. Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new interactive segmentation decoder that routes computation to boundary regions, using binary quantization attention and mixture-of-experts, achieves state-of-the-art accuracy with CPU-friendly latency.

  3. GeoSAM-Lite: A Lightweight Foundation Model for Onboard Remote Sensing Segmentation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A lightweight RS segmenter using domain distillation from an RS teacher plus spatial/frequency fusion layers matches other light SAMs and approaches a heavy teacher at 92.8% fewer parameters.

  4. IRS: Incremental Relationship-guided Segmentation for Digital Pathology

    eess.IV 2025-05 conditional novelty 5.0 of 10

    IRS uses anatomical relationships between old and new classes to guide knowledge distillation in a prompt-driven mixture-of-experts network, improving class-incremental segmentation of kidney pathology compared to baselines.

Pith tools