Pith. sign in

REVIEW 3 cited by

Foreground-Covering Prototype Generation and Matching for SAM-Aided Few-Shot Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.00752 v1 pith:TQS37553 submitted 2025-01-01 cs.CV

classification cs.CV
keywords queryfeaturesprototypessupportprototypepseudo-maskregionsresnet
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose Foreground-Covering Prototype Generation and Matching to resolve Few-Shot Segmentation (FSS), which aims to segment target regions in unlabeled query images based on labeled support images. Unlike previous research, which typically estimates target regions in the query using support prototypes and query pixels, we utilize the relationship between support and query prototypes. To achieve this, we utilize two complementary features: SAM Image Encoder features for pixel aggregation and ResNet features for class consistency. Specifically, we construct support and query prototypes with SAM features and distinguish query prototypes of target regions based on ResNet features. For the query prototype construction, we begin by roughly guiding foreground regions within SAM features using the conventional pseudo-mask, then employ iterative cross-attention to aggregate foreground features into learnable tokens. Here, we discover that the cross-attention weights can effectively alternate the conventional pseudo-mask. Therefore, we use the attention-based pseudo-mask to guide ResNet features to focus on the foreground, then infuse the guided ResNet feature into the learnable tokens to generate class-consistent query prototypes. The generation of the support prototype is conducted symmetrically to that of the query one, with the pseudo-mask replaced by the ground-truth mask. Finally, we compare these query prototypes with support ones to generate prompts, which subsequently produce object masks through the SAM Mask Decoder. Our state-of-the-art performances on various datasets validate the effectiveness of the proposed method for FSS. Our official code is available at https://github.com/SuhoPark0706/FCP

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Repurposing CLIP to Localize at Pixel Level

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CLIPix repurposes CLIP by tracing classification activations, applying noise-resistant correction, and localization embedding to reach SOTA zero-shot binary open-set segmentation on PASCAL-5i and COCO-20i.

  2. DFR: A Decompose-Fuse-Reconstruct Framework for Multi-Modal Few-Shot Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DFR integrates visual, textual, and audio guidance in a SAM-based framework and reports mIoU gains over state-of-the-art few-shot segmentation methods on PASCAL-5i and AVS-V3.

  3. Vision and Language Reference Prompt into SAM for Few-shot Segmentation

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Using a frozen vision-language model, VLP-SAM injects text-label semantics into SAM's prompt encoder and raises one-shot segmentation mIoU by 6.3 points on PASCAL-5i and 9.5 on COCO-20i.

Pith tools