Pith. sign in

REVIEW 11 cited by

Segment Everything Everywhere All at Once

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.06718 v4 pith:NIY6GFRM submitted 2023-04-13 cs.CV

classification cs.CV
keywords segmentationseemimagepromptstaskstextacrossdecoder
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this work, we present SEEM, a promptable and interactive model for segmenting everything everywhere all at once in an image, as shown in Fig.1. In SEEM, we propose a novel decoding mechanism that enables diverse prompting for all types of segmentation tasks, aiming at a universal segmentation interface that behaves like large language models (LLMs). More specifically, SEEM is designed with four desiderata: i) Versatility. We introduce a new visual prompt to unify different spatial queries including points, boxes, scribbles and masks, which can further generalize to a different referring image; ii) Compositionality. We learn a joint visual-semantic space between text and visual prompts, which facilitates the dynamic composition of two prompt types required for various segmentation tasks; iii) Interactivity. We further incorporate learnable memory prompts into the decoder to retain segmentation history through mask-guided cross-attention from decoder to image features; and iv) Semantic-awareness. We use a text encoder to encode text queries and mask labels into the same semantic space for open-vocabulary segmentation. We conduct a comprehensive empirical study to validate the effectiveness of SEEM across diverse segmentation tasks. Notably, our single SEEM model achieves competitive performance across interactive segmentation, generic segmentation, referring segmentation, and video object segmentation on 9 datasets with minimum 1/100 supervision. Furthermore, SEEM showcases a remarkable capacity for generalization to novel prompts or their combinations, rendering it a readily universal image segmentation interface.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 152 citations worldwide. Full citation record

  1. Discovering and using Spelke segments

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SpelkeNet, a self-supervised video world model, discovers Spelke segments in static images by aggregating motion correlations across imagined pokes.

  2. Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    DETRIS uses dense mixtures of convolutions and cross-attention adapters to tune a frozen DINOv2/CLIP pair, achieving top reported IoU on three referring image segmentation benchmarks while updating only a small fracti...

  3. On Moving Object Segmentation from Monocular Video with Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    M3Former fuses appearance and motion streams in a Mask2Former-style transformer and, when trained on a diverse mix that includes KITTI and DAVIS train splits, reaches state-of-the-art motion segmentation scores on tho...

  4. Segment Any Class (SAC): Multi-Class Few-Shot Semantic Segmentation via Class Region Proposals

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A training-free prompting method that adapts SAM to multi-class few-shot segmentation and reports higher mIoU than trained baselines on COCO-20i as the number of classes grows.

  5. From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The paper presents GEST, an event-graph representation of videos that is converted automatically into natural language and is also used as a teacher to pre-train end-to-end video captioning models.

  6. INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    INT improves task-generic promptable segmentation by progressively mining negative candidates, using VLM output differences after masking to select and refine instance-specific prompts.

  7. Efficient MedSAMs: Segment Anything in Medical Images on Laptop

    eess.IV 2024-12 conditional novelty 5.0 of 10

    A medical imaging competition produced lightweight segmentation models that match a large foundation model's accuracy on a new hidden test set while running on laptop CPUs with substantially less compute.

  8. SPT: Sequence Prompt Transformer for Interactive Image Segmentation

    cs.CV 2024-12 reject novelty 5.0 of 10

    A sequence-aware transformer for interactive image segmentation that uses previous images and clicks as prompts, claiming state-of-the-art results without reporting them.

  9. DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes

    cs.CV 2025-08 conditional novelty 4.0 of 10

    DrivingGaussian++ reconstructs dynamic surround-view driving scenes and performs training-free multi-task editing (weather, texture, object manipulation) using Gaussians, diffusion models, and LLM-generated trajectories.

  10. Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model

    cs.CV 2025-05 conditional novelty 3.0 of 10

    SegVLM reports 53.87 IoU on PhraseCut referring segmentation by adding SE blocks, deformable convolutions, residual shortcuts, and a fused BCE-Focal-Dice loss to CRIS.

  11. Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A structured survey of prompt engineering methods for the Segment Anything Model, covering geometric, textual, and multimodal prompts and their applications.

Pith tools