Pith. sign in

REVIEW 12 cited by

SegGPT: Segmenting Everything In Context

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03284 v1 pith:WRIRWEJ7 submitted 2023-04-06 cs.CV

classification cs.CV
keywords segmentationseggpttaskscontextin-contextsegmentingdataeverything
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present SegGPT, a generalist model for segmenting everything in context. We unify various segmentation tasks into a generalist in-context learning framework that accommodates different kinds of segmentation data by transforming them into the same format of images. The training of SegGPT is formulated as an in-context coloring problem with random color mapping for each data sample. The objective is to accomplish diverse tasks according to the context, rather than relying on specific colors. After training, SegGPT can perform arbitrary segmentation tasks in images or videos via in-context inference, such as object instance, stuff, part, contour, and text. SegGPT is evaluated on a broad range of tasks, including few-shot semantic segmentation, video object segmentation, semantic segmentation, and panoptic segmentation. Our results show strong capabilities in segmenting in-domain and out-of-domain targets, either qualitatively or quantitatively.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Functionalization via Structure Completion and Motion Rectification

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Object functionalization is cast as neural graph completion over a functional graph of parts, contacts, and motions, followed by geometry realization that also rectifies erroneous motions, demonstrated on furniture wi...

  2. GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents

    q-bio.QM 2025-10 unverdicted novelty 7.0 of 10

    GenCellAgent deploys a planner-executor-evaluator LLM agent loop to automatically select, adapt, and refine segmentation tools for diverse cellular microscopy images, matching or exceeding specialist performance on 4,...

  3. LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

    cs.CV 2023-03 conditional novelty 7.0 of 10

    LLaMA-Adapter turns frozen LLaMA 7B into a capable instruction follower using only 1.2M new parameters and zero-init attention, matching Alpaca while extending to image-conditioned reasoning on ScienceQA and COCO.

  4. Probing Intrinsic Medical Task Relationships: A Contrastive Learning Perspective

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    TaCo contrastively embeds semantic, generative, and transformation tasks from medical imaging into a joint space to reveal which tasks cluster, blend, or remain distinct.

  5. Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A two-stage pure RL method with an information-gap global view and hierarchical grounding loss makes MLLMs truly rely on precise crops and sets SOTA on high-res VQA under tight token budgets.

  6. Stable Diffusion Models are Secretly Good at Visual In-Context Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free attention recomputation inside Stable Diffusion self-attention enables visual in-context learning across six vision tasks.

  7. Decouple before Align: Visual Disentanglement Enhances Prompt Tuning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Decoupling images into foreground and background before aligning them with text improves CLIP prompt tuning on few-shot and generalization benchmarks.

  8. CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    CheXanatomy trains VLMs to generate 2D anatomical masks via next-token prediction on synthetic CXRs from CT, matching U-Net performance with better domain-shift robustness and sample efficiency.

  9. Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs

    cs.CV 2026-03 unverdicted novelty 5.0 of 10

    A two-stage RL method with information gaps and grounding loss trains MLLMs to focus on and precisely crop relevant image regions, yielding SOTA results on high-resolution VQA benchmarks.

  10. DOMR: Establishing Cross-View Segmentation via Dense Object Matching

    cs.CV 2025-08 conditional novelty 5.0 of 10

    DOMR jointly matches and refines multiple object masks across ego and exo views, reaching 49.7% and 55.2% mean IoU on Ego-Exo4D.

  11. UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    UnAC improves LMM performance on visual reasoning benchmarks by combining adaptive visual prompting, image abstraction, and gradual self-checking.

  12. Comparison Study: Glacier Calving Front Delineation in Synthetic Aperture Radar Images With Deep Learning

    cs.CV 2025-01 unverdicted novelty 4.0 of 10

    Deep learning systems for glacier calving front delineation in SAR imagery exhibit errors up to 221 m while human annotators deviate by only 38 m.

Pith tools