REVIEW 14 cited by
Segment Any Anomaly without Training via Hybrid Prompt Regularization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present a novel framework, i.e., Segment Any Anomaly + (SAA+), for zero-shot anomaly segmentation with hybrid prompt regularization to improve the adaptability of modern foundation models. Existing anomaly segmentation models typically rely on domain-specific fine-tuning, limiting their generalization across countless anomaly patterns. In this work, inspired by the great zero-shot generalization ability of foundation models like Segment Anything, we first explore their assembly to leverage diverse multi-modal prior knowledge for anomaly localization. For non-parameter foundation model adaptation to anomaly segmentation, we further introduce hybrid prompts derived from domain expert knowledge and target image context as regularization. Our proposed SAA+ model achieves state-of-the-art performance on several anomaly segmentation benchmarks, including VisA, MVTec-AD, MTD, and KSDD2, in the zero-shot setting. We will release the code at \href{https://github.com/caoyunkang/Segment-Any-Anomaly}{https://github.com/caoyunkang/Segment-Any-Anomaly}.
Forward citations
Cited by 14 Pith papers
-
IQE-CLIP: Instance-aware Query Embedding for Zero-/Few-shot Anomaly Detection in Medical Domain
IQE-CLIP improves zero- and few-shot medical anomaly detection by building query embeddings that combine text prompts with visual features from each test image, beating prior CLIP-based methods on six BMAD datasets.
-
ONER: Online Experience Replay for Incremental Anomaly Detection
ONER combines decomposed prompts and semantic prototypes to achieve state-of-the-art incremental anomaly detection on MVTec AD and VisA without replaying raw images.
-
Closed form perturbative relativistic modifications to wave-packet dynamics in the quantum harmonic oscillator
Closed-form O(1/c²) relativistic corrections to QHO wave-packet widths, variances, and uncertainty products leave minimum-uncertainty saturation intact and become percent-level for 1–10 keV electron confinement.
-
Toward Long-Tailed Online Anomaly Detection through Class-Agnostic Concepts
The paper proposes the long-tailed online anomaly detection (LTOAD) benchmark and a class-agnostic concept-based framework that outperforms class-aware baselines in most offline settings and in the online setting.
-
Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
Supervised finetuning of a CLIP image encoder with multi-scale features and text prompts detects anomaly objects in steel scrap at 28.6% pixel-level average precision, outperforming tested baselines on a private dataset.
-
MIAS-SAM: Medical Image Anomaly Segmentation without thresholding
MIAS-SAM generates medical anomaly segmentations by memory-bank matching of SAM encoder features, using the anomaly map's center of gravity as a point prompt for the SAM decoder, without thresholding the anomaly map.
-
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
OmniAD unifies industrial anomaly detection and understanding in a single multimodal model using text-encoded masks and reinforcement learning, reporting 79.1 on MMAD and strong detection scores.
-
Zero-Shot Industrial Anomaly Segmentation with Image-Aware Prompt Generation
A zero-shot anomaly segmentation pipeline that generates image-specific defect prompts using an image tagger and an LLM improves F1-max by up to 10% over fixed prompts.
-
Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning
A self-drafting and relation-aware fine-tuning scheme lifts SAM's anomaly segmentation mIoU by 1.5 to 2.0 points over PEFT baselines across six industrial datasets.
-
SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM
SAM-MI improves open-vocabulary segmentation by injecting aggregated SAM masks as low- and high-frequency guidance into CLIP cost maps, with sparse text-guided point prompts for speed.
-
MCL-AD: Multimodal Collaboration Learning for Zero-Shot 3D Anomaly Detection
MCL-AD fuses point cloud, RGB, and text modalities with learnable decoupled prompts and a contrastive loss, reporting top scores on two 3D anomaly detection benchmarks.
-
StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection
Stacking multiple category names in a CLIP text prompt, along with cluster-specific alignment layers, improves zero-shot industrial defect detection and localization.
-
Learning to Detect Multi-class Anomalies with Just One Normal Image Prompt
OneNIP reconstructs features with one normal image prompt and a supervised refiner, achieving state-of-the-art unified anomaly detection on MVTec, BTAD, and VisA.
-
Exploring Aleatoric Uncertainty in Object Detection via Vision Foundation Models
A Mahalanobis distance computed in SAM feature space is used as an aleatoric uncertainty score for object instances, and filtering and reweighting by this score yields modest AP gains on COCO and BDD100K.
Discussion (0). Continue with ORCID to comment.