Pith. sign in

REVIEW 4 cited by

Bootstrap Fine-Grained Vision-Language Alignment for Unified Zero-Shot Anomaly Localization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.15939 v2 pith:FUCDI4CR submitted 2023-08-30 cs.CV

classification cs.CV
keywords anomalyvisualanoclipcliplocalizationzero-shotfine-grainedtext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive Language-Image Pre-training (CLIP) models have shown promising performance on zero-shot visual recognition tasks by learning visual representations under natural language supervision. Recent studies attempt the use of CLIP to tackle zero-shot anomaly detection by matching images with normal and abnormal state prompts. However, since CLIP focuses on building correspondence between paired text prompts and global image-level representations, the lack of fine-grained patch-level vision to text alignment limits its capability on precise visual anomaly localization. In this work, we propose AnoCLIP for zero-shot anomaly localization. In the visual encoder, we introduce a training-free value-wise attention mechanism to extract intrinsic local tokens of CLIP for patch-level local description. From the perspective of text supervision, we particularly design a unified domain-aware contrastive state prompting template for fine-grained vision-language matching. On top of the proposed AnoCLIP, we further introduce a test-time adaptation (TTA) mechanism to refine visual anomaly localization results, where we optimize a lightweight adapter in the visual encoder using AnoCLIP's pseudo-labels and noise-corrupted tokens. With both AnoCLIP and TTA, we significantly exploit the potential of CLIP for zero-shot anomaly localization and demonstrate the effectiveness of AnoCLIP on various datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Anomalous Decision Discovery using Inverse Reinforcement Learning

    cs.AI 2025-07 conditional novelty 6.0 of 10

    TRAP combines trajectory-ranked IRL rewards with variable-horizon prefix sampling and DistilBERT to detect anomalous driving, reaching 0.90 AUC in simulation.

  2. IQE-CLIP: Instance-aware Query Embedding for Zero-/Few-shot Anomaly Detection in Medical Domain

    cs.CV 2025-06 conditional novelty 6.0 of 10

    IQE-CLIP improves zero- and few-shot medical anomaly detection by building query embeddings that combine text prompts with visual features from each test image, beating prior CLIP-based methods on six BMAD datasets.

  3. Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Anomaly-OV, trained on the new Anomaly-Instruct-125k dataset, improves zero-shot detection of image anomalies and their textual explanations over generalist MLLMs.

  4. StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Stacking multiple category names in a CLIP text prompt, along with cluster-specific alignment layers, improves zero-shot industrial defect detection and localization.

Pith tools