Supervised finetuning of a CLIP image encoder with multi-scale features and text prompts detects anomaly objects in steel scrap at 28.6% pixel-level average precision, outperforming tested baselines on a private dataset.
Setup Dataset and Metrics.The images used in this study were collected from a steel scrap recycling site and provided by anonymized
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
Supervised finetuning of a CLIP image encoder with multi-scale features and text prompts detects anomaly objects in steel scrap at 28.6% pixel-level average precision, outperforming tested baselines on a private dataset.