Supervised finetuning of a CLIP image encoder with multi-scale features and text prompts detects anomaly objects in steel scrap at 28.6% pixel-level average precision, outperforming tested baselines on a private dataset.
Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Recycling steel scrap can reduce carbon dioxide (CO2) emissions from the steel industry. However, a significant challenge in steel scrap recycling is the inclusion of impurities other than steel. To address this issue, we propose vision-language-model-based anomaly detection where a model is finetuned in a supervised manner, enabling it to handle niche objects effectively. This model enables automated detection of anomalies at a fine-grained level within steel scrap. Specifically, we finetune the image encoder, equipped with multi-scale mechanism and text prompts aligned with both normal and anomaly images. The finetuning process trains these modules using a multiclass classification as the supervision.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
Supervised finetuning of a CLIP image encoder with multi-scale features and text prompts detects anomaly objects in steel scrap at 28.6% pixel-level average precision, outperforming tested baselines on a private dataset.