A new smart-home video anomaly benchmark and a taxonomy-driven reflective LLM chain that improves MLLM anomaly detection accuracy by 11.62 percentage points over zero-shot prompting.
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We consider the problem of zero-shot one-class visual classification, extending traditional one-class classification to scenarios where only the label of the target class is available. This method aims to discriminate between positive and negative query samples without requiring examples from the target class. We propose a two-step solution that first queries large language models for visually confusing objects and then relies on vision-language pre-trained models (e.g., CLIP) to perform classification. By adapting large-scale vision benchmarks, we demonstrate the ability of the proposed method to outperform adapted off-the-shelf alternatives in this setting. Namely, we propose a realistic benchmark where negative query samples are drawn from the same original dataset as positive ones, including a granularity-controlled version of iNaturalist, where negative samples are at a fixed distance in the taxonomy tree from the positive ones. To our knowledge, we are the first to demonstrate the ability to discriminate a single category from other semantically related ones using only its label.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models
A new smart-home video anomaly benchmark and a taxonomy-driven reflective LLM chain that improves MLLM anomaly detection accuracy by 11.62 percentage points over zero-shot prompting.