An unverifiable report that a YOLO variant with ResC2Net, SPPF, and PConv modules detects machine-part defects at mAP 0.91 without comparing to any baseline.
Robust Domain Generalization for Multi-modal Object Recognition
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In multi-label classification, machine learning encounters the challenge of domain generalization when handling tasks with distributions differing from the training data. Existing approaches primarily focus on vision object recognition and neglect the integration of natural language. Recent advancements in vision-language pre-training leverage supervision from extensive visual-language pairs, enabling learning across diverse domains and enhancing recognition in multi-modal scenarios. However, these approaches face limitations in loss function utilization, generality across backbones, and class-aware visual fusion. This paper proposes solutions to these limitations by inferring the actual loss, broadening evaluations to larger vision-language backbones, and introducing Mixup-CLIPood, which incorporates a novel mix-up loss for enhanced class-aware visual fusion. Our method demonstrates superior performance in domain generalization across multiple datasets.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Detecting and Classifying Defective Products in Images Using YOLO
An unverifiable report that a YOLO variant with ResC2Net, SPPF, and PConv modules detects machine-part defects at mAP 0.91 without comparing to any baseline.