The paper introduces CRFI and CPRM, which align region-level visual features with CLIP text embeddings and mix proposals from clean and augmented images, achieving state-of-the-art on Cityscapes-C and DWD.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction
The paper introduces CRFI and CPRM, which align region-level visual features with CLIP text embeddings and mix proposals from clean and augmented images, achieving state-of-the-art on Cityscapes-C and DWD.