SAViL-Det combines CLIP, an asymptotic feature pyramid, and cross-modal attention to report F-scores of 84.8 on MLT-2019 and 90.2 on CTW1500, claiming state-of-the-art multi-script and curved text detection.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection
SAViL-Det combines CLIP, an asymptotic feature pyramid, and cross-modal attention to report F-scores of 84.8 on MLT-2019 and 90.2 on CTW1500, claiming state-of-the-art multi-script and curved text detection.