A pipeline that plans, places, and renders scene-coherent typographic adversarial text fools vision-language models more often than prior center or margin text attacks, but its success metric and naturalness evaluation are methodologically weak.
Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.arXiv
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
A pipeline that plans, places, and renders scene-coherent typographic adversarial text fools vision-language models more often than prior center or margin text attacks, but its success metric and naturalness evaluation are methodologically weak.