Introduces ViTextCaps dataset and PhonoSTFG phonological graph fusion framework for Vietnamese scene-text image captioning, showing cross-modal graph edges harm performance.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CV 2years
2026 2verdicts
UNVERDICTED 2roles
method 1polarities
background 1representative citing papers
Watermark removal leaves statistical artifacts that allow classifiers to detect the attempt at 10^{-3} FPR across tested methods, establishing forensic stealthiness as a required property.
citing papers explorer
-
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
Introduces ViTextCaps dataset and PhonoSTFG phonological graph fusion framework for Vietnamese scene-text image captioning, showing cross-modal graph edges harm performance.
-
The Forensic Cost of Watermark Removal
Watermark removal leaves statistical artifacts that allow classifiers to detect the attempt at 10^{-3} FPR across tested methods, establishing forensic stealthiness as a required property.