AVA-VLM reduces visual-token usage by 69% while improving PPE-violation F1 by 13 points over direct-QA baselines by training a VLM to adaptively crop high-resolution local regions from a downsampled global image.
Visual question answering-based referring expression segmentation for construction safety analysis
1 Pith paper cite this work, alongside 5 external citations. Polarity classification is still indexing.
1
Pith paper citing it
5
external citations · external index
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring
AVA-VLM reduces visual-token usage by 69% while improving PPE-violation F1 by 13 points over direct-QA baselines by training a VLM to adaptively crop high-resolution local regions from a downsampled global image.