A self-supervised training method that selects and merges instruction-relevant vision tokens, cutting LVLM compute while claiming stable VQA accuracy and better dense perception.
Few-shot adversarial prompt learning on vision-language models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
A self-supervised training method that selects and merges instruction-relevant vision tokens, cutting LVLM compute while claiming stable VQA accuracy and better dense perception.