CoME-VL fuses contrastive and self-supervised vision encoders via entropy-guided multi-layer aggregation and RoPE cross-attention to improve vision-language model performance on benchmarks.
Fiaz, Al- ham Fikri Aji, and Hisham Cholakkal
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.CV 2years
2026 2roles
dataset 1polarities
use dataset 1representative citing papers
AdaVFM trains a DINOv2-distilled, CLIP-aligned NAS supernet and uses a cloud LLM to select the cheapest subnet per scene, reducing edge FLOPs by up to 77.9% at similar accuracy.
citing papers explorer
-
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
CoME-VL fuses contrastive and self-supervised vision encoders via entropy-guided multi-layer aggregation and RoPE cross-attention to improve vision-language model performance on benchmarks.
-
AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference
AdaVFM trains a DINOv2-distilled, CLIP-aligned NAS supernet and uses a cloud LLM to select the cheapest subnet per scene, reducing edge FLOPs by up to 77.9% at similar accuracy.