TuringViT claims a new ViT design with linear attention and curated data that matches SOTA performance using 10% of typical pretraining data while supporting dynamic resolutions and improving VLM integration.
Mobileclip: Fast image-text models through multi-modal reinforced training
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 3years
2026 3representative citing papers
VisionAId is an offline-first Android application that combines six on-device models for depth, segmentation, embeddings, face detection and banknote recognition with a few-shot pipeline that lets users teach the system their personal objects and then guides them to those objects via AR, audio and h
A competition report covering 2025 LPCVC's Qualcomm AI Hub evaluation system and the winning solutions in classification, open-vocabulary segmentation, and depth estimation.
citing papers explorer
-
TuringViT: Making SOTA Vision Transformers Accessible to All
TuringViT claims a new ViT design with linear attention and curated data that matches SOTA performance using 10% of typical pretraining data while supporting dynamic resolutions and improving VLM integration.
-
VisionAId: An Offline-First Multimodal Android Assistant for People with Visual Impairment, Featuring Personalized Object Retrieval
VisionAId is an offline-first Android application that combines six on-device models for depth, segmentation, embeddings, face detection and banknote recognition with a few-shot pipeline that lets users teach the system their personal objects and then guides them to those objects via AR, audio and h
-
Evaluation of Winning Solutions of 2025 Low Power Computer Vision Challenge
A competition report covering 2025 LPCVC's Qualcomm AI Hub evaluation system and the winning solutions in classification, open-vocabulary segmentation, and depth estimation.