Active lifting of data distributions via latent-sequence sampling and max-rate uncertainty reduction formally derives slow-thinking LLMs and places them on representation and sampler hierarchies that can be climbed.
AdaptVision: Efficient vision-language models via adaptive visual acquisition.arXiv preprint arXiv:2512.03794, 2025
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
AVA-VLM reduces visual-token usage by 69% while improving PPE-violation F1 by 13 points over direct-QA baselines by training a VLM to adaptively crop high-resolution local regions from a downsampled global image.
A roadmap that defines architectural nativity for multimodal models and categorizes them into Multi-to-Text, Multi-to-Target, and Multi-to-Multi types while outlining an industrial pipeline toward unified transformer-based native multimodal modeling.
citing papers explorer
-
A First-Principles Theory of Slow Thinking and Active Perception
Active lifting of data distributions via latent-sequence sampling and max-rate uncertainty reduction formally derives slow-thinking LLMs and places them on representation and sampler hierarchies that can be climbed.
-
AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring
AVA-VLM reduces visual-token usage by 69% while improving PPE-violation F1 by 13 points over direct-QA baselines by training a VLM to adaptively crop high-resolution local regions from a downsampled global image.
-
Toward Native Multimodal Modeling: A Roadmap
A roadmap that defines architectural nativity for multimodal models and categorizes them into Multi-to-Text, Multi-to-Target, and Multi-to-Multi types while outlining an industrial pipeline toward unified transformer-based native multimodal modeling.