Mosaic uses cross-modal clusters as the unit for KVCache organization in VLMs to achieve up to 1.38x speedup in streaming long-video understanding.
Drivegpt4: Interpretable end-to-end autonomous driving via large language model
2 Pith papers cite this work, alongside 316 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 2roles
background 1polarities
background 1representative citing papers
A vision-language model classifies road surface and visibility conditions from front-camera images, parametrizing a context-adaptive safety envelope that couples braking and steering through a shared friction budget, achieving 100% trial success in CARLA simulation across adverse conditions.
citing papers explorer
-
Mosaic: Cross-Modal Clustering for Efficient Video Understanding
Mosaic uses cross-modal clusters as the unit for KVCache organization in VLMs to achieve up to 1.38x speedup in streaming long-video understanding.
-
VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving
A vision-language model classifies road surface and visibility conditions from front-camera images, parametrizing a context-adaptive safety envelope that couples braking and steering through a shared friction budget, achieving 100% trial success in CARLA simulation across adverse conditions.