Collaborative Edge-to-Server Inference for Vision-Language Models

Soochang Song; Yongjune Kim

arxiv: 2512.16349 · v2 · pith:QJWS37DJnew · submitted 2025-12-18 · 💻 cs.CV · cs.AI

Collaborative Edge-to-Server Inference for Vision-Language Models

Soochang Song , Yongjune Kim This is my paper

classification 💻 cs.CV cs.AI

keywords inferencecommunicationserveraccuracyframeworkcollaborativecostedge

0 comments

read the original abstract

We propose a collaborative edge-to-server inference framework for vision-language models (VLMs) that reduces communication cost while maintaining inference accuracy. In typical deployments, visual data captured at edge devices (clients) is transmitted to the server for VLM inference. However, transmitting full-resolution images incurs high communication cost. Conversely, aggressive downsizing or excessive compression to mitigate communication overhead can discard fine-grained details, leading to accuracy degradation. To overcome this limitation, we design a communication-efficient two-stage framework. In the first stage, the server performs inference on the downsized thumbnail (global image) and quantifies the min-entropy of the output tokens. If the min-entropy exceeds a predefined threshold, the server identifies a region of interest (RoI) using the VLM's internal attention and requests the edge device to send a detail-preserved local image of the RoI. The server then refines its inference by jointly leveraging the global and local images. This selective retransmission strategy ensures that only essential visual content is additionally transmitted. Experimental results consistently confirm that the proposed framework substantially reduces communication overhead while maintaining inference accuracy across diverse VQA benchmarks.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Progressive Semantic Communication for Efficient Edge-Cloud Vision-Language Models
cs.LG 2026-04 unverdicted novelty 5.0

A Meta AutoEncoder framework enables adaptive, progressive compression of visual features for low-latency edge-cloud VLM inference without model fine-tuning.