FedVLMBench systematically benchmarks federated fine-tuning of vision-language models and finds that a 2-layer MLP connector with joint connector-LLM training is optimal for encoder-based models, while vision-centric tasks are more sensitive to non-IID data.
FedMBridge: Bridgeable multimodal federated learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models
FedVLMBench systematically benchmarks federated fine-tuning of vision-language models and finds that a 2-layer MLP connector with joint connector-LLM training is optimal for encoder-based models, while vision-centric tasks are more sensitive to non-IID data.