Pith. sign in

REVIEW 5 cited by

Multimodal Federated Learning via Contrastive Representation Ensemble

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.08888 v3 pith:EXQAJIT6 submitted 2023-02-17 cs.LG cs.AI

classification cs.LGcs.AI
keywords multimodalmodeldataclientslearningmodalityensemblefederated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the increasing amount of multimedia data on modern mobile systems and IoT infrastructures, harnessing these rich multimodal data without breaching user privacy becomes a critical issue. Federated learning (FL) serves as a privacy-conscious alternative to centralized machine learning. However, existing FL methods extended to multimodal data all rely on model aggregation on single modality level, which restrains the server and clients to have identical model architecture for each modality. This limits the global model in terms of both model complexity and data capacity, not to mention task diversity. In this work, we propose Contrastive Representation Ensemble and Aggregation for Multimodal FL (CreamFL), a multimodal federated learning framework that enables training larger server models from clients with heterogeneous model architectures and data modalities, while only communicating knowledge on public dataset. To achieve better multimodal representation fusion, we design a global-local cross-modal ensemble strategy to aggregate client representations. To mitigate local model drift caused by two unprecedented heterogeneous factors stemming from multimodal discrepancy (modality gap and task gap), we further propose two inter-modal and intra-modal contrasts to regularize local training, which complements information of the absent modality for uni-modal clients and regularizes local clients to head towards global consensus. Thorough evaluations and ablation studies on image-text retrieval and visual question answering tasks showcase the superiority of CreamFL over state-of-the-art FL methods and its practical value.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. STAGE: Tackling Semantic Drift in Multimodal Federated Graph Learning

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    STAGE builds a shared semantic space through feature translation and controlled graph propagation to reduce semantic drift in multimodal federated graph learning, delivering state-of-the-art results with lower communi...

  2. FedTaste: Topology-Aware Structural Transfer for Multimodal Federated Learning with Missing Modalities

    cs.MM 2026-07 reject novelty 6.0 of 10

    FedTaste aligns missing-modality clients to a server-built semantic graph distilled from full-modality CLIP clients by updating only lightweight prompts, reporting state-of-the-art federated retrieval numbers.

  3. FedLAB: Traceable Semantic Codebooks for Federated Multimodal Graph Foundation Learning

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    FedLAB organizes multimodal graph knowledge into typed hierarchical codebooks for modality evidence, node semantics, and topology context via federated semantic barycenter pre-training, improving performance by up to ...

  4. Conditional Imputation for Within-Modality Missingness in Multi-Modal Federated Learning

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    CondI applies conditional diffusion models in a two-phase federated pipeline to impute within-modality missing data, then trains extractors on the completed inputs for downstream tasks on clinical datasets.

  5. Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    RCSR is a personalization-friendly federated framework that improves cross-modal retrieval accuracy and stability under missing modalities via semantic routing and adapters.

Pith tools