Pith. sign in

REVIEW 6 cited by

An Empirical Study of Multimodal Model Merging

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.14933 v2 pith:YS63IM62 submitted 2023-04-28 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords mergingmodeltaskstrainedarchitecturedifferentinitializationmodality-agnostic
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Model merging (e.g., via interpolation or task arithmetic) fuses multiple models trained on different tasks to generate a multi-task solution. The technique has been proven successful in previous studies, where the models are trained on similar tasks and with the same initialization. In this paper, we expand on this concept to a multimodal setup by merging transformers trained on different modalities. Furthermore, we conduct our study for a novel goal where we can merge vision, language, and cross-modal transformers of a modality-specific architecture to create a parameter-efficient modality-agnostic architecture. Through comprehensive experiments, we systematically investigate the key factors impacting model performance after merging, including initialization, merging mechanisms, and model architectures. We also propose two metrics that assess the distance between weights to be merged and can serve as an indicator of the merging outcomes. Our analysis leads to an effective training recipe for matching the performance of the modality-agnostic baseline (i.e., pre-trained from scratch) via model merging. Our method also outperforms naive merging significantly on various tasks, with improvements of 3% on VQA, 7% on COCO retrieval, 25% on NLVR2, 14% on Flickr30k and 3% on ADE20k. Our code is available at https://github.com/ylsung/vl-merging

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

    cs.AI 2026-08 conditional novelty 6.0 of 10

    AgentPatch is a training-free, coarse-to-fine repair method that merges agentic multimodal LLMs into one static checkpoint while recovering weak-task and behavior-critical capabilities.

  2. Olympus: A Universal Task Router for Computer Vision Tasks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Olympus is a trained MLLM router that delegates 20 vision tasks to specialist models and supports chain-of-action execution of up to five tasks per instruction.

  3. How to Merge Your Multimodal Models Over Time?

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A systematic study of temporal model merging shows that initialization and deployment choices matter far more than the merging technique, with EMA-style weight interpolation as the best practice.

  4. LoBAM: LoRA-Based Backdoor Attack on Model Merging

    cs.CR 2024-11 conditional novelty 6.0 of 10

    LoBAM constructs a malicious model for model merging by scaling the difference between a LoRA fine-tuned poisoned model and a benign model, achieving high backdoor attack success under low-resource assumptions.

  5. Rethinking Weight-Averaged Model-merging

    cs.LG 2024-11 conditional novelty 4.0 of 10

    Weight-averaged model merging is reinterpreted as template matching and implicit regularization, with systematic experiments showing logits ensembling generally outperforms weight averaging and ViTs degrade sharply un...

  6. Simplified Swarm Learning Framework for Robust and Scalable Diagnostic Services in Cancer Histopathology

    cs.DC 2025-04 reject novelty 3.0 of 10

    A blockchain-free peer-to-peer swarm learning framework shows modest accuracy on a three-class histopathology task, but missing details and weak baselines undermine the claim of parity with centralized models.

Pith tools