Pith. sign in

REVIEW 9 cited by

Evolutionary Optimization of Model Merging Recipes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.13187 v2 pith:EYKK2WKN submitted 2024-03-19 cs.NE

classification cs.NE
keywords modelsjapaneseapproachmodelmergingdatadevelopmenteven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have become increasingly capable, but their development often requires substantial computational resources. While model merging has emerged as a cost-effective promising approach for creating new models by combining existing ones, it currently relies on human intuition and domain knowledge, limiting its potential. Here, we propose an evolutionary approach that overcomes this limitation by automatically discovering effective combinations of diverse open-source models, harnessing their collective intelligence without requiring extensive additional training data or compute. Our approach operates in both parameter space and data flow space, allowing for optimization beyond just the weights of the individual models. This approach even facilitates cross-domain merging, generating models like a Japanese LLM with Math reasoning capabilities. Surprisingly, our Japanese Math LLM achieved state-of-the-art performance on a variety of established Japanese LLM benchmarks, even surpassing models with significantly more parameters, despite not being explicitly trained for such tasks. Furthermore, a culturally-aware Japanese VLM generated through our approach demonstrates its effectiveness in describing Japanese culture-specific content, outperforming previous Japanese VLMs. This work not only contributes new state-of-the-art models back to the open-source community, but also introduces a new paradigm for automated model composition, paving the way for exploring alternative, efficient approaches to foundation model development.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FlexOlmo: Open Language Models for Flexible Data Use

    cs.CL 2025-07 conditional novelty 7.0 of 10

    FlexOlmo merges independently trained language-model experts, trained on private data, into a single mixture-of-experts model without joint training.

  2. SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text

    cs.AI 2026-05 conditional novelty 6.0 of 10

    A 24-dataset benchmark for inducing schema graphs from raw text, plus an auditable LLM-based pipeline that reports the highest scores on the benchmark's four schema-similarity metrics.

  3. Competition and Attraction Improve Model Fusion

    cs.AI 2025-08 conditional novelty 6.0 of 10

    M2N2 evolves merging boundaries, uses resource competition for diversity and attraction-based pairing, achieving from-scratch evolution and specialized model fusion.

  4. GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching

    cs.CL 2025-06 conditional novelty 6.0 of 10

    GPTailor searches over layer removal, layer selection, and layer merging across fine-tuned model variants to produce smaller LLMs that retain more benchmark performance than single-model pruning.

  5. Training-free LLM Merging for Multi-task Learning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Hi-Merging merges two task-specialized LLMs by pruning and scaling delta vectors at model and layer level, reporting gains over prior merging and multi-task fine-tuning on English and Chinese MCQA and QA tasks.

  6. Tunable MAGMAX: Preference-Aware Model Merging for Continual Learning

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Tunable MAGMAX gives each task an element budget during model merging, letting users shift task-wise accuracy and auto-setting the budget from target-environment label/feature similarity.

  7. PSO-Merging: Merging Models Based on Particle Swarm Optimization

    cs.LG 2025-08 conditional novelty 5.0 of 10

    PSO-Merging applies particle swarm optimization over model weight space, seeded with original and sparsified experts, to build multitask models that outperform existing merging baselines on several language benchmarks.

  8. Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Sparse adapters trained with max connection sensitivity outperform LoRA and full fine-tuning both alone and after merging 20 task experts, but still lag multitask training on unseen tasks.

  9. Contrasting Cognitive Styles in Vision-Language Models: Holistic Attention in Japanese Versus Analytical Focus in English

    cs.CL 2025-07 reject novelty 5.0 of 10

    Japanese-prompted vision-language models produce more background-first captions than English-prompted ones, but the effect is confounded by the evaluator and by language grammar.

Pith tools