Published January 1, 2025 | Version v1
Journal article Open

KAFA-Merge: Model Merging Method With Layer-Based Bayes plus Linear Search

  • 1. Yildiz Tech Univ, Dept Comp Engn, TR-34220 Istanbul, Turkiye

Description

In this paper, we consider the model merging process for large language models (LLMs) under a two-stage optimization framework. Traditional merging methods usually apply fixed blending rates to all layers, which ignores the different levels of contribution of layers to the output. To overcome this limitation, we first group layers with similar functionality and use Bayesian optimization to determine the optimal blending rates for each group. Bayesian optimization offers a globally effective discovery strategy in high-dimensional and computationally expensive search spaces. However, an approach based solely on global search can lead to local artifacts being missed. Therefore, in the second stage, we perform Linear Search by applying controlled variations around the ratios obtained with Bayesian optimization and obtain more precise, locally optimized results. Extensive experiments on the Turkish versions of the Arc, Hellaswag, MMLU and GSM8K datasets show that our proposed Bayes+Linear strategy outperforms existing methods such as SLERP, Linear, Ties and Breadcrumbs. In addition to improving model accuracy and generalization capacity, the KAFA approach works without any additional fine-tuning, making it possible to effectively reuse large language models in different tasks.

Files

bib-c2c5afc1-47e0-42dd-b5be-e8db1f851789.txt

Files (138 Bytes)

Name Size Download all
md5:78e263bf08f90d979f1bcd9fdaee3bf0
138 Bytes Preview Download