mergekit: Combining Language Models Without Retraining
Tools for merging pretrained large language models.
At a glance
- What is it?
- mergekit lets you blend trained language models in weight space to create specialized variants without needing training data or the computational cost of ensemble inference. The toolkit supports many merge methods and runs on CPU or commodity GPUs.
- Who is it for?
- Use mergekit if you have multiple pretrained models and want to combine their strengths or create specialized variants without access to training data. Do not use it if you need to train models on new data or if your models use incompatible architectures.
- Can I use it commercially?
- Yes, with conditions. LGPL-3.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 18 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Why merge instead of ensemble or retrain
Training a language model from scratch requires weeks of compute on large GPU clusters and access to training data. Ensemble inference runs multiple models in parallel, multiplying inference latency and cost. Model merging operates directly in weight space to combine two or more models into a single artifact that keeps the inference cost of one model but draws from the strengths of all inputs.
Mergekit specifically targets cases where training is unavailable or unaffordable. The README states that merging can combine multiple specialized models into a single versatile model, transfer capabilities between models without access to training data, and find optimal trade-offs between different model behaviors. The merged result maintains the same inference speed as a single model while often matching or exceeding the performance of ensembles.
Mergekit-yaml merges models from YAML config files
mergekit requires Python 3.10 or later. Install it with pip in development mode so the command-line tools become available:
git clone https://github.com/arcee-ai/mergekit.git
cd mergekit
pip install -e .If the install fails with an error about setup.py missing, upgrade pip first with `python3 -m pip install --upgrade pip`.
The main entry point is `mergekit-yaml`. It takes a YAML configuration file and an output directory:
mergekit-yaml path/to/your/config.yml ./output-model-directory [--cuda] [--lazy-unpickle] [--allow-crimes] [... other options]The configuration file specifies which models to merge, which merge method to use, and the parameter weights for each input. The README provides examples in the examples/ directory: linear merges with specific weight distributions, SLERP-based merges for spherical interpolation, and TIES merges that resolve conflicts between model weights. Run `mergekit-yaml --help` for the full list of arguments.
How mergekit combines models
mergekit merges models by operating on their weight tensors rather than on activations or hidden states. The core operation takes a list of tensors, a merge method, and method-specific parameters, then produces a single output tensor.
For example, the linear method takes a list of weights and a list of interpolation coefficients; the output is the weighted sum of inputs. The SLERP (spherical linear interpolation) method interpolates on the surface of a high-dimensional sphere. TIES resolves conflicts when merging weights: if multiple input models assign very different values to the same weight, TIES drops the outliers and averages the remaining values.
mergekit supports Llama, Mistral, GPT-NeoX, and StableLM architectures. The tool also handles specialized merging: LoRA extraction breaks trained LoRA adapters back into the model weights, Mixture of Experts merging combines expert networks from different models, and multi-stage merging chains multiple merge operations together. The in-memory Python API exposes `merge_tensors()` for individual tensors and `merge_state_dicts()` for entire model state dictionaries, so you can integrate merging into Python code without the YAML interface.
Resource constraints and the lazy-loading approach
A typical 70 billion parameter model requires roughly 140 GB of RAM if stored as float32 weights. Loading three such models for merging would demand over 400 GB of memory, far beyond most hardware. mergekit uses an out-of-core approach with lazy loading: it loads tensor slices on demand rather than the entire model at once.
The README states that merges can run entirely on CPU or be accelerated with as little as 8 GB of VRAM. The `--lazy-unpickle` flag enables this behavior by deferring tensor loading until the merge actually needs them. This approach lets you merge large models on machines that would otherwise run out of memory.
Piecewise layer assembly and Frankenmerging
mergekit supports a technique called Frankenmerging: piecewise assembly of language models from different layer selections. Rather than interpolating all weights uniformly, you can specify that layers 0-20 come from Model A, layers 21-40 from Model B, and layers 41-80 from Model C. This creates a model with no training but with intentionally mismatched internal representations across layer boundaries.
The approach works because transformer models compartmentalize somewhat: early layers often learn tokenization and syntax, middle layers learn semantic features, and late layers specialize on the task. Frankenmerging can succeed or fail depending on the architecture and the specific models; the README does not provide success rates or guidance on predicting outcomes.
When merging fails
Merging requires all input models to have the same architecture, tokenizer, and embedding dimensions. If you attempt to merge a Llama model with a Mistral model, the operation will fail because their layer configurations differ. The token count must also match: a model with a 32,000-token vocabulary cannot merge with one using 128,000 tokens.
The README does not document how to detect these incompatibilities before running a merge, nor does it provide recovery procedures once a merge fails partway through. Large merges can run for hours, so an incompatibility discovered late in the process means lost compute time.
Uploading merged models to Hugging Face Hub
Once a merge completes successfully, mergekit generates a README.md file for the merged model with basic information for a model card. You can edit this file to include more details about your merge, such as explaining what the merged model is good at or giving it a descriptive name; rewrite it entirely; or use the generated card as-is. It is also possible to edit the README online after uploading to the Hub.
To share a merged model, use the huggingface_hub Python library. First, log in with `huggingface-cli login` using an access token that has write permission. Then upload with `huggingface-cli upload your_hf_username/my-cool-model ./output-model-directory .`. This command uploads the merged model directory to your Hugging Face Hub namespace. The huggingface_hub documentation covers additional upload options and configurations.
Multi-stage merging and PyTorch-level merging
mergekit supports multi-stage merging through the `mergekit-multi` command, enabling complex workflows that chain multiple merge operations together. This is useful when you need to merge three models but want to merge the first two, then merge the result with the third, rather than attempting a three-way merge directly. Each stage produces an intermediate model that feeds into the next stage.
For users needing to merge raw PyTorch models outside the configuration-file approach, `mergekit-pytorch` exposes lower-level primitives. Similarly, `mergekit-tokensurgeon` handles tokenizer transplantation, allowing you to swap the tokenizer of one model with that of another while preserving weights.
Evolutionary merge methods are also available through `mergekit-evolve`, which uses techniques like genetic algorithms to discover optimal merge parameters rather than requiring manual specification. These tools extend mergekit beyond simple two-or-three-way merges into larger, more flexible workflows.
Community tools and alternatives
FrankensteinAI provides a browser-based merging interface powered by mergekit, removing the need for local setup. The hosted platform includes a community gallery where users share merged models and a leaderboard for comparing results.
Llama.cpp can load some merged models if they use supported architectures. Ollama serves models over HTTP. Neither tool merges models; they run them. For merging specifically, mergekit is the primary open-source option documented in this material.
Editorial conclusion
Use mergekit if you have multiple pretrained models and want to combine their strengths or create specialized variants without access to training data. Do not use it if you need to train models on new data or if your models use incompatible architectures. Before starting, verify that all source models use the same tokenizer and architecture; mismatches will fail at merge time.
Frequently asked questions
What is model merging for LLMs?
Model merging combines multiple trained language models into a single model by operating on their weight tensors. The result preserves the inference cost of a single model while drawing from the capabilities of all inputs.
How do I use mergekit?
Write a YAML configuration file specifying the models, merge method, and parameter weights, then run `mergekit-yaml path/to/config.yml ./output`. The tool supports linear, SLERP, TIES, and other merge methods. A Python API is available for integration into larger workflows.
What is Frankenmerging?
Frankenmerging assembles a model by taking different layer ranges from different input models. You can specify that layers 0-20 come from Model A and layers 21-80 from Model B, creating a model with no training but with intentionally mixed internal representations.
Can I merge models on CPU?
Yes. mergekit uses lazy loading and out-of-core techniques to defer tensor loading until the merge needs them. Merges can run entirely on CPU or accelerated with as little as 8 GB of VRAM.
What models can mergekit merge?
mergekit supports Llama, Mistral, GPT-NeoX, and StableLM architectures. All input models must use the same architecture, tokenizer, and embedding dimensions; incompatible models will fail at merge time.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/arcee-ai-mergekit)