Downloads · 30 days
17
11% of all-time downloads
MatteoKhan/Mistral-LLaMA-Fusion
Mistral-LLaMA-Fusion is a text generation model from MatteoKhan. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Mistral-LLaMA-Fusion is an experimental merged language model combining the strengths of Mistral-7B-v0.1 and LLaMA-2-7B using the Linear Merge method via MergeKit. This hybrid model aims to balance Mistral’s efficienc…
Downloads · 30 days
17
11% of all-time downloads
All-time downloads
159
Public
Parameters
7.2B
14.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors14.5 GB · 100%
From the Hugging Face model README
Mistral-LLaMA-Fusion is an experimental merged language model combining the strengths of Mistral-7B-v0.1 and LLaMA-2-7B using the Linear Merge method via MergeKit. This hybrid model aims to balance Mistral’s efficiency and architecture with LLaMA-2’s robustness in reasoning and instruction following.
🔗 Created by: [Matteo Khan]
🎓 Affiliation: Apprentice at TW3 Partners (Generative AI Research)
📍 License: MIT
🔗 Connect on LinkedIn
🔗 Model on Hugging Face
This model is suited for research in model merging and hybridization, and can be used for:
As with all merged models, this fusion may inherit and combine weaknesses from both parents:
merge_method: linear
dtype: float16
models:
- model: mistralai/Mistral-7B-v0.1
parameters:
t: 1.0
weight: 0.6
- model: meta-llama/Llama-2-7b-hf
parameters:
t: 1.0
weight: 0.4
parameters:
normalize: true
int8_mask: false
layers:
- pattern: "model.*"
📌 Note: No additional fine-tuning was performed. This is a straight merge using MergeKit.
🌱 Why Merging?
Merging allows rapid experimentation with existing checkpoints while reducing the computational cost and carbon footprint compared to training from scratch.
🚀 How to Use
python
Copier
Modifier
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "MatteoKhan/Mistral-LLaMA-Fusion"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto")
prompt = "Explain the benefits of merging language models."
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_length=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))