Downloads · 30 days
20
15% of all-time downloads
MatteoKhan/MistralGemma-7B-Merged
MistralGemma-7B-Merged is a text generation model from MatteoKhan. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
MistralGemma-Hybrid-7B is an experimental hybrid language model that blends the strengths of Mistral-7B and Gemma-7B using the Spherical Linear Interpolation (slerp) merging technique. Designed to optimize both effici…
Downloads · 30 days
20
15% of all-time downloads
All-time downloads
131
Public
Parameters
7.2B
14.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors14.5 GB · 100%
From the Hugging Face model README
MistralGemma-Hybrid-7B is an experimental hybrid language model that blends the strengths of Mistral-7B and Gemma-7B using the Spherical Linear Interpolation (slerp) merging technique. Designed to optimize both efficiency and performance, this model offers robust text generation capabilities while leveraging the advantages of both parent models.
🔗 Created by: [Matteo Khan]
🎓 Affiliation: Apprentice at TW3 Partners (Generative AI Research)
📍 License: MIT
🔗 Connect with me on LinkedIn
🔗 Model on Hugging Face
This model is intended for research and experimentation in hybrid model optimization. Potential applications include:
While MistralGemma-Hybrid-7B offers enhanced capabilities, it also inherits limitations from its parent models:
This is not a newly trained model, but rather a merge of existing models using the following configuration:
merge_method: slerp # Using slerp instead of linear
dtype: float16
models:
- model: "mistralai/Mistral-7B-v0.1"
parameters:
weight: 0.5
- model: "google/gemma-7b"
parameters:
weight: 0.5
parameters:
normalize: true
int8_mask: false
rescale: true # Helps with different model scales
layers:
- pattern: ".*"
layer_range: [0, -1]
📊 No formal evaluation has been conducted yet. Users are encouraged to benchmark and share feedback!
By utilizing model merging rather than training from scratch, MistralGemma-Hybrid-7B significantly reduces computational and environmental costs.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "YourProfile/MistralGemma-Hybrid-7B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
# Example usage
prompt = "Write a short story about the future of AI."
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_length=200)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
📝 Citation
@misc{mistralgemma2025,
title={MistralGemma: A Hybrid Open-Source Language Model},
author={Your Name},
year={2025},
eprint={arXiv:XXXX.XXXXX},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
📩 Feedback & Contact: Reach out via Hugging Face.
🎉 Happy Experimenting! 🚀