Downloads · 30 days
37
10% of all-time downloads
ykarout/Mixtral-8x7B-DeepSeek-R1-Distill
Mixtral-8x7B-DeepSeek-R1-Distill is a text generation model from ykarout. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A reasoning-enhanced version of Mixtral-8x7B-Instruct-v0.1, fine-tuned on reasoning responses generated by DeepSeek's reasoning model.
Downloads · 30 days
37
10% of all-time downloads
All-time downloads
370
Public
Parameters
46.7B
93.4 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors93.4 GB · 100%
From the Hugging Face model README
A reasoning-enhanced version of Mixtral-8x7B-Instruct-v0.1, fine-tuned on reasoning responses generated by DeepSeek's reasoning model.
This model is a fine-tuned version of Mixtral-8x7B-Instruct-v0.1 that has been trained on reasoning-rich datasets to improve its step-by-step thinking and problem-solving capabilities. The model learns to generate explicit reasoning traces similar to those produced by advanced reasoning models like DeepSeek-R1.
This model is designed for tasks requiring explicit reasoning and step-by-step problem solving, including:
The model can be further fine-tuned for domain-specific reasoning tasks or integrated into applications requiring transparent AI reasoning processes.
Users should validate reasoning outputs, especially for critical applications. The model works best when prompted to "think step by step" or "show your reasoning."
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
tokenizer = AutoTokenizer.from_pretrained("ykarout/Mixtral-8x7B-DeepSeek-R1-Distill-16bit")
model = AutoModelForCausalLM.from_pretrained(
"ykarout/Mixtral-8x7B-DeepSeek-R1-Distill-16bit",
torch_dtype=torch.bfloat16,
device_map="auto"
)
# Example reasoning prompt
prompt = """<s>[INST] Solve this step by step: If a train travels 120 km in 2 hours, and then 180 km in 3 hours, what is its average speed for the entire journey? [/INST]"""
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(
**inputs,
max_new_tokens=512,
temperature=0.7,
do_sample=True
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
The model was fine-tuned on the open-r1/Mixture-of-Thoughts dataset, which contains reasoning responses generated by DeepSeek's reasoning model across various domains including mathematics, science, coding, and logical reasoning.
Evaluation pending on standard reasoning benchmarks including:
Training Metrics:
Comprehensive evaluation results on reasoning benchmarks will be updated post-training completion.
The model exhibits improved reasoning capabilities compared to the base Mixtral model, generating explicit step-by-step thinking processes. Analysis of attention patterns and reasoning trace quality is ongoing.
Estimated Training Impact:
BibTeX:
@model{mixtral-deepseek-r1-distill,
title={Mixtral-8x7B-DeepSeek-R1-Distill: Reasoning-Enhanced Mixture of Experts},
author={ykarout},
year={2025},
publisher={Hugging Face},
url={https://huggingface.co/ykarout/Mixtral-8x7B-DeepSeek-R1-Distill-16bit}
}
For questions or issues, please contact through Hugging Face