Downloads · 30 days
19
18% of all-time downloads
olusegunola/DeepSeek-R1-Distill-Merge-Qwen-Math-1.5Bb
DeepSeek-R1-Distill-Merge-Qwen-Math-1.5Bb is a machine learning model from olusegunola. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This model is a high-performance merge designed to bridge Mathematical Logic and Reasoning. It was constructed using MergeKit with the DARE-TIES method to preserve specialized weights from both source models.
Downloads · 30 days
19
18% of all-time downloads
All-time downloads
105
Public
Parameters
1.8B
7.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.6 GB · 100%
From the Hugging Face model README
This model is a high-performance merge designed to bridge Mathematical Logic and Reasoning. It was constructed using MergeKit with the DARE-TIES method to preserve specialized weights from both source models.
The merge integrates the logical foundations of Qwen2.5-Math with the distilled reasoning capabilities of DeepSeek-R1. This combination aims to improve accuracy in structured tasks such as USMLE-style Q&A and ICD-10 clinical coding.
dare_tiesQwen/Qwen2.5-1.5BQwen/Qwen2.5-Math-1.5B-Instruct (Weight: 0.5)DeepSeek-AI/DeepSeek-R1-Distill-Qwen-1.5B (Weight: 0.5)This model is intended for research purposes in the medical domain. It excels at tasks requiring Chain-of-Thought (CoT) explanation before providing a final medical answer.
If using this model for research, please cite the merge methodology and the source models accordingly.