Downloads · 30 days
26
5% of all-time downloads
gbueno86/QwQ-R1-Distill-Merge-32B
QwQ-R1-Distill-Merge-32B is a text generation model from gbueno86. Use it when you need the model to write or continue text. It is set up for transformers.
Testing locally it behaved very well for math problems. It usually starts a problem without the <think tag, but ends by closing it when using chatml template.
Downloads · 30 days
26
5% of all-time downloads
All-time downloads
491
Public
Parameters
32.8B
65.5 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors65.5 GB · 100%
From the Hugging Face model README
Testing locally it behaved very well for math problems. It usually starts a problem without the <think> tag, but ends by closing it when using chatml template.
This is a merge of pre-trained language models created using mergekit.
This model was merged using the SLERP merge method.
The following models were included in the merge:
The following YAML configuration was used to produce this model:
base_model: /models/Qwen/QwQ-32B
dtype: bfloat16
merge_method: slerp
parameters:
t:
- filter: self_attn
value: [0.0, 0.5, 0.3, 0.7, 1.0]
- filter: mlp
value: [1.0, 0.5, 0.7, 0.3, 0.0]
- value: 0.5
slices:
- sources:
- layer_range: [0, 64]
model: /models/Qwen/QwQ-32B
- layer_range: [0, 64]
model: /models/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B