Downloads · 30 days
18
8% of all-time downloads
doubleblind/DeepSeek-R1-Distill-QweNSA-1.5B
DeepSeek-R1-Distill-QweNSA-1.5B is a machine learning model from doubleblind. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository contains remote code and weights for a Native Sparse Attention distillation of DeepSeek-R1-Distill-Qwen-1.5B, distilled on mathematical reasoning data. Our parameter naming scheme refers to the paramet…
Downloads · 30 days
18
8% of all-time downloads
All-time downloads
222
Public
Parameters
2.7B
11 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors11 GB · 100%
From the Hugging Face model README
This repository contains remote code and weights for a Native Sparse Attention distillation of DeepSeek-R1-Distill-Qwen-1.5B, distilled on mathematical reasoning data. Our parameter naming scheme refers to the parameter count of the teacher model
To use this model, please ensure the following dependencies are installed:
pip install git+https://github.com/fnite1604/native-sparse-attention-pytorch.git
pip install transformers torch ...
Note: We recommend using the latest stable of Pytorch (currently 2.7.0) with CUDA 12.6 and the latest available version of Transformers
A quick_start.py script is included to help you get started with inference:
python quick_start.py
This will load the model and generate text based on a predefined prompt ("What is 1 + 1?") using our Native Sparse Attention enabled reasoning model.