Downloads · 30 days
19
4% of all-time downloads
emredeveloper/DeepSeek-R1-Distill-Qwen-1.5B-4bit
DeepSeek-R1-Distill-Qwen-1.5B-4bit is a machine learning model from emredeveloper. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This is a 4-bit quantized version of the deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B model, optimized for efficient inference with reduced memory usage. The quantization was performed using the bitsandbytes library.
Downloads · 30 days
19
4% of all-time downloads
All-time downloads
474
Public
Parameters
1.8B
1.6 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors1.6 GB · 99%
How the weights are stored.
U81.4B · 74%
From the Hugging Face model README
This is a 4-bit quantized version of the deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B model, optimized for efficient inference with reduced memory usage. The quantization was performed using the bitsandbytes library.
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5BThis model is intended for research and practical applications where memory efficiency is critical. It can be used for:
This model can be fine-tuned for specific tasks such as:
This model is not suitable for:
The model may inherit biases present in the training data. Users should be cautious when deploying the model in sensitive applications.
Users should evaluate the model's performance on their specific tasks and datasets before deployment. Consider fine-tuning the model for better alignment with your use case.
Use the code below to get started with the model:
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
import torch
# Quantization configuration
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True
)
# Load the model and tokenizer
tokenizer = AutoTokenizer.from_pretrained("emredeveloper/DeepSeek-R1-Distill-Qwen-1.5B-4bit", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
"emredeveloper/DeepSeek-R1-Distill-Qwen-1.5B-4bit",
quantization_config=quantization_config,
device_map="auto",
trust_remote_code=True
)
# Generate text
input_text = "Hello, how are you?"
inputs = tokenizer(input_text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))