Downloads · 30 days
2
4% of all-time downloads
anthonymikinka/DeepSeek-R1-Distill-Llama-8B-Stateful-CoreML
DeepSeek-R1-Distill-Llama-8B-Stateful-CoreML is a text generation model from anthonymikinka. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This repository contains a CoreML conversion of the DeepSeek-R1-Distill-Llama-8B model optimized for Apple Silicon devices. This conversion features stateful key-value caching for efficient text generation.
Downloads · 30 days
2
4% of all-time downloads
All-time downloads
52
Public
Repo size
16.1 GB
Likes
3
Public
Click a slice to open those files.
.bin16.1 GB · 100%
From the Hugging Face model README
This repository contains a CoreML conversion of the DeepSeek-R1-Distill-Llama-8B model optimized for Apple Silicon devices. This conversion features stateful key-value caching for efficient text generation.
DeepSeek-R1-Distill-Llama-8B is a distilled 8 billion parameter language model from the DeepSeek-AI team. The model is built on the Llama architecture and has been distilled to maintain performance while reducing the parameter count.
This CoreML conversion provides:
This conversion utilizes:
The model can be loaded and used with CoreML in your Swift or Python projects:
import coremltools as ct
# Load the model
model = ct.models.MLModel("DeepSeek-R1-Distill-Llama-8B.mlpackage")
# Prepare inputs for inference
# ...
# Run inference
output = model.predict({
"inputIds": input_ids,
"causalMask": causal_mask
})
The model was converted using CoreML Tools with the following steps:
To use this model:
This model conversion inherits the license of the original DeepSeek-R1-Distill-Llama-8B model.
If you use this model in your research, please cite both the original DeepSeek model and this conversion.