Downloads · 30 days
333
100% of all-time downloads
tokenfires/Qwen2.5-Coder-3B-Instruct-MLX-4bit
Qwen2.5-Coder-3B-Instruct-MLX-4bit is a text generation model from tokenfires. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
4-bit MLX quantization of Qwen/Qwen2.5-Coder-3B-Instruct for Apple Silicon. Quantization: affine, 4 bits, group size 64.
Downloads · 30 days
333
100% of all-time downloads
All-time downloads
333
Public
Parameters
3.1B
1.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.7 GB · 99%
How the weights are stored.
U323.1B · 100%
From the Hugging Face model README
4-bit MLX quantization of Qwen/Qwen2.5-Coder-3B-Instruct for Apple Silicon. Quantization: affine, 4 bits, group size 64.
Search for tokenfires/Qwen2.5-Coder-3B-Instruct-MLX-4bit in the LM Studio model downloader, or open this page and choose Use this model → LM Studio.
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("tokenfires/Qwen2.5-Coder-3B-Instruct-MLX-4bit")
messages = [{"role": "user", "content": "Write a Ruby method that reverses a string."}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
response = generate(model, tokenizer, prompt=prompt, verbose=True)