Downloads · 30 days
21
5% of all-time downloads
schaye/llama-3.2-1b-instruct-mlx-quantized
llama-3.2-1b-instruct-mlx-quantized is a text generation model from schaye. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.2.
The Model schaye/llama-3.2-1b-instruct-mlx-quantized was converted to MLX format from meta-llama/Llama-3.2-1B-Instruct using mlx-lm version 0.19.0.
Downloads · 30 days
21
5% of all-time downloads
All-time downloads
440
Public
Parameters
1.2B
695 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors695 MB · 99%
How the weights are stored.
U321.2B · 100%
From the Hugging Face model README
The Model schaye/llama-3.2-1b-instruct-mlx-quantized was converted to MLX format from meta-llama/Llama-3.2-1B-Instruct using mlx-lm version 0.19.0.
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("schaye/llama-3.2-1b-instruct-mlx-quantized")
prompt="hello"
if hasattr(tokenizer, "apply_chat_template") and tokenizer.chat_template is not None:
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
response = generate(model, tokenizer, prompt=prompt, verbose=True)