Downloads · 30 days
115
22% of all-time downloads
cyboghostginx/Llama3.1-8B-mlx-4Bit
Llama3.1-8B-mlx-4Bit is a text generation model from cyboghostginx. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.1.
Built with Llama. 4-bit MLX conversion of Meta's Llama 3.1 8B base model for Apple silicon. 4.5 GB, single shard, converted with mlx-lm 0.26.4 from cyboghostginx/Llama3.1-8B. Weights only, no fine-tuning.
Downloads · 30 days
115
22% of all-time downloads
All-time downloads
527
Public
Parameters
8B
4.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.5 GB · 100%
How the weights are stored.
U328B · 100%
From the Hugging Face model README
Built with Llama. 4-bit MLX conversion of Meta's Llama 3.1 8B base model for Apple silicon. 4.5 GB, single shard, converted with mlx-lm 0.26.4 from cyboghostginx/Llama3.1-8B. Weights only, no fine-tuning.
This is the base model, not Instruct. It ships no chat template, so it completes text rather than answering turns. For chat, quantize an Instruct checkpoint instead.
pip install mlx-lm
mlx_lm.generate --model cyboghostginx/Llama3.1-8B-mlx-4Bit \
--prompt "The three laws of robotics are" --max-tokens 256
from mlx_lm import load, generate
model, tokenizer = load("cyboghostginx/Llama3.1-8B-mlx-4Bit")
print(generate(model, tokenizer, prompt="The three laws of robotics are", verbose=True))
Higher precision: 8-bit MLX, 8.5 GB.
Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved. The full agreement is reproduced in the gate above.