Downloads · 30 days
17
5% of all-time downloads
Rafii/f1llama
f1llama is a text generation model from Rafii. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as other.
The Model Rafii/f1llama was converted to MLX format from mlx-community/Meta-Llama-3-8B-Instruct-8bit using mlx-lm version 0.20.1 and Finetuned using Low-rank adaptation (LoRA) on M3.
Downloads · 30 days
17
5% of all-time downloads
All-time downloads
329
Public
Parameters
8B
9.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9 GB · 100%
How the weights are stored.
U327.5B · 93%
From the Hugging Face model README
The Model Rafii/f1llama was converted to MLX format from mlx-community/Meta-Llama-3-8B-Instruct-8bit using mlx-lm version 0.20.1 and Finetuned using Low-rank adaptation (LoRA) on M3.
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("Rafii/f1llama")
prompt="hello"
if hasattr(tokenizer, "apply_chat_template") and tokenizer.chat_template is not None:
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
response = generate(model, tokenizer, prompt=prompt, verbose=True)