Downloads · 30 days
85
100% of all-time downloads
Zexiry/Zera-24B-4bit
Zera-24B-4bit is a text generation model from Zexiry. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
Zera is a 24B-parameter, 4-bit MLX language model fine-tuned for coding, debugging, technical explanations, general conversation, grammar, and vocabulary. It is a standalone fused model: users do not need a separate a…
Downloads · 30 days
85
100% of all-time downloads
All-time downloads
85
Public
Parameters
23.6B
13.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors13.3 GB · 100%
How the weights are stored.
U3223.6B · 100%
From the Hugging Face model README
Zera is a 24B-parameter, 4-bit MLX language model fine-tuned for coding, debugging, technical explanations, general conversation, grammar, and vocabulary. It is a standalone fused model: users do not need a separate adapter.
pip install mlx-lm
from mlx_lm import generate, load
model, tokenizer = load("Zexiry/Zera-24B-4bit")
messages = [{"role": "user", "content": "Introduce yourself, then write a Python trie."}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
print(generate(model, tokenizer, prompt=prompt, max_tokens=800, verbose=False))
Zera's default chat template supplies its assistant identity when an application does not provide a system prompt. Applications can still provide their own system prompt normally.
The base_model metadata above is retained for reproducibility, attribution, and
license compliance. In conversation, the assistant identity is Zera.