Downloads · 30 days
99
5% of all-time downloads
NexVeridian/GLM-4.7-Flash-3bit
GLM-4.7-Flash-3bit is a text generation model from NexVeridian. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as mit.
This model NexVeridian/GLM-4.7-Flash-3bit was converted to MLX format from zai-org/GLM-4.7-Flash using mlx-lm version 0.30.6.
Downloads · 30 days
99
5% of all-time downloads
All-time downloads
1.8K
Public
Parameters
29.9B
26.3 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors13.1 GB · 100%
How the weights are stored.
U3229.9B · 100%
From the Hugging Face model README
This model NexVeridian/GLM-4.7-Flash-3bit was converted to MLX format from zai-org/GLM-4.7-Flash using mlx-lm version 0.30.6.
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("NexVeridian/GLM-4.7-Flash-3bit")
prompt = "hello"
if tokenizer.chat_template is not None:
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_dict=False,
)
response = generate(model, tokenizer, prompt=prompt, verbose=True)