Downloads · 30 days
1.2K
10% of all-time downloads
spicyneuron/Kimi-K2.7-Code-MLX-3.6bit
Kimi-K2.7-Code-MLX-3.6bit is a text generation model from spicyneuron. Use it when you need the model to write or continue text. It is set up for mlx.
moonshotai/Kimi-K2.7-Code optimized for running on a Mac Studio M3 Ultra.
Downloads · 30 days
1.2K
10% of all-time downloads
All-time downloads
12.2K
Public
Parameters
1T
459 GB on disk
Likes
9
Public
Click a slice to open those files.
.safetensors459 GB · 100%
How the weights are stored.
U321T · 100%
From the Hugging Face model README
moonshotai/Kimi-K2.7-Code optimized for running on a Mac Studio M3 Ultra.
# Start server at http://localhost:8080/v1/chat/completions
uvx --from mlx-lm mlx_lm.server \
--host 127.0.0.1 \
--port 8080 \
--model spicyneuron/Kimi-K2.7-Code-MLX-3.6bit
| metric | this model |
|---|---|
| bpw | 3.578 |
| base memory | 427.579 |
| peak memory (1024/512) | 460.444 |
| prompt tok/s (1024) | 218.851 ± 0.208 |
| gen tok/s (512) | 21.035 ± 0.049 |
| perplexity | 4.462 ± 0.037 |
| arc_challenge | 0.692 ± 0.021 |
| hellaswag | 0.780 ± 0.019 |
Quantized with a mlx-lm fork. MLX quantization options differ than llama.cpp, but the principles are the same: