Downloads · 30 days
47
9% of all-time downloads
gittyeric/MiniMax-M2.7-oQ8e
MiniMax-M2.7-oQ8e is a text generation model from gittyeric. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as other.
Even as of July 2026 on M3 Mac 512Gb RAM I can't seem to find something better than Minimax 2.7, so here's the OMLX 8-bit oQ8e optimized quant, on paper it should be better than a regular 8-bit quant available elsewhe…
Downloads · 30 days
47
9% of all-time downloads
All-time downloads
507
Public
Parameters
229B
486 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors243 GB · 100%
How the weights are stored.
U32229B · 100%
From the Hugging Face model README
Even as of July 2026 on M3 Mac 512Gb RAM I can't seem to find something better than Minimax 2.7, so here's the OMLX 8-bit oQ8e optimized quant, on paper it should be better than a regular 8-bit quant available elsewhere on HF.
The optimized 8-bit quant runs quite fast and with turboquant you can pretty easily fit other models on a 512Gb or at least get a huge KV cache size while gaining up to 20 TPS vs 15 on the unquantized model; the difference between annoying and quite practical.
Below is the original base model card...
This model mlx-community/MiniMax-M2.7 was converted to MLX format from MiniMaxAI/MiniMax-M2.7 using mlx-lm version 0.31.3.
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("mlx-community/MiniMax-M2.7")
prompt = "hello"
if tokenizer.chat_template is not None:
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_dict=False,
)
response = generate(model, tokenizer, prompt=prompt, verbose=True)