Downloads · 30 days
136
0% of all-time downloads
fno2010/MiniMax-M2.7-TQ3
MiniMax-M2.7-TQ3 is a machine learning model from fno2010. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
A TurboQuant 3-bit quantized version of MiniMax-M2.7, optimized for inference with turboquant-vllm.
Downloads · 30 days
136
0% of all-time downloads
All-time downloads
40.9K
Public
Parameters
88.3B
94.9 GB on disk
Likes
14
Public
Click a slice to open those files.
.safetensors94.9 GB · 100%
How the weights are stored.
U885.3B · 97%
From the Hugging Face model README
A TurboQuant 3-bit quantized version of MiniMax-M2.7, optimized for inference with turboquant-vllm.
This quantized model is designed to work with the turboquant-vllm inference engine. Please refer to the turboquant-vllm repository for installation and usage instructions.
# Please refer to turboquant-vllm for proper model loading
The model uses a Jinja chat template with support for:
<minimax:tool_call> / </minimax:tool_call> delimiters)<think> / </minimax:tool_call> delimiters)The default model identity is: "You are a helpful assistant. Your name is MiniMax-M2.7 and is built by MiniMax."
This is a 3-bit quantized checkpoint intended for efficient inference. The quantization was applied using the TurboQuant method via the turboquant-vllm project.
This is a third-party quantized version of the original MiniMax-M2.7 model. Please refer to the original model card for base model details and licensing.