Downloads · 30 days
21
2% of all-time downloads
AlexGS74/MiniMax-M2.1-REAP-50-mlx-4bit
MiniMax-M2.1-REAP-50-mlx-4bit is a text generation model from AlexGS74. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
This is a 4-bit quantized version of the MiniMax REAP-50 model optimized for Apple Silicon using MLX.
Downloads · 30 days
21
2% of all-time downloads
All-time downloads
1.4K
Public
Parameters
116B
65.5 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors65.4 GB · 100%
How the weights are stored.
U32116B · 100%
From the Hugging Face model README
This is a 4-bit quantized version of the MiniMax REAP-50 model optimized for Apple Silicon using MLX.
from mlx_lm import load, generate
model, tokenizer = load("minimax-reap50-mlx-4bit")
response = generate(
model,
tokenizer,
prompt="Write a function to calculate fibonacci numbers",
max_tokens=500,
verbose=True
)
print(response)
Start the server:
mlx_lm.server --model minimax-reap50-mlx-4bit --port 8080
Make requests:
curl -X POST http://localhost:8080/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "default_model",
"prompt": "Write a function to calculate fibonacci numbers",
"max_tokens": 500
}'
Or use the chat endpoint:
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "default_model",
"messages": [
{"role": "user", "content": "Write a function to calculate fibonacci numbers"}
],
"max_tokens": 500
}'