Downloads · 30 days
16
17% of all-time downloads
rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32
FastContext-1.0-4B-SFT-mlx-4bit-g32 is a text generation model from rubybear. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as mit.
4-bit MLX quantization of microsoft/FastContext-1.0-4B-SFT with groupsize=32 for Apple Silicon.
Downloads · 30 days
16
17% of all-time downloads
All-time downloads
94
Public
Parameters
4B
2.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.5 GB · 100%
How the weights are stored.
U324B · 100%
From the Hugging Face model README
4-bit MLX quantization of microsoft/FastContext-1.0-4B-SFT with group_size=32 for Apple Silicon.
Tested on 10 SWE-bench Multilingual instances against other quantization variants:
| Model | Bits/Wt | Size | File F1 | Line F1 |
|---|---|---|---|---|
| affine 8-bit g64 | 8.5 | 4.0G | 0.507 | 0.140 |
| affine 4-bit g32 (this model) | 5.0 | 2.4G | 0.300 | 0.090 |
| affine 3-bit g64 | 3.5 | 1.7G | 0.100 | 0.000 |
| affine 4-bit g64 | 4.5 | 2.1G | 0.050 | 0.005 |
| mattrobenolt 4-bit g64 | 4.5 | 2.1G | 0.025 | 0.008 |
The finer group_size=32 delivers 12x better File F1 than standard 4-bit g64 quantization with only 300MB additional size.
from mlx_lm import load, generate
model, tokenizer = load("rubybear-lgtm/FastContext-1.0-4B-SFT-mlx-4bit-g32")
Or with fastcontext-mcp for Claude Code integration.