Downloads · 30 days
65
21% of all-time downloads
Micklavin/Mistral-Nemo-Instruct-2407-4bit
Mistral-Nemo-Instruct-2407-4bit is a text generation model from Micklavin. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
A 4-bit quantization of mistralai/Mistral-Nemo-Instruct-2407 for MLX on Apple Silicon.
Downloads · 30 days
65
21% of all-time downloads
All-time downloads
309
Public
Parameters
12.2B
6.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6.9 GB · 100%
How the weights are stored.
U3212.2B · 100%
From the Hugging Face model README
A 4-bit quantization of mistralai/Mistral-Nemo-Instruct-2407 for MLX on Apple Silicon.
Produced for Triad, a local three-model token-fusion experiment. These are the exact weights that project was built and validated against, published so its results can be reproduced.
Size on disk: 6.4 GB (down from ~24.0 GB at bf16)
mlx-lm >= 0.31.3pip install "mlx-lm>=0.31.3"
The PyPI release is sufficient — this model's model_type: mistral is resolved via
mlx-lm's MODEL_REMAPPING (mistral → llama), which is present in the published
wheel. No extra packages needed.
from mlx_lm import generate, load
model, tokenizer = load("Micklavin/Mistral-Nemo-Instruct-2407-4bit")
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "What is the chemical symbol for gold?"}],
tokenize=False,
add_generation_prompt=True,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=256))
Loading this tokenizer emits a transformers warning stating the regex pattern is incorrect and "will lead to incorrect tokenization", with a suggestion to pass fix_mistral_regex=True.
That flag is honoured — it silences the warning — but for this checkpoint it changes nothing. Measured directly: with and without the flag, the pre-tokenizer serialises byte-identically and tokenization is unchanged across 24 strings spanning digits, unicode, contractions, whitespace, code and chemical formulae.
So the warning appears to be a false positive here. If you want it gone:
model, tokenizer = load(
"Micklavin/Mistral-Nemo-Instruct-2407-4bit",
tokenizer_config={"fix_mistral_regex": True},
)
This was verified on this checkpoint only. If you depend on exact tokenization, verify it yourself rather than taking the above on trust.
Converted with mlx_lm.convert:
convert(
hf_path="mistralai/Mistral-Nemo-Instruct-2407",
mlx_path="models/nemo-4bit",
quantize=True,
q_bits=4,
q_group_size=64,
dtype="bfloat16",
)
Resulting config: {"group_size": 64, "bits": 4, "mode": "affine"}, model_type: mistral (which mlx-lm remaps to llama).
| Component | Version |
|---|---|
| mlx | 0.32.0 |
| mlx-lm | 0.31.3 |
| transformers | 5.12.1 |
The reproduction script is scripts/quantize_models.py.
None beyond a smoke test. No benchmark comparison against the bf16 original was run, so the quantization's quality cost is unmeasured. Treat it as an untested 4-bit conversion rather than a validated one.
The upstream repository carries an extra_gated_description referring to Mistral AI's privacy policy and their processing of your personal data. That has been removed here because it describes Mistral's data collection, not this re-upload's — carrying it over would misrepresent who is collecting what. If you want the upstream terms, get the model from Mistral directly.
Apache 2.0, inherited from the base model. Quantization does not alter the license or your obligations under it.