Downloads · 30 days
287
30% of all-time downloads
mlx-community/humanizer-1B-OptiQ-4bit
humanizer-1B-OptiQ-4bit is a text generation model from mlx-community. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). Try the Lab · All OptiQ quants · Docs
Downloads · 30 days
287
30% of all-time downloads
All-time downloads
949
Public
Parameters
1.1B
1.1 GB on disk
Likes
6
Trending 1
Click a slice to open those files.
.safetensors1.1 GB · 99%
How the weights are stored.
U321.1B · 100%
From the Hugging Face model README
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). Try the Lab · All OptiQ quants · Docs
A 1B model that scores the same as the human reference set on the RADAR AI detector. Stacked SFT + DPO LoRA adapters on top of mlx-community/MiniCPM5-1B-OptiQ-4bit close 100% of the gap to human writing on a 200-draft held-out evaluation.
| P(AI) (RADAR-Vicuna-7B) | |
|---|---|
| Source AI drafts (Qwen3.5-4B + Gemma-4-e4b output) | 0.51 |
humanizer-1B-OptiQ-4bit (SFT + DPO stacked) | 0.37 |
| Human reference (EditLens ICLR 2026, n=200) | 0.37 |
Build, recipe, and discussion: https://mlx-optiq.com/blog/humanizer-stacked-lora
humanizer-1B-OptiQ-4bit/
model.safetensors, config.json, tokenizer* base MiniCPM5-1B-OptiQ-4bit
optiq_metadata.json per-layer bit assignments
adapters/
humanizer-sft/ SFT humanizer LoRA
adapters.safetensors
adapter_config.json
optiq_lora_config.json
humanizer-dpo/ DPO continuation LoRA
adapters.safetensors
adapter_config.json
optiq_lora_config.json
mlx-community/MiniCPM5-1B-OptiQ-4bit. OptiQ mixed-precision quant of openbmb/MiniCPM5-1B. 875 MB on disk, Capability Score 30.28.--preset large (ranks 32 and 64, with the by_bits overlay), 600 iters, mask_prompt=True.optiq lora train --method dpo --mount-adapter. The reference KL is anchored against base + SFT (the textbook SFT then DPO continuation), so the saved adapter contains only the DPO delta. 300 iters, beta 0.1, LR 5e-5 with linear warmup then cosine decay (the OptiQ DPO defaults).The DPO adapter is meaningful only when applied alongside the SFT adapter. It is a delta from the SFT distribution, not a standalone LoRA. Apply both at inference for the headline result.
You need mlx-optiq >= 0.1.4 for the multi-LoRA serving and stacking syntax:
pip install 'mlx-optiq>=0.1.4'
# Download the repo
huggingface-cli download mlx-community/humanizer-1B-OptiQ-4bit \
--local-dir ./humanizer-1B-OptiQ-4bit
# Serve with both adapters mounted
optiq serve \
--model ./humanizer-1B-OptiQ-4bit \
--adapter ./humanizer-1B-OptiQ-4bit/adapters/humanizer-sft \
--adapter ./humanizer-1B-OptiQ-4bit/adapters/humanizer-dpo \
--port 8080
Send requests with both adapters active via the + stacking syntax in the request body:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "./humanizer-1B-OptiQ-4bit",
"adapter": "humanizer-sft+humanizer-dpo",
"messages": [
{"role": "system", "content": "Rewrite AI-generated drafts into natural human-style prose, preserving meaning, facts, names, numbers, citations, URLs, quotes, and formatting."},
{"role": "user", "content": "STYLE: direct technical blog\nTONE: analytical, clear, non-corporate\nLENGTH: preserve within 15%\n\nDraft to rewrite:\n\n[your AI-generated draft here]"}
],
"temperature": 0.4,
"max_tokens": 1600,
"chat_template_kwargs": {"enable_thinking": false}
}'
The OpenAI-compatible endpoint is a drop-in for Open WebUI, Continue, Cursor, your own scripts. Send "adapter": "humanizer-sft" to use SFT alone, or "adapter": "base" to bypass adapters entirely (useful for A/B comparisons).
200 AI-generated drafts from the EditLens ICLR 2026 held-out set, rewritten by each system and scored by RADAR-Vicuna-7B. Lower P(AI) is more human-like.
| Pipeline | P(AI) | Delta vs source | Slop / 1K tokens |
|---|---|---|---|
| Source AI draft (Qwen3.5-4B + Gemma-4-e4b) | 0.51 | , | 0.6 |
| SFT humanizer alone | 0.50 | -0.01 | 0.2 |
| SFT + DPO stacked (this repo) | 0.37 | -0.14 | 0.0 |
| Human reference (target) | 0.37 | -0.14 | 0.1 |
The stacked pipeline produces fewer slop phrases per 1K tokens (0.0) than the human reference set itself (0.1).
"adapter": "base" for general MiniCPM5-1B inference.This quant was produced by mlx-optiq. Point it at any Hugging Face model to get the same sensitivity-aware mixed precision:
pip install mlx-optiq
optiq convert <hf-model-id> --target-bpw 5.0 --candidate-bits 4,8
optiq lab # full local workbench: chat, compare, quantize, fine-tune
openbmb/MiniCPM5-1B (Apache-2.0).