Downloads · 30 days
143
100% of all-time downloads
GreenBitAI/Qwen3.8-27B-4bit
Qwen3.8-27B-4bit is a image-text-to-text model from GreenBitAI. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as apache-2.0.
Qwen3.8-27B quantized to four bits for MLX, carrying the model's own multi-token prediction head in mtp/. Speculative decoding therefore works from this repository alone -- nothing else to fetch, no environment variab…
Downloads · 30 days
143
100% of all-time downloads
All-time downloads
143
Public
Parameters
27.4B
16.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.3 GB · 100%
How the weights are stored.
U3226.9B · 98%
From the Hugging Face model README
Qwen3.8-27B quantized to four bits for MLX, carrying the model's own multi-token
prediction head in mtp/. Speculative decoding therefore works from this
repository alone -- nothing else to fetch, no environment variable pointing
somewhere else.
Generation runs 1.8-2.3x faster with the head on, and says the same thing. Every token it proposes is checked by the model itself, so the reply is the model's own either way; the head only saves passes over the weights.
Measured on a 48 GB MacBook Pro (M4 Pro), greedy decoding, 96 tokens, with
gbx_lm. Decode is timed from the first token, so prefill is not in it.
| context | decode, head off | decode, head on | draft acceptance |
|---|---|---|---|
| 1,024 | 14.7 tok/s | 33.8 tok/s | 0.95 |
| 4,096 | 14.3 | 27.3 | 0.78 |
| 16,384 | 13.5 | 24.8 | 0.78 |
| macOS | 15.0 or later |
| chip | Apple Silicon (arm64). There is no Intel build. |
| Python | none -- the binary carries what it needs |
These weights are held resident, not paged: on a 512 GB Mac Studio the model settles at about 16 GB, and a machine needs room for that much plus the conversation's cache.
# upgrading? clear the previous version's unpack directory first
rm -rf ~/.libra/cache/onefile/gbx_lm
curl -fL -o gbx_lm-darwin-arm64.tar.gz 'https://github.com/GreenBitAI/gbx-lm/releases/latest/download/gbx_lm-darwin-arm64.tar.gz' \
&& tar -xzf gbx_lm-darwin-arm64.tar.gz gbx_lm \
&& mkdir -p "$HOME/.local/bin" \
&& mv gbx_lm "$HOME/.local/bin/gbx_lm" \
&& chmod +x "$HOME/.local/bin/gbx_lm"
gbx_lm -h
The build is signed with a Developer ID and notarised, so macOS runs it without the usual detour for a downloaded binary.
command not found -- $HOME/.local/bin is not on your PATH:
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.zshrc && source ~/.zshrc # zsh
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bash_profile && source ~/.bash_profile # bash
Killed: 9 -- a previous version's files are still in the unpack directory,
and macOS refuses to mix two builds. Run the rm -rf line above, then try again.
gbx_lm --model GreenBitAI/Qwen3.8-27B-4bit
# the draft head is off unless asked for, and found in `mtp/` without a path
GBX_QWEN35_MTP=on gbx_lm --model GreenBitAI/Qwen3.8-27B-4bit
That serves an OpenAI-compatible API on port 11688, which is its default. The
weights download on first use into ~/.libra/cache/models; set HF_HOME to put
them elsewhere, and HF_TOKEN if you meet the Hub's rate limits for anonymous
downloads.
curl http://127.0.0.1:11688/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"GreenBitAI/Qwen3.8-27B-4bit","messages":[{"role":"user","content":"Hello"}]}'
From gbx_lm v0.7.1, the binary also runs the model without a server -- one
prompt, or an interactive chat:
GBX_QWEN35_MTP=on gbx_lm generate --model GreenBitAI/Qwen3.8-27B-4bit --prompt "Hello" --max-tokens 2048
GBX_QWEN35_MTP=on gbx_lm chat --model GreenBitAI/Qwen3.8-27B-4bit --max-tokens 2048
gbx_lm generate -h and gbx_lm chat -h list their options. Keep
--max-tokens generous: the default is 100 for generate and 256 for chat,
and the model's reasoning before it answers counts against it. Leave out
GBX_QWEN35_MTP=on to run without the draft head.
The server speaks three wire protocols on the same port, so the tools that expect a hosted API can be pointed at this one:
| path | for |
|---|---|
/v1/chat/completions | anything written against the OpenAI API |
/v1/responses | Codex |
/v1/messages | Claude Code |
Codex -- a provider in ~/.codex/config.toml:
[model_providers.gbx]
name = "gbx-lm"
base_url = "http://127.0.0.1:11688/v1"
wire_api = "responses"
and a profile in ~/.codex/gbx.config.toml:
model_provider = "gbx"
model = "GreenBitAI/Qwen3.8-27B-4bit"
model_context_window = 262144
Claude Code -- ~/.claude/gbx.settings.json:
{
"env": {
"ANTHROPIC_BASE_URL": "http://127.0.0.1:11688",
"ANTHROPIC_AUTH_TOKEN": "local",
"ANTHROPIC_MODEL": "GreenBitAI/Qwen3.8-27B-4bit",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "GreenBitAI/Qwen3.8-27B-4bit"
}
}
Both clients ask for a small model for their own background work, so every name in the settings has to be one this server is serving.
mtp/mtp.safetensors is built from the draft head
Qwen/Qwen3.8-27B ships under mtp.,
quantized to match these weights. Apache 2.0, as the original is.