Downloads · 30 days
273
51% of all-time downloads
ToPo-ToPo/Qwen3.8-Flash-Next-mlx-8bit
Qwen3.8-Flash-Next-mlx-8bit is a image-text-to-text model from ToPo-ToPo. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as other.
Qwen/Qwen3.8-Flash-Next を MLX 形式へ変換したもの(8bit 量子化)。
Downloads · 30 days
273
51% of all-time downloads
All-time downloads
539
Public
Parameters
177B
205 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors200 GB · 100%
How the weights are stored.
U32177B · 100%
From the Hugging Face model README
Qwen/Qwen3.8-Flash-Next を MLX 形式へ変換したもの(8bit 量子化)。
Qwen/Qwen3.8-Flash-Next(bf16 公式重み, revision de4b8e4)python -m mlx_vlm convert --hf-path Qwen/Qwen3.8-Flash-Next \
--mlx-path Qwen3.8-Flash-Next-mlx-8bit \
-q --q-bits 8 --q-group-size 32
--q-group-size 32 は必須。n-gram 埋め込みの最終次元が 160 で、既定の 64 では
weight.shape[-1] % group_size != 0 により 128 シャード(51B パラメータ)が
量子化対象から外れ、bf16 のまま残る。
pip install -U mlx-vlm
python -m mlx_vlm generate --model ToPo-ToPo/Qwen3.8-Flash-Next-mlx-8bit \
--prompt "この画像を説明してください" --image path/to/image.jpg --max-tokens 512
ドラフターは別リポジトリ
Qwen3.8-Flash-Next-MTP-bf16。
mlx-vlm 0.7.0 以降が必要(qwen4_exp_mtp はリリース版 0.6.17 に未収録):
pip install "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm@main"
python -m mlx_vlm generate --model ToPo-ToPo/Qwen3.8-Flash-Next-mlx-8bit \
--draft-model ToPo-ToPo/Qwen3.8-Flash-Next-MTP-bf16 \
--draft-kind mtp --draft-block-size 2 --prompt "..." --max-tokens 512