Downloads · 30 days
251
49% of all-time downloads
ToPo-ToPo/Qwen3.8-Flash-Next-mlx-bf16
Qwen3.8-Flash-Next-mlx-bf16 is a image-text-to-text model from ToPo-ToPo. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as other.
Qwen/Qwen3.8-Flash-Next を MLX 形式へ変換したもの(bf16・非量子化)。
Downloads · 30 days
251
49% of all-time downloads
All-time downloads
517
Public
Parameters
177B
355 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors355 GB · 100%
How the weights are stored.
BF16177B · 100%
From the Hugging Face model README
Qwen/Qwen3.8-Flash-Next を MLX 形式へ変換したもの(bf16・非量子化)。
Qwen/Qwen3.8-Flash-Next(bf16 公式重み, revision de4b8e4)python -m mlx_vlm convert --hf-path Qwen/Qwen3.8-Flash-Next \
--mlx-path Qwen3.8-Flash-Next-mlx-bf16 --dtype bfloat16
pip install -U mlx-vlm
python -m mlx_vlm generate --model ToPo-ToPo/Qwen3.8-Flash-Next-mlx-bf16 \
--prompt "この画像を説明してください" --image path/to/image.jpg --max-tokens 512
推論には 355 GB 超のメモリが要る。量子化版は 4bit(111 GB)/ 8bit(200 GB)。
MTP ヘッド(mtp.*)は変換時に除外される。投機デコードに使う場合は公式 bf16 から
別途切り出すこと。