Downloads · 30 days
151
18% of all-time downloads
mlx-community/LLaDA2.2-flash-OptiQ-2bit
LLaDA2.2-flash-OptiQ-2bit is a text generation model from mlx-community. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). Try the Lab · All OptiQ quants · Docs
Downloads · 30 days
151
18% of all-time downloads
All-time downloads
854
Public
Parameters
103B
38.9 GB on disk
Likes
4
Public
Click a slice to open those files.
.safetensors38.9 GB · 100%
How the weights are stored.
U32103B · 100%
From the Hugging Face model README
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). Try the Lab · All OptiQ quants · Docs
An OptiQ mixed-precision MLX quant of LLaDA2.2-flash, a ~100B diffusion language model with a 256-routed-expert sparse MoE. This is an extreme 2-bit build.
static build — per-layer bit-widths assigned to a 2.5
target bits-per-weight.Needs optiq >= 0.4.4, which ships the vendored llada2_moe decoder (the
256-expert diffusion MoE) and the block-diffusion decode loop. Stock mlx-lm
has no llada2_moe arch and cannot load or generate from this repo.
pip install -U optiq
LLaDA2 is a masked-diffusion model, not autoregressive — it denoises a canvas
block by block. optiq serve detects the arch and routes it through OptiQ's
vendored decoder and the block-diffusion decode loop automatically:
optiq serve --model mlx-community/LLaDA2.2-flash-OptiQ-2bit
Then call the OpenAI-compatible endpoint at http://localhost:8000/v1, or use
it from the OptiQ Lab and optiq code.