Downloads · 30 days
1.7K
11% of all-time downloads
mudler/Step-3.5-Flash-APEX-GGUF
Step-3.5-Flash-APEX-GGUF is a machine learning model from mudler. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
<div style="background-color: f59e0b; color: white; padding: 20px; border-radius: 10px; text-align: center; margin: 20px 0;" <h2 style="color: white; margin: 0 0 10px 0;"⚡ Each donation = another big MoE quantized</h2…
Downloads · 30 days
1.7K
11% of all-time downloads
All-time downloads
15.2K
Public
Repo size
795 GB
Likes
7
Public
Click a slice to open those files.
.gguf795 GB · 100%
From the Hugging Face model README
APEX (Adaptive Precision for EXpert Models) quantizations of Step-3.5-Flash.
Brought to you by the LocalAI team | APEX Project | Technical Report
Benchmarks coming soon. For reference APEX benchmarks on the Qwen3.5-35B-A3B architecture, see mudler/Qwen3.5-35B-A3B-APEX-GGUF.
APEX is a quantization strategy for Mixture-of-Experts (MoE) models. It classifies tensors by role (routed expert, shared expert, attention) and applies a layer-wise precision gradient -- edge layers get higher precision, middle layers get more aggressive compression. I-variants use diverse imatrix calibration (chat, code, reasoning, tool-calling, agentic traces, Wikipedia).
See the APEX project for full details, technical report, and scripts.
local-ai run mudler/[email protected]
APEX is brought to you by the LocalAI team. Developed through human-driven, AI-assisted research. Built on llama.cpp.