Downloads · 30 days
81K
100% of all-time downloads
Edge0/Edge0-35B-A3B-preview
Edge0-35B-A3B-preview is a text generation model from Edge0. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
Downloads · 30 days
81K
100% of all-time downloads
All-time downloads
81K
Public
Parameters
34.7B
19.7 GB on disk
Likes
3.6K
Public
Click a slice to open those files.
.safetensors19.7 GB · 100%
How the weights are stored.
U3234.7B · 100%
From the Hugging Face model README
A 35B-class sparse MoE that runs in phone-class memory.
3 GiB active memory · 15 tok/s · 4-bit
Native engines on iOS · macOS · Android · Windows.
</div>Edge0-35b-a3b — a 35B MoE LLM that runs at viable speed in under 3 GiB of active memory, via the edge0 streaming inference framework.
<div align="center">Preview status: this is an early preview release of the edge0 pipeline. The checkpoint ships as int4 quantization plus LoRA and prerouter adapters trained for this framework.
<video src="https://huggingface.co/Edge0/Edge0-35B-A3B-preview/resolve/main/20260910-105854.mp4" controls playsinline preload="metadata" width="80%"> </video>
</div>On 2026-09-30 we released the edge0 inference engines for four platforms — so users get the best inference experience across architectures and platforms. The source is open-sourced in the edge0 repo.
| Platform | Native engine (open source) |
|---|---|
| 📱 iOS | edge0/ios |
| 🖥️ macOS | edge0/macos |
| 🤖 Android | edge0/android |
| 🪟 Windows | edge0/windows |
This checkpoint is built for all of them: one model directory — the same int4 base plus LoRA / prerouter adapters — runs unchanged on every platform, so what you download here is what ships on a phone, a desktop and a laptop alike.
edge0.Three mechanisms make this work:
| Base model | Qwen3.6-35B-A3B |
| Quantization | 4-bit |
| Layers | 40 |
| Experts / active per token | 256 / 4 (K=4) |
| Hidden size | 2048 |
| License | Apache 2.0 |
| Framework | edge0 (MLX backend) |
| Contents | base checkpoint + lora_edge0_35b.safetensors + prerouter_edge0_35b.safetensors |
The LoRA and prerouter adapters are co-located with the base checkpoint
and load automatically — this repository is a complete, ready-to-run
model directory for edge0.
All benchmarks were run by us with OpenCompass under identical settings and parameters for both models. The loss of the edge0 pipeline (int4 + adapters) relative to the fp16 base model is small: 3.9 points on average. Max 100:
| Benchmark | edge0-35b (int4) | Qwen3.6-35B-A3B (fp16) |
|---|---|---|
| AIME 2026 | 86.6 | 92.7 |
| HumanEval | 90.9 | 95.1 |
| GPQA-Diamond | 79.8 | 81.8 |
| MMLU-Pro | 81.0 | 84.6 |
| IFBench | 57.9 | 61.7 |
| Average | 79.2 | 83.2 |
Measured with examples/bench.py on a Mac mini M4 Pro, 24 GB:
| Decode speed | Prefill throughput (cold / warm) | Peak active memory* |
|---|---|---|
| 14.9–17.7 tok/s | 113 / 140 tok/s | 2.9 GiB |
*Short contexts; long contexts add KV cache. Expert weights stream from SSD on demand and are not resident.
pip install -e 'git+https://github.com/Edge0-AI/edge0.git#egg=edge0[fetch]'
# Download this repository into a local directory
huggingface-cli download Edge0/Edge0-35b-a3b-preview --local-dir ./Edge0-35b-a3b-preview
# Run it
export EDGE0_35B_MODEL=$PWD/Edge0-35b-a3b-preview
edge0 chat --name edge0-35b --prompt "Introduce yourself"
# Or serve an OpenAI-compatible HTTP API
edge0 serve --name edge0-35b --port 8085
For full usage (Python API, streaming options, prerouter details), see the edge0 documentation.
Apache 2.0. See LICENSE.
If you find Edge0 useful in your research, please cite our paper:
@misc{lin2026halfmemorywallserving,
title={The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction},
author={Yu Lin and Yiming Wang and Runyuan Cai and Hanze Liu and Xiaodong Zeng},
year={2026},
eprint={2609.18063},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2609.18063},
}