Downloads · 30 days
6.2K
100% of all-time downloads
Edge0/Edge0-8B-A1B-preview
Edge0-8B-A1B-preview is a text generation model from Edge0. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
Downloads · 30 days
6.2K
100% of all-time downloads
All-time downloads
6.2K
Public
Parameters
7.9B
4.6 GB on disk
Likes
89
Trending 7
Click a slice to open those files.
.safetensors4.6 GB · 99%
How the weights are stored.
U327.9B · 100%
From the Hugging Face model README
An 8B-class sparse MoE that runs in phone-class memory.
1 GiB active memory · 25 tok/s · 4-bit
Native engines on iOS · macOS · Android · Windows.
</div>Edge0-8b-a1b — an 8B MoE LLM that runs at viable speed in under 1 GiB of active memory, via the edge0 streaming inference framework.
<div align="center">Preview status: this is an early preview release of the edge0 pipeline. The checkpoint ships as int4 quantization plus LoRA and prerouter adapters trained for this framework.
<video src="https://huggingface.co/Edge0/Edge0-35B-A3B-preview/resolve/main/20260910-105854.mp4" controls playsinline preload="metadata" width="80%"> </video>
</div>On 2026-09-30 we released the edge0 inference engines for four platforms — so users get the best inference experience across architectures and platforms. The source is open-sourced in the edge0 repo.
| Platform | Native engine (open source) |
|---|---|
| 📱 iOS | edge0/ios |
| 🖥️ macOS | edge0/macos |
| 🤖 Android | edge0/android |
| 🪟 Windows | edge0/windows |
This checkpoint is built for all of them: one model directory — the same int4 base plus LoRA / prerouter adapters — runs unchanged on every platform, so what you download here is what ships on a phone, a desktop and a laptop alike.
edge0.Three mechanisms make this work:
| Base model | inclusionAI Ling 3.0 tiny (bailing hybrid, MLA + MoE, ≈7.9B total / ≈1.2B active) |
| Quantization | 4-bit |
| Layers | 24 |
| Experts / active per token | 128 / 8 (K=8) |
| Hidden size | 1536 |
| Context | 128k |
| Thinking mode | yes (chat template) |
| License | Apache 2.0 |
| Framework | edge0 (MLX backend) |
| Contents | base checkpoint + lora_edge0_8b.safetensors + prerouter_edge0_8b.safetensors |
The LoRA and prerouter adapters are co-located with the base checkpoint
and load automatically — this repository is a complete, ready-to-run
model directory for edge0.
All benchmarks were run by us with OpenCompass under identical settings and parameters for both models. The loss of the edge0 pipeline (int4 + adapters) relative to the fp16 base model is small: 2.8 points on average, with MMLU-Pro above the base. Max 100:
| Benchmark | edge0-8b (int4) | Ling 3.0 tiny (fp16) |
|---|---|---|
| AIME 2026 | 63.3 | 73.3 |
| HumanEval | 91.5 | 92.7 |
| GPQA-Diamond | 70.7 | 71.2 |
| MMLU-Pro | 70.1 | 65.8 |
| IFBench | 53.9 | 60.6 |
| Average | 69.9 | 72.7 |
Measured with examples/bench.py on a Mac mini M4 Pro, 24 GB:
| Decode speed | Prefill throughput (cold / warm) | Peak active memory |
|---|---|---|
| 23.9–25.3 tok/s | 500 / 1428 tok/s | 1.0 GiB |
pip install -e 'git+https://github.com/Edge0-AI/edge0.git#egg=edge0[fetch]'
# Download this repository into a local directory
huggingface-cli download Edge0/Edge0-8b-a1b-preview --local-dir ./Edge0-8b-a1b-preview
# Run it
export EDGE0_8B_MODEL=$PWD/Edge0-8b-a1b-preview
edge0 chat --name edge0-8b --prompt "Introduce yourself"
# Or serve an OpenAI-compatible HTTP API
edge0 serve --name edge0-8b --port 8083
For full usage (Python API, streaming options, prerouter details), see the edge0 documentation.
Apache 2.0. See LICENSE.
If you find Edge0 useful in your research, please cite our paper:
@misc{lin2026halfmemorywallserving,
title={The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction},
author={Yu Lin and Yiming Wang and Runyuan Cai and Hanze Liu and Xiaodong Zeng},
year={2026},
eprint={2609.18063},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2609.18063},
}