Downloads · 30 days
62
13% of all-time downloads
ichsanlook/qwen3-eagle
qwen3-eagle is a machine learning model from ichsanlook. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for gguf. The card lists the license as apache-2.0.
EAGLE-3 draft model trained for speculative decoding with Qwen3-0.6B as the target model.
Downloads · 30 days
62
13% of all-time downloads
All-time downloads
488
Public
Repo size
1.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors940 MB · 56%
From the Hugging Face model README
EAGLE-3 draft model trained for speculative decoding with Qwen3-0.6B as the target model.
Trained with a 60K Fortran-optimized vocabulary for 30 epochs on CPU (PyTorch), achieving:
| File | Format | Size | Description |
|---|---|---|---|
eagle3_draft_f16.gguf | GGUF F16 | 470 MB | Draft model in F16 precision |
eagle3_draft_q8.gguf | GGUF Q8_0 | 238 MB | Draft model quantized to Q8_0 |
hf_model/ | safetensors | 896 MB | Hugging Face format for fine-tuning |
# Basic usage with speculative decoding
llama-cli \
-m Qwen3-0.6B-target.gguf \
-md eagle3_draft_q8.gguf \
--spec-type draft-eagle3 \
--spec-draft-n-max 5 \
-p "Hello, how are you?" \
-n 200
# Server mode (enables Ollama-like API)
llama-server \
-m Qwen3-0.6B-target.gguf \
-md eagle3_draft_q8.gguf \
--spec-type draft-eagle3 \
--spec-draft-n-max 5 \
--port 8080
ollama pull qwen3:0.6b
Download the draft model GGUF from this repo
Create a Modelfile:
FROM qwen3:0.6b
DRAFT_MODEL ./eagle3_draft_q8.gguf
PARAMETER speculative_type draft-eagle3
PARAMETER speculative_draft_max 5
ollama create qwen3-eagle -f Modelfile
ollama run qwen3-eagle
| Config | Speed (t/s) | vs Baseline |
|---|---|---|
| Baseline (no draft) | 33.2 | 1.00x |
| Ngram (len=3) | 34.6 | 1.04x |
| EAGLE-3 (len=3) | 15.6 | 0.47x |
| EAGLE-3 (len=5) | 11.8 | 0.36x |
Note: EAGLE-3 speedup requires GPU acceleration. On CPU-only, the dual-model overhead outweighs speculation benefits for small models (0.6B). Speedup expected with GPU offloading or larger target models.
EAGLE-3 extracts hidden states from target model layers [2, 14, 25], projects them through a feature compressor, and feeds them into a lightweight transformer block for draft prediction.