Downloads · 30 days
197
0% of all-time downloads
unkmaster/EAGLE3-LLaMA3.1-Instruct-8B
EAGLE3-LLaMA3.1-Instruct-8B is a machine learning model from unkmaster. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<img src="figs/logo.png" alt="EAGLE" width="220" align="left"<div align="center"<h1 EAGLE</h1</div
Downloads · 30 days
197
0% of all-time downloads
All-time downloads
115K
Public
Parameters
425M
865 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors850 MB · 98%
How the weights are stored.
F16425M · 100%
From the Hugging Face model README
<img src="figs/logo.png" alt="EAGLE" width="220" align="left"><div align="center"><h1> EAGLE</h1></div>
<p align="center"> | <a href="https://arxiv.org/pdf/2401.15077.pdf"><b>EAGLE</b></a> | <a href="https://arxiv.org/pdf/2406.16858"><b>EAGLE-2</b></a> | <a href="https://arxiv.org/pdf/2503.01840"><b>EAGLE-3</b></a> | <a href="https://sites.google.com/view/ eagle-llm"><b>Blog</b></a> | </p> <p align="center"> <a href=""> <img src="https://img.shields.io/badge/Version-v3.0.0-orange.svg" alt="Version"> </a> <a href="https://opensource.org/licenses/Apache-2.0"> <img src="https://img.shields.io/badge/License-Apache_2.0-blue.svg" alt="License"> </a> <a href="https://github.com/SafeAILab/EAGLE/issues"> <img src="https://img.shields.io/badge/Maintained%3F-yes-green.svg" alt="Maintenance"> </a> <a href="https://github.com/SafeAILab/EAGLE/pulls"> <img src="https://img.shields.io/badge/Contributions-welcome-brightgreen.svg?style=flat" alt="Contributions welcome"> </a> </p> <p align="center"> <img src="./figs/eagle3r.jpg" alt="benchmark" width="790"> </p>EAGLE (Extrapolation Algorithm for Greater Language-model Efficiency) is a new baseline for fast decoding of Large Language Models (LLMs) with provable performance maintenance. This approach involves extrapolating the second-top-layer contextual feature vectors of LLMs, enabling a significant boost in generation efficiency.
EAGLE-2 uses the confidence scores from the draft model to approximate acceptance rates, dynamically adjusting the draft tree structure, which further enhances performance.
EAGLE-3 removes the feature prediction constraint in EAGLE and simulates this process during training using training-time testing. Considering that top-layer features are limited to next-token prediction, EAGLE-3 replaces them with a fusion of low-, mid-, and high-level semantic features. EAGLE-3 further improves generation speed while ensuring lossless performance.
Inference is conducted on 2x RTX 3090 GPUs at fp16 precision using the Vicuna 13B model.
EAGLE has been merged in the following mainstream LLM serving frameworks (listed in alphabetical order).
For technical details and full experimental results, please check the paper of EAGLE, the paper of EAGLE-2, and the paper of EAGLE-3.
@inproceedings{li2024eagle,
author = {Yuhui Li and Fangyun Wei and Chao Zhang and Hongyang Zhang},
title = {{EAGLE}: Speculative Sampling Requires Rethinking Feature Uncertainty},
booktitle = {International Conference on Machine Learning},
year = {2024}
}
@inproceedings{li2024eagle2,
author = {Yuhui Li and Fangyun Wei and Chao Zhang and Hongyang Zhang},
title = {{EAGLE-2}: Faster Inference of Language Models with Dynamic Draft Trees},
booktitle = {Empirical Methods in Natural Language Processing},
year = {2024}
}
@misc{li2025eagle3scalinginferenceacceleration,
title={{EAGLE-3}: Scaling up Inference Acceleration of Large Language Models via Training-Time Test},
author={Yuhui Li and Fangyun Wei and Chao Zhang and Hongyang Zhang},
year={2025},
eprint={2503.01840},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2503.01840},
}