Downloads · 30 days
21
7% of all-time downloads
RedHatAI/SparseLlama-2-7b-evolcodealpaca-pruned_50.2of4
SparseLlama-2-7b-evolcodealpaca-pruned_50.2of4 is a text generation model from RedHatAI. Use it when you need the model to write or continue text. It is set up for transformers.
- Model Architecture: Llama-2 - Input: Text - Output: Text - Model Optimizations: - Pruned: 50% 2:4 - Release Date: 7/2/2024 - Version: 1.0 - Model Developers: Neural Magic
Downloads · 30 days
21
7% of all-time downloads
All-time downloads
293
Public
Parameters
6.7B
27 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors27 GB · 100%
From the Hugging Face model README
Compressed version of Llama-2-7b specialized for code-generation. This model was obtained by fine-tuning the Sparse Foundational model SparseLlama-2-7b-pruned_50.2of4 on the evol-codealpaca-v1 dataset. SquareHead knowledge distillation was used with Llama-2-7b-evolcodealpaca as teacher. It achieves HumanEval pass@1 of 34.58%, whereas the dense Llama-2-7b-evolcodealpaca model achieves 32.03%.
This model was produced as part if Neural Magic's Sparse Foundational Models initiative, and demostrates the capability of Sparse Foundational Models to transfer to the code-generation domain.
This model is derived from the Sparse Foundational model Sparse-Llama-2-7b-pruned_50.2of4, which was obtained by applying the SparseGPT algorithm to prune Llama-2-7b to 50% sparsity with a 2:4 mask. This optimization reduces the number of parameters by 50%, reducing the disk size and FLOPs by the same level.
This model was evaluated in the HumanEval benchmark using the bigcode-evaluation-harness.
| Model | HumanEval pass@1 | Recovery |
|---|---|---|
| Llama-2-7b-evolcodealpaca | 32.03% | -- |
| SparseLlama-2-7b-evolcodealpaca-pruned_50.2of4 | 34.58% | 108% |