Downloads · 30 days
0
Vinnnf/LLaMA-2-7B-MaskLLM-C4
LLaMA-2-7B-MaskLLM-C4 is a machine learning model from Vinnnf. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
<div align="center" <figure <img src="https://github.com/NVlabs/MaskLLM/blob/main/assets/teaser.png?raw=true" style="width:70%; display:block; margin-left:auto; margin-right:auto;" </figure </div
Downloads · 30 days
0
Access
Public
Updated Dec 7, 2024
Repo size
548 MB
Likes
1
Public
Click a slice to open those files.
.npz548 MB · 100%
From the Hugging Face model README
This work introduces MaskLLM, a learnable pruning method that establishes Semi-structured (or ``N:M'') Sparsity in LLMs, aimed at reducing computational overhead during inference. The proposed method is scalable and stands to benefit from larger training datasets.
We provide pre-computed masks for Huggingface Models such as Llama-2 7B and Llama-3 8B with the minimum requirements. It will not involve docker, Megatron or data preprocessing.
pip install transformers accelerate datasets SentencePiece
The following masks were trained and provided by @VainF. We use huggingface_hub to automatically download those masks and apply them to offcical LLMs for evaluation. Those mask files were compressed using numpy.savez_compressed. More results for baselines (SparseGPT, Wanda) can be found in the appendix.
| Model | Pattern | Training Data | Training/Eval SeqLen | PPL (Dense) | PPL (SparseGPT) | PPL (MaskLLM) | Link |
|---|---|---|---|---|---|---|---|
| LLaMA-2 7B | 2:4 | C4 (2B Tokens) | 4096 | 5.12 | 10.42 | 6.78 | HuggingFace |
| LLaMA-3 8B | 2:4 | C4 (2B Tokens) | 4096 | 5.75 | 17.64 | 8.49 | HuggingFace |
| LLaMA-3.1 8B | 2:4 | C4 (2B Tokens) | 4096 | - | - | - | Coming Soon |
Please see NVlabs/MaskLLM.