Downloads · 30 days
38
8% of all-time downloads
INC4AI/gpt-j-6b-sparse
gpt-j-6b-sparse is a text generation model from INC4AI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
The sparse version of GPT-J 6B is a pruned variant derived from the original GPT-J 6B model and the vast majority of linear layers maintain a 40% unstructured sparsity (except for the 'lmhead').
Downloads · 30 days
38
8% of all-time downloads
All-time downloads
500
Public
Repo size
48.4 GB
Likes
1
Public
Click a slice to open those files.
.bin24.2 GB · 100%
From the Hugging Face model README
The sparse version of GPT-J 6B is a pruned variant derived from the original GPT-J 6B model and the vast majority of linear layers maintain a 40% unstructured sparsity (except for the 'lm_head').
<figure>| Hyperparameter | Value |
|---|---|
| \(n_{parameters}\) | 6053381344 |
| \(n_{layers}\) | 28* |
| \(d_{model}\) | 4096 |
| \(d_{ff}\) | 16384 |
| \(n_{heads}\) | 16 |
| \(d_{head}\) | 256 |
| \(n_{ctx}\) | 2048 |
| \(n_{vocab}\) | 50257/50400† (same tokenizer as GPT-2/3) |
| Positional Encoding | Rotary Position Embedding RoPE |
| RoPE Dimensions | 64 |
The model consists of 28 layers with a model dimension of 4096, and a feedforward dimension of 16384. The model dimension is split into 16 heads, each with a dimension of 256. Rotary Position Embedding (RoPE) is applied to 64 dimensions of each head. The model is trained with a tokenization vocabulary of 50257, using the same set of BPEs as GPT-2/GPT-3.
Evaluating the accuracy of the sparse model of gpt-j-6b using the lambada_openai dataset in lm_eval, providing the accuracy fluctuation under two precisions: FP32 and BF16.
<figure>| Sparsity | Dataset | Precision | Dense Acc ↑ | Sparse Acc ↑ | Acc fluctuations |
|---|---|---|---|---|---|
| 40% | Lambada_openai | FP32 | 0.6831 | 0.6922 | +1.33% |
| 40% | Lambada_openai | BF16 | 0.6771 | 0.6874 | +0.63% |