Downloads · 30 days
11
39% of all-time downloads
distily/distily_verify_update8
distily_verify_update8 is a machine learning model from distily. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for Distily. The card lists the license as creativeml-openrail-m.
Distilled with Distily library using teacher model gpt2 on dataset wikimedia/wikipedia.
Downloads · 30 days
11
39% of all-time downloads
All-time downloads
28
Public
Parameters
81.9M
164 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors164 MB · 98%
From the Hugging Face model README
Distilled with Distily library using teacher model gpt2 on dataset wikimedia/wikipedia.
<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. # Model description More information needed # Intended uses & limitations More information needed -->GPT2LMHeadModelGPT2LMHeadModel(
(transformer): GPT2Model(
(wte): Embedding(50257, 768)
(wpe): Embedding(1024, 768)
(drop): Dropout(p=0.1, inplace=False)
(h): ModuleList(
(0-5): 6 x GPT2Block(
(ln_1): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
(attn): GPT2FlashAttention2(
(c_attn): Conv1D()
(c_proj): Conv1D()
(attn_dropout): Dropout(p=0.1, inplace=False)
(resid_dropout): Dropout(p=0.1, inplace=False)
)
(ln_2): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
(mlp): GPT2MLP(
(c_fc): Conv1D()
(c_proj): Conv1D()
(act): NewGELUActivation()
(dropout): Dropout(p=0.1, inplace=False)
)
)
)
(ln_f): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
)
(lm_head): Linear(in_features=768, out_features=50257, bias=False)
)
</details>
<br/>
GPT2LMHeadModel -> GPT2LMHeadModel--- teacher model modules
+++ student model modules
@@ -4,7 +4,7 @@
(wpe): Embedding(1024, 768)
(drop): Dropout(p=0.1, inplace=False)
(h): ModuleList(
- (0-11): 12 x GPT2Block(
+ (0-5): 6 x GPT2Block(
(ln_1): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
(attn): GPT2FlashAttention2(
(c_attn): Conv1D()
</details>
<br/>
Trained on 3,221,571 tokens from the wikimedia/wikipedia dataset.
3,96020231101.entrainDistillationObjective(logits_loss_component=LossComponent(label=logits, weight=1, loss_fn=kl), attn_loss_component=LossComponent(label=attn, weight=5.0, loss_fn=raw_mse, layer_mapper=layer-2, norm=layernorm_teacher_only_affine, projector=orthogonal))
The following hyperparameters were used during training:
<details> <summary>Expand</summary>0.000216842Adam with betas=(0.9,0.999) and epsilon=1e-08polynomial1.0DistillationObjective(logits_loss_component=LossComponent(label=logits, weight=1, loss_fn=kl), attn_loss_component=LossComponent(label=attn, weight=5.0, loss_fn=raw_mse, layer_mapper=layer-2, norm=layernorm_teacher_only_affine, projector=orthogonal))<torch.optim.lr_scheduler.LambdaLR object at 0x7feff1761870>Nonedistilbert/distilgpt2NoneNone[('lm_head', False)]Falsegpt2FalseFalsewikimedia/wikipedia20231101.entraintext40000.0110.01.00.00True