Downloads · 30 days
28
7% of all-time downloads
HenryHHHH/DistilLlamaV1
DistilLlamaV1 is a text generation model from HenryHHHH. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This model is a distilled version of LLaMA 2, containing approximately 80 million parameters. It was trained using a mix of OpenWebText and WikiText Raw V1 datasets. Knowledge distillation was employed to transfer kno…
Downloads · 30 days
28
7% of all-time downloads
All-time downloads
419
Public
Parameters
87.3M
699 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors349 MB · 100%
From the Hugging Face model README
This model is a distilled version of LLaMA 2, containing approximately 80 million parameters. It was trained using a mix of OpenWebText and WikiText Raw V1 datasets. Knowledge distillation was employed to transfer knowledge from a larger "teacher" model—Meta’s 7B LLaMA 2—to help this smaller model mimic the behavior of the teacher. This version is the latest version of DistilLlama, which has gone through 5 days of training using two Nvidia A100 80G GPU.
30 out of 300 checkpoints were examined, and the one with the best performance in semantic and factual accuracy has now been updated in this repository.
The architecture is based on LLaMA 2, with the following parameters:
| Parameter | Value |
|---|---|
| Hidden Dimension | 512 |
| Intermediate Dimension | 1536 |
| Max Positional Embeddings | 128 |
| Attention Heads | 8 |
| Transformer Layers | 16 |
Cosine Similarity using Word Embeddings
Exact Match (EM)
ROUGE Score
| Model Name | Duration (s) | Emissions (kgCO₂e) | Avg. EM | Avg. Cosine Similarity | Avg. ROUGE Score |
|---|---|---|---|---|---|
| LLaMA-2-7B-HF | 18215.61 | 1.84e-01 | 0.715 | 0.7257 | 0.0821 |
| baby-llama-58m | 57.20 | 2.73e-06 | 0.025 | 0.6556 | 0.0097 |
| DistilLlama | 77.12 | 7.79e-04 | 0.02 | 0.6623 | 0.0115 |
| DistilLlamaV1 | 78.46 | 8.49e-04 | 0.065 | 0.6776 | 0.0135 |
Note: CodeCarbon was used to track carbon emission. Allocated 80GB memory, 32 cores, Intel(R) Xeon(R) Gold 6448H for the evaluation
@misc{timiryasov2023babyllamaknowledgedistillation, title={Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty}, author={Inar Timiryasov and Jean-Loup Tastet}, year={2023}, eprint={2308.02019}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2308.02019}, }
Note: The repository will be updated as training progresses. Last update 2024-11-06