Downloads · 30 days
251
11% of all-time downloads
ericflo/Llama-3.1-8B-ContinuedTraining
Llama-3.1-8B-ContinuedTraining is a text generation model from ericflo. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
This model is a custom-trained language model based on the Meta-Llama-3.1-8B architecture. Unlike most instruction-tuned models, it was trained directly on a mixture of high-quality datasets for general text and code…
Downloads · 30 days
251
11% of all-time downloads
All-time downloads
2.4K
Public
Parameters
8B
260 GB on disk
Likes
0
Public
Click a slice to open those files.
.gguf67.3 GB · 68%
From the Hugging Face model README
This model is a custom-trained language model based on the Meta-Llama-3.1-8B architecture. Unlike most instruction-tuned models, it was trained directly on a mixture of high-quality datasets for general text and code completion tasks, as well as instruction-following. A high-rank adapter (rank 128) is used to enhance learning capacity while mitigating catastrophic forgetting, which distinguishes this model from common low-rank fine-tuning methods.
Instead of fine-tuning an instruction-tuned model, the base Meta-Llama-3.1-8B model was trained with a diverse set of high-quality pretraining and instruction datasets. The training focused on both text completion/prediction and instruction-following tasks.
Key features of the training process:
The model was trained for 650 steps using the datasets mentioned above. During this process, the focus was on ensuring a balanced learning process across different task types (text completion, code completion, instruction-following). The high-rank adapter plays a significant role in maintaining model capacity while reducing computational complexity.
This model is designed for a variety of natural language processing tasks, including:
While this model was trained on a mix of high-quality datasets, it may still exhibit biases present in the training data, especially in domains with limited or skewed representation. Users should:
| Tasks | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| tinyBenchmarks | N/A | |||||||
| - tinyArc | 0 | none | 25 | acc_norm | ↑ | 0.6056 | ± | N/A |
| - tinyGSM8k | 0 | flexible-extract | 5 | exact_match | ↑ | 0.4793 | ± | N/A |
| strict-match | 5 | exact_match | ↑ | 0.4793 | ± | N/A | ||
| - tinyHellaswag | 0 | none | 10 | acc_norm | ↑ | 0.8261 | ± | N/A |
| - tinyMMLU | 0 | none | 0 | acc_norm | ↑ | 0.6358 | ± | N/A |
| - tinyTruthfulQA | 0 | none | 0 | acc | ↑ | 0.5098 | ± | N/A |
| - tinyWinogrande | 0 | none | 5 | acc_norm | ↑ | 0.7447 | ± | N/A |
For inquiries about this model, please contact Eric Florenzano through the model repository.