Downloads · 30 days
45
10% of all-time downloads
sequelbox/Llama3.1-8B-PlumMath
Llama3.1-8B-PlumMath is a text generation model from sequelbox. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.1.
This is a merge of pre-trained language models created using mergekit.
Downloads · 30 days
45
10% of all-time downloads
All-time downloads
443
Public
Parameters
8B
16.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
This is a merge of pre-trained language models created using mergekit.
This model was merged using the della merge method using meta-llama/Llama-3.1-8B-Instruct as a base.
The following models were included in the merge:
The following YAML configuration was used to produce this model:
merge_method: della
dtype: bfloat16
parameters:
normalize: true
models:
- model: ValiantLabs/Llama3.1-8B-ShiningValiant2
parameters:
density: 0.5
weight: 0.3
- model: ValiantLabs/Llama3.1-8B-Cobalt
parameters:
density: 0.5
weight: 0.2
base_model: meta-llama/Llama-3.1-8B-Instruct
Detailed results can be found here
| Metric | Value |
|---|---|
| Avg. | 13.80 |
| IFEval (0-Shot) | 22.42 |
| BBH (3-Shot) | 16.45 |
| MATH Lvl 5 (4-Shot) | 3.93 |
| GPQA (0-shot) | 9.06 |
| MuSR (0-shot) | 8.98 |
| MMLU-PRO (5-shot) | 21.95 |