Downloads · 30 days
18
2% of all-time downloads
srinivasbilla/tinymix-8x1b
tinymix-8x1b is a text generation model from srinivasbilla. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
18
2% of all-time downloads
All-time downloads
1.1K
Public
Parameters
6.4B
12.9 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors12.9 GB · 100%
From the Hugging Face model README
This is a MoE-ification of TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T using the Mixtral branch of mergekit
The Goal was to MoE-fy the TinyLlama model and then use this as a base model to further train from. The intuition being finetuning 8x1b should give better performance than finetuning 1b by itself.
More work coming!
This is a merge of the base model, so treat it like a completion.
llm.generate('Quantum Tunneling is')
base_model: TinyLlama/TinyLlama-1.1B-Chat-v1.0
gate_mode: hidden
dtype: bfloat16
experts:
- source_model: /TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
positive_prompts: [""]
- source_model: /TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
positive_prompts: [""]
- source_model: /TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
positive_prompts: [""]
- source_model: /TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
positive_prompts: [""]
- source_model: /TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
positive_prompts: [""]
- source_model: /TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
positive_prompts: [""]
- source_model: /TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
positive_prompts: [""]
- source_model: /TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
positive_prompts: [""]
Thanks to u/mhenrichsen for thr HellaSwag score
| Tasks |Version|Filter|n-shot| Metric |Value | |Stderr|
|---------|-------|------|-----:|--------|-----:|---|-----:|
|hellaswag|Yaml |none | 0|acc |0.4659|± |0.0050|
| | |none | 0|acc\_norm|0.6044|± |0.0049|