Downloads · 30 days
117
1% of all-time downloads
chargoddard/MixtralRPChat-ZLoss
MixtralRPChat-ZLoss is a text generation model from chargoddard. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
<img src="https://raw.githubusercontent.com/OpenAccess-AI-Collective/axolotl/main/image/axolotl-badge-web.png" alt="Built with Axolotl" width="200" height="32"/
Downloads · 30 days
117
1% of all-time downloads
All-time downloads
9.6K
Public
Parameters
46.7B
93.4 GB on disk
Likes
26
Public
Click a slice to open those files.
.safetensors93.4 GB · 100%
From the Hugging Face model README
QLoRA tuned from mistralai/Mixtral-8x7B-v0.1.
My main reason for training this model was to investigate using an altered router balancing loss combined with the z-loss introduced in ST-MoE: Designing Stable and Transferable Sparse Expert Models. The result is pretty decent, I think! It does a good job of respecting character information in system prompts and performed adequately on a few simple coding tasks.
To train this I used a custom branch of Transformers that adds z-loss and reimplements the router balancing loss based on the version in MegaBlocks. The config used with my custom hacked-up branch of axolotl is available here.
Uses my favorite non-ChatML token-economic chat prompt format. Messages should be prefixed with " ***System:", " ***Query:", or " ***Response:" for system, user, and model messages respectively. No newlines are necessary but the space before the triple asterisk is mandatory.