Downloads · 30 days
71
1% of all-time downloads
M4-ai/Hercules-Mini-1.8B
Hercules-Mini-1.8B is a text generation model from M4-ai. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
We fine-tuned Qwen1.5-1.8B on Locutusque's Hercules-v4.
Downloads · 30 days
71
1% of all-time downloads
All-time downloads
8.1K
Public
Parameters
1.8B
5.4 GB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors3.7 GB · 100%
From the Hugging Face model README
We fine-tuned Qwen1.5-1.8B on Locutusque's Hercules-v4.
This model has capabilities in math, coding, function calling, roleplay, and more. We fine-tuned it using 700,000 examples of Hercules-v4.
General purpose assistant, question answering, chain-of-thought, etc..
The eos token was not setup properly, so to prevent infinite generation you'll need to implement a stopping criteria when the model generates the <|im_end|> token.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
Coming soon
https://huggingface.co/datasets/Locutusque/hercules-v4.0
We used 8 Kaggle TPUs, and we trained at a global batch size of 256 and sequence length of 1536
Thanks to @Tonic, @aloobun, @fhai50032, and @Locutusque for their contributions to this model.