Downloads · 30 days
75
0% of all-time downloads
Locutusque/Hercules-3.1-Mistral-7B
Hercules-3.1-Mistral-7B is a text generation model from Locutusque. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
75
0% of all-time downloads
All-time downloads
49.5K
Public
Parameters
7.2B
14.5 GB on disk
Likes
16
Public
Click a slice to open those files.
.safetensors14.5 GB · 100%
From the Hugging Face model README

Hercules-3.1-Mistral-7B is a fine-tuned language model derived from Mistralai/Mistral-7B-v0.1. It is specifically designed to excel in instruction following, function calls, and conversational interactions across various scientific and technical domains. The dataset used for fine-tuning, also named Hercules-v3.0, expands upon the diverse capabilities of OpenHermes-2.5 with contributions from numerous curated datasets. This fine-tuning has hercules-v3.0 with enhanced abilities in:
Hercules-3.1-Mistral-7B is well-suited to the following applications:
Important Note: Although Hercules-v3.0 is carefully constructed, it's important to be aware that the underlying data sources may contain biases or reflect harmful stereotypes. Use this model with caution and consider additional measures to mitigate potential biases in its responses.
Hercules-3.1-Mistral-7B is fine-tuned from the following sources:
cognitivecomputations/dolphinEvol Instruct 70K & 140Kteknium/GPT4-LLM-Cleanedjondurbin/airoboros-3.2AlekseyKorshuk/camel-chatmlCollectiveCognition/chats-data-2023-09-22Nebulous/lmsys-chat-1m-smortmodelsonlyglaiveai/glaive-code-assistant-v2glaiveai/glaive-code-assistantglaiveai/glaive-function-calling-v2garage-bAInd/Open-Platypusmeta-math/MetaMathQAteknium/GPTeacher-General-InstructGPTeacher roleplay datasetsBI55/MedTextpubmed_qa labeled subsetUnnatural InstructionsM4-ai/LDJnr_combined_inout_formatCollectiveCognition/chats-data-2023-09-27CollectiveCognition/chats-data-2023-10-16NobodyExistsOnTheInternet/sharegptPIPPAyuekai/openchat_sharegpt_v3_vicuna_formatise-uiuc/Magicoder-Evol-Instruct-110Ksablo/oasst2_curatedThe bluemoon dataset was filtered from the training data as it showed to cause performance degradation.
<|im_start|>system\n{message}<|im_end|>\n<|im_start|>user\n{user message}<|im_end|>\n<|im_start|>call\n{function call message}<|im_end|>\n<|im_start|>function\n{function response message}<|im_end|>\n<|im_start|>assistant\n{assistant message}</s>This model was fine-tuned using the TPU-Alignment repository. https://github.com/Locutusque/TPU-Alignment
ExLlamaV2 by bartowski https://huggingface.co/bartowski/Hercules-3.1-Mistral-7B-exl2
Detailed results can be found here
| Metric | Value |
|---|---|
| Avg. | 62.09 |
| AI2 Reasoning Challenge (25-Shot) | 61.18 |
| HellaSwag (10-Shot) | 83.55 |
| MMLU (5-Shot) | 63.65 |
| TruthfulQA (0-shot) | 42.83 |
| Winogrande (5-shot) | 79.01 |
| GSM8k (5-shot) | 42.30 |