Downloads · 30 days
60
0% of all-time downloads
zhengchenphd/Mistral-Plus-7B
Mistral-Plus-7B is a text generation model from zhengchenphd. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Mistral-Plus is a chat assistant trained by Reinforcement Learning from Human Feedback (RLHF) using the Mistral-7B base model as the backbone.
Downloads · 30 days
60
0% of all-time downloads
All-time downloads
13.4K
Public
Repo size
29 GB
Likes
4
Public
Click a slice to open those files.
.bin14.5 GB · 100%
From the Hugging Face model README
Mistral-Plus is a chat assistant trained by Reinforcement Learning from Human Feedback (RLHF) using the Mistral-7B base model as the backbone.
Paper (Mistral-Plus): https://arxiv.org/abs/2403.02513
Mistral-Plus is primarily utilized for research in the areas of large language models and chatbots. It is intended chiefly for use by researchers and hobbyists specializing in natural language processing, machine learning, and artificial intelligence.
Mistral-Plus not only preserves the Mistral base model's general capabilities, but also significantly enhances its conversational abilities and notably reduces the generation of toxic outputs as human preference.

To the best of knowledge, this is the first academic endeavor to bypass supervised fine-tuning and directly apply reinforcement learning from human feedback. More importantly, Mistral-Plus is publicly available through HuggingFace for promoting collaborative research and innovation. This initiative to open-source Mistral-Plus seeks to empower researchers worldwide, enabling them to delve deeper into and build upon Mistral-Plus work, with a particular focus on conversational tasks, such as customer service, intelligent assistant, etc.

![]() | ![]() |
|---|---|
![]() | ![]() |
![]() | ![]() |
|---|
Bad word generation probablity on Mistral-Instruct and Mistral-Plus. The x-axis represents different intermittent layers, y-axis shows token probability.

