Downloads · 30 days
86
1% of all-time downloads
MexIvanov/zephyr-python-ru-merged
zephyr-python-ru-merged is a text generation model from MexIvanov. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
- Developed by: C.B. Pronin, A.V. Volosova, A.V. Ostroukh, Yu.N. Strogov, V.V. Kurbatov, A.S. Umarova. - Model type: Base model HuggingFaceH4/zephyr-7b-beta merged with LoRA (Peft) adapter model MexIvanov/zephyr-pytho…
Downloads · 30 days
86
1% of all-time downloads
All-time downloads
12.8K
Public
Parameters
7.2B
14.5 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors14.5 GB · 100%
From the Hugging Face model README
An experimental finetune of Zephyr-7b-beta, aimed at improving coding performance and support for coding-related instructions written in Russian language.
Instruction-based coding in Python, based of instructions written in natural language (English or Russian)
Prompt template - Zephyr:
<|system|>
</s>
<|user|>
{prompt}</s>
<|assistant|>
This adapter model is intended (but not limited) for research usage only. It was trained on a code based instruction set and it does not have any moderation mechanisms. Use at your own risk, we are not responsible for any usage or output of this model.
Quote from Zephyr (base-model) repository: "Zephyr-7B-β has not been aligned to human preferences for safety within the RLHF phase or deployed with in-the-loop filtering of responses like ChatGPT, so the model can produce problematic outputs (especially when prompted to do so). It is also unknown what the size and composition of the corpus was used to train the base model (mistralai/Mistral-7B-v0.1), however it is likely to have included a mix of Web data and technical sources like books and code. See the Falcon 180B model card for an example of this."
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
The following bitsandbytes quantization config was used during training:
The following bitsandbytes quantization config was used during training: