Downloads · 30 days
32
5% of all-time downloads
datatab/Yugo55A-4bit
Yugo55A-4bit is a text generation model from datatab. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
32
5% of all-time downloads
All-time downloads
703
Public
Parameters
7.5B
4.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors4.1 GB · 100%
How the weights are stored.
U87.2B · 96%
From the Hugging Face model README
4bit<table> <tr> <th>MODEL</th> <th>ARC-E</th> <th>ARC-C</th> <th>Hellaswag</th> <th>BoolQ</th> <th>Winogrande</th> <th>OpenbookQA</th> <th>PiQA</th> </tr> <tr> <td><a href="https://huggingface.co/datatab/Yugo55-GPT-v4-4bit/">*Yugo55-GPT-v4-4bit</a></td> <td>51.41</td> <td>36.00</td> <td>57.51</td> <td>80.92</td> <td><strong>65.75</strong></td> <td>34.70</td> <td><strong>70.54</strong></td> </tr> <tr> <td><a href="https://huggingface.co/datatab/Yugo55A-GPT/">Yugo55A-GPT</a></td> <td><strong>51.52</strong></td> <td><strong>37.78</strong></td> <td><strong>57.52</strong></td> <td><strong>84.40</strong></td> <td>65.43</td> <td><strong>35.60</strong></td> <td>69.43</td> </tr> </table>Results obtained through the Serbian LLM evaluation, released by Aleksa Gordić: serbian-llm-eval
- Evaluation was conducted on a 4-bit version of the model due to hardware resource constraints.
This is a merge of pre-trained language models created using mergekit. This model was merged using the linear merge method.
The following models were included in the merge:
The following YAML configuration was used to produce this model:
models:
- model: datatab/Yugo55-GPT-v4
parameters:
weight: 1.0
- model: datatab/Yugo55-GPT-DPO-v1-chkp-300
parameters:
weight: 1.0
- model: mlabonne/AlphaMonarch-7B
parameters:
weight: 0.5
- model: NousResearch/Nous-Hermes-2-Mistral-7B-DPO
parameters:
weight: 0.5
merge_method: linear
dtype: float16
!pip -q install git+https://github.com/huggingface/transformers # need to install from github
!pip install -q datasets loralib sentencepiece
!pip -q install bitsandbytes accelerate
from IPython.display import HTML, display
def set_css():
display(HTML('''
<style>
pre {
white-space: pre-wrap;
}
</style>
'''))
get_ipython().events.register('pre_run_cell', set_css)
import torch
import transformers
from transformers import AutoTokenizer, AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
"datatab/Yugo55A-GPT", torch_dtype="auto"
)
tokenizer = AutoTokenizer.from_pretrained(
"datatab/Yugo55A-GPT", torch_dtype="auto"
)
from typing import Optional
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
def generate(
user_content: str, system_content: Optional[str] = ""
) -> str:
system_content = "Ispod je uputstvo koje opisuje zadatak, upareno sa unosom koji pruža dodatni kontekst. Napišite odgovor koji na odgovarajući način kompletira zahtev."
messages = [
{
"role": "system",
"content": system_content,
},
{"role": "user", "content": user_content},
]
tokenized_chat = tokenizer.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True, return_tensors="pt"
).to("cuda")
text_streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
output = model.generate(
tokenized_chat,
streamer=text_streamer,
max_new_tokens=2048,
temperature=0.1,
repetition_penalty=1.11,
top_p=0.92,
top_k=1000,
pad_token_id=tokenizer.pad_token_id,
eos_token_id=tokenizer.eos_token_id,
do_sample=True,
)
generated_text = tokenizer.decode(output[0], skip_special_tokens=True)
generate("Nabroj mi sve planete suncevog sistemai reci mi koja je najveca planeta")
generate("Koja je razlika između lame, vikune i alpake?")
generate("Napišite kratku e-poruku Semu Altmanu dajući razloge za GPT-4 otvorenog koda")