Downloads · 30 days
68
7% of all-time downloads
Gen2B/HyGPT-10b-it
HyGPT-10b-it is a text generation model from Gen2B. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
HyGPT-10b-it is an instruction-tuned version of HyGPT-10b, the first Armenian large language model that was pretrained on a corpus of Armenian text data. This model has been fine-tuned on a diverse instruction dataset…
Downloads · 30 days
68
7% of all-time downloads
All-time downloads
1K
Public
Parameters
10.2B
20.4 GB on disk
Likes
9
Public
Click a slice to open those files.
.safetensors20.3 GB · 100%
From the Hugging Face model README
HyGPT-10b-it is an instruction-tuned version of HyGPT-10b, the first Armenian large language model that was pretrained on a corpus of Armenian text data. This model has been fine-tuned on a diverse instruction dataset to enhance its ability to follow instructions, engage in multi-turn conversations, and perform various language tasks in Armenian, Russian, and English.
HyGPT-10b-it is a decoder-only language model based on the HyGPT-10b base model that was first pretrained on 10B tokens of Armenian text and then instruction-tuned (SFT) on a diverse dataset of 50,000 instruction samples.
First, install the Transformers library with:
pip install -U transformers
Then, run this example:
from transformers import AutoModelForCausalLM, AutoTokenizer, GenerationConfig
import torch
model_path = "Gen2B/HyGPT-10b-it"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype=torch.float16,
device_map="auto",
)
# Example of a single-turn conversation
chat = [
{"role": "user", "content": "Ինչու է խոտը Կանաչ:"}
]
# Example of a multi-turn conversation
# chat = [
# {"role": "user", "content": "Բարև, ինչպե՞ս ես:"},
# {"role": "assistant", "content": "Բարև, ես լավ եմ: Ինչով կարող եմ օգնել քեզ այսօր:"},
# {"role": "user", "content": "Ինչու է խոտը Կանաչ:"}
# ]
PROMPT = tokenizer.apply_chat_template(chat, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(
PROMPT,
return_tensors="pt",
)
print("Generating...")
generation_output = model.generate(
input_ids=inputs["input_ids"].cuda(),
generation_config=GenerationConfig(
temperature=0.0001,
repetition_penalty=1.1,
do_sample=True
),
return_dict_in_generate=True,
output_scores=True,
max_new_tokens=1024,
)
for s in generation_output.sequences:
print(tokenizer.decode(s))
# Խոտի մեջ կան պիգմենտներ, որոնք կլանում են լույսի միայն կարճ ալիքները և անդրադարձնում երկար ալիքները։ Դրանք նաև բաց են թողնում ուլտրամանուշակագույն և ինֆրակարմիր ալիքները։ Մարդու աչքերը զգայուն չեն այս ալիքների նկատմամբ, ուստի դրանք տեսանելի չեն։ Այսպիսով, երբ արևի լույսը հարվածում է խոտին, այն անդրադարձնում է երկար ալիքները՝ առաջացնելով կանաչ գույնը, որը մենք տեսնում ենք:
HyGPT-10b-it can be used directly for:
The base model (HyGPT-10b) was pretrained on a diverse corpus of Armenian text data comprising approximately 10 billion tokens, including:
The model was fine-tuned on a diverse instruction dataset consisting of 50,000 samples with the following characteristics:
Dataset Composition:
Task Types:
The instruction tuning data underwent several preprocessing steps:
The model was fine-tuned from the HyGPT-10b base model using supervised fine-tuning (SFT) techniques. The training focused on teaching the model to:
The model was evaluated on several standard benchmarks that were translated into Armenian to accurately assess its performance in the target language. The benchmarks include:
Below is a table of accuracy of different models on 4 benchmarks. The results demonstrate significant improvements over the base model across these tasks:
| Gen2B/HyGPT-10b-it | google/gemma-3-12b-it | mistralai/Mistral-Small-3.1-24B-Instruct-2503 | google/gemma-2-9b-it | mistralai/Mistral-Nemo-Instruct-2407 | meta-llama/Llama-3.1-8B-Instruct | |
|---|---|---|---|---|---|---|
| Flores | 79.33 | 80.59 | 80.62 | 78.61 | 79.1 | 77.67 |
| ARC | 76.1 | 79.42 | 81.76 | 72.54 | 73.2 | 58.91 |
| Truthful QA | 72.83 | 65.52 | 67.98 | 67.49 | 39.9 | 39.41 |
| GSM8K | 68.0 | 65.8 | 41.07 | 38.0 | 44.19 | 17.22 |
| avg | 74.06 | 72.83 | 67.86 | 64.16 | 59.1 | 48.3 |
The instruction-tuned model demonstrates significantly improved capabilities in following instructions and engaging in conversations compared to the base model. It shows enhanced abilities in:
HyGPT-10b-it builds upon the strong foundation of HyGPT-10b to provide a more interactive and instruction-following Armenian language model. It is particularly well-suited for conversational applications, educational tools, and multilingual assistance systems that require Armenian language support.
This model is based on Gemma and is distributed according to the Gemma Terms of Use.
Notice: Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms.
This model is a modified version of the original Gemma-2-9b model. The modifications include:
According to the Gemma Terms of Use, the model should not be used:
UNLESS REQUIRED BY APPLICABLE LAW, THE GEMMA SERVICES, AND OUTPUTS, ARE PROVIDED ON AN "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, EITHER EXPRESS OR IMPLIED, INCLUDING ANY WARRANTIES OR CONDITIONS OF TITLE, NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE. YOU ARE SOLELY RESPONSIBLE FOR DETERMINING THE APPROPRIATENESS OF USING, REPRODUCING, MODIFYING, PERFORMING, DISPLAYING OR DISTRIBUTING ANY OF THE GEMMA SERVICES OR OUTPUTS AND ASSUME ANY AND ALL RISKS ASSOCIATED WITH YOUR USE OR DISTRIBUTION OF ANY OF THE GEMMA SERVICES OR OUTPUTS AND YOUR EXERCISE OF RIGHTS AND PERMISSIONS UNDER THIS AGREEMENT.