Downloads · 30 days
0
alexanderpl/gemma2_2b_gec_v1
gemma2_2b_gec_v1 is a text generation model from alexanderpl. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This model is a fine-tuned version of unsloth/gemma-2-2b-bnb-4bit on the p1746-lingua/ru-gec-v1 dataset. It is designed for Grammatical Error Correction (GEC) for Russian texts, generating corrected versions of input…
Downloads · 30 days
0
Access
Public
Updated Mar 8, 2026
Repo size
122 MB
Likes
0
Public
Click a slice to open those files.
.safetensors83.1 MB · 68%
From the Hugging Face model README
This model is a fine-tuned version of unsloth/gemma-2-2b-bnb-4bit on the p1746-lingua/ru-gec-v1 dataset. It is designed for Grammatical Error Correction (GEC) for Russian texts, generating corrected versions of input sentences with grammatical, spelling, and punctuation errors.
google/gemma-2-2b (via Unsloth's 4-bit quantized version)p1746-lingua/ru-gec-v1 dataset, consisting of approximately 707,000 sentence pairs (erroneous → corrected).| Parameter | Value |
|---|---|
| Batch Size | 32 |
| Learning Rate | 1e-5 |
| Total Epochs | 10,000 |
| Warmup Steps | 100 |
| Optimizer | adamw_bnb_8bit |
The model operates in a standard text-to-text way. Here are examples using the Hugging Face transformers pipeline and direct inference.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
peft_model_id = "p1746-lingua/gemma2-2b-gec-v1"
base_model_id = "unsloth/gemma-2-2b-bnb-4bit"
tokenizer = AutoTokenizer.from_pretrained(peft_model_id)
# Load base model
base_model = AutoModelForCausalLM.from_pretrained(
base_model_id,
torch_dtype=torch.float16,
device_map="auto"
)
model = PeftModel.from_pretrained(base_model, peft_model_id)
examples = [
"Я будуш делать задание завтра.",
"Она купила три яблоки.",
"Это моя лучшая друзья.",
"Мы ходили в кино вчерашний день."
]
def correct_sentence(sentence, max_length=200):
prompt = f"Correct this Russian sentence: {sentence}\nCorrected:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=max_length,
do_sample=False,
num_return_sequences=1,
pad_token_id=tokenizer.pad_token_id,
eos_token_id=tokenizer.eos_token_id,
)
generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
if "Corrected:" in generated_text:
corrected = generated_text.split("Corrected:")[1].strip()
else:
corrected = generated_text.replace(prompt, "").strip()
return corrected
for sentence in examples:
corrected = correct_sentence(sentence)
print(f"Input: {sentence}")
print(f"Output: {corrected}\n")
# Я буду делать задание завтра.
# Она купила три яблока.
# Это моя лучшая подруга.
# Мы ходили в кино вчера.
Disclaimer: This model is a tool to assist with writing. Its output should be reviewed by a human, especially in critical or formal contexts.
Модель дообучена на unsloth/gemma-2-2b-bnb-4bit на датасете p1746-lingua/ru-gec-v1. Модель предназначена для исправления грамматических ошибок (GEC) в русских текстах, она генерирует исправленные версии входных предложений с грамматическими, орфографическими и пунктуационными ошибками.
google/gemma-2-2b (4-битная квантизированная версия от Unsloth)p1746-lingua/ru-gec-v1, состоящий приблизительно из 707 000 пар предложений (с ошибкой → исправленное).| Параметр | Значение |
|---|---|
| Размер батча | 32 |
| Скорость обучения | 1e-5 |
| Всего эпох | 10 000 |
| Шагов warmup | 100 |
| Оптимизатор | adamw_bnb_8bit |
Модель работает стандартным способом. Ниже приведены примеры использования пайплайна Hugging Face transformers.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
peft_model_id = "p1746-lingua/gemma2-2b-gec-v1"
base_model_id = "unsloth/gemma-2-2b-bnb-4bit"
tokenizer = AutoTokenizer.from_pretrained(peft_model_id)
# Load base model
base_model = AutoModelForCausalLM.from_pretrained(
base_model_id,
torch_dtype=torch.float16,
device_map="auto"
)
model = PeftModel.from_pretrained(base_model, peft_model_id)
examples = [
"Я будуш делать задание завтра.",
"Она купила три яблоки.",
"Это моя лучшая друзья.",
"Мы ходили в кино вчерашний день."
]
def correct_sentence(sentence, max_length=200):
prompt = f"Correct this Russian sentence: {sentence}\nCorrected:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=max_length,
do_sample=False,
num_return_sequences=1,
pad_token_id=tokenizer.pad_token_id,
eos_token_id=tokenizer.eos_token_id,
)
generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
if "Corrected:" in generated_text:
corrected = generated_text.split("Corrected:")[1].strip()
else:
corrected = generated_text.replace(prompt, "").strip()
return corrected
for sentence in examples:
corrected = correct_sentence(sentence)
print(f"Input: {sentence}")
print(f"Output: {corrected}\n")
# Я буду делать задание завтра.
# Она купила три яблока.
# Это моя лучшая подруга.
# Мы ходили в кино вчера.
Дисклеймер: эта модель является инструментом для помощи в написании текстов. Ее вывод должен проверяться человеком, особенно в критически важных областях.