Downloads · 30 days
5
29% of all-time downloads
RimasZzz/agriculture_text_transform_model
agriculture_text_transform_model is a machine learning model from RimasZzz. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Модель на базе cointegrated/rut5-base-multitask, дообученная на паре input → output из Excel-файла, предназначена для трансформации аграрных или иных текстов: исправления, нормализации, переформулирования и т.п.
Downloads · 30 days
5
29% of all-time downloads
All-time downloads
17
Public
Parameters
244M
978 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors977 MB · 100%
From the Hugging Face model README
Модель на базе cointegrated/rut5-base-multitask, дообученная на паре input → output из Excel-файла, предназначена для трансформации аграрных или иных текстов: исправления, нормализации, переформулирования и т.п.
cointegrated/rut5-base-multitaskagriculture_text_transform_model/
├── config.json
├── generation_config.json
├── model.safetensors
├── special_tokens_map.json
├── spiece.model
├── tokenizer_config.json
Обучение производилось на кастомном датасете в формате Excel (train.xlsx) без заголовков, где:
input — исходный текст,output — целевой текст.cointegrated/rut5-base-multitask9516128Adam (lr=1e-5)import torch
from transformers import T5ForConditionalGeneration, T5Tokenizer
model = T5ForConditionalGeneration.from_pretrained("RimasZzz/agriculture_text_transform_model").cuda()
tokenizer = T5Tokenizer.from_pretrained("RimasZzz/agriculture_text_transform_model")
def transform_text(text, **kwargs):
inputs = tokenizer(text, return_tensors='pt', truncation=True, max_length=128).to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_length=128,
num_beams=5,
repetition_penalty=2.5,
**kwargs
)
return tokenizer.decode(outputs[0], skip_special_tokens=True)
# Пример:
transform_text("Пахота зяби под мн тр")
# → "Пахота зяби под Многолетние травы"
license: apache-2.0 language: