Downloads · 30 days
93
0% of all-time downloads
maximxls/text-normalization-ru-terrible
text-normalization-ru-terrible is a machine learning model from maximxls. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
Normalization for Russian text. Couldn't find any existing solutions (besides algorithms, don't like those) so made this.
Downloads · 30 days
93
0% of all-time downloads
All-time downloads
29.2K
Public
Parameters
8.4M
67.9 MB on disk
Likes
4
Public
Click a slice to open those files.
.bin33.6 MB · 49%
From the Hugging Face model README
Normalization for Russian text. Couldn't find any existing solutions (besides algorithms, don't like those) so made this.
Tiny T5 trained from scratch for normalizing Russian texts:
Useful in TTS, for example with Silero to make it read numbers and English words (even if not perfectly, it's at least not ignoring)
from transformers import (
T5ForConditionalGeneration,
PreTrainedTokenizerFast,
)
model_path = "maximxls/text-normalization-ru-terrible"
tokenizer = PreTrainedTokenizerFast.from_pretrained(model_path)
model = T5ForConditionalGeneration.from_pretrained(model_path)
example_text = "Я ходил в McDonald's 10 июля 2022 года."
inp_ids = tokenizer(
example_text,
return_tensors="pt",
).input_ids
out_ids = model.generate(inp_ids, max_new_tokens=128)[0]
out = tokenizer.decode(out_ids, skip_special_tokens=True)
print(out)
я ходил в макдоналд'эс десятого июля две тысячи двадцать второго года.
Very much unreliable:
Data from this Kaggle challenge (761435 sentences) aswell as a bit of extra data written by me.
See preprocessing.py
See train.py
I have reset lr manually several times during training, see metrics.
See README on github for a step-by-step overview of the training procedure.
Couple tens of hours of RTX 3090 Ti compute on my personal PC (21.65 epochs)