Downloads · 30 days
67
1% of all-time downloads
Ahmad/parsT5-base
parsT5-base is a machine learning model from Ahmad. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
A monolingual T5 model for Persian trained on OSCAR 21.09 (https://oscar-corpus.com/) corpus with self-supervised method. 35 Gig deduplicated version of Persian data was used for pre-training the model.
Downloads · 30 days
67
1% of all-time downloads
All-time downloads
9.4K
Public
Repo size
2 GB
Likes
6
Public
Click a slice to open those files.
.bin990 MB · 100%
From the Hugging Face model README
A monolingual T5 model for Persian trained on OSCAR 21.09 (https://oscar-corpus.com/) corpus with self-supervised method. 35 Gig deduplicated version of Persian data was used for pre-training the model.
It's similar to the English T5 model but just for Persian. You may need to fine-tune it on your specific task.
Example code:
from transformers import T5ForConditionalGeneration,AutoTokenizer
import torch
model_name = "Ahmad/parsT5-base"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = T5ForConditionalGeneration.from_pretrained(model_name)
input_ids = tokenizer.encode('دانش آموزان به <extra_id_0> میروند و <extra_id_1> میخوانند.', return_tensors='pt')
with torch.no_grad():
hypotheses = model.generate(input_ids)
for h in hypotheses:
print(tokenizer.decode(h))
Steps: 725000
Accuracy: 0.66
To train the model further please refer to its github repository at: