Downloads · 30 days
329
8% of all-time downloads
arm-on/BERTweet-FA
BERTweet-FA is a fill-mask model from arm-on. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
BERTweet-FA: A pre-trained language model for Persian (a.k.a Farsi) Tweets ---
Downloads · 30 days
329
8% of all-time downloads
All-time downloads
4K
Public
Repo size
871 MB
Likes
6
Public
Click a slice to open those files.
.bin435 MB · 100%
From the Hugging Face model README
BERTweet-FA is a transformer-based model trained on 20665964 Persian tweets. The model has been trained on the data only for 1 epoch (322906 steps), and yet it has the ability to recognize the meaning of most of the conversational sentences used in Farsi. Note that the architecture of this model follows the original BERT [Devlin et al.].
from transformers import BertForMaskedLM, BertTokenizer, pipeline
model = BertForMaskedLM.from_pretrained('arm-on/BERTweet-FA')
tokenizer = BertTokenizer.from_pretrained('arm-on/BERTweet-FA')
fill_sentence = pipeline('fill-mask', model=model, tokenizer=tokenizer)
fill_sentence('اینجا جمله مورد نظر خود را بنویسید و کلمه موردنظر را [MASK] کنید')
The first version of the model was trained on the "Large Scale Colloquial Persian Dataset" containing more than 20 million tweets in Farsi, gathered by Khojasteh et al., and published on 2020.
| Training Loss | Epoch | Step |
|---|---|---|
| 0.0036 | 1.0 | 322906 |