Downloads · 30 days
40
1% of all-time downloads
sdadas/xlm-roberta-large-twitter
xlm-roberta-large-twitter is a fill-mask model from sdadas. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as mit.
This is a XLM-RoBERTa-large model tuned on a corpus of over 156 million tweets in ten languages: English, Spanish, Italian, Portuguese, French, Chinese, Hindi, Arabic, Dutch and Korean. The model has been trained from…
Downloads · 30 days
40
1% of all-time downloads
All-time downloads
3.4K
Public
Parameters
560M
4.5 GB on disk
Likes
2
Public
Click a slice to open those files.
.bin2.2 GB · 50%
How the weights are stored.
F32560M · 100%
From the Hugging Face model README
This is a XLM-RoBERTa-large model tuned on a corpus of over 156 million tweets in ten languages: English, Spanish, Italian, Portuguese, French, Chinese, Hindi, Arabic, Dutch and Korean. The model has been trained from the original XLM-RoBERTA-large checkpoint for 2 epochs with a batch size of 1024.
For best results, preprocess the tweets using the following method before passing them to the model:
def preprocess(text):
new_text = []
for t in text.split(" "):
t = '@user' if t.startswith('@') and len(t) > 1 else t
t = 'http' if t.startswith('http') else t
new_text.append(t)
return " ".join(new_text)