Downloads · 30 days
14
2% of all-time downloads
Nuri-Tas/roberturk-base
roberturk-base is a feature extraction model from Nuri-Tas. Use it when you need embeddings to search or compare text. It is set up for transformers.
RoBERTurk is pretrained on Oscar Turkish Split (27GB) and a small chunk of C4 Turkish Split (1GB) with sentencepiece BPE tokenizer that is trained on randomly selected 30M sentences from the training data, which is co…
Downloads · 30 days
14
2% of all-time downloads
All-time downloads
746
Public
Parameters
125M
2.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors499 MB · 98%
From the Hugging Face model README
RoBERTurk is pretrained on Oscar Turkish Split (27GB) and a small chunk of C4 Turkish Split (1GB) with sentencepiece BPE tokenizer that is trained on randomly selected 30M sentences from the training data, which is composed of 90M sentences. The training data in total contains 5.3B tokens and the vocabulary size is 50K. The learning rate is warmed up to the peak value of 1e-5 for the first 10K updates and linearly decayed at $0.01$ rate. The model is pretrained for maximum 600K updates only with sequences of at most T=256 length.
Load the pretrained tokenizer as follows:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Nuri-Tas/roberturk-base")
Get the pretrained model with:
from transformers import RobertaModel
model = RobertaModel.from_pretrained("Nuri-Tas/roberturk-base")
There is a slight mismatch between our tokenizer and the default tokenizer used by RobertaTokenizer, which results in some underperformance. I'm working on the issue and will update the tokenizer/model accordingly.
Additional TODOs are (although some of them can take some time and I may include them on different repositories):