Downloads · 30 days
29
50% of all-time downloads
haryads/kurdish-roberta
kurdish-roberta is a fill-mask model from haryads. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
A RoBERTa language model for Central Kurdish (Sorani), trained from scratch on a 1-million-sentence Kurdish corpus with a custom 32k Kurdish BPE tokenizer. It's a masked-language model meant as a base for fine-tuning…
Downloads · 30 days
29
50% of all-time downloads
All-time downloads
58
Public
Parameters
35.7M
143 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors143 MB · 98%
From the Hugging Face model README
A RoBERTa language model for Central Kurdish (Sorani), trained from scratch on a 1-million-sentence Kurdish corpus with a custom 32k Kurdish BPE tokenizer. It's a masked-language model meant as a base for fine-tuning on Kurdish tasks (classification, NER, embeddings, and so on).
It's the base model behind haryads/kurdish-sentence-embeddings.
from transformers import pipeline
fill = pipeline("fill-mask", model="haryads/kurdish-roberta")
fill("زمانی کوردی زۆر <mask> ە")
Or load it directly to fine-tune:
from transformers import AutoModel, AutoTokenizer
tok = AutoTokenizer.from_pretrained("haryads/kurdish-roberta")
model = AutoModel.from_pretrained("haryads/kurdish-roberta")
| Type | RoBERTa (masked LM) |
| Layers | 6 |
| Hidden size | 512 |
| Attention heads | 8 |
| Max sequence length | 128 |
| Vocab size | 32,000 |
| Parameters | ~35M |
A base model, not an instruct or chat model. Best used fine-tuned for a specific Kurdish task. Sorani only, and the corpus is news-heavy, so very colloquial or domain-specific text may need in-domain fine-tuning.
Apache-2.0.