Downloads · 30 days
13
3% of all-time downloads
aviadrom/HeArBERT
HeArBERT is a feature extraction model from aviadrom. Use it when you need embeddings to search or compare text. It is set up for transformers.
A bilingual BERT for Arabic and Hebrew, pretrained on the respective parts of the OSCAR corpus.
Downloads · 30 days
13
3% of all-time downloads
All-time downloads
410
Public
Repo size
873 MB
Likes
4
Public
Click a slice to open those files.
.bin436 MB · 100%
From the Hugging Face model README
A bilingual BERT for Arabic and Hebrew, pretrained on the respective parts of the OSCAR corpus.
In order to process Arabic with this model, one would have to transliterate it to Hebrew script. The code for doing so is available on the preprocessing file and can be used as follows:
from transformers import AutoTokenizer
from preprocessing import transliterate_arabic_to_hebrew
tokenizer = AutoTokenizer.from_pretrained("aviadrom/HeArBERT")
text_ar = "مرحبا"
text_he = transliterate_arabic_to_hebrew(text_ar)
tokenizer(text_he)
If you find our work useful in your research, please consider citing:
@article{rom2024training,
title={Training a Bilingual Language Model by Mapping Tokens onto a Shared Character Space},
author={Rom, Aviad and Bar, Kfir},
journal={arXiv preprint arXiv:2402.16065},
year={2024}
}