Downloads · 30 days
0
0% of all-time downloads
IsmaelMousa/arabic-bpe-tokenizer
arabic-bpe-tokenizer is a summarization model from IsmaelMousa. Use it when you need a shorter version of a longer text. It is set up for tokenizers. The card lists the license as mit.
Byte Level Tokenizer for Arabic, a robust tokenizer designed to handle Arabic text with precision and efficiency. This tokenizer utilizes a Byte-Pair Encoding (BPE) approach to create a vocabulary of 50,000 tokens, ca…
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
1
Public
Repo size
—
Likes
1
Public
Click a slice to open those files.
.json3.6 MB · 100%
From the Hugging Face model README
Byte Level Tokenizer for Arabic, a robust tokenizer designed to handle Arabic text with precision and efficiency.
This tokenizer utilizes a Byte-Pair Encoding (BPE) approach to create a vocabulary of 50,000 tokens, catering specifically to the intricacies of the Arabic language.
This tokenizer was created as part of the development of an Arabic BART transformer model for summarization from scratch using PyTorch.
In adherence to the configurations outlined in the official BART paper, which specifies the use of BPE tokenization, I sought a BPE tokenizer specifically tailored for Arabic.
While there are Arabic-only tokenizers and multilingual BPE tokenizers, a dedicated Arabic BPE tokenizer was not available. This gap inspired the creation of a BPE tokenizer focused solely on Arabic, ensuring alignment with BART's recommended configurations and enhancing the effectiveness of Arabic NLP tasks.
IsmaelMousa/arabic-bpe-tokenizer50,000The Byte Level Tokenizer is optimized to manage Arabic text, which often includes a range of diacritics, different forms of the same word, and various prefixes and suffixes. This tokenizer addresses these challenges by breaking down text into byte-level tokens, ensuring that it can effectively process and understand the nuances of the Arabic language.
tokenizers library for seamless tokenization.To use this tokenizer, you need to install the tokenizers library. If you haven’t installed it yet, you can do so using pip:
pip install tokenizers
Here is an example of how to use the Byte Level Tokenizer with the tokenizers library.
This example demonstrates tokenization of the Arabic sentence "لاشيء يعجبني, أريد أن أبكي":
from tokenizers import Tokenizer
tokenizer = Tokenizer.from_pretrained("IsmaelMousa/arabic-bpe-tokenizer")
text = "لاشيء يعجبني, أريد أن أبكي"
encoded = tokenizer.encode(text)
decoded = tokenizer.decode(encoded.ids)
print("Encoded Tokens:", encoded.tokens)
print("Token IDs:", encoded.ids)
print("Decoded Text:", decoded)
output:
Encoded Tokens: ['<s>', 'ÙĦا', 'ĠØ´ÙĬØ¡', 'ĠÙĬع', 'جب', 'ÙĨÙĬ', ',', 'ĠأرÙĬد', 'ĠØ£ÙĨ', 'Ġأب', 'ÙĥÙĬ', '</s>']
Token IDs: [0, 419, 1773, 667, 2281, 489, 16, 7578, 331, 985, 1344, 2]
Decoded Text: لا شيء يعجبني, أريد أن أبكي
This project is licensed under the MIT License.