Downloads · 30 days
0
BeitTigreAI/tigre-spm-tokenizer
tigre-spm-tokenizer is a machine learning model from BeitTigreAI. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
A standalone SentencePiece (Unigram/BPE) subword tokenizer built and optimized for the Tigre language (ትግሬ) written in Ge'ez script.
Downloads · 30 days
0
Access
Public
Updated Aug 9, 2026
Repo size
585 KB
Likes
0
Public
Click a slice to open those files.
.json1.1 MB · 54%
From the Hugging Face model README
A standalone SentencePiece (Unigram/BPE) subword tokenizer built and optimized for the Tigre language (ትግሬ) written in Ge'ez script.
| Metric | Value | Status |
|---|---|---|
| Vocabulary Size | 16,384 | ✅ Standard |
| Round-Trip Integrity | 100% Lossless | ✅ Verified |
Unknown Token (<unk>) Rate | 0.00% | ✅ Verified |
| Punctuation Isolation | Clean | ✅ Isolated |
Ge'ez Wordspace (፡) Handling | Atomic Token | ✅ Verified |
| Base Architecture | SentencePiece | ✅ Native |
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("BeitTigreAI/tigre-spm-tokenizer")
text = "ሰላም፡ እሊ ናይ ሂጋ ትግሬ ክታበት ቱ።"
# Encode to tokens & IDs
tokens = tokenizer.tokenize(text)
input_ids = tokenizer.encode(text)
print("Tokens:", tokens)
print("IDs :", input_ids)
# Lossless Decoding
decoded_text = tokenizer.decode(input_ids, clean_up_tokenization_spaces=False)
print("Decoded:", decoded_text)
assert text == decoded_text
tokenizer.model: Native SentencePiece binary model file.tokenizer.json: Serialized fast tokenizer representation for Python/Rust environments.tokenizer_config.json: Metadata and special token configuration mapping.special_tokens_map.json: Explicit <pad>, <s>, </s>, and <unk> assignments.