Downloads · 30 days
0
eligapris/kirundi-tokenizer
kirundi-tokenizer is a machine learning model from eligapris. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This is a SentencePiece-based tokenizer model trained for the Kirundi language. It can be used for tokenizing text in Kirundi for NLP tasks.
Downloads · 30 days
0
Access
Public
Updated Dec 6, 2024
Repo size
523 KB
Likes
0
Public
Click a slice to open those files.
.model523 KB · 63%
From the Hugging Face model README
This is a SentencePiece-based tokenizer model trained for the Kirundi language. It can be used for tokenizing text in Kirundi for NLP tasks.
The tokenizer was trained on a diverse corpus of Kirundi text collected from various sources. The data was preprocessed to remove any unwanted characters and cleaned for tokenization.
import sentencepiece as spm
# Load the tokenizer
sp = spm.SentencePieceProcessor(model_file='kirundi.model')
# Tokenize text
text = "Ndakunda igihugu canje."
tokens = sp.encode(text, out_type=str)
print(tokens)
# Detokenize text
decoded_text = sp.decode(tokens)
print(decoded_text)