Downloads · 30 days
0
billingsmoore/getok-v0
getok-v0 is a machine learning model from billingsmoore. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This is a custom Byte Pair Encoding (BPE) tokenizer specifically for Tibetan Buddhist texts. It was trained using the SentencePiece / Hugging Face Tokenizers library. It is designed to tokenize text data efficiently f…
Downloads · 30 days
0
Access
Public
Updated Jul 17, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json2.8 MB · 100%
From the Hugging Face model README
This is a custom Byte Pair Encoding (BPE) tokenizer specifically for Tibetan Buddhist texts. It was trained using the SentencePiece / Hugging Face Tokenizers library. It is designed to tokenize text data efficiently for downstream NLP tasks. The tokenizer supports Unicode text in both Tibetan and English and was trained on a domain-specific corpus of Tibetan Buddhist texts.
This model was developed as part of the MLotsawa project. More information can be found here.
Getok was introduced in the paper Optimizing T5 for Lightweight Tibetan-English Translation
(Moore & Lauren, 2025), where it is used as the tokenizer for the mlotsawa-ground-small and
mlotsawa-ground-base translation models. The paper's code, including how this tokenizer was
trained and evaluated, is available at
optimizing-t5-tibetan-english-mt.
Special thanks to Andres Montano for suggesting the name of this tokenizer.
Tokenizer Type: BPE (Byte Pair Encoding)
Vocabulary Size: 32,000
Normalization: None
Special Tokens: "[PAD]", "[BOS]", "[EOS]", "<unk>"
Tokenization Level: Subword
Languages: Tibetan, English
The tokenizer was trained on a corpus consisting of:
from transformers import PreTrainedTokenizerFast
tokenizer = PreTrainedTokenizerFast.from_pretrained('billingsmoore/getok-v0')
tokenizer.encode('འཇམ་དཔལ་གཞོན་ནུར་གྱུར་པ་ལ་ཕྱག་འཚལ་ལོ༔')
The tokenizer currently supports unicode text in Tibetan or Latin script. However, it was only trained on Tibetan and English texts and should not be expected to perform well on other languages that use those scripts (i.e. Dzongkha, French)
This tokenizer is not suitable for languages that are written in other scripts (i.e. Greek, Russian)
Finetuning a pretrained model using this tokenizer should be expected to take longer than finetuning using the model's own tokenizer because the model will need to adapt to the new encodings.
If you use this tokenizer, please cite the paper it was introduced in:
@article{moore2025optimizing,
title = {Optimizing T5 for Lightweight Tibetan-English Translation},
author = {Moore, Jacob and Lauren, Paula},
year = {2025},
journal = {Research Square},
doi = {10.21203/rs.3.rs-7409829/v1},
url = {https://doi.org/10.21203/rs.3.rs-7409829/v1},
note = {Preprint}
}
Author: billingsmoore
Contact: billingsmoore [at] gmail [dot] com