Downloads · 30 days
14
15% of all-time downloads
raptorkwok/pyctokenizer
pyctokenizer is a translation model from raptorkwok. Use it when you need text moved from one language to another. The card lists the license as apache-2.0.
This is a Cantonese sentence tokenizer based on BART Chinese. It can be used along with our CCPC Parallel Corpus dataset.
Downloads · 30 days
14
15% of all-time downloads
All-time downloads
95
Public
Parameters
225M
939 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors899 MB · 95%
From the Hugging Face model README
This is a Cantonese sentence tokenizer based on BART Chinese. It can be used along with our CCPC Parallel Corpus dataset.
from transformers import BertTokenizer
tokenizer = BertTokenizer.from_pretrained("raptorkwok/pyctokenizer")
print(tokenizer.tokenize("我哋去咗尖沙咀睇醫生呀!"))
# Output: ['我哋', '去', '咗', '尖沙咀', '睇醫生', '呀', '!']