Downloads · 30 days
0
ocisd4/openllama_tokenizer_ext_zh
openllama_tokenizer_ext_zh is a machine learning model from ocisd4. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
note: - The first token might be a whitespace in LLamaTokenizer. - Open LlaMa的tokenizer is incompatible with original LlaMa - This tokenizer will encode continuous spaces to ONE space
Downloads · 30 days
0
Access
Public
Updated Jun 6, 2023
Repo size
6.6 MB
Likes
0
Public
Click a slice to open those files.
.model828 KB · 100%
From the Hugging Face model README
from transformers import LlamaTokenizer
tokenizer = LlamaTokenizer.from_pretrained(
'ocisd4/openllama_tokenizer_ext_zh',
add_bos_token=True,
add_eos_token=False,
use_auth_token='True',
)
print('vocab size:',tokenizer.vocab_size)
#vocab size: 52928
text = '今天天氣真好!'
print(tokenizer.tokenize(text))
#['▁', '今天', '天氣', '真', '好', '<0xEF>', '<0xBC>', '<0x81>']
print(tokenizer.encode(text))
#[1, 31822, 32101, 32927, 45489, 45301, 242, 191, 132]
print(tokenizer.decode(tokenizer.encode(text)))
# 今天天氣真好!</s>
** note: **