Downloads · 30 days
0
maritaca-ai/sabia-2-tokenizer-small
sabia-2-tokenizer-small is a machine learning model from maritaca-ai. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This is the tokenizer used by the Sabia-2 Small model.
Downloads · 30 days
0
Access
Public
Updated Feb 23, 2024
Repo size
—
Likes
1
Public
Click a slice to open those files.
.json1.8 MB · 100%
From the Hugging Face model README
This is the tokenizer used by the Sabia-2 Small model.
Sabiá-2 Small is a proprietary LLM that can be used through an API endpoint, which we refer to as the "MariTalk API", or a downloadable version that can be used locally and is encrypted, known as "MariTalk Local".
The purpose of including this tokenizer is to allow you to estimate the number of tokens in your prompts and, therefore, the cost of using the model.
For example:
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("maritaca-ai/sabia-2-tokenizer-small")
prompt = "Com quantos paus se faz uma canoa?"
tokens = tokenizer.encode(prompt)
print(f'O prompt "{prompt}" contém {len(tokens)} tokens.') # It should print 11 tokens.
For more information on how to use the model, please refer to our documentation at this link.