Downloads · 30 days
0
yhavinga/dutch-llama-tokenizer
dutch-llama-tokenizer is a machine learning model from yhavinga. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
The Dutch-Llama Tokenizer is a versatile tokenizer trained to handle a variety of languages and formats, including Dutch, English, Python code, Markdown, and general text. It's based on a dataset consisting of diverse…
Downloads · 30 days
0
Access
Public
Updated Jan 4, 2024
Repo size
—
Likes
2
Public
Click a slice to open those files.
.json2 MB · 100%
From the Hugging Face model README
The Dutch-Llama Tokenizer is a versatile tokenizer trained to handle a variety of languages and formats, including Dutch, English, Python code, Markdown, and general text. It's based on a dataset consisting of diverse sources, which ensures its capability to tokenize a wide range of text inputs effectively.
The tokenizer was trained on a comprehensive dataset, including:
The tokenizer was trained using the spm_train command with the following settings:
To use the Dutch-Llama Tokenizer, ensure you have Python 3.10.12 or later installed. Then, install the Transformers library from Hugging Face:
pip install transformers
First, import the AutoTokenizer from the Transformers library and load the Dutch-Llama Tokenizer:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("yhavinga/dutch-llama-tokenizer")
To tokenize text, use the tokenizer.tokenize method. For converting tokens to IDs and decoding them back to text, use tokenizer.convert_tokens_to_ids and tokenizer.decode respectively:
# Example text
text = "Steenvliegen of oevervliegen[2] (Plecoptera) 华为发布Mate60手机"
# Tokenization and decoding
tokens = tokenizer.tokenize(text)
token_ids = tokenizer.convert_tokens_to_ids(tokens)
decoded_text = tokenizer.decode(token_ids)
print(decoded_text)
Compare the effectiveness of this tokenizer on different inputs at the Hugging Face Space: Dutch Tokenizer Arena.
The following table shows the number of tokens produced by the Dutch-Llama Tokenizer, the Mistral Tokenizer, the GroNLP GPT-2 Dutch Tokenizer, and the UL2 Dutch Tokenizer on a variety of inputs.
| Input Type | Dutch LLama (32k) | Mistral (32k) | GroNLP GPT-2 Dutch (40k) | UL2 Dutch (32k) |
|---|---|---|---|---|
| Dutch news | 440 | 658 | 408 | 410 |
| English news | 414 | 404 | 565 | 402 |
| Code python | 566 | 582 | 767 | 639 (no newlines) |
| LaTeX math | 491 | 497 | 717 | 666 (no newlines) |
| Total | 1911 | 2141 | 2457 | 2117 |
🇳🇱 🇧🇪🐍📐