Downloads · 30 days
0
Sigurdur/ice-tokenizer
ice-tokenizer is a machine learning model from Sigurdur. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This BPE (Byte Pair Encoding) tokenizer is designed for the Icelandic GPT model, available at Sigurdur/ice-gpt. Trained on the Icelandic Gigaword Corpus ({IGC}-2022) - annotated version, it excels in accurately segmen…
Downloads · 30 days
0
Access
Public
Updated Nov 21, 2023
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json1.6 MB · 88%
From the Hugging Face model README
This BPE (Byte Pair Encoding) tokenizer is designed for the Icelandic GPT model, available at Sigurdur/ice-gpt. Trained on the Icelandic Gigaword Corpus ({IGC}-2022) - annotated version, it excels in accurately segmenting Icelandic text into meaningful tokens.
Integrate this tokenizer into your NLP pipeline for preprocessing Icelandic text. The following example demonstrates basic usage:
from transformers import GPT2Tokenizer
# Load the tokenizer
tokenizer = GPT2Tokenizer.from_pretrained("Sigurdur/ice-tokenizer")
tokenizer.pad_token_id = tokenizer.eos_token_id
tokenizer("Halló heimur!")["input_ids"]
If you use this tokenizer in your work, please cite the original source of the training data:
@misc{20.500.12537/254,
title = {Icelandic Gigaword Corpus ({IGC}-2022) - annotated version},
author = {Barkarson, Starkaður and Steingrímsson, Steinþór and Andrésdóttir, Þórdís Dröfn and Hafsteinsdóttir, Hildur and Ingimundarson, Finnur Ágúst and Magnússon, Árni Davíð},
url = {http://hdl.handle.net/20.500.12537/254},
note = {{CLARIN}-{IS}},
year = {2022}
}
We welcome user feedback to enhance the tokenizer's functionality. Feel free to reach out with your insights and suggestions.
Happy tokenizing!
Sigurdur Haukur Birgisson
(readme created with chatgpt)