Downloads · 30 days
0
KateMajzel/tokenizer-pl-32k
tokenizer-pl-32k is a machine learning model from KateMajzel. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
Byte-level BPE for Polish. 32,568 vocabulary entries + 200 special tokens = 32,768 (2¹⁵, convenient for GPU kernels and tensor parallelism).
Downloads · 30 days
0
Access
Public
Updated Sep 7, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json2.4 MB · 100%
From the Hugging Face model README
Byte-level BPE for Polish. 32,568 vocabulary entries + 200 special tokens = 32,768 (2¹⁵, convenient for GPU kernels and tensor parallelism).
Trained on a Polish corpus and validated experimentally in the gollem-pl project — a controlled ablation comparing it against the GPT-2 tokenizer under an identical training budget.
Model trained with this tokenizer: KateMajzel/GoLLeM-45M-PL.
| this tokenizer | GPT-2 | |
|---|---|---|
| bytes/token — mixed PL corpus (2.96 GB) | 4.050 | 2.066 |
| bytes/token — held-out (2.77 MB) | 4.103 | 2.134 |
| tokens from 2.96 GB of text | 731M | 1,434M |
| relative density | 1.96× | 1.00× |
Fertility on Polish prose: ~1.35 tokens per word (5.80 characters/token). Note: on a real corpus — with URLs, numbers and leftover formatting — the figure comes out at 4.05 bytes/token, which is markedly worse. Fertility should be reported from a corpus, not from hand-picked sentences.
A practical example of the difference: the word niejednoznaczny is 3 tokens here and 9 under GPT-2.
| type | byte-level BPE (tokenizers, tokenizer.json format) |
| normalization | none — lossless round-trip |
| pre-tokenizer | GPT-4 style: Isolated + ByteLevel(use_regex=false, add_prefix_space=false) |
| digits | split into groups of at most 3 (\p{N}{1,3}) |
byte_fallback | not needed (byte-level cannot produce UNK) |
| tokens with diacritics | 27.9% of the vocabulary |
| tokens unreachable through merges | 0 (256 bytes + 32,312 merges = 32,568) |
| whitespace-only tokens longer than 2 chars | 18 |
| purely numeric tokens | 409 |
Round-trip verified on, among others, Zażółć gęślą jaźń — «cytat» … 😀\ttab, as well as
repeated spaces and newlines.
<|endoftext|> (32568), <|begin_of_text|>, <|pad|>, <|unk|>, chat, FIM and tool-call
tokens, plus 186 <|reserved_N|> slots — headroom for future extensions without changing
the shape of the embedding matrix.
Important when training from scratch: if you use only <|endoftext|> as the document
separator, the remaining special tokens never receive a gradient. In that case it is worth
blocking them at generation time (bad_words_ids) and setting
bos_token_id = eos_token_id = 32568.
from tokenizers import Tokenizer
tok = Tokenizer.from_pretrained("KateMajzel/tokenizer-pl-32k")
ids = tok.encode("Zażółć gęślą jaźń").ids
print(len(ids), tok.decode(ids))
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("KateMajzel/tokenizer-pl-32k")
An ablation on 42–51M parameter models (3 GB corpus, 3 random seeds) found:
The Bielik v3 PL team reached the same conclusion independently at 11B scale (arXiv 2604.10799): swapping in a dedicated tokenizer preserved quality and nearly doubled representation density.
Ġniepełnosprawności, Ġzagospodarowania). For a model meant to sound
colloquial, this is worth accounting for in the data mix.'(?i:[sdmt]|ll|ve|re)) — dead on
Polish, but harmless.post_processor — BOS/EOS are not added automatically and offsets are not trimmed.
Irrelevant for language modeling, relevant for NER/QA.d_model = 512 the embedding matrix is 16.8M parameters. In a
42.5M model that is 40% — worth computing before matching a vocabulary size to a small
model.MIT. Methodology and code: https://github.com/KateMajzel/gollem-pl
@misc{tokenizerpl32k,
title = {tokenizer-pl-32k: a Polish BPE tokenizer and its ablation},
author = {Majzel-Pośpiech, Katarzyna},
year = {2026},
url = {https://huggingface.co/KateMajzel/tokenizer-pl-32k}
}
Byte-level BPE dla języka polskiego. 32 568 pozycji słownika + 200 tokenów specjalnych = 32 768 (2¹⁵, wygodne dla kerneli GPU i tensor-parallel).
Tokenizer wytrenowany na polskim korpusie i zwalidowany eksperymentalnie w projekcie gollem-pl — kontrolowanej ablacji porównującej go z tokenizerem GPT-2 przy identycznym budżecie treningowym.
Model wytrenowany z tym tokenizerem: KateMajzel/GoLLeM-45M-PL.
| ten tokenizer | GPT-2 | |
|---|---|---|
| bajty/token — korpus mieszany PL (2,96 GB) | 4,050 | 2,066 |
| bajty/token — held-out (2,77 MB) | 4,103 | 2,134 |
| tokenów z 2,96 GB tekstu | 731 mln | 1 434 mln |
| gęstość względna | 1,96× | 1,00× |
Fertility na polskiej prozie: ~1,35 tokena na słowo (5,80 znaku/token). Uwaga: na realnym korpusie — z URL-ami, liczbami, resztkami formatowania — wychodzi 4,05 bajtu/token, czyli wyraźnie gorzej. Fertility należy podawać z korpusu, nie z wyselekcjonowanych zdań.
Przykład różnicy w praktyce — słowo niejednoznaczny to 3 tokeny tutaj i 9 u GPT-2.
| typ | byte-level BPE (tokenizers, format tokenizer.json) |
| normalizacja | brak — round-trip bezstratny |
| pre-tokenizer | w stylu GPT-4: Isolated + ByteLevel(use_regex=false, add_prefix_space=false) |
| cyfry | cięte po maks. 3 (\p{N}{1,3}) |
byte_fallback | nie jest potrzebny (byte-level nie może wyprodukować UNK) |
| tokeny z diakrytykami | 27,9% słownika |
| tokeny nieosiągalne przez merge | 0 (256 bajtów + 32 312 merge'ów = 32 568) |
| tokeny „same białe znaki" dłuższe niż 2 zn. | 18 |
| tokeny czysto liczbowe | 409 |
Round-trip zweryfikowany m.in. na Zażółć gęślą jaźń — «cytat» … 😀\ttab oraz
wielokrotnych spacjach i znakach nowej linii.
<|endoftext|> (32568), <|begin_of_text|>, <|pad|>, <|unk|>, tokeny czatowe, FIM
i tool-call oraz 186 pozycji <|reserved_N|> — zapas na przyszłe rozszerzenia bez zmiany
kształtu tablicy embeddingów.
Ważne przy trenowaniu od zera: jeśli korzystasz tylko z <|endoftext|> jako separatora
dokumentów, pozostałe tokeny specjalne nigdy nie dostaną gradientu. Warto je wtedy
zablokować przy generowaniu (bad_words_ids) i ustawić
bos_token_id = eos_token_id = 32568.
from tokenizers import Tokenizer
tok = Tokenizer.from_pretrained("KateMajzel/tokenizer-pl-32k")
ids = tok.encode("Zażółć gęślą jaźń").ids
print(len(ids), tok.decode(ids))
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("KateMajzel/tokenizer-pl-32k")
Ablacja na modelach 42–51 M parametrów (3 GB korpusu, 3 ziarna losowe) wykazała:
Ten sam wniosek uzyskał niezależnie zespół Bielika v3 PL na skali 11B (arXiv 2604.10799): wymiana tokenizera na dedykowany zachowała jakość i niemal podwoiła gęstość reprezentacji.
Ġniepełnosprawności, Ġzagospodarowania). Przy modelu, który
ma brzmieć potocznie, warto to uwzględnić w miksie danych.'(?i:[sdmt]|ll|ve|re)) — na polskim
martwa, ale nieszkodliwa.post_processor — BOS/EOS nie są dodawane automatycznie, offsety nie są trymowane.
Bez znaczenia dla modelowania języka, istotne przy NER/QA.d_model = 512 tablica embeddingów to 16,8 M parametrów.
W modelu 42,5 M stanowi to 40% — warto to policzyć, zanim dobierze się rozmiar słownika
do małego modelu.MIT. Metodologia i kod: https://github.com/KateMajzel/gollem-pl
@misc{tokenizerpl32k,
title = {tokenizer-pl-32k: polski tokenizer BPE i jego ablacja},
author = {Majzel-Pośpiech, Katarzyna},
year = {2026},
url = {https://huggingface.co/KateMajzel/tokenizer-pl-32k}
}