Downloads · 30 days
4
17% of all-time downloads
flexitok/supertokenizer-mod_tokenizers_zero_padded
supertokenizer-mod_tokenizers_zero_padded is a machine learning model from flexitok. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A merged super-vocabulary built from 9 tokenizer(s).
Downloads · 30 days
4
17% of all-time downloads
All-time downloads
24
Public
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json1.8 MB · 100%
From the Hugging Face model README
A merged super-vocabulary built from 9 tokenizer(s).
Vocab size: 110931
flexitok/mod-tokenizers-zero-padded-individualflexitok/mod-tokenizers-zero-padded-ltr_3digitflexitok/mod-tokenizers-zero-padded-ltr_2digitflexitok/mod-tokenizers-zero-padded-ltr_4digitflexitok/mod-tokenizers-zero-padded-ltr_5digitflexitok/mod-tokenizers-zero-padded-rtl_2digitflexitok/mod-tokenizers-zero-padded-rtl_3digitflexitok/mod-tokenizers-zero-padded-rtl_4digitflexitok/mod-tokenizers-zero-padded-rtl_5digitsuper_vocab.json — merged vocabulary mapping token string → super indexconfig.yaml — model config with vocab_sizeparticipating_tokenizers.json — list of tokenizer names included<tokenizer>_super_mapping.json — per-tokenizer index → super index mapping<tokenizer>_vocab.json — per-tokenizer vocabulary<tokenizer>_info.json / .yaml — tokenizer metadata