Downloads · 30 days
0
mkd-ai/keural-tokenizer
keural-tokenizer is a machine learning model from mkd-ai. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Keural Tokenizer is the official tokenizer used for training the Keural Foundation Model, a large-scale Mixture-of-Experts (MoE) language model architecture designed for enterprise AI, long-context reasoning, and mult…
Downloads · 30 days
0
Access
Public
Updated Jun 24, 2026
Repo size
2.7 MB
Likes
0
Public
Click a slice to open those files.
.model2.7 MB · 51%
From the Hugging Face model README
Keural Tokenizer is the official tokenizer used for training the Keural Foundation Model, a large-scale Mixture-of-Experts (MoE) language model architecture designed for enterprise AI, long-context reasoning, and multilingual language understanding.
This repository provides the tokenizer used during the pretraining stage of the Keural model, including configuration files, vocabulary, and metadata required to reproduce tokenization behavior during training and inference.
Large Language Models rely heavily on efficient tokenization. The Keural tokenizer was designed with the following goals:
The tokenizer was trained using the SentencePiece Unigram model on a curated multilingual corpus.
| Property | Value |
|---|---|
| Tokenizer Type | SentencePiece Unigram |
| Vocabulary Size | 131072 tokens |
| Normalization | NFKC |
| Byte Fallback | Enabled |
| Digit Splitting | Enabled |
| Unknown Token | <unk> |
| Padding Token | <pad> |
| BOS Token | <bos> |
| EOS Token | <eos> |
The tokenizer supports multilingual text including:
The tokenizer was trained on a 54.77 GB multilingual corpus consisting of multiple domains to ensure robust token coverage.
| Domain | Description |
|---|---|
| Web Text | Large-scale English web corpus |
| Scientific Papers | ArXiv and PubMed datasets |
| Literature | PG19 and BookCorpus |
| Wikipedia | Clean Korean Wikipedia |
| Source Code | Large-scale code repositories |
| Korean Web Data | Korean web text corpora |
| Multilingual Corpus | CC100 Korean |
The dataset pipeline was designed to reduce noise while preserving linguistic diversity across domains.
This repository contains the following tokenizer artifacts:
keural_tokenizer.model
keural_tokenizer.vocab
tokenizer_config.json
tokenizer_metadata.json
tokenizer.sha256
keural_tokenizer.model Binary SentencePiece tokenizer model used for tokenization.
keural_tokenizer.vocab Vocabulary mapping tokens to IDs.
tokenizer_config.json Tokenizer configuration used during model training.
tokenizer_metadata.json Metadata including training corpus information.
tokenizer.sha256 Checksum file for verifying tokenizer integrity.
import sentencepiece as spm
sp = spm.SentencePieceProcessor()
sp.load("keural_tokenizer.model")
tokens = sp.encode("Keural is a foundation model.", out_type=int)
print(tokens)
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("mkd-ai/keural-tokenizer")
tokens = tokenizer("Keural foundation model tokenizer example")
print(tokens)
This tokenizer is used for training the Keural Foundation Model, which uses the following architecture:
| Parameter | Value |
|---|---|
| Architecture | Transformer Mixture-of-Experts |
| Hidden Size | 4096 |
| Layers | 32 |
| Attention Heads | 32 |
| Experts per Layer | 32 |
| Active Experts per Token | 4 |
| Context Length | 4096 (scalable) |
| Vocabulary Size | 131072 |
Estimated model capacity:
The Keural model is designed to scale context length progressively using YaRN positional scaling.
| Stage | Context Length |
|---|---|
| Stage 1 | 4096 |
| Stage 2 | 8192 |
| Stage 3 | 32768 |
| Stage 4 | 131072 |
| Stage 5 | 262144 |
| Stage 6 | 524288 |
| Stage 7 | 1,048,576 |
This staged context expansion enables efficient training while supporting ultra-long context inference.
The tokenizer was trained as part of the Keural dataset pipeline, which includes:
The dataset preparation pipeline is available in the Keural model repository.
The Keural project roadmap includes the following stages.
Tokenizer development and dataset processing were performed on a high-performance server environment:
This tokenizer is part of the Keural Foundation Model project.
Developed by
MKD Corp AI Research
Republic of Korea
https://github.com/MKD-CORP/keural-tokenizer-model
If you use the Keural tokenizer in research, please cite the Keural project repository.
@misc{keural_tokenizer,
title={Keural Tokenizer},
author={MKD Corp AI Research, Md. Najmul Hossain},
year={2026},
url={https://huggingface.co/mkd-ai/keural-tokenizer}
}