Downloads · 30 days
4
17% of all-time downloads
Zarinaaa/punctuator_model
punctuator_model is a token classification model from Zarinaaa. Use it when you need labels on individual words, such as names. The card lists the license as apache-2.0.
The first punctuation restoration model for the Kyrgyz language, achieving 94.1% precision and 90.3% F1-score — surpassing benchmarks for other low-resource languages.
Downloads · 30 days
4
17% of all-time downloads
All-time downloads
24
Public
Repo size
1.1 GB
Likes
0
Public
Click a slice to open those files.
.onnx1.1 GB · 99%
From the Hugging Face model README
The first punctuation restoration model for the Kyrgyz language, achieving 94.1% precision and 90.3% F1-score — surpassing benchmarks for other low-resource languages.
📄 Published research: "AI-Based Punctuation Restoration using Transformer Model for Kyrgyz Language" — Uvalieva Z., Muhametjanova G. (SCOPUS-indexed)
| Metric | Score |
|---|---|
| Precision | 94.1% |
| Recall | 86.8% |
| F1-Score | 90.3% |
| Model | Language | F1-Score |
|---|---|---|
| Ours (XLM-RoBERTa) | Kyrgyz | 90.3% |
| Alam et al. (2020) | English (clean) | 87.0% |
| Alam et al. (2020) | Bangla | 69.5% |
| Nagy et al. (2021) | Hungarian | ~82.0% |
The model demonstrates strong performance on frequent punctuation marks (periods, commas) with reduced accuracy on rare marks (question marks, exclamation points) due to class imbalance.
| Parameter | Value |
|---|---|
| Base model | XLM-RoBERTa-base |
| Parameters | ~270M |
| Transformer layers | 12 |
| Hidden dimensions | 768 |
| Attention heads | 12 |
| Export format | ONNX |
A custom-built 200 MB Kyrgyz text corpus, collected over 2 months:
| Source | Size | Description |
|---|---|---|
| Kyrgyz-Turkish Manas University Library | 135 MB | Books (literature, math, physics) |
| Kyrgyz Wikipedia | 40 MB | Encyclopedia articles |
| News portals | 25 MB | Journalistic text |
Preprocessing pipeline: PDF → EasyOCR text extraction → manual cleaning → JSON formatting with punctuation labels.
Specialized augmentation techniques designed for Kyrgyz agglutinative morphology:
| Parameter | Value |
|---|---|
| Batch size | 32 |
| Epochs | 10 |
| Optimizer | Adam |
| Learning rate | 5e-5 |
| Regularization | Dropout |
| Hardware | Google Colab TPU |
| Training time | 42 hours |
import onnxruntime as ort
import numpy as np
# Load the ONNX model
session = ort.InferenceSession("model.onnx")
# Prepare input (see config.yaml for tokenizer settings)
# The model predicts punctuation labels for each token:
# O (no punctuation), COMMA, PERIOD, QUESTION, EXCLAMATION
# Example inference
input_text = "бул кыргыз тилиндеги текст"
# Tokenize and run inference (see main.py for full pipeline)
├── model.onnx # Trained model in ONNX format (1.11 GB)
├── main.py # Inference pipeline
├── env.py # Environment configuration
├── config.yaml # Hyperparameters and model config
├── requirements.txt # Python dependencies
└── Files/ # Additional model files
| Use Case | Description |
|---|---|
| ASR post-processing | Restore punctuation in speech-to-text output for Kyrgyz |
| Text normalization | Clean and format raw Kyrgyz text with proper punctuation |
| NLP preprocessing | Improve downstream task performance (NER, MT, summarization) |
| Accessibility | Enhance readability of automatically generated Kyrgyz content |
@article{uvalieva2024punctuation,
author = {Uvalieva, Zarina and Muhametjanova, Gulshat},
title = {AI-Based Punctuation Restoration using Transformer Model for Kyrgyz Language},
year = {2024},
institution = {Kyrgyz-Turkish Manas University}
}
Zarina Uvalieva — ML Engineer specializing in NLP and Speech Technologies for low-resource languages.