Downloads · 30 days
13
3% of all-time downloads
latofat/uzpostagger-cyrillic-3
uzpostagger-cyrillic-3 is a token classification model from latofat. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as mit.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
13
3% of all-time downloads
All-time downloads
499
Public
Repo size
1.7 GB
Likes
1
Public
Click a slice to open those files.
.bin434 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of coppercitylabs/uzbert-base-uncased on uzbekpos dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Precision | Recall | F1 | Accuracy |
|---|---|---|---|---|---|---|---|
| No log | 1.0 | 25 | 0.8765 | 0.6558 | 0.5477 | 0.5969 | 0.7485 |
| No log | 2.0 | 50 | 0.4086 | 0.8496 | 0.8237 | 0.8364 | 0.9004 |
| No log | 3.0 | 75 | 0.3133 | 0.8615 | 0.8552 | 0.8583 | 0.9142 |
| No log | 4.0 | 100 | 0.2806 | 0.8730 | 0.8657 | 0.8693 | 0.9193 |
| No log | 5.0 | 125 | 0.2715 | 0.8763 | 0.8699 | 0.8731 | 0.9219 |
@inproceedings{bobojonova-etal-2025-bbpos,
title = "{BBPOS}: {BERT}-based Part-of-Speech Tagging for {U}zbek",
author = "Bobojonova, Latofat and
Akhundjanova, Arofat and
Ostheimer, Phil Sidney and
Fellenz, Sophie",
editor = "Hettiarachchi, Hansi and
Ranasinghe, Tharindu and
Rayson, Paul and
Mitkov, Ruslan and
Gaber, Mohamed and
Premasiri, Damith and
Tan, Fiona Anting and
Uyangodage, Lasitha",
booktitle = "Proceedings of the First Workshop on Language Models for Low-Resource Languages",
month = jan,
year = "2025",
address = "Abu Dhabi, United Arab Emirates",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.loreslm-1.23/",
pages = "287--293",
abstract = "This paper advances NLP research for the low-resource Uzbek language by evaluating two previously untested monolingual Uzbek BERT models on the part-of-speech (POS) tagging task and introducing the first publicly available UPOS-tagged benchmark dataset for Uzbek. Our fine-tuned models achieve 91{\%} average accuracy, outperforming the baseline multi-lingual BERT as well as the rule-based tagger. Notably, these models capture intermediate POS changes through affixes and demonstrate context sensitivity, unlike existing rule-based taggers."
}