Downloads · 30 days
49
46% of all-time downloads
shah-bakhsh/BalPOS
BalPOS is a token classification model from shah-bakhsh. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
49
46% of all-time downloads
All-time downloads
106
Public
Parameters
277M
2.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
| Accuracy | Weighted F1 | MCC | Macro F1 |
|---|---|---|---|
| 87.26 % | 87.22 % | 0.8526 | 78.99 % |
BalPOS v2 is a transformer-based token-classification model fine-tuned from shahbakhsh/BalBERT specifically for Universal Dependencies (UD) Part-of-Speech tagging on the Balochi language (bal).
Developed by Shah Bakhsh, BalPOS v2 provides the foundational syntactic layer for the Balochi NLP research ecosystem.
shahbakhsh/BalBERT)bal)42), cuDNN deterministic execution enabledThe evaluation results below reflect performance on the 15% final hold-out test set, which was never seen during hyperparameter optimization or fold selection:
<div align="center">| Metric | Exact Score | Percentage | Status |
|---|---|---|---|
| Accuracy | 0.8726 | 87.26% | 🥇 Primary Benchmark |
| Weighted F1 | 0.8722 | 87.22% | ⚡ Distribution Weighted |
| Weighted Precision | 0.8741 | 87.41% | ⚡ Distribution Weighted |
| Weighted Recall | 0.8726 | 87.26% | ⚡ Distribution Weighted |
| Matthews Correlation Coefficient (MCC) | 0.8526 | 0.8526 | 📐 Multi-Class Quality |
| Macro F1 | 0.7899 | 78.99% | ⚖️ Unweighted Class Mean |
| Macro Precision | 0.7956 | 79.56% | ⚖️ Unweighted Class Mean |
| Macro Recall / Balanced Accuracy | 0.7881 | 78.81% | ⚖️ Unweighted Class Mean |
Model selection was performed via 5-fold cross-validation over the development split following randomized search hyperparameter optimization:
{
"learning_rate": 1e-05,
"batch_size": 8,
"epochs": 15,
"weight_decay": 0.01,
"warmup_ratio": 0.05,
"gradient_accumulation_steps": 1,
"seed": 42,
"best_fold": "Fold 5"
}
pipelinefrom transformers import pipeline
# Load BalPOS v2 pipeline from Hugging Face
tagger = pipeline(
task="token-classification",
model="shahbakhsh/BalPOS",
aggregation_strategy="simple"
)
# Tag Balochi sentence
text = "وتی فلسفہ"
results = tagger(text)
for entity in results:
print(f"Token: {entity['word']:<15} | Tag: {entity['entity_group']:<8} | Score: {entity['score']:.4f}")
import torch
from transformers import AutoTokenizer, AutoModelForTokenClassification
model_id = "shahbakhsh/BalPOS"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForTokenClassification.from_pretrained(model_id)
model.eval()
sentence = "وتی فلسفہ"
inputs = tokenizer(sentence, return_tensors="pt")
with torch.no_grad():
logits = model(**inputs).logits
predictions = torch.argmax(logits, dim=-1)[0]
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
for token, pred_id in zip(tokens, predictions):
if token not in [tokenizer.cls_token, tokenizer.sep_token, tokenizer.pad_token]:
tag = model.config.id2label[pred_id.item()]
print(f"{token:<15} -> {tag}")
BalPOS v2 maps Balochi tokens to all 16 UPOS categories:
| Tag | Name | Description | Example Tokens |
|---|---|---|---|
| ADJ | Adjective | Modifies a noun | گران, شَر |
| ADP | Adposition | Preposition / Postposition | گۆں, ماں |
| ADV | Adverb | Modifies verb/adj | انّۆ, سک |
| AUX | Auxiliary | Auxiliary or copular verb | اِنت, بوت |
| CCONJ | Coordinating Conjunction | Connects words or clauses | ءُ, یا |
| DET | Determiner | Demonstrative or article | اے, آ |
| INTJ | Interjection | Exclamation | واہ, ھۆ |
| NOUN | Noun | Common noun | مارچ, کتاب |
| NUM | Numeral | Cardinal/ordinal number | یک, دۆ |
| PART | Particle | Function word | مئے |
| PRON | Pronoun | Personal/possessive pronoun | من, تئو, وتی |
| PROPN | Proper Noun | Name of entity | بلوچستان, کراچی |
| PUNCT | Punctuation | Punctuation marks | . , ؟ ! |
| SCONJ | Subordinating Conjunction | Subordinating connector | کہ, پرچا کہ |
| VERB | Verb | Action/state verb | رَوت, گوشیت |
| X | Other | Foreign or unclassified token | — |
INTJ or X) have higher variance in per-class evaluation.BalPOS v2 is part of an ongoing open-source initiative for Balochi NLP led by Shah Bakhsh:
shahbakhsh/BalBERT — Pretrained Masked Language Model (Released)shahbakhsh/BalPOS — Universal POS Tagger (Released)If you use BalPOS v2 or BalBERT, please cite this repository:
@misc{balpos_v2_2026,
title = {{BalPOS v2: Balochi Universal Part-of-Speech Tagger}},
author = {Shah Bakhsh},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/shahbakhsh/BalPOS}},
note = {Fine-tuned from shahbakhsh/BalBERT on Balochi UD Corpus}
}
Distributed under the Apache License 2.0.