Downloads · 30 days
0
boffire/kabyle-stanza-tokenizer
kabyle-stanza-tokenizer is a token classification model from boffire. Use it when you need labels on individual words, such as names. It is set up for stanza. The card lists the license as apache-2.0.
Sentence tokenizer for Kabyle (Taqbaylit) language, trained on Tatoeba corpus designed to be used on MiniSBD.
Downloads · 30 days
0
Access
Public
Updated Jul 10, 2026
Repo size
1.3 MB
Likes
0
Public
Click a slice to open those files.
.onnx651 KB · 50%
From the Hugging Face model README
Sentence tokenizer for Kabyle (Taqbaylit) language, trained on Tatoeba corpus designed to be used on MiniSBD.
| Property | Value |
|---|---|
| Language | Kabyle (kab) |
| Type | Sentence tokenizer |
| Architecture | CNN + BiLSTM |
| Training data | Tatoeba Kabyle sentences (~789K sentences) |
| Dev F1 | 99.19% |
| Token F1 | 99.96% |
| Sentence F1 | 98.43% |
| ONNX size | 0.62 MB |
| Vocab size | 223 characters |
kab.onnx — ONNX runtime modelkab.pt — PyTorch checkpointvocab.json — Character vocabulary (223 entries)config.json — Model hyperparameterstokenizer_config.json — HF tokenizer configimport stanza
nlp = stanza.Pipeline(
lang="kab",
processors="tokenize",
tokenize_model_path="kab.pt"
)
doc = nlp("Amcic ha-t-an deg uxxam-nneɣ. Teciḍ fell-as?")
for sent in doc.sentences:
print(sent.text)
# Amcic ha-t-an deg uxxam-nneɣ.
# Teciḍ fell-as?
import onnxruntime as ort
import numpy as np
sess = ort.InferenceSession("kab_tokenizer.onnx")
# units: (batch, seq_len) int64 — char IDs from vocab.json
# features: (batch, seq_len, 5) float32 — Stanza features
units = np.zeros((1, 100), dtype=np.int64)
features = np.zeros((1, 100, 5), dtype=np.float32)
outputs = sess.run(None, {"units": units, "features": features})
# outputs[0] shape: (batch, seq_len, 3) — logits for B/I/O
| Input | Tokens |
|---|---|
Ad tseddumt ɣer Taskriwt. | ['Ad', 'tseddumt', 'ɣer', 'Taskriwt', '.'] |
Aweḍ ɣer Tezmalt. | ['Aweḍ', 'ɣer', 'Tezmalt', '.'] |
Tettawḍem ɣer Kendira. | ['Tettawḍem', 'ɣer', 'Kendira', '.'] |
Efk-asen tizwal-nni. | ['Efk-asen', 'tizwal-nni', '.'] |
Melmi ara ad d-taɣeḍ lmitra? | ['Melmi', 'ara', 'ad', 'd-taɣeḍ', 'lmitra', '?'] |
The tokenizer handles standard Kabyle Latin characters including:
ɛ / Ɛ (open e)ɣ / Ɣ (voiced velar fricative)ṭ, ḍ, č, ǧ (emphatic and palatal consonants)If you use this model, please cite:
Apache-2.0