Downloads · 30 days
8
0% of all-time downloads
hamzabouajila/distilled_tunbert
distilled_tunbert is a fill-mask model from hamzabouajila. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as mit.
A distilled, efficient version of TunBERT for Tunisian Arabic. This model is faster, smaller, and fully reproducible thanks to an open Tunisian corpus and transparent distillation pipeline.
Downloads · 30 days
8
0% of all-time downloads
All-time downloads
6.1K
Public
Parameters
67M
268 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
A distilled, efficient version of TunBERT for Tunisian Arabic. This model is faster, smaller, and fully reproducible thanks to an open Tunisian corpus and transparent distillation pipeline.
distilbert-base-uncased)distilbert-base-uncasedBias: Model inherits cultural/linguistic biases present in the Tunisian corpus.
Limitations:
cosine ≈ 0.02) due to tokenizer mismatch and lack of hidden-state loss.Risk: Misuse in contexts requiring semantic alignment (e.g., search, embeddings).
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("hamzabouajila/distilled_tunbert")
model = AutoModel.from_pretrained("hamzabouajila/distilled_tunbert")
text = "نحب النموذج هذا يخدم بسرعه"
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
| Metric | Original TunBERT | Distilled TunBERT | Notes |
|---|---|---|---|
| Perplexity | 34838.7 | 4.26 | Strong LM performance. Teacher likely uninitialized. |
| Inference Time (s) | 0.106 | 0.058 | 1.83× faster |
| Parameters | 109M | 66M | 1.65× smaller |
| Embedding Similarity | — | 0.02 | Near-zero due to tokenizer mismatch |
| Training Data | Unknown | Open corpus | Fully reproducible |
The distilled model is faster, lighter, and trained on open data. It performs competitively on classification tasks but embeddings should not be used for similarity-based applications.
BibTeX:
@misc{bouajila2025distilledtunbert,
title={Distilled TunBERT: Efficient Tunisian Arabic BERT via Knowledge Distillation},
author={Bouajila Hamza},
year={2025},
publisher={HuggingFace},
howpublished={\url{https://huggingface.co/hamzabouajila/distilled_tunbert}}
}
👉 This version positions your model as efficient, open, and reproducible — while honestly stating limitations (embeddings, risks).
Do you want me to also draft a shorter, lightweight Hugging Face card (2–3 sections only) for quick readers, in addition to this full professional one?