Downloads · 30 days
148
27% of all-time downloads
medyas/FarukSTT
FarukSTT is a automatic speech recognition model from medyas. Use it when you need speech turned into text. It is set up for transformers.
FarukSTT is the best publicly available ASR model for Tunisian Arabic (Derja), fine-tuned from Whisper Large v3. It handles real-world Tunisian speech including natural code-switching between Arabic, French, and English.
Downloads · 30 days
148
27% of all-time downloads
All-time downloads
540
Public
Parameters
1.5B
132 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors18.5 GB · 84%
From the Hugging Face model README
FarukSTT is the best publicly available ASR model for Tunisian Arabic (Derja), fine-tuned from Whisper Large v3. It handles real-world Tunisian speech including natural code-switching between Arabic, French, and English.
Whisper Large v3 out of the box scores above 50% WER on Tunisian Derja. This model brings that down to ~32% through targeted fine-tuning on real Tunisian conversational data.
| Model | WER |
|---|---|
| Whisper Large v3 (baseline) | >50% |
| FarukSTT v1 | 33.26% |
| FarukSTT v2 | 31.99% |
Evaluated on FARUKxAUTO/tunisian-asr-cleaned validation split.
Fine-tuned on FARUKxAUTO/tunisian-asr-cleaned — 54,156 audio-transcription pairs of real Tunisian speech including natural code-switching.
| Step | Validation Loss | WER |
|---|---|---|
| 500 | 0.3228 | 34.02% |
| 2000 | 0.3121 | 32.79% |
| 4000 | 0.3110 | 33.26% |
| Step | Validation Loss | WER |
|---|---|---|
| 250 | 0.3114 | 32.90% |
| 500 | 0.3060 | 31.99% |
from transformers import pipeline
pipe = pipeline("automatic-speech-recognition", model="FARUKxAUTO/FarukSTT")
result = pipe("audio.wav")
print(result["text"])
If you use this model, please credit:
FarukSTT by FARUK BATTIKH — Tunisian Derja ASR, 2024