Downloads · 30 days
75
30% of all-time downloads
Quran-Lab/mfa-quran-hafs
mfa-quran-hafs is a machine learning model from Quran-Lab. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for montreal-forced-aligner. The card lists the license as apache-2.0.
A Montreal Forced Aligner acoustic model purpose-built for Quranic recitation in the Hafs riwaya, with a pronunciation dictionary derived from a rule-verified phonetic script. Built for phone-level tajweed measurement…
Downloads · 30 days
75
30% of all-time downloads
All-time downloads
251
Public
Repo size
61.7 MB
Likes
4
Public
Click a slice to open those files.
.zip61.7 MB · 85%
From the Hugging Face model README
A Montreal Forced Aligner acoustic model purpose-built for Quranic recitation in the Hafs riwaya, with a pronunciation dictionary derived from a rule-verified phonetic script. Built for phone-level tajweed measurement: madd durations, ghunna, qalqala, not just word timestamps.
Generic Arabic aligners fail on recitation: multi-second madd vowels, ghunna nasals, melismatic (mujawwad) style, and mosque reverb are far outside normal speech. This model was trained and evaluated specifically against those:
ا+ = long vowel of any prescribed length,
vs. the source convention of repeating characters, which is degenerate for
HMM alignment). 67 symbols; tajweed-bearing segments are single intervals.| Metric | Result |
|---|---|
| Phone-boundary jitter vs. signal landmarks (geminate stop releases, n=533, 4 reciters incl. mujawwad) | 10-20 ms median |
| Speech uncovered by any phone (murattal, held-out reciters) | 0.0-0.4% |
| Speech uncovered (mujawwad, Abdul Basit held-out) | 5.0% |
| Madd duration vs. prescription (tempo-normalized, anchored): normal (2) / monfasel (4) | 2.1 / 4.8 |
| Word boundary on a human-labelled mujawwad elongation (58:20) | within ~100 ms |
Notes for measurement use: word-onset comparisons against QUL word timings are biased: QUL starts are ~245 ms early vs. physical burst landmarks (playback convention). Trust obstruent-anchored spans; treat boundaries inside sonorant runs as untrustworthy and measure rule segments between obstruent anchors. Pre-pause madds benefit from an energy/voicing end-trim (breath and room decay otherwise attach to the final phone).
| File | Purpose |
|---|---|
quran_hafs_acoustic.zip | MFA acoustic model (train with MFA 3.4) |
quran_hafs.dict | pronunciation dictionary (20,967 entries, contextual) |
tokens.txt | phone symbol table (67 + blank) for CTC integrations |
rule_index.jsonl | per-ayah tajweed rule annotations: 93,430 positions with the governing rule and its prescribed length (golden_len); join with alignments to measure tajweed |
mfa align CORPUS_DIR quran_hafs.dict quran_hafs_acoustic.zip OUT_DIR \
--beam 40 --retry_beam 160
Corpus: one wav (16 kHz mono) + one .lab per ayah clip, .lab containing
the Uthmani ayah text (whole words; the dictionary handles phonetization).
Keep all paths ASCII (OpenFST on Windows fails on non-ASCII paths).
Built by the Quran-Lab effort on a 59,000-hour multi-site crawl of public
recitation audio (verified per-ayah against canonical text before any
training; label source is always the canonical Uthmani text, never ASR
output). Phone representation and rule index from the companion
quran-phones package.