Downloads · 30 days
49
39% of all-time downloads
jankoko/PALF-Whisper-small
PALF-Whisper-small is a automatic speech recognition model from jankoko. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of openai/whisper-small on the wTIMIT-US dataset using the F0-Mask augmentation method. It was evaluated on both normal and whispered speech subsets, with Word Error Rate (WER) as th…
Downloads · 30 days
49
39% of all-time downloads
All-time downloads
126
Public
Parameters
242M
2.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors967 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of openai/whisper-small on the wTIMIT-US dataset using the F0-Mask augmentation method. It was evaluated on both normal and whispered speech subsets, with Word Error Rate (WER) as the primary metric.
The results below highlight performance improvements over the baseline for whispered speech, validating the effectiveness of phoneme-aware low-frequency masking (PALF-Mask).
| Setup | Training Data | Augmentation | WER (Normal) | WER (Whispered) |
|---|---|---|---|---|
| No Fine-tuning | Zero-shot | None | 5.0 | 13.7 |
| Baseline | Both modes | None | 5.8 | 11.7 |
| SpecAugment | Both modes | SpecAugment (LD) | 5.2 | 12.3 |
| F0-Mask (Ours) | Both modes | F0-based Masking | 5.0 (ns, p=0.144) | 11.5 (★, p=0.002) |
★ = Statistically significant improvement over SpecAugment (paired MAPSSWE)
ns = No significant difference (not statistically significant)
Compared to the SpecAugment baseline, F0-Mask achieved a statistically significant improvement in whispered speech recognition (↓0.8% absolute WER, p=0.002), while maintaining comparable performance on normal speech (p=0.144).
Notably, the whispered WER of 11.5% matches the best result previously reported on this dataset by Marchenko (2024).
Kokowski, J. (2025). F0-Based Masking Policies for Self-Supervised Whispered Speech Recognition. Master’s Thesis, University of Groningen, Campus Fryslân.
Available at: https://campus-fryslan.studenttheses.ub.rug.nl/view/degree_programme/voice=5Ftechnology.html
If you use this model or build upon this work, please cite the thesis above.
Model: Whisper-small
Augmentation: F0-Mask
Evaluation toolkit: SCTK (sclite)
Notes: For complete results, including MAPSSWE and CER scores, refer to Section 5 of the thesis.