Downloads · 30 days
16
3% of all-time downloads
Scrya/whisper-medium-id-augmented
whisper-medium-id-augmented is a automatic speech recognition model from Scrya. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
16
3% of all-time downloads
All-time downloads
521
Public
Repo size
33.6 GB
Likes
2
Public
Click a slice to open those files.
.bin3.1 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of openai/whisper-medium on the following datasets:
It achieves the following results on the evaluation set (Common Voice 11.0):
More information needed
More information needed
Training:
Evaluation:
Datasets were augmented on-the-fly using audiomentations via PitchShift, AddGaussianNoise and TimeStretch transformations at p=0.3.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Wer | Cer |
|---|---|---|---|---|---|
| 0.3002 | 1.9 | 1000 | 0.1659 | 8.1850 | 2.5333 |
| 0.0514 | 3.8 | 2000 | 0.1818 | 8.0559 | 2.5244 |
| 0.0145 | 5.7 | 3000 | 0.2150 | 7.8945 | 2.5281 |
| 0.0037 | 7.6 | 4000 | 0.2248 | 7.7100 | 2.3738 |
| 0.0016 | 9.51 | 5000 | 0.2402 | 7.6224 | 2.3591 |
| 0.0009 | 11.41 | 6000 | 0.2525 | 7.7654 | 2.3952 |
| 0.0005 | 13.31 | 7000 | 0.2609 | 7.5994 | 2.3487 |
| 0.0008 | 15.21 | 8000 | 0.2682 | 7.5855 | 2.3347 |
| 0.0002 | 17.11 | 9000 | 0.2756 | 7.6178 | 2.3288 |
| 0.0002 | 19.01 | 10000 | 0.2788 | 7.6132 | 2.3332 |