Downloads · 30 days
14
2% of all-time downloads
deepdml/whisper-medium-mix-es
whisper-medium-mix-es is a automatic speech recognition model from deepdml. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
14
2% of all-time downloads
All-time downloads
901
Public
Repo size
18.3 GB
Likes
2
Public
Click a slice to open those files.
.bin3.1 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of openai/whisper-medium on the mozilla-foundation/common_voice_11_0, google/fleurs, facebook/multilingual_librispeech and facebook/voxpopuli datasets. It achieves the following results on the evaluation set:
Using the evaluation script provided in the Whisper Sprint the model achieves these results on the test sets (WER):
google/fleurs: 4.0266 %
(python run_eval_whisper_streaming.py --model_id="deepdml/whisper-medium-mix-es" --dataset="google/fleurs" --config="es_419" --device=0 --language="es")
facebook/multilingual_librispeech: 4.6644 %
(python run_eval_whisper_streaming.py --model_id="deepdml/whisper-medium-mix-es" --dataset="facebook/multilingual_librispeech" --config="spanish" --device=0 --language="es")
facebook/voxpopuli: 8.3668 %
(python run_eval_whisper_streaming.py --model_id="deepdml/whisper-medium-mix-es" --dataset="facebook/voxpopuli" --config="es" --device=0 --language="es")
More information needed
More information needed
Training data used:
Evaluating over test split from mozilla-foundation/common_voice_11_0 dataset.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Wer |
|---|---|---|---|---|
| 0.266 | 0.2 | 1000 | 0.1657 | 8.0395 |
| 0.1394 | 0.4 | 2000 | 0.1539 | 7.3937 |
| 0.1316 | 0.6 | 3000 | 0.1452 | 6.9656 |
| 0.1165 | 0.8 | 4000 | 0.1392 | 6.5765 |
| 0.2816 | 1.0 | 5000 | 0.1344 | 6.3465 |