Downloads · 30 days
30
4% of all-time downloads
devasheeshG/whisper_medium_fp16_transformers
whisper_medium_fp16_transformers is a automatic speech recognition model from devasheeshG. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
- CUDA: 12.1 - cuDNN Version: 8.9.2.261.0-1amd64
Downloads · 30 days
30
4% of all-time downloads
All-time downloads
790
Public
Repo size
3.1 GB
Likes
2
Public
Click a slice to open those files.
.bin1.5 GB · 100%
From the Hugging Face model README
RAM: 2.8 GB (Original_Model: 5.5GB)
VRAM: 1812 MB (Original_Model: 6GB)
test.wav: 23 s (Multilingual Speech i.e. English+Hindi)
| Device Name | float32 (Original) | float16 | CudaCores | TensorCores |
|---|---|---|---|---|
| 3060 | 1.7 | 1.1 | 3,584 | 112 |
| 1660 Super | OOM | 3.3 | 1,408 | N/A |
| Collab (Tesla T4) | 2.8 | 2.2 | 2,560 | 320 |
| Collab (CPU) | 35 | N/A | N/A | N/A |
| M1 (CPU) | - | - | - | - |
| M1 (GPU -> 'mps') | - | - | - | - |
Punchuation: True
Test done on RTX 3060 on 2557 Samples
| WER | MER | WIL | WIP | CER | |
|---|---|---|---|---|---|
| Original_Model (54 min) | 52.02 | 47.86 | 66.82 | 33.17 | 23.76 |
| This_Model (38 min) | 54.97 | 47.86 | 66.83 | 33.16 | 30.23 |
Test done on RTX 3060 on 1000 Samples
| WER | MER | WIL | WIP | CER | |
|---|---|---|---|---|---|
| Original_Model (30 min) | - | - | - | - | - |
| This_Model (20 min) | - | - | - | - | - |
Test done on RTX 3060 on __ Samples
| WER | MER | WIL | WIP | CER | |
|---|---|---|---|---|---|
| Original_Model | - | - | - | - | - |
| This_Model | - | - | - | - | - |
Test done on RTX 3060 on __ Samples
| WER | MER | WIL | WIP | CER | |
|---|---|---|---|---|---|
| Original_Model | - | - | - | - | - |
| This_Model | - | - | - | - | - |
A file __init__.py is contained inside this repo which contains all the code to use this model.
Firstly, clone this repo and place all the files inside a folder.
git lfs install
git clone https://huggingface.co/devasheeshG/whisper_medium_fp16_transformers
Please try in jupyter notebook
# Import the Model
from whisper_medium_fp16_transformers import Model, load_audio, pad_or_trim
# Initilise the model
model = Model(
model_name_or_path='whisper_medium_fp16_transformers',
cuda_visible_device="0",
device='cuda',
)
# Load Audio
audio = load_audio('whisper_medium_fp16_transformers/test.wav')
audio = pad_or_trim(audio)
# Transcribe (First transcription takes time)
model.transcribe(audio)
It is fp16 version of openai/whisper-medium