Downloads · 30 days
27
2% of all-time downloads
devasheeshG/whisper_large_v2_fp16_transformers
whisper_large_v2_fp16_transformers is a automatic speech recognition model from devasheeshG. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
- CUDA: 12.1 - cuDNN Version: 8.9.2.261.0-1amd64
Downloads · 30 days
27
2% of all-time downloads
All-time downloads
1.4K
Public
Repo size
6.2 GB
Likes
2
Public
Click a slice to open those files.
.bin3.1 GB · 100%
From the Hugging Face model README
RAM: 3 GB (Original_Model: 6GB)
VRAM: 3.7 GB (Original_Model: 11GB)
test.wav: 23 s (Multilingual Speech i.e. English+Hindi)
| Device Name | float32 (Original) | float16 | CudaCores | TensorCores |
|---|---|---|---|---|
| 3060 | 2.2 | 1.3 | 3,584 | 112 |
| 1660 Super | OOM | 6 | 1,408 | N/A |
| Collab (Tesla T4) | - | - | 2,560 | 320 |
| Collab (CPU) | - | N/A | N/A | N/A |
| M1 (CPU) | - | - | N/A | N/A |
| M1 (GPU -> 'mps') | - | - | N/A | N/A |
Punchuation: Sometimes False ('I don't know the exact reason why this is happening')
Test done on RTX 3060 on 1000 Samples
| WER | MER | WIL | WIP | CER | |
|---|---|---|---|---|---|
| Original_Model (30 min) | 43.99 | 41.65 | 59.47 | 40.52 | 16.23 |
| This_Model (20 min) | 44.64 | 41.69 | 59.53 | 40.46 | 16.80 |
Test done on RTX 3060 on 1000 Samples
| WER | MER | WIL | WIP | CER | |
|---|---|---|---|---|---|
| Original_Model (30 min) | - | - | - | - | - |
| This_Model (20 min) | - | - | - | - | - |
Test done on RTX 3060 on ___ Samples
| WER | MER | WIL | WIP | CER | |
|---|---|---|---|---|---|
| Original_Model | - | - | - | - | - |
| This_Model | - | - | - | - | - |
Test done on RTX 3060 on ___ Samples
| WER | MER | WIL | WIP | CER | |
|---|---|---|---|---|---|
| Original_Model | - | - | - | - | - |
| This_Model | - | - | - | - | - |
A file __init__.py is contained inside this repo which contains all the code to use this model.
Firstly, clone this repo and place all the files inside a folder.
git lfs install
git clone https://huggingface.co/devasheeshG/whisper_large_v2_fp16_transformers
Please try in jupyter notebook
# Import the Model
from whisper_large_v2_fp16_transformers import Model, load_audio, pad_or_trim
# Initilise the model
model = Model(
model_name_or_path='whisper_large_v2_fp16_transformers',
cuda_visible_device="0",
device='cuda',
)
# Load Audio
audio = load_audio('whisper_large_v2_fp16_transformers/test.wav')
audio = pad_or_trim(audio)
# Transcribe (First transcription takes time)
model.transcribe(audio)
It is fp16 version of openai/whisper-large-v2