Downloads · 30 days
15
7% of all-time downloads
AfroLogicInsect/whisper-finetuned-float32
whisper-finetuned-float32 is a automatic speech recognition model from AfroLogicInsect. Use it when you need speech turned into text. The card lists the license as apache-2.0.
Fine-tuned Whisper model (float32 version) for speech recognition
Downloads · 30 days
15
7% of all-time downloads
All-time downloads
228
Public
Parameters
242M
967 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors967 MB · 100%
From the Hugging Face model README
Fine-tuned Whisper model (float32 version) for speech recognition
from transformers import WhisperProcessor, WhisperForConditionalGeneration
import librosa
# Load model and processor
processor = WhisperProcessor.from_pretrained("AfroLogicInsect/whisper-finetuned-float32")
model = WhisperForConditionalGeneration.from_pretrained("AfroLogicInsect/whisper-finetuned-float32")
# Load audio
audio, sr = librosa.load("path/to/audio.wav", sr=16000)
# Process
input_features = processor(audio, sampling_rate=16000, return_tensors="pt").input_features
# Generate transcription
with torch.no_grad():
predicted_ids = model.generate(input_features)
transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)[0]
print(transcription)
If you use this model, please cite:
@misc{whisper-finetuned,
author = {Daniel AMAH},
title = {Fine-tuned Whisper Model},
year = {2024},
publisher = {Hugging Face},
url = {https://huggingface.co/AfroLogicInsect/whisper-finetuned-float32}
}