Downloads · 30 days
7
2% of all-time downloads
AventIQ-AI/whisper-speech-text
whisper-speech-text is a machine learning model from AventIQ-AI. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository hosts a fine-tuned version of the OpenAI Whisper-Base model optimized for speech-to-text tasks using the Mozilla Common Voice 13.0 dataset. The model is designed to efficiently transcribe speech into t…
Downloads · 30 days
7
2% of all-time downloads
All-time downloads
283
Public
Parameters
72.6M
145 MB on disk
Likes
4
Public
Click a slice to open those files.
.safetensors145 MB · 99%
From the Hugging Face model README
This repository hosts a fine-tuned version of the OpenAI Whisper-Base model optimized for speech-to-text tasks using the Mozilla Common Voice 13.0 dataset. The model is designed to efficiently transcribe speech into text while maintaining high accuracy.
pip install transformers torch
from transformers import WhisperProcessor, WhisperForConditionalGeneration
import torch
device = "cuda" if torch.cuda.is_available() else "cpu"
model_name = "AventIQ-AI/whisper-speech-text"
model = WhisperForConditionalGeneration.from_pretrained(model_name).to(device)
processor = WhisperProcessor.from_pretrained(model_name)
import torchaudio
# Load and process audio file
def transcribe(audio_path):
waveform, sample_rate = torchaudio.load(audio_path)
inputs = processor(waveform, sampling_rate=sample_rate, return_tensors="pt").input_features.to(device)
# Generate transcription
with torch.no_grad():
predicted_ids = model.generate(inputs)
transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)[0]
return transcription
# Example usage
audio_file = "sample_audio.wav"
print(transcribe(audio_file))
After fine-tuning the Whisper-Base model for speech-to-text, we evaluated the model's performance on the validation set from the Common Voice 13.0 dataset. The following results were obtained:
| Metric | Score | Meaning |
|---|---|---|
| WER | 8.2% | Word Error Rate: Measures transcription accuracy |
| CER | 4.5% | Character Error Rate: Measures character-level accuracy |
The Mozilla Common Voice 13.0 dataset, containing diverse multilingual speech samples, was used for fine-tuning the model.
Post-training quantization was applied using PyTorch's built-in quantization framework to reduce the model size and improve inference efficiency.
.
├── model/ # Contains the quantized model files
├── tokenizer_config/ # Tokenizer configuration and vocabulary files
├── model.safetensors/ # Quantized Model
├── README.md # Model documentation
Contributions are welcome! Feel free to open an issue or submit a pull request if you have suggestions or improvements.