Downloads · 30 days
12
24% of all-time downloads
bartelds/whisper-dro-set2-baseline
whisper-dro-set2-baseline is a machine learning model from bartelds. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository contains a baseline automatic speech recognition (ASR) model fine-tuned from openai/whisper-large-v3. The model was trained on balanced training data from set 2 (eng, fas, hrv, ita, slk, yue).
Downloads · 30 days
12
24% of all-time downloads
All-time downloads
51
Public
Parameters
1.5B
6.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6.2 GB · 100%
From the Hugging Face model README
This repository contains a baseline automatic speech recognition (ASR) model fine-tuned from openai/whisper-large-v3.
The model was trained on balanced training data from set 2 (eng, fas, hrv, ita, slk, yue).
This model is intended for multilingual ASR. Users can run inference using the HuggingFace Transformers library:
import torch
import librosa
from transformers import WhisperForConditionalGeneration, WhisperProcessor
model = WhisperForConditionalGeneration.from_pretrained("bartelds/whisper-dro-set2-baseline")
processor = WhisperProcessor.from_pretrained("bartelds/whisper-dro-set2-baseline")
model.eval()
audio, sr = librosa.load("input.wav", sr=16000)
inputs = processor.feature_extractor(audio, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
generated = model.generate(input_features=inputs.input_features)
text = processor.tokenizer.batch_decode(generated, skip_special_tokens=True)[0]
print("Recognized text:", text)
pip install transformers torch librosafrom_pretrained() as shown above.openai/whisper-large-v3