Downloads Β· 30 days
5
28% of all-time downloads
khouloudCh15/multitaskwav2vec2
multitaskwav2vec2 is a machine learning model from khouloudCh15. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This model is fine-tuned on the Wav2Vec2 architecture for speech emotion recognition. It can classify speech into 8 different emotions with corresponding confidence scores.
Downloads Β· 30 days
5
28% of all-time downloads
All-time downloads
18
Public
Parameters
94.6M
378 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors378 MB Β· 100%
From the Hugging Face model README
This model is fine-tuned on the Wav2Vec2 architecture for speech emotion recognition. It can classify speech into 8 different emotions with corresponding confidence scores.
The model was trained with the following configuration:
For detailed training process, check out the Fine-tuning Notebook
https://huggingface.co/spaces/Dpngtm/Audio-Emotion-Recognition
For issues and questions, feel free to:
from transformers import Wav2Vec2ForSequenceClassification, Wav2Vec2Processor
import torch
import torchaudio
# Load model and processor
model = Wav2Vec2ForSequenceClassification.from_pretrained("Dpngtm/wav2vec2-emotion-recognition")
processor = Wav2Vec2Processor.from_pretrained("Dpngtm/wav2vec2-emotion-recognition")
# Load and preprocess audio
speech_array, sampling_rate = torchaudio.load("path_to_audio.wav")
if sampling_rate != 16000:
resampler = torchaudio.transforms.Resample(orig_freq=sampling_rate, new_freq=16000)
speech_array = resampler(speech_array)
# Convert to mono if stereo
if speech_array.shape[0] > 1:
speech_array = torch.mean(speech_array, dim=0, keepdim=True)
speech_array = speech_array.squeeze().numpy()
# Process through model
inputs = processor(speech_array, sampling_rate=16000, return_tensors="pt", padding=True)
with torch.no_grad():
outputs = model(**inputs)
predictions = torch.nn.functional.softmax(outputs.logits, dim=-1)
# Get predicted emotion
emotion_labels = ["angry", "calm", "disgust", "fearful", "happy", "neutral", "sad", "surprised"]
predicted_emotion = emotion_labels[predictions.argmax().item()]