Downloads · 30 days
10
2% of all-time downloads
asadullah797/ssl-semi-multitask
ssl-semi-multitask is a audio classification model from asadullah797. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for fairseq. The card lists the license as apache-2.0.
Downloads · 30 days
10
2% of all-time downloads
All-time downloads
461
Public
Parameters
94.7M
33 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors379 MB · 100%
From the Hugging Face model README
Multitask Speech Model with Wav2Vec2
This repository contains a multitask learning pipeline built on top of Wav2Vec2 , designed to jointly perform:
Automatic Speech Recognition (ASR) (character-level CTC loss)
Speaker Identification
Emotion Recognition
The system is trained on a combination of training dataset with parallel data from speech transcriptions, speaker identification and emotion recognition labels.
📌 Features
Multitask model (Wav2Vec2MultiTasks) with shared Wav2Vec2 encoder and separate heads for:
Speech Recognition (CTC)
Speaker classification
Emotion classification
Custom data preprocessing:
Cleans transcripts (removes punctuation & special characters)
Converts numbers into words
Builds a vocabulary and tokenizer
Filters short/invalid audio
Training, validation, and test splits with collators for CTC.
Evaluation metrics:
Character Error Rate (CER) for character recognition
Accuracy for speaker and emotion classification sh