Downloads ยท 30 days
0
saadmannan/speech-emotion-recognition
speech-emotion-recognition is a audio classification model from saadmannan. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
[](https://www.python.org/downloads/release/python-3100/) [](https://pytorch.org/) [](https://opensource.org/licenses/MIT)
Downloads ยท 30 days
0
Access
Public
Updated Nov 13, 2025
Repo size
143 MB
Likes
0
Public
Click a slice to open those files.
.pth143 MB ยท 100%
From the Hugging Face model README
A production-ready deep learning system for detecting emotions from speech using the RAVDESS dataset. Achieved 75% validation accuracy through enhanced CNN architecture with residual connections, attention mechanisms, and comprehensive data augmentation.
โ
Primary Goal Met: 75% validation accuracy (66.2% test accuracy)
โ
Enhanced Features: 196-dimensional feature vectors
โ
Advanced Architecture: 11.8M parameter CNN with residual blocks and attention
โ
Production Ready: Complete pipeline from data to deployment
| Metric | Baseline Model | Enhanced Model | Improvement |
|---|---|---|---|
| Validation Accuracy | 38.89% | 75.00% | +36.11% |
| Test Accuracy | 39.81% | 66.20% | +26.39% |
| Parameters | 536K | 11.8M | 22x larger |
| Features | 143 | 196 | +37% richer |
| Emotion | Baseline | Enhanced | Improvement | Status |
|---|---|---|---|---|
| Neutral | 78.57% | 71.43% | -7.14% | โ Good |
| Calm | 85.71% | 85.71% | +0.00% | โ Excellent |
| Happy | 6.90% | 58.62% | +51.72% | ๐ Huge gain |
| Sad | 0.00% | 51.72% | +51.72% | ๐ Huge gain |
| Angry | 31.03% | 68.97% | +37.94% | โ Major gain |
| Fearful | 13.79% | 41.38% | +27.59% | โ Good gain |
| Disgust | 68.97% | 75.86% | +6.89% | โ Improved |
| Surprised | 55.17% | 79.31% | +24.14% | โ Major gain |
# Clone the repository
git clone https://github.com/yourusername/speech-emotion-recognition.git
cd speech-emotion-recognition
# Create conda environment
conda create -n voice_ai python=3.10
conda activate voice_ai
# Install dependencies
pip install -r requirements.txt
python data/download_dataset.py
python data/prepare_data.py
python models/train_v2.py
python models/evaluate_v2.py
streamlit run deployment/app.py
import torch
from models.emotion_cnn_v2 import ImprovedEmotionCNN
from data.prepare_data import extract_features
# Load model
model = ImprovedEmotionCNN(num_classes=8)
checkpoint = torch.load('results/best_model_v2.pth')
model.load_state_dict(checkpoint['model_state_dict'])
model.eval()
# Extract features from audio
features = extract_features('path/to/audio.wav')
features_tensor = torch.FloatTensor(features).unsqueeze(0).unsqueeze(0)
# Predict
with torch.no_grad():
output = model(features_tensor)
probs = torch.softmax(output, dim=1)
predicted = output.argmax(1)
emotions = ['neutral', 'calm', 'happy', 'sad', 'angry', 'fearful', 'disgust', 'surprised']
print(f"Predicted emotion: {emotions[predicted]}")
print(f"Confidence: {probs[0][predicted]:.2%}")
Features (196 dimensions):
Model Architecture:
Input (1, 196, 128)
โ
Conv2d 7ร7, stride 2 โ 64 channels
โ
Residual Block ร 2 (64 channels) + Channel Attention
โ
Residual Block ร 2 (128 channels) + Channel Attention
โ
Residual Block ร 2 (256 channels) + Channel Attention
โ
Residual Block ร 2 (512 channels) + Channel Attention
โ
Dual Global Pooling (Avg + Max) โ 1024 features
โ
FC 1024 โ 512 โ 256 โ 8 (emotions)
Total Parameters: 11,873,480
Key Improvements:
Features (143 dimensions):
Model Architecture:
speech-emotion-recognition/
โโโ data/
โ โโโ download_dataset.py # RAVDESS dataset downloader
โ โโโ prepare_data.py # Enhanced feature extraction (196 features)
โ โโโ dataset.py # PyTorch Dataset with train/val/test splits
โ โโโ augmentation.py # Data augmentation (SpecAugment, noise, etc.)
โ
โโโ models/
โ โโโ emotion_cnn.py # Baseline CNN (536K params)
โ โโโ emotion_cnn_v2.py # Enhanced CNN (11.8M params) โญ
โ โโโ train.py # Baseline training script
โ โโโ train_v2.py # Enhanced training script โญ
โ โโโ evaluate.py # Baseline evaluation
โ โโโ evaluate_v2.py # Enhanced evaluation โญ
โ
โโโ deployment/
โ โโโ app.py # Streamlit demo application
โ โโโ requirements.txt # Deployment dependencies
โ
โโโ notebooks/
โ โโโ emotion_eda.ipynb # Exploratory analysis + model comparison
โ
โโโ results/
โ โโโ best_model.pth # Baseline model weights
โ โโโ best_model_v2.pth # Enhanced model weights โญ
โ โโโ confusion_matrix_v2.png # Confusion matrix visualization
โ โโโ per_class_accuracy_v2.png # Per-class performance chart
โ โโโ model_comparison.png # Baseline vs Enhanced comparison
โ
โโโ runs/ # TensorBoard logs
โโโ README.md # This file
โโโ requirements.txt # Python dependencies
โโโ LICENSE # MIT License
Ryerson Audio-Visual Database of Emotional Speech and Song
config = {
'batch_size': 24,
'learning_rate': 0.001,
'epochs': 150,
'optimizer': 'AdamW',
'weight_decay': 1e-4,
'loss': 'CrossEntropyLoss + Label Smoothing (0.1)',
'lr_scheduler': 'ReduceLROnPlateau (patience=8, factor=0.5)',
'early_stopping': 'patience=20',
'mixed_precision': 'FP16',
'gradient_clipping': 'max_norm=1.0',
'data_augmentation': True
}
tensorboard --logdir=runs/
View real-time training metrics:
The trained model is available on Hugging Face:
from huggingface_hub import hf_hub_download
model_path = hf_hub_download(
repo_id="yourusername/speech-emotion-recognition",
filename="best_model_v2.pth"
)
Live demo: [Coming Soon]
streamlit run deployment/app.py
Features:
precision recall f1-score support
neutral 0.667 0.714 0.690 14
calm 0.686 0.857 0.762 28
happy 0.531 0.586 0.557 29
sad 0.500 0.517 0.508 29
angry 0.769 0.690 0.727 29
fearful 0.706 0.414 0.522 29
disgust 0.688 0.759 0.721 29
surprised 0.793 0.793 0.793 29
accuracy 0.662 216
macro avg 0.667 0.666 0.660 216
weighted avg 0.667 0.662 0.658 216
# Test model architecture
python models/emotion_cnn_v2.py
# Test dataset loading
python data/dataset.py
# Check environment
python quick_start.py
# Complete pipeline
./run_pipeline.sh
# Or step by step:
python data/download_dataset.py
python data/prepare_data.py
python models/train_v2.py
python models/evaluate_v2.py
RAVDESS Dataset: Livingstone SR, Russo FA (2018) The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS). PLoS ONE 13(5): e0196391.
SpecAugment: Park et al. (2019) "SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition"
ResNet: He et al. (2016) "Deep Residual Learning for Image Recognition"
Channel Attention: Hu et al. (2018) "Squeeze-and-Excitation Networks"
Contributions are welcome! Please feel free to submit a Pull Request.
git checkout -b feature/AmazingFeature)git commit -m 'Add some AmazingFeature')git push origin feature/AmazingFeature)This project is licensed under the MIT License - see the LICENSE file for details.
For questions or feedback, please open an issue on GitHub.
Built with โค๏ธ using PyTorch, librosa, and Streamlit