Downloads · 30 days
8
13% of all-time downloads
TiMauzi/EraClassifierBiLSTM-4.76M
EraClassifierBiLSTM-4.76M is a audio classification model from TiMauzi. Use it for the audio classification task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-sa-4.0.
This model is a compact bidirectional LSTM neural network designed for musical era classification from MIDI data. It achieves the following results on the evaluation set: - Loss: 1.0935 - Accuracy: 0.5852 - F1: 0.4299
Downloads · 30 days
8
13% of all-time downloads
All-time downloads
62
Public
Parameters
4.8M
38.6 MB on disk
Likes
0
Public
Click a slice to open those files.
.bin38.1 MB · 50%
From the Hugging Face model README
This model is a compact bidirectional LSTM neural network designed for musical era classification from MIDI data. It achieves the following results on the evaluation set:
The EraClassifierBiLSTM-4.76M is a custom bidirectional LSTM neural network specifically designed for classifying musical compositions into historical eras based on MIDI data analysis. This compact model variant (~4.76M parameters) offers a good balance between performance and computational efficiency. For higher accuracy at the cost of compute, see the larger EraClassifierBiLSTM-134M model.
The model processes 8 key MIDI features per message, automatically selected as the most frequent features across the dataset:
Numerical Features (7):
Categorical Features (1):
All numerical features are normalized using dataset statistics (mean and standard deviation), while categorical features are encoded using learned ID mappings.
The model uses a sliding window approach to capture temporal patterns in musical structure that are characteristic of different historical periods. Each MIDI file is processed into multiple overlapping sequences, allowing the model to learn both local and global musical patterns.
Below is the confusion matrix for the best-performing checkpoint, visually highlighting these misclassifications (click to enlarge):
<img src="confusion_matrix_best.png" alt="Confusion Matrix" width="500"/>
The numbers 0 through 5 correspond to each era's index during inference.
The model was trained on 6,992 MIDI files from the IMSLP dataset with the following era distribution:
Era thresholding was applied (minimum 150 samples per era), with rare eras like "Early 20th century" (125 samples) and "Medieval" (5 samples) mapped to the "Other" category to maintain classification stability.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Accuracy | F1 |
|---|---|---|---|---|---|
| 1.2797 | 0.1031 | 2000 | 1.3522 | 0.4608 | 0.2486 |
| 1.1521 | 0.2063 | 4000 | 1.2422 | 0.4987 | 0.3139 |
| 1.0887 | 0.3094 | 6000 | 1.2189 | 0.5056 | 0.3223 |
| 1.0432 | 0.4126 | 8000 | 1.1715 | 0.5252 | 0.3479 |
| 1.019 | 0.5157 | 10000 | 1.2021 | 0.5150 | 0.3304 |
| 0.9963 | 0.6188 | 12000 | 1.1789 | 0.5252 | 0.3487 |
| 0.976 | 0.7220 | 14000 | 1.1151 | 0.5759 | 0.3983 |
| 0.9544 | 0.8251 | 16000 | 1.1800 | 0.5299 | 0.3529 |
| 0.9455 | 0.9283 | 18000 | 1.1866 | 0.5415 | 0.3662 |
| 0.9276 | 1.0314 | 20000 | 1.1744 | 0.5350 | 0.3792 |
| 0.9167 | 1.1345 | 22000 | 1.1032 | 0.5774 | 0.4120 |
| 0.9084 | 1.2377 | 24000 | 1.1312 | 0.5553 | 0.3818 |
| 0.8758 | 1.3408 | 26000 | 1.1042 | 0.5667 | 0.4109 |
| 0.859 | 1.4440 | 28000 | 1.1065 | 0.5733 | 0.4125 |
| 0.8607 | 1.5471 | 30000 | 1.1104 | 0.5695 | 0.4115 |
| 0.8526 | 1.6503 | 32000 | 1.1011 | 0.5830 | 0.4255 |
| 0.8559 | 1.7534 | 34000 | 1.1083 | 0.5765 | 0.4136 |
| 0.8501 | 1.8565 | 36000 | 1.1113 | 0.5752 | 0.4163 |
| 0.8497 | 1.9597 | 38000 | 1.0935 | 0.5775 | 0.4220 |
| 0.8473 | 2.0628 | 40000 | 1.1092 | 0.5745 | 0.4181 |
| 0.8441 | 2.1660 | 42000 | 1.1095 | 0.5733 | 0.4164 |
| 0.8396 | 2.2691 | 44000 | 1.0935 | 0.5852 | 0.4299 |
| 0.8391 | 2.3722 | 46000 | 1.1054 | 0.5744 | 0.4160 |
| 0.8401 | 2.4754 | 48000 | 1.1008 | 0.5755 | 0.4198 |
| 0.8327 | 2.5785 | 50000 | 1.1097 | 0.5712 | 0.4132 |
| 0.838 | 2.6817 | 52000 | 1.1055 | 0.5720 | 0.4143 |
| 0.8329 | 2.7848 | 54000 | 1.1055 | 0.5728 | 0.4165 |
| 0.8346 | 2.8879 | 56000 | 1.1038 | 0.5743 | 0.4172 |
| 0.8353 | 2.9911 | 58000 | 1.1090 | 0.5728 | 0.4167 |
| 0.8385 | 3.0942 | 60000 | 1.1013 | 0.5755 | 0.4201 |
| 0.8337 | 3.1974 | 62000 | 1.1088 | 0.5733 | 0.4163 |
| 0.8256 | 3.3005 | 64000 | 1.1076 | 0.5748 | 0.4177 |
| 0.8367 | 3.4036 | 66000 | 1.1066 | 0.5730 | 0.4159 |
| 0.831 | 3.5068 | 68000 | 1.1083 | 0.5732 | 0.4164 |
| 0.8283 | 3.6099 | 70000 | 1.1067 | 0.5744 | 0.4173 |
| 0.8349 | 3.7131 | 72000 | 1.1058 | 0.5747 | 0.4180 |
| 0.8313 | 3.8162 | 74000 | 1.1058 | 0.5741 | 0.4171 |
| 0.8313 | 3.9193 | 76000 | 1.1065 | 0.5735 | 0.4169 |
| 0.8309 | 4.0225 | 78000 | 1.1067 | 0.5736 | 0.4171 |
| 0.8331 | 4.1256 | 80000 | 1.1055 | 0.5744 | 0.4174 |
| 0.8371 | 4.2288 | 82000 | 1.1058 | 0.5735 | 0.4167 |
| 0.8344 | 4.3319 | 84000 | 1.1060 | 0.5734 | 0.4166 |
| 0.8291 | 4.4350 | 86000 | 1.1049 | 0.5747 | 0.4185 |
| 0.8343 | 4.5382 | 88000 | 1.1053 | 0.5735 | 0.4171 |
| 0.8293 | 4.6413 | 90000 | 1.1056 | 0.5736 | 0.4174 |
| 0.8294 | 4.7445 | 92000 | 1.1056 | 0.5736 | 0.4174 |
| 0.8316 | 4.8476 | 94000 | 1.1055 | 0.5736 | 0.4174 |
| 0.8264 | 4.9508 | 96000 | 1.1056 | 0.5736 | 0.4174 |
Below is the full training metrics plot, showing loss, accuracy, and F1-score trends over the entire training process (click to enlarge):
<img src="training_metrics.png" alt="Training Metrics" width="500"/>
The training shows stable convergence with the model reaching its best performance around step 44,000 (epoch 2.27). The training loss decreases steadily while validation metrics stabilize, indicating good generalization without severe overfitting. The model achieves its peak F1 score of 0.4299 at step 44,000, which was selected as the best checkpoint.