Downloads · 30 days
0
KalineZephyr/music-lstm-midi-codealpha
music-lstm-midi-codealpha is a machine learning model from KalineZephyr. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A Stacked LSTM + Attention model for symbolic piano music generation, trained on 46 million musical events from ~12,000 MIDI files.
Downloads · 30 days
0
Access
Public
Updated Jun 27, 2026
Repo size
264 MB
Likes
0
Public
Click a slice to open those files.
.pt264 MB · 100%
From the Hugging Face model README
A Stacked LSTM + Attention model for symbolic piano music generation, trained on 46 million musical events from ~12,000 MIDI files.
Built from scratch as part of the CodeAlpha AI Internship (Task 3) — no pre-trained model, no external API, just raw deep learning on MIDI data.
The model treats music generation as a next-token prediction problem — the same principle behind language models like GPT, applied to piano music.
Each musical event is encoded as a single token representing three attributes simultaneously:
The model learns to predict the next token given a sequence of 64 past tokens, then generates music autoregressively.
Token sequence (64 tokens)
│
▼
Embedding (vocab_size=6543, dim=256)
│
▼
LSTM Layer 1 (hidden=1024) ← local patterns: intervals, rhythm
│
▼
LSTM Layer 2 (hidden=1024) ← higher-level patterns: phrases, motifs
│
▼
Attention (additive, Bahdanau-style)
│
▼
Linear → Softmax (6543 classes)
│
▼
Next token prediction
| Parameter | Value |
|---|---|
| Total parameters | 22,030,480 |
| Vocabulary size | 6,543 tokens |
| Sequence length | 64 tokens |
| Hidden size | 1,024 |
| LSTM layers | 2 |
| Embedding dim | 256 |
| Dropout | 0.3 |
| Dataset | Files | Events |
|---|---|---|
| Maestro v3.0.0 | 1,276 | ~6M |
| GiantMIDI-Piano v1.21 | 10,841 | ~41M |
| Total | 12,109 | ~46M |
Both datasets consist of professional and semi-professional solo piano recordings in classical style, ensuring a consistent musical domain.
Preprocessing used symusic (C++ MIDI parser) for fast extraction of pitch, duration, and velocity attributes from raw MIDI files.
Trained on Kaggle 2×T4 GPUs (30GB VRAM total) using torch.nn.DataParallel.
| Hyperparameter | Value |
|---|---|
| Optimizer | Adam (lr=0.001) |
| Scheduler | ReduceLROnPlateau (factor=0.5, patience=2) |
| Batch size | 256 |
| Gradient clipping | 5.0 |
| Early stopping patience | 5 |
| Epoch | Train Loss | Val Loss |
|---|---|---|
| 1 | 6.834 | 6.407 |
| 4 | 5.906 | 5.900 |
| 8 | 5.505 | 5.806 ✓ best |
| 13 | 5.041 | 5.857 — early stop |
Best checkpoint: epoch 8, val_loss = 5.805
Random baseline: ln(6543) ≈ 8.78 — the model significantly outperforms random prediction.
git clone https://github.com/Tahsine/CodeAlpha_Music_Generation
cd CodeAlpha_Music_Generation
pip install -r requirements.txt
The model weights are downloaded automatically from this repo on first run:
# Generate MIDI (256 notes, temperature=0.9)
python generate.py
# Full options
python generate.py \
--n_tokens 512 \
--temperature 0.9 \
--bpm 120 \
--output artifacts/my_music.mid \
--device cpu
# Install FluidSynth
sudo apt-get install fluidsynth # Linux
brew install fluidsynth # macOS
# Generate audio — soundfont (~30MB) downloads automatically
python generate.py --n_tokens 512 --temperature 0.9 --audio
| Temperature | Effect |
|---|---|
0.7 | Conservative — coherent but repetitive |
0.9 | Balanced — musical and varied (recommended) |
1.1 | Creative — surprising but less coherent |
import torch
from huggingface_hub import hf_hub_download
# Download weights
model_path = hf_hub_download(
repo_id="KalineZephyr/music-lstm-midi-codealpha",
filename="best_model.pt",
)
# Load checkpoint
ckpt = torch.load(model_path, map_location="cpu")
print(f"Best epoch : {ckpt['epoch']}")
print(f"Val loss : {ckpt['val_loss']:.4f}")
print(f"Config : {ckpt['config']}")
These limitations are inherent to LSTM-based sequence models. A Transformer architecture with full self-attention would address the long-range coherence issue.
GitHub: Tahsine/CodeAlpha_Music_Generation
MIT