Downloads · 30 days
0
projectlosangeles/midisimx
midisimx is a machine learning model from projectlosangeles. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Sep 24, 2026
Repo size
1.5 GB
Likes
2
Public
Click a slice to open those files.
.pth1.1 GB · 99%
From the Hugging Face model README

| Feature / Change | midisimx | midisim |
|---|---|---|
| Model Architecture | ⭐ One unified larger model | Two smaller models |
| Model Dimension | 🔥 768 | 512 |
| Model Depth | 🔥 16 layers | 16 + 8 layers |
| Attention Heads | 🔥 12 heads | 8 heads |
| Training Corpus Size | 🌍 3M+ filtered & processed MIDIs | 1M+ raw MIDIs |
| MIDI Event Representation | 🎼 start-time · note/chord · pitch · duration | start-time · duration · pitch |
| Codebase Quality | 💎 Improved, extended, modernized | Older original codebase |
| Overall Quality | ✅ Major upgrade | Baseline |
midisimx_trained_model_14391_steps_0.255_loss_0.9036_acc.pth - Unified and fast large model for a nuanced embeddings generation. Download checkpoint from Hugging Facediscover_midi_dataset_3267574_clean_midis_embeddings_1_2_1_2_weighted_cc_by_nc_sa.npy - 3267574 clean MIDIs weighted embeddings from Discover MIDI Dataset for large scale similarity search and analysis tasks
lakh_midi_dataset_17203_clean_midis_embeddings_1_2_1_2_weighted_cc_by_nc_sa.npy - 17203 LAKH clean_midi subset weighted embeddings tailored primarily for artist/song identification tasks
midisimx-similarity-search-output-samples-1-2-1-2-weighted-CC-BY-NC-SA.zip - ~182k+ MIDIs filtered by weighted midisimx music discovery pipeline
!pip install -U midisimx
!pip install x-transformers==2.3.1
# ================================================================================================
# Initalize midisimx
# ================================================================================================
# Import main midisimx module
import midisimx
# ================================================================================================
# Prepare midisimx embeddings
# ================================================================================================
# Option 1: Download sample pre-computed embeddings corpus from Hugging Face
emb_path = midisimx.download_embeddings()
# Option 2: use custom pre-computed embeddings corpus
# See custom embeddings generation section of this README for details
# emb_path = './custom_midis_embeddings_corpus.npy'
# Load downloaded embeddings corpus
corpus_midi_names, corpus_emb = midisimx.load_embeddings(emb_path)
# ================================================================================================
# Prepare midisimx model
# ================================================================================================
# Option 1: Download main pre-trained midisimx model from Hugging Face
model_path = midisimx.download_model()
# Option 2: Use main pre-trained midisimx model included in midisimx PyPI package
# model_path = midisimx.get_package_models()[0]['path']
# Load midisimx model
model, ctx, dtype = midisimx.load_model(model_path)
# ================================================================================================
# Prepare source MIDI
# ================================================================================================
# Load source MIDI
input_toks_seqs = midisimx.midi_to_tokens('Come To My Window.mid')
# ================================================================================================
# Calculate and analyze embeddings
# ================================================================================================
# Compute source/query embeddings
query_emb = midisimx.get_embeddings_bf16(model,
input_toks_seqs,
device=torch.device('cuda'),
pooling='weighted_mean',
# The following arg is optional but recommended if
# you want to make an emphasis on music
# Remove it for overall/general similarity searches
# PLEAE NOTE: You must enable it if you are using
# included pre-computed weighted embeddings
token_type_weights={(128, 256): 2, # Pitches weight
(384, 718): 2 # Chords weight
},
)
# Calculate cosine similarity between source/query MIDI embeddings and embeddings corpus
idxs, sims = midisimx.cosine_similarity_topk(query_emb, corpus_emb)
# ================================================================================================
# Processs, print and save results
# ================================================================================================
# Convert the results to sorted list with transpose values
idxs_sims_tvs_list = midisimx.idxs_sims_to_sorted_list(idxs, sims)
# Print corpus matches (and optionally) convert the final result to a handy list for further processing
corpus_matches_list = midisimx.print_sorted_idxs_sims_list(idxs_sims_tvs_list, corpus_midi_names, return_as_list=True)
# ================================================================================================
# Copy matched MIDIs from the MIDI corpus for listening and further evaluation and analysis
# ================================================================================================
# Copy matched corpus MIDI to a desired directory for easy evaluation and analysis
out_dir_path = midisimx.copy_corpus_files(corpus_matches_list)
# ================================================================================================
import torch
from x_transformers import TransformerWrapper, Encoder
# Original model hyperparameters
SEQ_LEN = 3072
MASK_IDX = 718 # Use this value for masked modelling
PAD_IDX = 719 # Model pad index
VOCAB_SIZE = 720 # Total vocab size
MASK_PROB = 0.15 # Original training mask probability value (use for masked modelling)
DEVICE = 'cuda' # You can use any compatible device or CPU
DTYPE = torch.bfloat16 # Original training dtype
# Official main midisimx model checkpoint name
MODEL_CKPT = 'midisimx_trained_model_14391_steps_0.255_loss_0.9036_acc.pth'
# Model architecture using x-transformers
model = TransformerWrapper(
num_tokens = VOCAB_SIZE,
max_seq_len = SEQ_LEN,
attn_layers = Encoder(
dim = 768,
depth = 16,
heads = 12,
rotary_pos_emb = True,
attn_flash = True,
),
)
model.load_state_dict(torch.load(MODEL_CKPT, map_location=DEVICE))
model.to(DEVICE)
model.eval()
# Original training autoxast setup
autocast_ctx = torch.amp.autocast(device_type=DEVICE, dtype=DTYPE)
# ================================================================================================
# Load main midisimx module
import midisimx
# Import helper modules
import os
import tqdm
# ================================================================================================
# Call included TMIDIX module through midisimx to create MIDI files list
custom_midi_corpus_file_names = midisimx.TMIDIX.create_files_list(['./custom_midi_corpus_dir/'])
# ================================================================================================
# Create two lists: one with MIDI corpus file names
# and another with MIDI corpus tokens representations suitable for embeddings generation
midi_corpus_file_names = []
midi_corpus_tokens = []
for midi_file in tqdm.tqdm(custom_midi_corpus_file_names):
midi_corpus_file_names.append(os.path.splitext(os.path.basename(midi_file))[0])
midi_tokens = midisimx.midi_to_tokens(midi_file, transpose_factor=0, verbose=False)[0]
midi_corpus_tokens.append(midi_tokens)
# It is highly recommended to sort the resulting corpus by tokens sequence length
# This greatly speeds up embeddings calculations
sorted_midi_corpus = sorted(zip(midi_corpus_file_names, midi_corpus_tokens), key=lambda x: len(x[1]))
midi_corpus_file_names, midi_corpus_tokens = map(list, zip(*sorted_midi_corpus))
# ================================================================================================
# Now you are ready to generate embeddings as follows:
# ================================================================================================
# Load main midisimx model
model, ctx, dtype = midisimx.load_model(verbose=False)
# Generate MIDI corpus embeddings
midi_corpus_embeddings = midisimx.get_embeddings_bf16(model, midi_corpus_tokens, verbose=False)
# ================================================================================================
# Save generated MIDI corpus embeddings and MIDI corpus file names in one handy NumPy file
midisimx.save_embeddings(midi_corpus_file_names,
midi_corpus_embeddings,
verbose=False
)
# ================================================================================================
# You now can use this saved custom MIDI corpus NumPy file with midisimx.load_embeddings()
# and the rest of the pipeline outlined in the general use section above
Here is a complete MIDI music discovery pipeline example using midisimx and Discover MIDI Dataset
!pip install -U midisimx
!pip install -U discovermidi
import discovermidi
from discovermidi import fast_parallel_extract
discovermidi.download_dataset()
fast_parallel_extract.fast_parallel_extract()
model_ckpt = 'midisimx_trained_model_14391_steps_0.255_loss_0.9036_acc.pth'
model_depth = 16
embeddings_file = 'discover_midi_dataset_3267574_clean_midis_embeddings_1_2_1_2_weighted_cc_by_nc_sa.npy'
import os
os.makedirs('./Master-MIDI-Dataset/', exist_ok=True)
# Import main midisimx module
import midisimx
# Download embeddings from Hugging Face
emb_path = midisimx.download_embeddings(filename=embeddings_file)
# Load downloaded embeddings corpus
corpus_midi_names, corpus_emb = midisimx.load_embeddings(embeddings_path=emb_path)
# Download midisimx model from Hugging Face
model_path = midisimx.download_model(filename=model_ckpt)
# Load midisimx model
model, ctx, dtype = midisimx.load_model(model_path,
depth=model_depth
)
filez = midisimx.TMIDIX.create_files_list(['./Master-MIDI-Dataset/'])
import os
import tqdm
for fa in tqdm.tqdm(filez):
# Load source MIDI
input_toks_seqs = midisimx.midi_to_tokens(fa, verbose=False)
if input_toks_seqs:
# ================================================================================================
# Calculate and analyze embeddings
# ================================================================================================
# Compute source/query embeddings
query_emb = midisimx.get_embeddings_bf16(model,
input_toks_seqs,
device=torch.device('cuda'),
pooling='weighted_mean',
# The following arg is optional but recommended if
# you want to make an emphasis on music
# Remove it for overall/general similarity searches
# PLEAE NOTE: You must enable it if you are using
# included pre-computed weighted embeddings
token_type_weights={(128, 256): 2, # Pitches weight
(384, 718): 2 # Chords weight
},
verbose=False,
show_progress_bar=False
)
# Calculate cosine similarity between source/query MIDI embeddings and embeddings corpus
idxs, sims = midisimx.cosine_similarity_topk(query_emb,
corpus_emb,
verbose=False
)
# ================================================================================================
# Processs, print and save results
# ================================================================================================
# Convert the results to sorted list with transpose values
idxs_sims_tvs_list = midisimx.idxs_sims_to_sorted_list(idxs, sims)
# Print corpus matches (and optionally) convert the final result to a handy list for further processing
corpus_matches_list = midisimx.print_sorted_idxs_sims_list(idxs_sims_tvs_list,
corpus_midi_names,
return_as_list=True
)
# ================================================================================================
# Copy matched MIDIs from the MIDI corpus for listening and further evaluation and analysis
# ================================================================================================
# Copy matched corpus MIDI to a desired directory for easy evaluation and analysis
out_dir_path = midisimx.copy_corpus_files(corpus_matches_list,
corpus_midis_dirs=['./Discover-MIDI-Dataset/MIDIs/'],
main_output_dir='Output-MIDI-Dataset',
sub_output_dir=os.path.splitext(os.path.basename(fa))[0],
verbose=False
)
# ================================================================================================
midisimx uses a compact, event‑structured token format that lets the model understand timing, harmony, melody, and rhythm with minimal overhead.
Each event is encoded in a strict order, and notes and chords share the same structure—chords simply contain multiple pitch–duration pairs.
| Token Type | Range | Meaning | Notes |
|---|---|---|---|
| Delta Start‑Time | 0–127 | Time since previous event | Encodes rhythmic spacing |
| Note/Chord Token | 384–717 | Semitone or chord class | 384–395 → 12 semitones; 396–716 → 321 chords |
| Pitch | 128–255 | MIDI pitch (0–127) | One per note; multiple for chords |
| Duration | 256–383 | Note length | One per pitch |
A note event always has four tokens:
[delta‑start, note-token, pitch, duration]
A chord event starts with the same two tokens, but then includes multiple (pitch, duration) pairs:
[delta‑start, chord-token, pitch, duration, pitch, duration, pitch, duration, ...]
This allows encoding triads, extended chords, clusters, or any multi‑note harmony.
Below is a real midisimx token sequence excerpt, formatted for readability.
Events are grouped to show how notes and chords appear:
[0, 643, 193, 321]
[186, 321, 179, 325]
[16, 391, 195, 265]
[9, 391, 195, 298]
[16, 387, 191, 272]
[16, 387, 191, 266]
[8, 689, 193, 323, 186, 321] ← chord (two pitch–duration pairs)
[1, 386, 178, 323]
[15, 391, 195, 265]
[9, 391, 195, 298]
[16, 387, 191, 283]
[24, 711, 196, 321, 186, 321] ← chord
[1, 384, 176, 320]
[15, 391, 195, 296]
[9, 389, 193, 297]
[23, 387, 191, 273]
[8, 391, 195, 315]
[9, 707, 183, 323, 174, 324] ← chord
[50, 391, 195, 274]
...
You can clearly see:
midisimx/ # Project root
├── LICENSE # Apache-2.0 license text
├── MANIFEST.in # Setuptools manifest — package-data inclusion rules for sdist/wheel
├── README.md # Main project README — features, usage guides, links, citations
├── midisimx/ # The installable Python package
│ ├── API_REFERENCE.md # This document — complete public API reference
│ ├── MIDI.py # LEGACY — original parent of TMIDIX; unused, kept for reference/posterity
│ ├── README.md # Package README (PyPI landing page)
│ ├── TMIDIX.py # TMIDIX MIDI parsing/processing suite; re-exported as midisimx.TMIDIX
│ ├── artwork/ # Project images
│ │ ├── Project-Los-Angeles.png # Project Los Angeles logo
│ │ ├── README.md # Artwork notes and credits
│ │ ├── Tegridy-Code-2026.png # Tegridy Code 2026 branding image
│ │ └── midisimx.png # Project banner (embedded in READMEs)
│ ├── crossmodal_mapper.py # Non-ML closed-form bi-directional cross-modal embedding mapper (procrustes/ridge/cca)
│ ├── embeddings/ # Bundled pre-computed embeddings
│ │ ├── README.md # Notes on bundled embeddings sets
│ │ └── lakh_midi_dataset_17209...... # Tiny 128-dim weighted (1-2-1-2) embeddings for 17 209 clean LAKH MIDIs — pairs with the bundled tiny model
│ ├── helpers.py # Utilities — bundled assets listing, MIDI normalization, file hashing, apt install
│ ├── instrumentation_similarity.py # Deterministic timbre-aware GM instrumentation similarity scoring
│ ├── ldmb.py # LDMB — mmap-backed binary storage for large lists of dicts (lazy reads, byte-level merges)
│ ├── memmap.py # Single-file memmap storage for paired names + float32 embeddings
│ ├── midi_to_colab_audio.py # AUX (optional) — renders MIDIs to audio via fluidsynth + SF2 soundfont banks
│ ├── midisimx.py # CORE — model/embeddings I/O, MIDI↔tokens, embedding computation, similarity search
│ ├── models/ # Bundled model checkpoints
│ │ ├── README.md # Notes on bundled models
│ │ └── midisimx_tiny_trained_model... # Tiny 6.49M-param Transformer checkpoint (14 401 steps · 0.5146 loss · 0.8202 acc)
│ ├── pca_reduce.py # Streaming, GPU-accelerated PCA reduction (PCAReductor, PCAReductionResult)
│ └── x_transformer_2_3_1.py # CORE (models) — vendored, stand-alone x-transformers v2.3.1 by lucidrains
└── pyproject.toml # PEP 621 packaging metadata — version, dependencies, PyPI URLs, classifiers
Legend:
fluidsynth (installable via midisimx.helpers.install_apt_package('fluidsynth')) and SF2 banks; audio rendering only.x_transformer_2_3_1.py makes the core pipeline independent of the PyPI x-transformers package (which is only needed for raw/custom tasks).@misc{project_los_angeles_2026,
author = { Project Los Angeles and Tegridy Code },
title = { midisimx (Revision cfed861) },
year = 2026,
url = { https://huggingface.co/projectlosangeles/midisimx },
doi = { 10.57967/hf/10032 },
publisher = { Hugging Face }
}
@misc{project_los_angeles_2026,
author = { Project Los Angeles and Tegridy Code },
title = { midisimx-embeddings (Revision 0af7bbc) },
year = 2026,
url = { https://huggingface.co/datasets/projectlosangeles/midisimx-embeddings },
doi = { 10.57967/hf/10082 },
publisher = { Hugging Face }
}
@misc{project_los_angeles_2026,
author = { Project Los Angeles and Tegridy Code },
title = { midisimx-samples (Revision 3c28df7) },
year = 2026,
url = { https://huggingface.co/datasets/projectlosangeles/midisimx-samples },
doi = { 10.57967/hf/10085 },
publisher = { Hugging Face }
}
@misc{project_los_angeles_2025,
author = { Project Los Angeles },
title = { Discover-MIDI-Dataset (Revision 0eaecb5) },
year = 2025,
url = { https://huggingface.co/datasets/projectlosangeles/Discover-MIDI-Dataset },
doi = { 10.57967/hf/7361 },
publisher = { Hugging Face }
}
@phdthesis{raffel2016learning,
author = { Colin Raffel },
title = { Learning-Based Methods for Comparing Sequences, with Applications to Audio-to-{MIDI} Alignment and Matching },
school = { Columbia University },
year = { 2016 },
url = { https://colinraffel.com/projects/lmd/ }
}