Downloads · 30 days
0
vincentamato/ARIA
ARIA is a machine learning model from vincentamato. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
ARIA is a multimodal AI model that generates MIDI music based on the emotional content of artwork. It uses a CLIP-based image encoder to extract emotional valence and arousal from images, then generates emotionally ap…
Downloads · 30 days
0
Access
Public
Updated Jul 23, 2025
Repo size
5.2 GB
Likes
2
Public
Click a slice to open those files.
.pt3.5 GB · 100%
From the Hugging Face model README
ARIA is a multimodal AI model that generates MIDI music based on the emotional content of artwork. It uses a CLIP-based image encoder to extract emotional valence and arousal from images, then generates emotionally appropriate music using conditional MIDI generation.
ARIA consists of two main components:
The model offers three different conditioning modes:
continuous_concat (Recommended)Creates a single vector from valence and arousal values, repeats it across the sequence, and concatenates it with every music token embedding. This approach gives the emotion information global influence throughout the entire generation process, allowing the transformer to access emotional context at every timestep. Research shows this method achieves the best performance in both note prediction accuracy and emotional coherence.
continuous_tokenConverts each emotion value (valence and arousal) into separate condition vectors with the same length as music token embeddings, then concatenates them in the sequence dimension. The emotion vectors are inserted at the beginning of the input sequence during generation. This treats emotions similarly to music tokens but can lose influence as the sequence grows longer.
discrete_tokenQuantizes continuous emotion values into 5 discrete bins (very low, low, moderate, high, very high) and converts them into control tokens. These tokens are placed before the music tokens in the sequence. While this represents the current state-of-the-art approach in conditional text generation, it suffers from information loss due to binning and can lose emotional context during longer generations when tokens are truncated.
The repository contains three variants of the MIDI generation model, each trained with a different conditioning strategy. Each variant includes:
model.pt: The trained model weightsmappings.pt: Token mappings for MIDI generationmodel_config.pt: Model configurationAdditionally, image_encoder.pt contains the CLIP-based image emotion encoder.
This model is designed for:
The model combines:
This project builds upon:
This model is released under the MIT License. However, usage of the midi-emotion component should comply with its GPL-3.0 license.