Downloads · 30 days
29
67% of all-time downloads
nishitanand/FIGMA
FIGMA is a feature extraction model from nishitanand. Use it when you need embeddings to search or compare text. It is set up for pytorch. The card lists the license as cc-by-nc-4.0.
Downloads · 30 days
29
67% of all-time downloads
All-time downloads
43
Public
Repo size
3.8 GB
Likes
0
Public
Click a slice to open those files.
.ckpt3.8 GB · 100%
From the Hugging Face model README
Accepted to ACL 2026.
FIGMA retrieves music from natural-language descriptions that specify fine-grained musical attributes — tempo, key, chord progression, beat/time signature — in addition to high-level attributes like genre and mood. This repository hosts the trained model checkpoint and the code to run inference.
pip install torch transformers librosa soundfile numpy
pip3 install muq
from huggingface_hub import snapshot_download
import os
local_dir = snapshot_download(repo_id="nishitanand/FIGMA")
ckpt = os.path.join(local_dir, "figma.ckpt")
import torch, librosa
from figma_model import Figma, get_tokenizer
device = "cuda"
model = Figma.from_checkpoint(ckpt, device=device) # ckpt from the download step
tok = get_tokenizer()
# Text embedding
t = tok(["a song in F minor at 120 BPM in 4/4 time"], return_tensors="pt",
padding="max_length", truncation=True, max_length=128).to(device)
text_emb = model.encode_text(t) # [1, 512], L2-normalized
# Audio embedding (24 kHz mono -> [B, 1, samples])
wav, _ = librosa.load("clip.wav", sr=24000)
wav = torch.tensor(wav)[None, None].to(device)
audio_emb = model.encode_audio(wav) # [1, 512], L2-normalized
similarity = (audio_emb @ text_emb.T).item()
Command-line text→audio retrieval over a folder:
python inference.py --checkpoint figma.ckpt \
--audio_dir ./clips --query "a jazzy track in C major at 90 bpm" --topk 5
@inproceedings{figma2026,
title = {FIGMA: Towards FIne-Grained Music retrievAl},
author = {Anand, Nishit and Seth, Ashish and Ghosh, Sreyan and Manocha, Dinesh and Duraiswami, Ramani},
booktitle = {Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics},
year = {2026},
url = {https://arxiv.org/abs/2606.06615}
}
Released for non-commercial research use under CC BY-NC 4.0. The checkpoint bundles the MuQ encoder (weights CC BY-NC 4.0); the E5 encoder is MIT. You must attribute MuQ and E5 and comply with CC BY-NC 4.0. The inference code is MIT-licensed in the code repository.