Downloads ยท 30 days
19
2% of all-time downloads
Narsil/riffusion
riffusion is a text-to-speech model from Narsil. Use it when you need text read aloud. It is set up for diffusers. The card lists the license as creativeml-openrail-m.
Riffusion is an app for real-time music generation with stable diffusion.
Downloads ยท 30 days
19
2% of all-time downloads
All-time downloads
770
Public
Repo size
21.8 GB
Likes
1
Public
Click a slice to open those files.
.ckpt14.6 GB ยท 67%
From the Hugging Face model README
Riffusion is an app for real-time music generation with stable diffusion.
Read about it at https://www.riffusion.com/about and try it at https://www.riffusion.com/.
This repository contains the model files, including:
Riffusion is a latent text-to-image diffusion model capable of generating spectrogram images given any text input. These spectrograms can be converted into audio clips.
The model was created by Seth Forsgren and Hayk Martiros as a hobby project.
You can use the Riffusion model directly, or try the Riffusion web app.
The Riffusion model was created by fine-tuning the Stable-Diffusion-v1-5 checkpoint. Read about Stable Diffusion here ๐ค's Stable Diffusion blog.
The model is intended for research purposes only. Possible research areas and tasks include
If you build on this work, please cite it as follows:
@software{Forsgren_Martiros_2022,
author = {Forsgren, Seth* and Martiros, Hayk*},
title = {{Riffusion - Stable diffusion for real-time music generation}},
url = {https://riffusion.com/about},
year = {2022}
}