Downloads · 30 days
0
ZeyuXie/PicoAudio
PicoAudio is a machine learning model from ZeyuXie. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Duplicate of github repo [](https://arxiv.org/abs/2407.02869v2) [](https://zeyuxie29.github.io/PicoAudio.github.io/)
Downloads · 30 days
0
Access
Public
Updated Jul 19, 2024
Repo size
1.4 MB
Likes
0
Public
Click a slice to open those files.
.json2.2 MB · 52%
From the Hugging Face model README
Duplicate of github repo
Bullet contribution:
You can see the demo on the website Huggingface Online Inference and Github Demo. Or you can use the "inference.py" script provided by website Huggingface Inference to generate. Huggingface Online Inference uses Gemini as a preprocessor, and we also provide a GPT preprocessing script consistent with the paper in "llm_preprocess.py"
<!-- <[GoogleDrive](https://drive.google.com/file/d/1oez7kzFFhqU9JZQhqJdDshXrRQczBmlp/view?usp=sharing) -->Simulated data can be downloaded from (1) HuggingfaceDataset or (2) BaiduNetDisk with the extraction code "pico".
The metadata is stored in "data/meta_data/{}.json", one instance is as follows:
{
"filepath": "data/multi_event_test/syn_1.wav",
"onoffCaption": "cat meowing at 0.5-2.0, 3.0-4.5 and whistling at 5.0-6.5 and explosion at 7.0-8.0, 8.5-9.5",
"frequencyCaption": "cat meowing two times and whistling one times and explosion two times"
}
where:
Download data into the "data" folder. The training and inference code can be found in the "picoaudio" folder.
cd picoaudio
pip install -r requirements.txt
To start traning:
accelerate launch runner/controllable_train.py
Our code referred to the AudioLDM and Tango. We appreciate their open-sourcing of their code.