Downloads · 30 days
0
nvidia/audio-flamingo
audio-flamingo is a machine learning model from nvidia. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, Bryan Catanzaro
Downloads · 30 days
0
Access
Public
Updated Oct 2, 2024
Repo size
35.3 GB
Likes
28
Public
Click a slice to open those files.
.pt35.3 GB · 100%
From the Hugging Face model README
Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, Bryan Catanzaro
This repo contains the model checkpoints of Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities (ICML 2024). Audio Flamingo is a novel audio-understanding language model with
We introduce a series of training techniques, architecture design, and data strategies to enhance our model with these abilities. Extensive evaluations across various audio understanding tasks confirm the efficacy of our method, setting new state-of-the-art benchmarks. Sound demos can be found in this website.

Our code is at https://github.com/NVIDIA/audio-flamingo
@article{kong2024audio,
title={Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities},
author={Kong, Zhifeng and Goel, Arushi and Badlani, Rohan and Ping, Wei and Valle, Rafael and Catanzaro, Bryan},
journal={arXiv preprint arXiv:2402.01831},
year={2024}
}