Downloads · 30 days
577
1% of all-time downloads
AVoCaDO-Captioner/AVoCaDO
AVoCaDO is a machine learning model from AVoCaDO-Captioner. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<p align="left" <a href="https://avocado-captioner.github.io/"<img src="https://img.shields.io/badge/Project%20webpage-558b2f?style=for-the-badge"</a <a href="https://github.com/AVoCaDO-Captioner/AVoCaDO"<img src="htt…
Downloads · 30 days
577
1% of all-time downloads
All-time downloads
39.1K
Public
Parameters
8.9B
17.9 GB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors17.9 GB · 100%
From the Hugging Face model README
Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. We introduce <b>AVoCaDO</b>, a powerful audiovisual video captioner driven by the temporal orchestration between audio and visual modalities. Experimental results demonstrate that AVoCaDO significantly outperforms existing open-source models across four audiovisual video captioning benchmarks, and also achieves competitive performance under visual-only settings.
Please refer to our Github repository for more details.
If you find our work helpful for your research, please consider giving a star ⭐ and citing our paper. We appreciate your support!
@article{chen2025avocado,
title={AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration},
author={Chen, Xinlong and Ding, Yue and Lin, Weihong and Hua, Jingyun and Yao, Linli and Shi, Yang and Li, Bozhou and Zhang, Yuanxing and Liu, Qiang and Wan, Pengfei and others},
journal={arXiv preprint arXiv:2510.10395},
year={2025}
}