Downloads · 30 days
0
ai-forever/fbc3_baseline
fbc3_baseline is a machine learning model from ai-forever. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This solution is inspired by the methodologies of FROMAGe and Kosmos-1. It primarily employs these approaches to fine-tune the linear mapping from visual and audio vector spaces into the language model-decoder's vecto…
Downloads · 30 days
0
Access
Public
Updated Sep 28, 2023
Repo size
269 MB
Likes
0
Public
Click a slice to open those files.
Other269 MB · 99%
From the Hugging Face model README
This solution is inspired by the methodologies of FROMAGe and Kosmos-1. It primarily employs these approaches to fine-tune the linear mapping from visual and audio vector spaces into the language model-decoder's vector space. Subsequently, the response is generated exclusively using the intact language model.
As a modality encoder, we utilize ImageBind. This encoder has been trained specifically for understanding images, audio, and text and other data formats in a shared embedding space.
During the training phase, the weights of the encoder and the language model remain frozen. The exceptions to this are the additional embeddings for two tokens marking the beginning and end of the respective modalities in the language model: <SOI>, <EOI> and <SOA>, <EOA> (S, E — Start, End; I,A — Image, Audio).
Training is leveraged using four datasets: VisualDialogues, COCO Captions, Clotho v. 2.1 and Clotho-AQA. The core training objective is next token prediction with CrossEntropy loss. The general architecture is illustrated here:
<p align="center"> <img alt="Baseline model architecture" src="./baseline.png" width="100%"> </p>To reproduce training, please run notebook after installing requirements:
pip install requirements.txt