Downloads · 30 days
42
7% of all-time downloads
sartifyllc/AViLaMa
AViLaMa is a zero-shot image classification model from sartifyllc. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Learning Visual Concepts Directly From African Languages Supervision. Paper is coming
Downloads · 30 days
42
7% of all-time downloads
All-time downloads
643
Public
Parameters
775M
43.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.bin3.1 GB · 50%
From the Hugging Face model README
Learning Visual Concepts Directly From African Languages Supervision. Paper is coming
AViLaMa is the large open-source text-vision alignment pre-training model in African languages. It brings a way to learn visual concepts directly from African languages supervision. Based on African languages to capture the nuances, cultural context, and social aspect use of our languages that are so impossible to get just from machine translation. It includes techniques like agnostic languages encoding, data filtering network etc... All for more than 12 African languages, trained on the #AViLaDa-2B datasets of filtered image-text pairs.
import torch
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("sartifyllc/AViLaMa")
tokenizer = AutoTokenizer.from_pretrained("sartifyllc/AViLaMa")
model = model.eval()
BibTeX:
AViLaMa paper
@article{sartifyllc2023africanvision,
title={AViLaMa: Learning Visual Concepts Directly From African Languages Supervision},
author={Sartify LLC Research Team},
journal={To be inserted},
year={2024}
}