Downloads · 30 days
11
1% of all-time downloads
hfl/vle-base-for-vqa
vle-base-for-vqa is a machine learning model from hfl. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
VLE (Visual-Language Encoder) is an image-text multimodal understanding model built on the pre-trained text and image encoders. It can be used for multimodal discriminative tasks such as visual question answering and…
Downloads · 30 days
11
1% of all-time downloads
All-time downloads
1.7K
Public
Repo size
3.1 GB
Likes
1
Public
Click a slice to open those files.
.bin1.5 GB · 100%
From the Hugging Face model README
VLE (Visual-Language Encoder) is an image-text multimodal understanding model built on the pre-trained text and image encoders. It can be used for multimodal discriminative tasks such as visual question answering and image-text retrieval. Especially on the visual commonsense reasoning (VCR) task, which requires high-level language understanding and reasoning skills, VLE achieves significant improvements.
For more details see https://github.com/iflytek/VLE.
Online VLE demo on Visual Question Answering: https://huggingface.co/spaces/hfl/VQA_VLE_LLM