Downloads · 30 days
11
31% of all-time downloads
YifanXu/libra-11b-base
libra-11b-base is a image-to-text model from YifanXu. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as apache-2.0.
Libra: Building Decoupled Vision System on Large Language Models
Downloads · 30 days
11
31% of all-time downloads
All-time downloads
36
Public
Parameters
11B
25.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors22 GB · 87%
From the Hugging Face model README
Libra: Building Decoupled Vision System on Large Language Models
This model was trained on image-text pairs for basic multi-modal understanding ability.
In addition to the pretrained weights in this repo, please download the pretrained CLIP model in huggingface and merge it into the path, as:
libra-base/
├── ...
└── openai-clip-vit-large-patch14-336/
└── ...
The CLIP model can be downloaded here.