Downloads · 30 days
21
4% of all-time downloads
Efficient-Large-Model/VILA-7b-4bit-awq
VILA-7b-4bit-awq is a text generation model from Efficient-Large-Model. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
Model type: VILA is a visual language model (VLM) pretrained with interleaved image-text data at scale, enabling multi-image VLM. VILA is deployable on the edge, including Jetson Orin and laptop by AWQ 4bit quantizati…
Downloads · 30 days
21
4% of all-time downloads
All-time downloads
523
Public
Repo size
4.6 GB
Likes
2
Public
Click a slice to open those files.
.pt4.6 GB · 100%
From the Hugging Face model README
Model type: VILA is a visual language model (VLM) pretrained with interleaved image-text data at scale, enabling multi-image VLM. VILA is deployable on the edge, including Jetson Orin and laptop by AWQ 4bit quantization through TinyChat framework. We find: (1) image-text pairs are not enough, interleaved image-text is essential; (2) unfreezing LLM during interleaved image-text pre-training enables in-context learning; (3)re-blending text-only instruction data is crucial to boost both VLM and text-only performance. VILA unveils appealing capabilities, including: multi-image reasoning, in-context learning, visual chain-of-thought, and better world knowledge.
Model date: VILA-7b-4bit-awq was trained in Feb 2024.
Paper or resources for more information: https://github.com/Efficient-Large-Model/VILA
@misc{lin2023vila,
title={VILA: On Pre-training for Visual Language Models},
author={Ji Lin and Hongxu Yin and Wei Ping and Yao Lu and Pavlo Molchanov and Andrew Tao and Huizi Mao and Jan Kautz and Mohammad Shoeybi and Song Han},
year={2023},
eprint={2312.07533},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
Where to send questions or comments about the model: https://github.com/Efficient-Large-Model/VILA/issues
Primary intended uses: The primary use of VILA is research on large multimodal models and chatbots.
Primary intended users: The primary intended users of the model are researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence.
See Dataset Preparation for more details.
A collection of 12 benchmarks, including 5 academic VQA benchmarks and 7 recent benchmarks specifically proposed for instruction-following LMMs.