Downloads · 30 days
74
0% of all-time downloads
mtgv/MobileVLM-3B
MobileVLM-3B is a text generation model from mtgv. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
MobileVLM is a competent multimodal vision language model (MMVLM) targeted to run on mobile devices. It is an amalgamation of a myriad of architectural designs and techniques that are mobile-oriented, which comprises…
Downloads · 30 days
74
0% of all-time downloads
All-time downloads
27.5K
Public
Repo size
12.1 GB
Likes
14
Public
Click a slice to open those files.
.bin6.1 GB · 100%
From the Hugging Face model README
MobileVLM is a competent multimodal vision language model (MMVLM) targeted to run on mobile devices. It is an amalgamation of a myriad of architectural designs and techniques that are mobile-oriented, which comprises a set of language models at the scale of 1.4B and 2.7B parameters, trained from scratch, a multimodal vision model that is pre-trained in the CLIP fashion, cross-modality interaction via an efficient projector. We evaluate MobileVLM on several typical VLM benchmarks. Our models demonstrate on par performance compared with a few much larger models. More importantly, we measure the inference speed on both a Qualcomm Snapdragon 888 CPU and an NVIDIA Jeston Orin GPU, and we obtain state-of-the-art performance of 21.5 tokens and 65.3 tokens per second, respectively.
The MobileVLM-3B was built on our MobileLLaMA-2.7B-Chat to facilitate the off-the-shelf deployment.
Inference examples can be found at Github.
Please refer to our paper: MobileVLM: A Fast, Strong and Open Vision Language Assistant for Mobile Devices