Downloads · 30 days
145
1% of all-time downloads
mtgv/MobileVLM_V2-7B
MobileVLM_V2-7B is a text generation model from mtgv. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
MobileVLM V2 is a family of significantly improved vision language models upon MobileVLM, which proves that a delicate orchestration of novel architectural design, an improved training scheme tailored for mobile VLMs,…
Downloads · 30 days
145
1% of all-time downloads
All-time downloads
15.2K
Public
Repo size
28.3 GB
Likes
5
Public
Click a slice to open those files.
.bin14.1 GB · 100%
From the Hugging Face model README
MobileVLM V2 is a family of significantly improved vision language models upon MobileVLM, which proves that a delicate orchestration of novel architectural design, an improved training scheme tailored for mobile VLMs, and rich high-quality dataset curation can substantially benefit VLMs’ performance. Specifically, MobileVLM V2 1.7B achieves better or on-par performance on standard VLM benchmarks compared with much larger VLMs at the 3B scale. Notably, MobileVLM_V2-3B model outperforms a large variety of VLMs at the 7B+ scale.
The MobileVLM_V2-7B was built on Vicuna-7B-v1.5 to facilitate the off-the-shelf deployment.
Inference examples can be found at Github.