Downloads · 30 days
23
26% of all-time downloads
zeromodels/qwen2-vl-2b-instruct
qwen2-vl-2b-instruct is a image-text-to-text model from zeromodels. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for zeromodels. The card lists the license as apache-2.0.
[](https://github.com/IMvision12/ZeroModels) [](https://imvision12.github.io/ZeroModels/qwen2vl/) [](https://huggingface.co/collections/zeromodels/qwen2-vl-6a8eae2d09e489fe58f3910c)
Downloads · 30 days
23
26% of all-time downloads
All-time downloads
89
Public
Repo size
4.4 GB
Likes
0
Public
Click a slice to open those files.
.h54.4 GB · 100%
From the Hugging Face model README
Pure-Keras 3 conversion of Qwen/Qwen2-VL-2B-Instruct for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX. This is the 2B variant, served here as image + text -> text via Qwen2VLProcessor; weights are stored in bfloat16.
For model details, license, and usage terms, see the upstream model card.
Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution (arXiv:2409.12191) · HF Papers
Paper: Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities (arXiv:2308.12966) · HF Papers
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
from zeromodels.models.qwen2_vl import Qwen2VLTextGenerate, Qwen2VLProcessor
model = Qwen2VLTextGenerate.from_weights("zeromodels/qwen2-vl-2b-instruct")
processor = Qwen2VLProcessor.from_weights("zeromodels/qwen2-vl-2b-instruct")
inputs = processor(conversation=[
{"role": "user", "content": [{"type": "text", "text": "Hello, who are you?"}]}
])
outputs = model.generate(**inputs, max_new_tokens=64)
print(processor.decode(outputs[0]))
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
from PIL import Image
from zeromodels.models.qwen2_vl import Qwen2VLConditionalGenerate, Qwen2VLProcessor
model = Qwen2VLConditionalGenerate.from_weights("zeromodels/qwen2-vl-2b-instruct")
processor = Qwen2VLProcessor.from_weights("zeromodels/qwen2-vl-2b-instruct")
inputs = processor(conversation=[
{"role": "user", "content": [
{"type": "image", "image": Image.open("photo.jpg")},
{"type": "text", "text": "Describe this image in one sentence."},
]}
])
outputs = model.generate(**inputs, max_new_tokens=64)
print(processor.decode(outputs[0]))
Load any Qwen2-VL variant the same way with from_weights("zeromodels/<variant>"):
| Variant | Hub |
|---|---|
qwen2-vl-2b | zeromodels/qwen2-vl-2b |
qwen2-vl-2b-instruct | zeromodels/qwen2-vl-2b-instruct |
qwen2-vl-7b | zeromodels/qwen2-vl-7b |
qwen2-vl-7b-instruct | zeromodels/qwen2-vl-7b-instruct |
qwen2-vl-72b | zeromodels/qwen2-vl-72b |
qwen2-vl-72b-instruct | zeromodels/qwen2-vl-72b-instruct |
A huge thank you to the Qwen team at Alibaba for creating and releasing these models.
License: Apache 2.0.