Downloads · 30 days
6
12% of all-time downloads
alpharomercoma/vqwen3-4b
vqwen3-4b is a image-text-to-text model from alpharomercoma. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
A ready-to-use vision-language model built by swapping Vicuna for Qwen3-4B in the LLaVA-1.5 recipe. Everything is pre-wired: drop in a LlavaForConditionalGeneration loader, pass an image + a prompt, get text out. No r…
Downloads · 30 days
6
12% of all-time downloads
All-time downloads
52
Public
Parameters
4.3B
8.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.7 GB · 100%
From the Hugging Face model README
A ready-to-use vision-language model built by swapping Vicuna for Qwen3-4B
in the LLaVA-1.5 recipe. Everything is pre-wired: drop in a LlavaForConditionalGeneration
loader, pass an image + a prompt, get text out. No rigging.
openai/clip-vit-large-patch14-336 (frozen)Qwen/Qwen3-4B with LoRA merged back into the weightsimport torch
from transformers import LlavaForConditionalGeneration, AutoProcessor
from PIL import Image
model_id = "alpharomercoma/vqwen3-4b"
model = LlavaForConditionalGeneration.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained(model_id)
image = Image.open("my_image.jpg").convert("RGB")
messages = [{
"role": "user",
"content": [
{"type": "image"},
{"type": "text", "text": "Describe this image in detail."},
],
}]
prompt = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = processor(text=prompt, images=image, return_tensors="pt").to(model.device)
inputs["pixel_values"] = inputs["pixel_values"].to(torch.bfloat16)
out = model.generate(**inputs, max_new_tokens=256, do_sample=False)
reply = processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(reply)
Two-stage reproduction of the LLaVA-1.5 recipe, both stages on a single H200 141 GB.
liuhaotian/LLaVA-Pretrain (558,128 BLIP caption pairs)liuhaotian/LLaVA-Instruct-150K (LLaVA-1.5 mix665k: COCO, GQA,
OCR-VQA, TextVQA, VisualGenome + ShareGPT text-only; 665,286 records after
filtering 12 dead image refs)[q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj]LengthGroupedSampler over mix665k for padding efficiencyThe stage-2 LoRA has been merged back into Qwen3's weights in this release,
so loading is a single .from_pretrained() call.
The model is the transformers-standard LlavaForConditionalGeneration with:
vision_config: CLIP ViT-L/14-336 (fixed)text_config: Qwen3-4B (with LoRA merged)image_seq_length: 576vision_feature_layer: −2 (penultimate hidden state)vision_feature_select_strategy: "default" (strips CLS)image_token_index: 151669 (the added <image> special token)projector_hidden_act: "gelu"Because these choices match the LLaVA class upstream, no custom code or
trust_remote_code=True is required.
LLaVA-Instruct-150K — inherits its distribution: English-heavy,
mostly natural-image QA, OCR-light. Don't expect SOTA on GUI / document /
chart tasks.Apache 2.0 for the projector weights and LoRA-merged Qwen3 delta. Base models retain their original licenses: OpenAI CLIP (MIT), Qwen3-4B (Apache 2.0).
Qwen/Qwen3-4B as the language backboneopenai/clip-vit-large-patch14-336 as the vision tower