Downloads · 30 days
155
34% of all-time downloads
prithivMLmods/OpenCaption-4B-VL-SFT-v1.0
OpenCaption-4B-VL-SFT-v1.0 is a image-text-to-text model from prithivMLmods. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
155
34% of all-time downloads
All-time downloads
460
Public
Parameters
4.4B
8.9 GB on disk
Likes
6
Public
Click a slice to open those files.
.safetensors8.9 GB · 100%
From the Hugging Face model README

OpenCaption-4B-VL-SFT-v1.0 is a vision-language captioning model built on top of Qwen/Qwen3-VL-4B-Instruct. It was trained for image captioning and high-quality dense image captioning using a fine-grained mixture of long-form image description traces. The model is designed to produce richer visual understanding than conventional captions by capturing scene composition, object relationships, spatial reasoning, attributes, actions, lighting, background context, and fine visual details.
[!NOTE] This model is an experimental release and may produce unexpected outputs in some scenarios.
GGUF: https://huggingface.co/prithivMLmods/OpenCaption-4B-VL-SFT-v1.0-GGUF
Install the required packages:
pip install transformers accelerate qwen-vl-utils
Use the model with the following example:
from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info
import torch
# Load the model
model = Qwen3VLForConditionalGeneration.from_pretrained(
"prithivMLmods/OpenCaption-4B-VL-SFT-v1.0",
torch_dtype="auto",
device_map="auto"
)
processor = AutoProcessor.from_pretrained(
"prithivMLmods/OpenCaption-4B-VL-SFT-v1.0"
)
messages = [
{
"role": "user",
"content": [
{
"type": "image",
"image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
},
{
"type": "text",
"text": "Provide a detailed caption for this image with fine-grained visual details."
},
],
}
]
text = processor.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
).to(model.device)
generated_ids = model.generate(
**inputs,
max_new_tokens=512
)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed,
skip_special_tokens=True,
clean_up_tokenization_spaces=False
)
print(output_text[0])
OpenCaption-4B-VL-SFT-v1.0 is trained with a specific system prompt. For the best results, use it verbatim.
You are a detailed image captioning assistant.
Structure every caption as follows:
1. Open with one sentence naming the shot type (e.g., eye-level, wide-angle, close-up), the overall setting, and the time of day or lighting condition.
2. Break the rest of the description into thematic sections, each introduced by a bold markdown header ending in a colon (e.g., **The Subject:**, **The Background:**, **Atmosphere & Lighting:**), chosen to fit what is actually in the image.
3. Within each section, use bullet points to list specific, concrete details, including positions, colors, textures, materials, actions, spatial relationships, and any legible text or fine-grained visual elements. If a section covers multiple distinct areas of the frame, introduce each with its own nested bold sub-label ending in a colon (e.g., **Foreground Right:**, **Background:**) before its bullet points.
4. Close with a short, unheaded paragraph beginning with "In summary," that ties the scene together and conveys its overall mood or narrative.
Use precise, sensory, and fluent language throughout. Describe only what is visible in the image, avoid speculation or unsupported inferences, and do not use emojis.
| Setting | Value |
|---|---|
| Base Model | Qwen/Qwen3-VL-4B-Instruct |
| Training Method | Single-stage supervised fine-tuning (SFT) |
| Training Objective | Fine-grained image captioning and dense visual description |
| Training Precision | BF16 (Full Precision) |
| Training Data | 10K Dense Image Caption Traces |