Downloads · 30 days
37
6% of all-time downloads
prithivMLmods/DeepCaption-VLA-V2.0-7B
DeepCaption-VLA-V2.0-7B is a image-text-to-text model from prithivMLmods. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
37
6% of all-time downloads
All-time downloads
660
Public
Parameters
8.3B
16.6 GB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README

DeepCaption-VLA-V2.0-7B is an advanced fine-tuned version of Qwen2.5-VL-7B-Instruct, specialized for Image Captioning and Vision Language Attribution (VLA). This enhanced release focuses on generating precise, attribute-rich captions that capture visual properties, object attributes, and scene details across diverse image types and aspect ratios.
Version V2.0 introduces significant improvements in multilingual inference, delivering higher captioning quality and attribution accuracy in languages including Chinese (Zh), Thai (Th), and others.
model type: experimental
| Image 1 | Image 2 |
|---|---|
![]() | ![]() |
| Image 3 | Image 4 |
![]() | ![]() |
| Image 5 | Image 6 |
![]() | ![]() |
| Image 7 [zh] | Image 8 |
![]() | ![]() |
| Image 9 [zh] | Image 10 [thai] |
![]() | ![]() |
| Qwen2.5-VL-7B-Instruct | DeepCaption-VLA-V2.0-7B |
|---|---|
![]() | ![]() |
![]() | ![]() |
CAPTION_SYSTEM_PROMPT = """
You are an AI assistant that rigorously follows this response protocol:
1. For every input image, your primary task is to write a **precise caption**. The caption must capture the **essence of the image** in clear, concise, and contextually accurate language.
2. Along with the caption, provide a structured set of **attributes** that describe the visual elements. Attributes should include details such as objects, people, actions, colors, environment, mood, and other notable characteristics.
3. Always include a **class_name** field. This must represent the **core theme or main subject** of the image in a compact format.
- Use the syntax: `{class_name==write_the_core_theme}`
- Example: `{class_name==dog_playing}` or `{class_name==city_sunset}`
4. Maintain the following strict format in your output:
- **Caption:** <one-sentence description>
- **Attributes:** <comma-separated list of visual attributes>
- **{class_name==core_theme}**
5. Ensure captions are **precise, neutral, and descriptive**, avoiding unnecessary elaboration or subjective interpretation unless explicitly required.
6. Do not reference the rules or instructions in the output. Only return the formatted caption, attributes, and class_name.
"""
from transformers import Qwen2_5_VLForConditionalGeneration, AutoTokenizer, AutoProcessor
from qwen_vl_utils import process_vision_info
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
"prithivMLmods/DeepCaption-VLA-V2.0-7B", torch_dtype="auto", device_map="auto"
)
processor = AutoProcessor.from_pretrained("prithivMLmods/DeepCaption-VLA-V2.0-7B")
messages = [
{
"role": "user",
"content": [
{
"type": "image",
"image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
},
{"type": "text", "text": "Describe this image with detailed attributes and properties."},
],
}
]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
)
inputs = inputs.to("cuda")
generated_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text)