Downloads · 30 days
243
100% of all-time downloads
Dibachain/Diba-Vision
Diba-Vision is a image-text-to-text model from Dibachain. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
243
100% of all-time downloads
All-time downloads
243
Public
Parameters
4.7B
9.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.3 GB · 100%
How the weights are stored.
BF164.7B · 100%
From the Hugging Face model README
A Persian-first vision-language model by Dibachain مدل بینایی-زبانی فارسیمحور، ساختهی دیباچین
🌐 dibachain.ir · 🤖 Agent · Chat demo (GPU) · Diba-Base · Diba-Embed
</div>Diba-Vision sees images and answers in Persian and English. It joins the vision understanding of a strong multimodal encoder with the Persian language ability of the Diba family, so it can look at a picture, a document, or a screenshot and talk about it fluently in Persian — where most open vision models are weak.
| Type | Vision‑language model (image + text → text) |
| Parameters | ~4B |
| Languages | Persian‑first, plus English |
| License | Apache 2.0 |
from transformers import AutoModelForImageTextToText, AutoProcessor
from PIL import Image
import torch
model = AutoModelForImageTextToText.from_pretrained(
"Dibachain/Diba-Vision", trust_remote_code=True, dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained("Dibachain/Diba-Vision", trust_remote_code=True)
messages = [{"role": "user", "content": [
{"type": "image", "image": "path/or/url/to/image.jpg"},
{"type": "text", "text": "این تصویر را به فارسی توضیح بده."},
]}]
inputs = processor.apply_chat_template(messages, add_generation_prompt=True,
tokenize=True, return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(processor.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])
Diba-Vision ships with the Diba model definition, so pass
trust_remote_code=Truewhen loading.
Diba-Vision is for understanding images and answering about them, not for generating images. Quality is strongest on everyday photos, documents, and screens; very small text, low‑quality scans, or highly specialized diagrams may be misread. It reflects biases present in its training data. For text‑only chat and code use Diba-Base; for semantic search use Diba-Embed.
دیبا-ویژن عکس را میبیند و به فارسی و انگلیسی پاسخ میدهد. این مدل، درک تصویری یک رمزگذار چندحالتهی قوی را با توانایی زبان فارسیِ خانوادهی دیبا ترکیب میکند؛ پس میتواند به یک عکس، سند یا اسکرینشات نگاه کند و روان دربارهاش فارسی حرف بزند، جایی که بیشتر مدلهای بینایی متنباز ضعیفاند.
| نوع | مدل بینایی-زبانی (تصویر + متن ← متن) |
| تعداد پارامتر | حدود ۴ میلیارد |
| زبانها | فارسیمحور، بههمراه انگلیسی |
| مجوز | Apache 2.0 |
از همان کد بخش انگلیسی استفاده کنید. هنگام بارگذاری، trust_remote_code=True را بدهید و برای پرسشهای جستوجو تصویر و متن را با هم بفرستید.
دیبا-ویژن برای درک تصویر و پاسخ دربارهی آن است، نه برای تولید تصویر. بهترین کیفیت روی عکسهای روزمره، اسناد و صفحههاست؛ متنهای بسیار ریز، اسکنهای بیکیفیت یا نمودارهای خیلی تخصصی ممکن است اشتباه خوانده شوند. مدل سوگیریهای دادهی خود را بازتاب میدهد. برای گفتگوی متنی و کد از Diba-Base و برای جستوجوی معنایی از Diba-Embed استفاده کنید.
خانوادهی دیبا · The Diba family — Diba-Base · Diba-Embed · Diba-Vision · Diba-Code · Diba-TTS · Diba-STT · Diba-Image · Diba-ImageEdit
</div>