Downloads · 30 days
4
4% of all-time downloads
zhaospei/Model_14
Model_14 is a machine learning model from zhaospei. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Mô hình BLIP (Bootstrapping Language–Image Pre‑training) sử dụng Vision Transformer (ViT) để tạo ra mô hình hiểu và mô tả hình ảnh một cách linh hoạt, bao gồm cả các tác vụ như image captioning, image–text retrieval v…
Downloads · 30 days
4
4% of all-time downloads
All-time downloads
95
Public
Repo size
3 GB
Likes
0
Public
Click a slice to open those files.
.h5990 MB · 50%
From the Hugging Face model README
Mô hình BLIP (Bootstrapping Language–Image Pre‑training) sử dụng Vision Transformer (ViT) để tạo ra mô hình hiểu và mô tả hình ảnh một cách linh hoạt, bao gồm cả các tác vụ như image captioning, image–text retrieval và visual question answering. Phiên bản base được fine‑tune trên tập dữ liệu COCO cho nhiệm vụ generate caption, hỗ trợ cả hai chế độ:
Sinh caption chất lượng cao theo ngữ cảnh hoặc không cần prompt. Chỉ cần ViT + Q‑Former + Text decoder (BLIP 0. Similarly BLIP‑2 use LLM) — hiệu quả mà vẫn mạnh mẽ. Chạy trên CPU hoặc GPU, hỗ trợ chế độ half‑precision (FP16) để tối ưu tốc độ.
Hình ảnh: RGB
Kích thước đầu vào: bất kỳ, vì BlipProcessor sẽ tự resize và crop center về 224×224
Prompt (tuỳ chọn): ví dụ "a photo of" cho image-to-text có định hướng
Caption ở dạng chuỗi văn bản (string), đã được decode qua tokenizer
Có thể lấy logits của từng token nếu cần
pip install torch torchvision transformers pillow
import requests
from PIL import Image
from transformers import BlipProcessor, BlipForConditionalGeneration
processor = BlipProcessor.from_pretrained("zhaospei/Model_14")
model = BlipForConditionalGeneration.from_pretrained("zhaospei/Model_14")
url = "https://storage.googleapis.com/sfr-vision-language-research/BLIP/demo.jpg"
raw_image = Image.open(requests.get(url, stream=True).raw).convert("RGB")
# -- Conditional captioning --
inputs = processor(raw_image, "a photography of", return_tensors="pt")
out = model.generate(**inputs)
print(processor.decode(out[0], skip_special_tokens=True))
# → "a photography of a woman and her dog"
# -- Unconditional captioning --
inputs = processor(raw_image, return_tensors="pt")
out = model.generate(**inputs)
print(processor.decode(out[0], skip_special_tokens=True))
# → "a woman sitting on the beach with her dog"
Tăng ~2.8 điểm CIDEr cho task captioning so với các baseline trước đó.
Mô hình cũng thể hiện khả năng zero-shot tốt trên video (inference có thể dùng chế độ freeze) .
Ứng dụng thực tế gồm: trợ năng dành cho người khiếm thị, caption sản phẩm/E‑commerce, social media metadata, v.v.