Downloads · 30 days
27
100% of all-time downloads
ftmdeveloperz006/SaShi-1.0-Vision
SaShi-1.0-Vision is a visual question answering model from ftmdeveloperz006. Use it for the visual question answering task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as apache-2.0.
Downloads · 30 days
27
100% of all-time downloads
All-time downloads
27
Public
Parameters
860M
2.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.1 GB · 100%
From the Hugging Face model README
SaShi 1.0 Vision is an all-in-one unified Visual Intelligence & Art engine developed by ftmdeveloperz006 in Varanasi (Banaras), Uttar Pradesh, India (Two Brothers).
</div>SaShi 1.0 Vision consolidates three essential visual AI capabilities into one unified framework:
┌─────────────────────────┐
│ 👁️ SaShi 1.0 Vision │
└────────────┬────────────┘
┌───────────────────────────┼───────────────────────────┐
▼ ▼ ▼
┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐
│ 🎨 Text-to-Image │ │ 🖼️ Image-to-Image│ │ 🔍 Image-to-Text │
│ (T2I) │ │ (I2I) │ │ (I2T/Vision) │
│ 8K Ultra-HD Art │ │ Style & Upscale │ │ OCR & Visual QA │
└──────────────────┘ └──────────────────┘ └──────────────────┘
import torch
from PIL import Image
from diffusers import StableDiffusionPipeline, StableDiffusionImg2ImgPipeline, EulerAncestralDiscreteScheduler
from transformers import pipeline
device = "cuda" if torch.cuda.is_available() else "cpu"
# 1. Text-to-Image (T2I)
t2i_pipe = StableDiffusionPipeline.from_pretrained(
"ftmdeveloperz006/SaShi-1.0-Vision",
torch_dtype=torch.float16
).to(device)
def sashi_t2i(prompt):
return t2i_pipe(prompt + ", 8k uhd photorealistic", num_inference_steps=30).images[0]
# 2. Image-to-Image (I2I)
i2i_pipe = StableDiffusionImg2ImgPipeline.from_pretrained(
"ftmdeveloperz006/SaShi-1.0-Vision",
torch_dtype=torch.float16
).to(device)
def sashi_i2i(init_image, prompt, strength=0.75):
return i2i_pipe(prompt=prompt, image=init_image, strength=strength).images[0]
# 3. Image-to-Text (I2T / Vision Analysis)
vqa_pipe = pipeline("image-to-text", model="Salesforce/blip-image-captioning-large", device=0 if device=="cuda" else -1)
def sashi_i2t(image):
return vqa_pipe(image)[0]["generated_text"]