Downloads · 30 days
6
7% of all-time downloads
Ransaka/TrOCR.si
TrOCR.si is a image-to-text model from Ransaka. Use it when you need a caption or text from an image. It is set up for transformers.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
6
7% of all-time downloads
All-time downloads
83
Public
Parameters
348M
23.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.4 GB · 100%
From the Hugging Face model README
See training metrics tab for performance details.
This model is finetuned version of Microsoft TrOCR Printed
More information needed
More information needed
from PIL import Image
import requests
from io import BytesIO
from transformers import TrOCRProcessor, VisionEncoderDecoderModel, AutoTokenizer
image_url = "https://datasets-server.huggingface.co/assets/Ransaka/sinhala_synthetic_ocr/--/bf7c8a455b564cd73fe035031e19a5f39babb73b/--/default/train/0/image/image.jpg"
response = requests.get(image_url)
img = Image.open(BytesIO(response.content))
processor = TrOCRProcessor.from_pretrained('Ransaka/TrOCR-Sinhala')
model = VisionEncoderDecoderModel.from_pretrained('Ransaka/TrOCR-Sinhala')
model.to("cuda:0")
pixel_values = processor(img, return_tensors="pt").pixel_values.to('cuda:0')
generated_ids = model.generate(pixel_values,num_beams=2,early_stopping=True)
generated_text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
generated_text #දිවයිනට බලයට ඇති ආපදා තත්ත්වය හමුවේ සබරගමුව පළාතේ