Downloads · 30 days
88
4% of all-time downloads
Ransaka/TrOCR-Sinhala
TrOCR-Sinhala is a image-to-text model from Ransaka. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as mit.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
88
4% of all-time downloads
All-time downloads
2.4K
Public
Parameters
315M
34 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
See training metrics tab for performance details.
This model is finetuned version of Microsoft TrOCR Printed
More information needed
More information needed
from PIL import Image
import requests
from io import BytesIO
from transformers import TrOCRProcessor, VisionEncoderDecoderModel, AutoTokenizer
image_url = "https://datasets-server.huggingface.co/assets/Ransaka/sinhala_synthetic_ocr/--/bf7c8a455b564cd73fe035031e19a5f39babb73b/--/default/train/0/image/image.jpg"
response = requests.get(image_url)
img = Image.open(BytesIO(response.content))
processor = TrOCRProcessor.from_pretrained('Ransaka/TrOCR-Sinhala')
model = VisionEncoderDecoderModel.from_pretrained('Ransaka/TrOCR-Sinhala')
model.to("cuda:0")
pixel_values = processor(img, return_tensors="pt").pixel_values.to('cuda:0')
generated_ids = model.generate(pixel_values,num_beams=2,early_stopping=True)
generated_text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
generated_text #දිවයිනට බලයට ඇති ආපදා තත්ත්වය හමුවේ සබරගමුව පළාතේ