Downloads · 30 days
21
1% of all-time downloads
Rodr16020/detr_handwriten_cursive_text_detection
detr_handwriten_cursive_text_detection is a object detection model from Rodr16020. Use it when you need objects located in an image. It is set up for transformers.
DETR allows to detect and generate the bounding boxes for handwritten and cursive text. This model was finetuned using the base model facebook/detr-resnet-101. The dataset used is still under development and possible…
Downloads · 30 days
21
1% of all-time downloads
All-time downloads
2.5K
Public
Parameters
60.7M
1.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.bin243 MB · 50%
From the Hugging Face model README
DETR allows to detect and generate the bounding boxes for handwritten and cursive text. This model was finetuned using the base model facebook/detr-resnet-101. The dataset used is still under development and possible released in future versions. Mainly, the model detects spanish text. Note: The default value of generated bounding boxes was used (num_queries: 100). Modifying this value when using the model could lead to unexpected behavior.
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
from transformers import DetrForObjectDetection, DetrImageProcessor
import torch
import cv2
import supervision as sv
# User defined constants
MODEL_CHECKPOINT = "Rodr16020/detr_handwriten_cursive_text_detection"
DEVICE = "cuda"
CONFIDENCE_TRESHOLD = 0.5 # This parameter allows to filter the generated boxes with a confidence score >= to this value
IOU_TRESHOLD = 0.5
TEST_IMAGE = "demo.jpeg" # Path to the test image
#Load the model and preprocessor
img_proc = DetrImageProcessor.from_pretrained(MODEL_CHECKPOINT)
detr_model = DetrForObjectDetection.from_pretrained(
pretrained_model_name_or_path=MODEL_CHECKPOINT,
ignore_mismatched_sizes=True
).to(DEVICE)
# Get the pixel values of the image (matrix)
image = cv2.imread(TEST_IMAGE)
# inference
with torch.no_grad():
# load image and predict
inputs = img_proc(images=image, return_tensors='pt').to(DEVICE)
outputs = detr_model(**inputs)
# post-process
# Resize the generated Bounding Boxes coords to the image original size
target_sizes = torch.tensor([image.shape[:2]]).to(DEVICE)
results = img_proc.post_process_object_detection(
outputs=outputs,
threshold=CONFIDENCE_TRESHOLD,
target_sizes=target_sizes
)[0]
# To extract all the generated bboxes
boxes = results["boxes"].tolist()[0]
# With supervision lib, use the generated coords to annotate the image and preview the boxes
box_annotator = sv.BoxAnnotator()
detections = sv.Detections.from_transformers(transformers_results=results).with_nms(threshold=0.1)
labels = [f"{confidence:.2f}" for _,_, confidence, class_id, _ in detections]
frame = box_annotator.annotate(scene=image.copy(), detections=detections, labels=labels)
sv.plot_image(frame, (16, 16))
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
Use the code below to get started with the model.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
[More Information Needed]
A simple and a tiny computer at CIC research lab.
When finetuning, the model and data used a total of
And possibly others
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]