Downloads · 30 days
141
4% of all-time downloads
prithivMLmods/Imgscope-OCR-2B-0527
Imgscope-OCR-2B-0527 is a image-text-to-text model from prithivMLmods. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
141
4% of all-time downloads
All-time downloads
3.8K
Public
Parameters
2.2B
4.4 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors4.4 GB · 100%
From the Hugging Face model README

The Imgscope-OCR-2B-0527 model is a fine-tuned version of Qwen2-VL-2B-Instruct, specifically optimized for messy handwriting recognition, document OCR, realistic handwritten OCR, and math problem solving with LaTeX formatting. This model is trained on custom datasets for document and handwriting OCR tasks and integrates a conversational approach with strong visual and textual understanding for multi-modal applications.
[!note] Colab Demo : https://huggingface.co/prithivMLmods/Imgscope-OCR-2B-0527/blob/main/Imgscope%20OCR%202B%200527%20Demo/Imgscope-OCR-2B-0527.ipynb
[!note] Video Understanding Demo : https://huggingface.co/prithivMLmods/Imgscope-OCR-2B-0527/blob/main/Imgscope-OCR-2B-05270-Video-Understanding/Imgscope-OCR-2B-0527-Video-Understanding.ipynb
SoTA Understanding of Images of Various Resolution & Ratio Imgscope-OCR-2B-0527 achieves state-of-the-art performance on visual understanding benchmarks such as MathVista, DocVQA, RealWorldQA, and MTVQA.
Enhanced Handwriting OCR Specifically optimized for recognizing and interpreting realistic and messy handwriting with high accuracy. Ideal for digitizing handwritten documents and notes.
Document OCR Fine-Tuning Fine-tuned with curated and realistic document OCR datasets, enabling accurate extraction of text from various structured and unstructured layouts.
Understanding Videos of 20+ Minutes Capable of processing long videos for video-based question answering, transcription, and content generation.
Device Control Agent Supports decision-making and control capabilities for integration with mobile devices, robots, and automation systems using visual-textual commands.
Multilingual OCR Support In addition to English and Chinese, the model supports OCR in multiple languages including European languages, Japanese, Korean, Arabic, and Vietnamese.
from transformers import Qwen2VLForConditionalGeneration, AutoTokenizer, AutoProcessor
from qwen_vl_utils import process_vision_info
# Load the model
model = Qwen2VLForConditionalGeneration.from_pretrained(
"prithivMLmods/Imgscope-OCR-2B-0527", # replace with updated model ID if available
torch_dtype="auto",
device_map="auto"
)
# Optional: Flash Attention for performance optimization
# model = Qwen2VLForConditionalGeneration.from_pretrained(
# "prithivMLmods/Imgscope-OCR-2B-0527",
# torch_dtype=torch.bfloat16,
# attn_implementation="flash_attention_2",
# device_map="auto",
# )
# Load processor
processor = AutoProcessor.from_pretrained("prithivMLmods/Imgscope-OCR-2B-0527")
messages = [
{
"role": "user",
"content": [
{
"type": "image",
"image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
},
{"type": "text", "text": "Recognize the handwriting in this image."},
],
}
]
# Prepare input
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
)
inputs = inputs.to("cuda")
# Generate output
generated_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text)


buffer = ""
for new_text in streamer:
buffer += new_text
buffer = buffer.replace("<|im_end|>", "")
yield buffer
Realistic Messy Handwriting OCR
Document OCR and Layout Understanding
Image and Text Multi-modal Reasoning
Math Problem Solving and LaTeX Rendering
Multi-turn Conversations
Video + Image + Text-to-Text Generation
Imgscope-OCR-2B-0527 is intended for: