Downloads · 30 days
10
11% of all-time downloads
ManvithReddy/Omniverse-v1
Omniverse-v1 is a image-to-text model from ManvithReddy. Use it when you need a caption or text from an image. The card lists the license as apache-2.0.
Omniverse-v1 is a powerful Vision-Language Model (VLM) customized for Intelligent Document Processing (IDP). It is designed to understand complex document layouts, such as dense mark sheets, invoices, and unstructured…
Downloads · 30 days
10
11% of all-time downloads
All-time downloads
93
Public
Parameters
959M
2.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.9 GB · 89%
From the Hugging Face model README
Omniverse-v1 is a powerful Vision-Language Model (VLM) customized for Intelligent Document Processing (IDP). It is designed to understand complex document layouts, such as dense mark sheets, invoices, and unstructured PDFs, and natively extract beautifully structured Markdown.
This model is intended to be used with the Omniverse Python/Gradio pipeline.
from transformers import AutoModelForCausalLM, AutoProcessor
import torch
model_id = "ManvithReddy/Omniverse-v1"
# Load the processor and model
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16
).eval()
This model is derived from the PaddleOCR-VL architecture and customized for local, offline document extraction tasks under the Omniverse initiative.