Downloads · 30 days
9
24% of all-time downloads
kylewhite0314/multimodal-vision-ocr
multimodal-vision-ocr is a machine learning model from kylewhite0314. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This model serves as a complete AI Agent utilizing a RAG workflow: 1. Input: An image containing text. 2. Retrieval: OCR is used to pull text from the image (Document AI). 3. Augmented Generation: The extracted text i…
Downloads · 30 days
9
24% of all-time downloads
All-time downloads
38
Public
Parameters
224M
896 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors896 MB · 100%
From the Hugging Face model README
This model serves as a complete AI Agent utilizing a RAG workflow:
To use this, simply load the model and run the Vision + OCR pipeline to get structured data.