Downloads · 30 days
66
17% of all-time downloads
polymer/dots.mocr-GGUF
dots.mocr-GGUF is a image-to-text model from polymer. Use it when you need a caption or text from an image. The card lists the license as apache-2.0.
DOTS.OCR is a vision-language document understanding model developed by RedNote HiLab, designed for OCR, document parsing, layout reasoning, and structured content understanding. This repository contains GGUF quantize…
Downloads · 30 days
66
17% of all-time downloads
All-time downloads
400
Public
Repo size
5.5 GB
Likes
0
Public
Click a slice to open those files.
.gguf5.5 GB · 100%
From the Hugging Face model README
DOTS.OCR is a vision-language document understanding model developed by RedNote HiLab, designed for OCR, document parsing, layout reasoning, and structured content understanding. This repository contains GGUF quantized variants of the model optimized for efficient local inference using llama.cpp.
Rather than functioning as a conventional OCR engine, DOTS.OCR is designed to understand complete document layouts while recognizing textual content. The model jointly interprets paragraphs, tables, mathematical expressions, figures, diagrams, forms, and hierarchical document structures, enabling accurate reconstruction of visually complex documents for downstream AI workflows.
The quantized formats significantly reduce memory requirements while preserving document understanding capability, making the model practical for local deployment, enterprise document intelligence, and large-scale document processing applications.
This repository provides various GGUF quantized versions of the DOTS.OCR model optimized for efficient local inference using llama.cpp.
DOTS.OCR is trained with an emphasis on document understanding, multimodal reasoning, layout analysis, and structured document interpretation across a diverse collection of visual documents.
Document Parsing Understands complete document layouts while preserving structural relationships.
Optical Character Recognition (OCR) Extracts textual information from scanned documents and visual content.
Layout Understanding Identifies logical organization across headings, paragraphs, tables, figures, and forms.
Structured Content Extraction Generates machine-readable representations while maintaining semantic document structure.
Complex Document Analysis Processes visually rich documents containing mixed layouts, equations, diagrams, and technical content.
Efficient Local Deployment Quantized variants enable practical document intelligence workloads on consumer hardware.
./llama-mtmd-cli \
-m SandlogicTechnologies/DOTS-OCR_IQ4_NL.gguf \
--mmproj SandlogicTechnologies/mmproj-dots.ocr-f16.gguf \
--image research_paper.png \
-p "Convert this document into structured Markdown while preserving headings, tables, equations, and figures."
Document Parsing Convert visually complex documents into structured machine-readable formats.
Enterprise Document Intelligence Automate understanding of reports, contracts, manuals, and technical documentation.
Layout-Aware Information Extraction Preserve document hierarchy and semantic relationships during extraction.
Knowledge Base Construction Prepare structured documents for enterprise search and Retrieval-Augmented Generation (RAG).
Technical Document Processing Parse scientific papers, books, forms, diagrams, and engineering documentation.
Research and Evaluation Benchmark multimodal document understanding and layout reasoning capabilities.
These quantized models are based on the original work by the RedNote HiLab development team.
Special thanks to:
llama.cpp open-source community for enabling efficient quantization and inference via the GGUF format.For questions, feedback, or support, please reach out at [email protected] or visit https://www.sandlogic.com/