Downloads · 30 days
21
44% of all-time downloads
enclavelabs/enclave-scribe-devanagari
enclave-scribe-devanagari is a image-text-to-text model from enclavelabs. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as mit.
A LoRA adapter for allenai/olmOCR-2-7B-1025 that adds Devanagari OCR capability. The base model can transcribe English documents well but cannot read Devanagari at all — this adapter fixes that.
Downloads · 30 days
21
44% of all-time downloads
All-time downloads
48
Public
Repo size
381 MB
Likes
0
Public
Click a slice to open those files.
.safetensors381 MB · 100%
From the Hugging Face model README
A LoRA adapter for allenai/olmOCR-2-7B-1025 that adds Devanagari OCR capability. The base model can transcribe English documents well but cannot read Devanagari at all — this adapter fixes that.
Built by Enclave Labs. MIT-licensed. Part of the EnclaveScribe project — a self-hostable, Indic-first document OCR system.
Given an image containing Devanagari text (Hindi, Marathi, Sanskrit, Nepali, Pali), returns the Unicode transcription. Best suited for word-level and short-line images. For full-page PDFs, use the EnclaveScribe agent pipeline which handles page rasterization and generation-config tuning.
Evaluated on a 500-sample held-out slice of himalaya-ai/devanagari_ocr_dataset (never seen during training):
| Metric | Base OLMoCR-2-7B | This adapter | Improvement |
|---|---|---|---|
| CER ↓ | 16.26 (1626%) | 0.174 (17.4%) | ~93× |
| WER ↓ | 22.64 | 0.468 | 48× |
| F1 ↑ | 0.013 | 0.534 | 41× |
| Latency ↓ | 1.75 s/sample | 0.84 s/sample | 2× faster |
Base is unable to read Devanagari — it hallucinates verbose English descriptions instead of transcribing, which is why it's also slower (more tokens generated).
import torch
from transformers import AutoProcessor, AutoModelForImageTextToText
from peft import PeftModel
from PIL import Image
BASE = "allenai/olmOCR-2-7B-1025"
ADAPTER = "enclavelabs/enclave-scribe-devanagari"
processor = AutoProcessor.from_pretrained(BASE)
model = AutoModelForImageTextToText.from_pretrained(
BASE, dtype=torch.bfloat16, device_map="auto"
)
model = PeftModel.from_pretrained(model, ADAPTER)
model.eval()
image = Image.open("hindi_word.png").convert("RGB")
messages = [{
"role": "user",
"content": [
{"type": "image", "image": image},
{"type": "text", "text": "Transcribe the Devanagari text from this image:"},
],
}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
with torch.no_grad():
out = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
repetition_penalty=1.1, # prevents generation loops on long/dense inputs
)
print(processor.batch_decode(
out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True
)[0].strip())
Important: use repetition_penalty=1.1 (or higher) in generation. Without it, the model can enter degenerate loops on long or ambiguous inputs. See "Limitations" below.
allenai/olmOCR-2-7B-1025 (Qwen2.5-VL-7B fine-tuned by Allen AI for English document OCR)repetition_penalty ≥ 1.1 in generation config, the model can enter loops that emit <tool_call> tokens (a Qwen2.5-VL quirk that survives skip_special_tokens=True) until max_new_tokens is exhausted. Always pass repetition_penalty=1.1.Iter-4 will add page-level Devanagari data (Nayana, IndicVisionBench) to close the page-level gap, and formally benchmark English regression. Follow github.com/Enclave-Labs-Inc/enclave-scribe for updates.
@software{enclavescribe_devanagari_2026,
title = {EnclaveScribe: Self-hostable Indic OCR — Devanagari adapter},
author = {Enclave Labs},
year = {2026},
url = {https://huggingface.co/enclavelabs/enclave-scribe-devanagari},
}
MIT — same as the base model and the EnclaveScribe repo.