Downloads · 30 days
0
Rivok/paddleocr-hebrew
paddleocr-hebrew is a image-to-text model from Rivok. Use it when you need a caption or text from an image. It is set up for onnxruntime. The card lists the license as apache-2.0.
Finetuned from PaddleOCR. Every model in this repository is a finetune of a PaddleOCR model (PaddleOCR v3.7.0, Apache-2.0) and is released under the same licence — see the per-model MODELCARD.md for each lineage. All…
Downloads · 30 days
0
Access
Public
Updated Aug 29, 2026
Repo size
387 MB
Likes
1
Public
Click a slice to open those files.
.onnx387 MB · 100%
From the Hugging Face model README
Finetuned from PaddleOCR. Every model in this repository is a finetune of a PaddleOCR model (PaddleOCR v3.7.0, Apache-2.0) and is released under the same licence — see the per-model
MODEL_CARD.mdfor each lineage.All reported CER figures are in-house evaluation on our own test sets: measured, but vendor-reported and not independently verified. The benchmark harness is not shipped; methodology is documented in the GitHub repo.
Training weights (
.pdparams) and training configs are not published. This is a release of inference-ready ONNX models, not a reproducible training pipeline.
Built by Rivok Labs — rivoklabs.com · [email protected] · code on GitHub
ONNX weights for Hebrew OCR (finetuned from PaddleOCR / SVTRv2). Runnable pipeline code, quickstart, benchmark, and docs are on GitHub: https://github.com/RivoksLab/paddleocr-hebrew.
All recognizers share one byte-identical 120-char charset
(charset_v2f.txt, md5 e17ce22e7b4ab8224a3dad9e4c85b6ae). Each folder has a
MODEL_CARD.md + md5sums.txt. Split-ONNX pairs (nrtr-encoder + nrtr-decstep)
are one logical model — the attention decode loop runs on the host.
Hebrew is RTL — output is LOGICAL Unicode order. Apply
python-bidiget_display()only when rendering, never before storing/scoring. This is the #1 way to get garbage out. See the GitHubdocs/charset.md.
| folder | role | files | size |
|---|---|---|---|
server-svtrv2/ | flagship server REC (SVTRv2, CTC + NRTR) | ctc.onnx, nrtr-encoder.onnx, nrtr-decstep.onnx | 77 + 72 + 27 MB |
light-svtrv2small/ | edge/CPU REC (NRTR-only KD student) | nrtr-encoder.onnx, nrtr-decstep.onnx | 28 + 27 MB |
server-v5/ | alt word-level server REC (PPHGNetV2-B4) | rec.onnx | 73 MB |
server-v6/ | alt word-level server REC (PPLCNetV4) | rec.onnx | 60 MB |
mobile-word/ | mobile word REC (PPLCNetV3 KD) | rec.onnx | 7.4 MB |
word-det/ | word detector (mobile DBNet) | det.onnx | 4.6 MB |
line-det/ | line detector (situational) | det.onnx | 4.6 MB |
pip install "git+https://github.com/RivoksLab/paddleocr-hebrew" huggingface_hub
hf download Rivok/paddleocr-hebrew \
--include "charset_v2f.txt" "word-det/*" "server-svtrv2/*" \
--local-dir hebrew-ocr-models
from ocr import HebrewOCR
ocr = HebrewOCR(models_dir="hebrew-ocr-models")
for line in ocr.read("page.png")["lines"]:
print(line["text"]) # logical order
Full tables, methodology, and the CTC-vs-attention finding: see GitHub.
Released as ONNX only (inference-ready, portable — CPU/CUDA/Jetson). Paddle training weights + configs are not published. To fine-tune on your own Hebrew data or collaborate, open an issue on the GitHub repo or reach out: [email protected].
Rivok Labs builds computational intelligence and automation tools, with a particular focus on Hebrew and other right-to-left languages that mainstream tooling handles badly.
We released these models because there was no open-source Hebrew OCR with a commercial-friendly licence that held up on real documents. If they are useful to you, or if they fail on your documents, we would like to hear about it.
Apache-2.0. Finetuned from PaddleOCR (Apache-2.0) — see GitHub NOTICE.