Downloads · 30 days
46
79% of all-time downloads
ueuegio/ITA-OCR
ITA-OCR is a image-text-to-text model from ueuegio. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for gguf. The card lists the license as mit.
A fine-tune of GLM-OCR for Italian handwriting, packaged as GGUF for llama.cpp. These are the weights used by the ITA-OCR desktop application, which runs recognition entirely on the user's machine — no cloud OCR servi…
Downloads · 30 days
46
79% of all-time downloads
All-time downloads
58
Public
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.gguf1.2 GB · 100%
From the Hugging Face model README
A fine-tune of GLM-OCR for Italian handwriting, packaged as GGUF for llama.cpp. These are the weights used by the ITA-OCR desktop application, which runs recognition entirely on the user's machine — no cloud OCR service, no API key, no telemetry.
| File | Role | Size |
|---|---|---|
glm-ocr-ocr-ita-v8-native-3ep-clean-q8_0.gguf | fine-tuned language model | 686 MiB |
glm-ocr-base-mmproj-q8_0.gguf | vision projector (multimodal) | 462 MiB |
Both are required: the projector alone cannot transcribe, the model alone cannot see the page. Quantised to Q8_0.
95c6a7b7f318293c4a5276a5dab773d0ef55d49ad357e24f89f9d4a701a5be8a glm-ocr-ocr-ita-v8-native-3ep-clean-q8_0.gguf
2f83f69e7e5268474c8e257606ace1c1c28ff3546e1308a63fb15d3dc2ebf459 glm-ocr-base-mmproj-q8_0.gguf
With llama.cpp. The prompt is GLM-OCR's own, Text Recognition:, with no
system prompt and no extra instruction — the fine-tune stays directly
comparable with the base model.
llama-mtmd-cli -m glm-ocr-ocr-ita-v8-native-3ep-clean-q8_0.gguf \
--mmproj glm-ocr-base-mmproj-q8_0.gguf \
--image page.png -p "Text Recognition:"
Pages are rendered onto a fixed 960×1248 canvas at 150 DPI, as in training.
With the application: put both files in the models folder next to the
executable, or point OCR_ITA_MODELS at the folder holding them. Instructions
in docs/en/MODEL.md.
The app starts llama-server on loopback and applies a cascade of retries when
a decode ends up incomplete, so command-line output can differ from the app's
on the same page.
LoRA rank 8 on the base model, 3 epochs, on Italian handwritten pages with human reference transcriptions. The split is writer-disjoint: no writer present in training appears in evaluation.
The work is documented in the technical report Teaching a Vision Model When to Stop, which describes the termination collapse observed after fine-tuning, the two stop tokens that caused it and the inference cascade that compensates for it.
The dataset remains private and is not published in any form.
Base model against the shipped system. Every row states the set it was measured on: a percentage without its set is meaningless. The measurement noise floor is 0.3 points, so the margins below are well clear of it.
| Set | Pages | CER | WER |
|---|---|---|---|
| Development holdout, writer-disjoint | 164 | 30.88 → 25.71 (−16.7%) | 54.68 → 45.54 (−16.7%) |
| Sealed benchmark, readable part | 67 | 29.28 → 18.78 (−35.9%) | 56.74 → 38.99 (−31.3%) |
| Cohort | 136 | −8.8% | −20.4% |
| Subset labelled easy a priori | 61 | 12.06 → 11.28 (−6.5%) | 36.97 → 30.67 (−17.0%) |
The holdout uses capped CER, the only fair statistic where pages are lost; on the other sets neither model loses a page, so raw and capped coincide. The benchmark row excludes one writer at the edge of legibility whom the base model already reads at 52% capped CER — the aggregate over all 103 pages says more about that hand than about either model. On the easy subset the shipped system wins on 43 pages, loses on 17 and ties on 1.
Fine-tuning induced a termination collapse: the model stops emitting the stop token reliably and runs to the context ceiling. Inference-time mitigations, measured on a 48-page panel of affected pages:
| Configuration | Pages lost / 48 | Recovers | Loops |
|---|---|---|---|
| Greedy, no mitigation | 24 | n/a | 0 |
| Frequency 0.35 + presence 0.20 | 5 | 20 / 24 | 1 |
| Frequency 0.35 | 5 | 21 / 24 | 5 |
| Presence 0.20 alone | 22 | 4 / 24 | 0 |
| DRY sampling | 22 | 2 / 24 | 0 |
The shipped cascade uses the first configuration as its retry stage. It is cheaper than not having it: on the holdout, a full cascaded run takes less wall clock than a single plain greedy pass, because runaway pages are cut at around 550 tokens instead of reaching the 4,096 ceiling. Cascade behaviour: 140 pages resolved greedily, 19 on retry 1, 3 on retry 2, 2 left incomplete on the holdout; 71 / 21 / 3 / 8 on the benchmark.
Both official stop tokens must be honoured — this is what the
--override-kv argument in the application sets up.
Paired per-page difference in character error against the bf16 reference, 126-page cohort. Positive is worse. Each configuration is the language tower plus the vision projector.
| Configuration | Mean diff. vs bf16 | Pages identical to bf16 |
|---|---|---|
| Q4_K_M + bf16 | −0.02 | 2 / 126 |
| Q8_0 + bf16 | +0.08 | 49 / 126 |
| Q8_0 + Q8_0 (shipped) | +0.23 | 22 / 126 |
| bf16 + Q8_0 | +0.55 | 24 / 126 |
| Q5_K_M + bf16 | +0.65 | 10 / 126 |
The mean alone is misleading — the count of identical pages is the discriminator. Q8_0 is the only quantisation genuinely indistinguishable from bf16: median difference of zero, better and worse pages close to a coin toss, and 49 pages out of 126 character-identical. Q4_K_M looks fine on the mean yet leaves almost no page untouched (2 / 126) with roughly 60% of pages worse: a small but directional degradation. Against bf16, Q8_0 uses about 30% less video memory and decodes about 35% faster, which is why it ships for both towers.
Quality is unchanged on the CPU; latency is not. On the test machine a page goes from roughly 1.6 s on the GPU to roughly 14 s on the CPU, and 11–12 s of that is image encoding rather than generation: decoding nearly doubles from bf16 to Q4_K_M (63.8 → 124.2 tok/s) while the page still costs about fourteen seconds. Absolute timings depend on the machine — processor, GPU, drivers, page complexity — and should be read as a ratio. On the CPU, quantisation buys memory rather than latency: 3,410 MB resident at bf16 against 2,402 MB at Q8_0. The lever for latency would be the canvas resolution, not weight precision.
Inference cost is unchanged by fine-tuning itself: about 1.61 s per page for both base and fine-tuned model on the easy subset, ~297 generated tokens per page against a 1,542-token prompt that is almost entirely the image.
MIT, like the GLM-OCR base weights. The application code is MIT; third-party components keep their own licences, listed in THIRD-PARTY.md.
The base model is GLM-OCR by Z.ai. "ITA-OCR" is an independent name and implies no affiliation with Z.ai or the GLM project.
The application code, part of the documentation and the graphic assets were produced with the help of artificial-intelligence tools. That is not a guarantee of correctness or accuracy: transcriptions must be verified.
Fine-tuning di GLM-OCR per la scrittura a mano italiana, in formato GGUF per llama.cpp. Sono i pesi usati dall'applicazione desktop ITA-OCR, che esegue il riconoscimento interamente sul computer dell'utente: nessun servizio OCR cloud, nessuna chiave API, nessuna telemetria.
| File | Ruolo | Dimensione |
|---|---|---|
glm-ocr-ocr-ita-v8-native-3ep-clean-q8_0.gguf | modello linguistico adattato | 686 MiB |
glm-ocr-base-mmproj-q8_0.gguf | proiettore visivo (multimodale) | 462 MiB |
Servono entrambi: il proiettore da solo non trascrive, il modello da solo non vede la pagina. Quantizzazione Q8_0. I checksum SHA-256 sono quelli riportati sopra.
Il prompt è quello di GLM-OCR, Text Recognition:, senza prompt di sistema e
senza istruzioni aggiuntive. Le pagine vanno rese su tela fissa 960×1248 a
150 DPI, come in addestramento.
llama-mtmd-cli -m glm-ocr-ocr-ita-v8-native-3ep-clean-q8_0.gguf \
--mmproj glm-ocr-base-mmproj-q8_0.gguf \
--image pagina.png -p "Text Recognition:"
Con l'applicazione: metti i due file nella cartella models accanto
all'eseguibile, oppure indica la cartella con OCR_ITA_MODELS. Istruzioni in
docs/MODELLO.md.
LoRA rank 8 sul modello base, 3 epoche, su pagine manoscritte italiane con trascrizione umana di riferimento. Lo split è per scrivente: le persone presenti nell'addestramento non compaiono nella valutazione. Il lavoro è documentato nel report tecnico. Il dataset resta privato e non è pubblicato in nessuna forma.
La trascrizione può omettere, ripetere o inventare testo, soprattutto con scrittura difficile, impaginazione irregolare, formule e tabelle: va sempre confrontata con l'originale. Le misure interne non sono un benchmark indipendente — metodo e limiti in BENCHMARKS.md.
MIT, come i pesi base di GLM-OCR. I componenti di terze parti mantengono le proprie licenze, elencate in TERZE-PARTI.md. «ITA-OCR» è un nome indipendente e non indica affiliazione con Z.ai o con il progetto GLM.
Il codice dell'applicazione, parte della documentazione e gli asset grafici sono stati realizzati con il supporto di strumenti di intelligenza artificiale. Non costituisce garanzia di correttezza o accuratezza: le trascrizioni vanno verificate.