Downloads · 30 days
10
100% of all-time downloads
TLDRKKU/burapha
burapha is a image-text-to-text model from TLDRKKU. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as cc-by-nc-4.0.
LoRA adapter that fine-tunes Qwen/Qwen2.5-VL-7B-Instruct to read single handwritten Thai characters. Trained on the BURAPHA-TH dataset only (experiment e1 of THIRA — Thai Handwriting Intelligence Recognition & Analysis).
Downloads · 30 days
10
100% of all-time downloads
All-time downloads
10
Public
Repo size
202 MB
Likes
0
Public
Click a slice to open those files.
.safetensors190 MB · 91%
From the Hugging Face model README
LoRA adapter that fine-tunes Qwen/Qwen2.5-VL-7B-Instruct to read single handwritten Thai characters. Trained on the BURAPHA-TH dataset only (experiment e1 of THIRA — Thai Handwriting Intelligence Recognition & Analysis).
Qwen/Qwen2.5-VL-7B-Instruct, revision cc594898137f460bfe9f0759e9844b3ce807cfb5q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj"Read the handwritten Thai character in the image. Answer with the character only."Trained on BURAPHA-TH: A Multi-Purpose Character, Digit, and Syllable Handwriting Dataset (release 2021-03-06, ~70.8k balanced train samples used).
From the project's internal eval harness (runs/e1, see results.json in this repo for full detail).
| Test set | N | CER | char F1 | Exact match |
|---|---|---|---|---|
| burapha (in-domain) | 13,600 | 0.0388 | 0.9612 | 96.12% |
| hme (out-of-domain math) | 24,607 | 0.4324 | 0.7490 | 12.47% |
| thai_sentence (out-of-domain prose) | 2,032 | 0.5041 | 0.6856 | 1.87% |
| combined | 38,207 | 0.4277 | 0.7516 | 42.24% |
Everything below is measured, from the runs recorded in this project. It is reported here because a model card that only shows the in-domain number is misleading.
1. It learned Thai glyphs, not Thai text. 96.12% exact match on isolated BURAPHA characters collapses to 1.87% on connected Thai sentences (CER 0.504). BURAPHA is a character dataset; nothing in this adapter's training teaches word or sentence structure. Do not deploy it on Thai prose — use a sentence-trained adapter for that.
2. It is unusable on handwritten mathematics (CER 0.432, 12.47% exact). Expected for a
single-dataset adapter, and the reason the project trained the joint e3.
3. Joint training dominates it. e3, trained on BURAPHA + HME100K together, scores
0.0400 on BURAPHA — statistically indistinguishable from this adapter's 0.0388 — while also
scoring 0.0421 on maths, where this adapter scores 0.4324. There is no measured task on
which e1 beats e3 by a meaningful margin. e1 is superseded by e3.
4. Weight-space merging does not recover this adapter's Thai skill. Merging e1 with
the math adapter e2 at several ratios was tried and all of it was worse than joint
training:
| merge (e1–e2 weight) | burapha CER | hme CER | thai_sentence CER |
|---|---|---|---|
| 0.7 – 0.3 | 0.0480 | 0.1134 | 0.4386 |
| 0.5 – 0.5 | 0.1215 | 0.0592 | 0.4133 |
| 0.3 – 0.7 | 0.5144 | 0.0483 | 0.4797 |
e3 (joint training) | 0.0400 | 0.0421 | 0.6753 |
Every merge ratio trades one skill for the other; joint training gets both. A three-way
merge (e7) was also tried and lands at burapha 0.109–0.171 and hme 0.369–0.443 — worse
than e3 on both. Do not merge these adapters; retrain jointly instead.
5. Not selected for the production pipeline. The Learnly MVP uses a routed two-adapter
configuration — a sentence-trained Thai adapter for Thai-script regions, e3 for symbolic
and mathematical regions. e1 is retained as the single-dataset baseline that establishes
the BURAPHA ceiling, not as a deployed component.
CER here is corpus-level (total Levenshtein edits ÷ total reference characters) computed
after a normalisation pass that strips outer math delimiters and applies latex_normalize.
That normalisation collapses runs of whitespace but keeps single spaces, which is fine for
this adapter but not comparable across decoders that emit space-separated tokens — see
the note in the e3 card. Character precision/recall/F1 are an order-agnostic
bag-of-characters overlap, not an alignment.
import torch
from peft import PeftModel
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
base = Qwen2_5_VLForConditionalGeneration.from_pretrained(
"Qwen/Qwen2.5-VL-7B-Instruct",
revision="cc594898137f460bfe9f0759e9844b3ce807cfb5",
torch_dtype=torch.bfloat16,
)
model = PeftModel.from_pretrained(base, "TLDRKKU/burapha")
processor = AutoProcessor.from_pretrained("TLDRKKU/burapha")
TLDRKKU/HMk100 — HME100K-only adapter (e2)TLDRKKU/burapha-HMk100 — joint adapter (e3), the one to preferPart of THIRA — Thai Handwriting Intelligence Recognition & Analysis.