Downloads · 30 days
50
100% of all-time downloads
DAIR-Group/ExpertHTR
ExpertHTR is a image-to-text model from DAIR-Group. Use it when you need a caption or text from an image. It is set up for experthtr.
<p align="center" <a href="https://github.com/DAIR-Group/ExpertHTR"<img src="https://img.shields.io/badge/💻%20Code-ExpertHTR-blue" alt="Code"</a <a href="https://huggingface.co/datasets/DAIR-Group/ExpertHTR-Dataset"<…
Downloads · 30 days
50
100% of all-time downloads
All-time downloads
50
Public
Repo size
2.3 GB
Likes
2
Public
Click a slice to open those files.
.pt2.2 GB · 99%
From the Hugging Face model README
Validation-selected page-level handwritten text recognition checkpoint from ExpertHTR. It is based on Qwen3.5-0.8B-Base and adds four routed experts plus one shared expert at layers 1, 5, 9, 13, 17 and 21.
The selected checkpoint is training step 1914:
| Metric | Value |
|---|---|
| Validation pages | 962 |
| Micro page CER | 16.439% |
| Micro page WER | 33.300% |
CER removes only the leading region marker ([Rk]:). <del> and <gap> are
kept as distinct OCR symbols.
The companion gated dataset is available at 🤗 ExpertHTR-Dataset. The current dataset card excludes HWDB/CASIA. This checkpoint was trained on the original seven-source experiment, including HWDB2.0, so the companion dataset is not an exact reproduction of its training data. Do not describe this checkpoint as Bentham-only or claim that the gated dataset alone produced these weights. Follow every upstream dataset term; this model card does not grant rights to any training data.
Clone the code, install dependencies, and follow the evaluation instructions:
git clone https://github.com/DAIR-Group/ExpertHTR
cd ExpertHTR
python -m pip install -e .
export HWVLM_DATA_ROOT=/path/to/dataset
export HWVLM_FINAL_TEST_CHECKPOINT=/path/to/this/model
experthtr final-test
The checkpoint uses ExpertHTR's custom full_model_state.pt format. Keep the
tokenizer, processor, checkpoint_manifest.json, and the repository version
together; do not load it as a generic Transformers checkpoint.
This release contains runtime files only: model state, config, tokenizer, processor, chat template, checkpoint manifest, metadata and metrics. Training state, optimizer state, predictions, email addresses, access tokens and local machine paths are excluded. Review the manifest before mirroring this model.
The code is Apache-2.0. Qwen3.5-0.8B-Base and the training data retain their upstream licenses. This checkpoint is intended for research on handwritten text recognition and may make transcription errors. Do not use it for high-stakes decisions without human review.