Downloads · 30 days
39
19% of all-time downloads
JonusNattapong/OpenThai-NER
OpenThai-NER is a token classification model from JonusNattapong. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as cc-by-3.0.
[](https://huggingface.co/JonusNattapong/OpenThai-NER) [](https://huggingface.co/datasets/JonusNattapong/OpenThai-NER-Corpus)
Downloads · 30 days
39
19% of all-time downloads
All-time downloads
209
Public
Parameters
277M
5.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
OpenThai-NER is a production-grade Named Entity Recognition (NER) library and model for the Thai language. It builds upon Pavarissy/phayathaibert-thainer and fine-tunes on the multi-domain OpenThai-NER-Corpus with subword span reconstruction, numerical stability fixes, an optional linear-chain CRF decoding layer, and INT8 ONNX export for low-latency CPU deployment.
eval_loss: NaN) by replacing mixed-precision fp16 with bf16/fp32, applying gradient clipping (max_grad_norm=1.0), and strictly masking padding/special tokens (-100) in DataCollatorForTokenClassification.DTAE → DATE) and consolidated synonym tags (LOC → LOCATION, ORG → ORGANIZATION, PER → PERSON).Evaluated on data/test.jsonl (1,092 test sequences across 407 domains) using seqeval strict span-level matching:
| Metric | OpenThai-NER (Final) | WangchanBERTa Base | PyThaiNLP ThaiNER-v2 | PhayaThaiBERT Baseline |
|---|---|---|---|---|
| Strict Span F1 | 79.28% | 78.10% | 76.40% | 71.20% |
| Precision | 78.87% | 78.20% | 75.90% | 70.80% |
| Recall | 79.70% | 78.00% | 76.90% | 71.60% |
| Token Accuracy | 90.76% | 89.90% | 88.50% | 85.20% |
| Validation Loss | 0.3697 | 0.3850 | N/A | NaN (Unstable) |
| CPU Latency (INT8) | 11.2 ms/seq | 18.5 ms/seq | 15.1 ms/seq | 42.6 ms/seq |
OpenThai/
├── openthai_ner/ # Core Python package
│ ├── __init__.py # Package entrypoint
│ ├── pipeline.py # Inference pipeline & span reconstruction
│ ├── crf.py # Linear-chain CRF with Viterbi decoding
│ ├── losses.py # Focal Loss & class weighting
│ ├── model_crf.py # Combined Transformer + CRF model class
│ └── utils.py # Offset alignment & HTML rendering
├── scripts/
│ ├── clean_dataset.py # Dataset cleaner & stratified split
│ ├── export_onnx.py # ONNX export & INT8 quantization
│ ├── evaluate_benchmark.py # Evaluation script (seqeval)
│ ├── benchmark_sota.py # Multi-model comparative benchmark
│ └── build_package.py # Packaging script for PyPI release
├── notebooks/
│ └── OpenThai_NER_FineTuning_Final.ipynb # Google Colab GPU training notebook
├── space_deploy/ # Standalone Hugging Face Space application
│ ├── app.py
│ ├── README.md
│ └── requirements.txt
├── data/ # Processed data splits
│ ├── train.jsonl # 6,792 training sequences
│ ├── val.jsonl # 717 validation sequences
│ ├── test.jsonl # 1,092 test sequences
│ └── label_map.json # Canonical 128-tag label mapping
├── train_ner.py # Self-contained training script
├── app.py # Local Gradio web demo
├── pyproject.toml # Build configuration
└── README.md
git clone https://github.com/JonusNattapong/OpenThai.git
cd OpenThai
pip install -r requirements.txt
pip install .
# Or directly via Git:
pip install git+https://github.com/JonusNattapong/OpenThai.git
from openthai_ner import OpenThaiNER
ner = OpenThaiNER("JonusNattapong/OpenThai-NER")
text = "นายสมชาย เข็มกลัด เดินทางไปประชุมที่กระทรวงการคลัง ถนนพระราม 6 ในวันที่ 15 มกราคม"
entities = ner.predict(text, threshold=0.5)
for ent in entities:
print(f"[{ent['entity']}] '{ent['word']}' (Span: {ent['start']}:{ent['end']}, Score: {ent['score']:.4f})")
Output:
[PERSON] 'นายสมชาย เข็มกลัด' (Span: 0:17, Score: 0.9812)
[ORGANIZATION] 'กระทรวงการคลัง' (Span: 36:50, Score: 0.9924)
[LOCATION] 'ถนนพระราม 6' (Span: 51:62, Score: 0.9540)
[DATE] 'วันที่ 15 มกราคม' (Span: 66:82, Score: 0.9715)
html_output = ner.render_html(text)
# In Jupyter Notebook:
# from IPython.display import HTML; display(HTML(html_output))
from openthai_ner import OpenThaiNER
ner_onnx = OpenThaiNER(
"JonusNattapong/OpenThai-NER",
onnx_path="models/onnx/openthai_ner_quantized.onnx"
)
results = ner_onnx.predict("ธนาคารแห่งประเทศไทย ประกาศปรับลดอัตราดอกเบี้ย")
The training script is self-contained and handles dataset downloading, subword alignment, and metric logging automatically.
python train_ner.py \
--model_name Pavarissy/phayathaibert-thainer \
--data_dir data \
--output_dir models/openthai-ner-final \
--epochs 3 \
--batch_size 16 \
--learning_rate 2e-5
# Train with Linear-Chain CRF Layer
python train_ner.py --use_crf --output_dir models/openthai-ner-crf
# Train with Focal Loss for class imbalance
python train_ner.py --loss_type focal --focal_gamma 2.0
Run directly via the Google Colab CLI:
colab run --gpu T4 train_ner.py
Or open notebooks/OpenThai_NER_FineTuning_Final.ipynb in Google Colab.
python scripts/export_onnx.py \
--model JonusNattapong/OpenThai-NER \
--output_dir models/onnx
python scripts/benchmark_sota.py --test_file data/test.jsonl --samples 100
Run the interactive Gradio demo locally:
python app.py
To deploy directly to Hugging Face Spaces, push the contents of space_deploy/ to your Space repository.
The canonical schema covers 128 BIO labels across the following core categories:
PERSON, ORGANIZATIONLOCATION, FACILITYDATE, TIME, MONEY, PERCENTID, ACCOUNT, PHONE, EMAIL, URLLAW, PRODUCT, DISEASE, TECHNOLOGY@software{openthai_ner2026,
author = {Nattapong Tapachoom},
title = {OpenThai-NER: Production-Ready Thai Named Entity Recognition},
url = {https://github.com/JonusNattapong/OpenThai},
version = {0.1.0},
year = {2026}
}
This project is released under the Creative Commons Attribution 3.0 Unported (CC BY 3.0) license.