Downloads · 30 days
83
49% of all-time downloads
bravend/bartpho-syllable-vi-spellcheck
bartpho-syllable-vi-spellcheck is a token classification model from bravend. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
Mô hình kiểm tra chính tả tiếng Việt: một encoder-decoder BARTpho-syllable với hai đầu ra dùng chung encoder — đầu sinh câu đã sửa (generation) và đầu phát hiện lỗi mức token (detection head). Đánh giá trên VSEC (9.34…
Downloads · 30 days
83
49% of all-time downloads
All-time downloads
168
Public
Parameters
397M
1.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.6 GB · 100%
From the Hugging Face model README
Mô hình kiểm tra chính tả tiếng Việt: một encoder-decoder BARTpho-syllable với hai đầu ra dùng chung encoder — đầu sinh câu đã sửa (generation) và đầu phát hiện lỗi mức token (detection head). Đánh giá trên VSEC (9.341 câu).
A Vietnamese spelling-error detector/corrector. One BARTpho-syllable encoder-decoder with two heads sharing the encoder: a seq2seq correction head and a token-level error-detection head. Evaluated on the VSEC benchmark (9,341 sentences).
| Model | DP | DR | DF | DF0.5 | CP | CR | CF | CF0.5 |
|---|---|---|---|---|---|---|---|---|
| N-gram (VSEC paper, Table 4) | 0.912 | 0.731 | 0.812 | – | 0.891 | 0.714 | 0.793 | – |
| VSEC subword Transformer (VSEC paper) | 0.931 | 0.813 | 0.868 | 0.905 | 0.874 | 0.763 | 0.815 | 0.849 |
| VinAI spelling correction system (Nguyen et al., IUI 2023) | 0.937 | 0.868 | 0.901 | 0.923 | 0.909 | 0.843 | 0.875 | 0.896 |
This model — generate, greedy | 0.966 | 0.855 | 0.907 | 0.942 | 0.911 | 0.806 | 0.855 | 0.887 |
This model — generate, beam 4 | 0.968 | 0.858 | 0.910 | 0.944 | 0.912 | 0.809 | 0.857 | 0.890 |
| This model — detection head only, threshold 0.4 | 0.935 | 0.833 | 0.881 | 0.913 | – | – | – | – |
Detection-head threshold sweep (encoder only, no decoding):
| threshold | 0.3 | 0.4 | 0.5 | 0.6 | 0.7 | 0.9 |
|---|---|---|---|---|---|---|
| DP | 0.915 | 0.935 | 0.951 | 0.962 | 0.972 | 0.987 |
| DR | 0.847 | 0.833 | 0.817 | 0.799 | 0.780 | 0.697 |
| DF | 0.880 | 0.881 | 0.879 | 0.873 | 0.865 | 0.817 |
import torch, difflib, unicodedata, re
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
repo = "bravend/bartpho-syllable-vi-spellcheck"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval().cuda()
# --- input normalization (the model was trained on text normalized this way) ---
_TONE = {"òa":"oà","óa":"oá","ỏa":"oả","õa":"oã","ọa":"oạ","òe":"oè","óe":"oé","ỏe":"oẻ",
"õe":"oẽ","ọe":"oẹ","ùy":"uỳ","úy":"uý","ủy":"uỷ","ũy":"uỹ","ụy":"uỵ"}
_TONE.update({k.capitalize(): v.capitalize() for k, v in list(_TONE.items())})
_TONE_RE = re.compile("|".join(map(re.escape, _TONE)))
def normalize(text):
text = unicodedata.normalize("NFC", text.strip())
text = text.translate(str.maketrans({"“":'"',"”":'"',"‘":"'","’":"'","–":"-","—":"-","…":"..."}))
return _TONE_RE.sub(lambda m: _TONE[m.group()], text)
text = normalize("Tôi đi hoc ở Học viện Công nghệ Bưu Chính Viễn Thông, chuyên nghành công nghệ thông tin.")
enc = tok(text, return_tensors="pt", truncation=True, max_length=1024).to("cuda")
with torch.no_grad():
# 1) correction
out = model.generate(**enc, num_beams=1, max_length=1024)
corrected = tok.decode(out[0], skip_special_tokens=True)
# 2) token-level error probability from the detection head (encoder only)
tok_probs = model.detect(enc.input_ids, enc.attention_mask)[0] # (seq_len,)
print(corrected) # Tôi đi học ở Học viện Công nghệ Bưu Chính Viễn Thông, chuyên ngành công nghệ thông tin.
def word_error_probs(words):
"""P(error) per word = max over the word's sub-word pieces (labels were assigned per word)."""
ids, spans = [], []
for w in words:
p = tok(w, add_special_tokens=False)["input_ids"]
spans.append((len(ids), len(ids) + len(p)))
ids.extend(p)
inp = torch.tensor([tok.build_inputs_with_special_tokens(ids)], device="cuda")
probs = model.detect(inp)[0]
off = 1 # one leading <s>
return [probs[off+a:off+b].max().item() if b > a else 0.0 for a, b in spans]
def gen_changed(src_words, pred_words):
flags = [False] * len(src_words)
for tag, i1, i2, _, _ in difflib.SequenceMatcher(a=src_words, b=pred_words).get_opcodes():
if tag in ("replace", "delete"):
for k in range(i1, i2): flags[k] = True
elif tag == "insert" and src_words:
flags[min(i1, len(src_words) - 1)] = True
return flags
src_words = text.split()
flags = [g or p >= 0.9 for g, p in zip(gen_changed(src_words, corrected.split()),
word_error_probs(src_words))]
print([w for w, f in zip(src_words, flags) if f]) # ['hoc', 'nghành']
Notes:
trust_remote_code=True is required: the repo ships the multi-task class
(modeling_spellcheck.py) and a tokenizer whose SentencePiece model is restricted to the
40K vocabulary (tokenization_spellcheck.py). Loading with the stock classes silently drops
the detection head and tokenizes rare strings differently from training.model.detect) is encoder-only and ~30× faster than generation; use it
for triage, and generation for the final decision.| Base model | vinai/bartpho-syllable (MBart, 12+12 layers, d=1024, ~400M params) |
| Extra parameters | detection head: Linear(1024,1024) → ReLU → Dropout → Linear(1024,2) on encoder outputs |
| Precision | float32 |
| Max length | 1024 sub-word pieces |
| Tokenizer | BartphoTokenizer with SentencePiece restricted to the 40,030-entry vocabulary |
Released under CC BY-NC 4.0: free to use, share and adapt for non-commercial purposes
with attribution. Any commercial or production use requires prior written permission from
the author — please contact the author at [email protected] or through this Hugging Face profile.
The base model vinai/bartpho-syllable is MIT-licensed.
@misc{bravend2026vispellcheck,
title = {Vietnamese Spell Checker: BARTpho-syllable multi-task correction + detection},
author = {Nguyen Duy Dung},
year = {2026},
url = {https://huggingface.co/bravend/bartpho-syllable-vi-spellcheck}
}
Base model:
@inproceedings{bartpho,
title = {{BARTpho: Pre-trained Sequence-to-Sequence Models for Vietnamese}},
author = {Nguyen Luong Tran and Duong Minh Le and Dat Quoc Nguyen},
booktitle = {Proceedings of INTERSPEECH},
year = {2022}
}
Benchmark dataset:
@inproceedings{do2021vsec,
title = {{VSEC: Transformer-based Model for Vietnamese Spelling Correction}},
author = {Do, Dinh-Truong and Nguyen, Ha Thanh and Bui, Thang Ngoc and Vo, Hieu Dinh},
booktitle = {PRICAI 2021: Trends in Artificial Intelligence},
year = {2021},
doi = {10.1007/978-3-030-89363-7_20}
}
Compared system:
@inproceedings{nguyen2023vietnamese,
title = {A Vietnamese Spelling Correction System},
author = {Nguyen, Thien Hai and Pham, Thinh and Le, Khoi Minh and Luong, Manh and Tran, Nguyen Luong and Man, Hieu and Nguyen, Dang Minh and Luu, Anh Tuan and Nguyen, Thien Huu and Bui, Hung and Phung, Dinh and Nguyen, Dat Quoc},
booktitle = {Companion Proceedings of the 28th International Conference on Intelligent User Interfaces (IUI '23)},
year = {2023},
doi = {10.1145/3581754.3584159}
}