Downloads · 30 days
0
yammdd/vietnamese-error-correction
vietnamese-error-correction is a machine learning model from yammdd. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as mit.
This model is a Fine-tuned version of vinai/bartpho-syllable using LoRA (Low-Rank Adaptation). It is specifically designed for Vietnamese Error Correction (VEC) tasks.
Downloads · 30 days
0
Access
Public
Updated Dec 26, 2025
Parameters
396M
1.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors792 MB · 99%
From the Hugging Face model README
This model is a Fine-tuned version of vinai/bartpho-syllable using LoRA (Low-Rank Adaptation). It is specifically designed for Vietnamese Error Correction (VEC) tasks.
Unlike simple diacritic restoration models, this model aims to correct:
The model was trained on a dataset of approximately 70,000 sentences across the training, validation, and test splits, which were automatically labeled using a large language model from crawled Vietnamese social media comments. Due to the nature of social media data, the dataset may contain noise or labeling imperfections; however, it is not intended to include any offensive content or to target any individual or organization.
vinai/bartpho-syllableThe model is designed for Vietnamese text error correction. It takes noisy Vietnamese text as input, including missing diacritics, spelling mistakes, and informal or teencode expressions, and produces grammatically correct and orthographically normalized Vietnamese text as output.
Example:
You can use this model with the transformers libraries.
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, pipeline
path = "yammdd/vietnamese-error-correction"
tokenizer = AutoTokenizer.from_pretrained(path)
model = AutoModelForSeq2SeqLM.from_pretrained(path)
pipe = pipeline("text2text-generation", model=model, tokenizer=tokenizer)
text = "hum ni a bùn wá bé iu ưi"
out = pipe(text, max_new_tokens=256)
print(out[0]["generated_text"])
# Output: hôm nay anh buồn quá bé yêu ơi
vinai/bartpho-syllableq_proj, v_proj, out_proj, fc1, fc2 (covering both attention and feed-forward layers).DataCollatorForSeq2Seq with label_pad_token_id=-100.Seq2SeqTrainer with predict_with_generate enabled for validation metrics.The model was evaluated on a held-out test set of 7,056 samples, covering a diverse range of Vietnamese sentence structures and lengths.
| Metric | Score | Note |
|---|---|---|
| BLEU | 86.34 | High linguistic and semantic fidelity |
| Word Accuracy | 93.28% | Robust word-level correction |
| Exact Match | 51.53% | Entire sentence perfectly restored |
| WER | 0.0838 | ~8.38% error rate per word |
| CER | 0.0360 | ~3.60% error rate per character |
Note: The Exact Match score reflects the inherent ambiguity in the Vietnamese language (e.g., "muon" could be "muốn", "mượn", or "muộn"), where multiple correct interpretations may exist without broader paragraph context.
The model's performance varies based on the complexity and length of the input:
| Category | Length (words) | Accuracy | Sample Count |
|---|---|---|---|
| Short | < 10 | 60.88% | 2,927 |
| Medium | 10 - 30 | 47.83% | 3,577 |
| Long | > 30 | 25.91% | 552 |
Analysis: The model performs exceptionally well on short to medium sentences. Accuracy declines on longer sequences (>30 words), likely due to the increased probability of cumulative errors and the 256-token limit.