Downloads · 30 days
15
2% of all-time downloads
uvegesistvan/huBERTPlain
huBERTPlain is a text classification model from uvegesistvan. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
Cased fine-tuned BERT model for Hungarian, trained on a dataset provided by National Tax and Customs Administration - Hungary (NAV): Public Accessibilty Programme.
Downloads · 30 days
15
2% of all-time downloads
All-time downloads
648
Public
Parameters
111M
1.8 GB on disk
Likes
1
Public
Click a slice to open those files.
.bin885 MB · 40%
How the weights are stored.
F32111M · 100%
From the Hugging Face model README
Cased fine-tuned BERT model for Hungarian, trained on a dataset provided by National Tax and Customs Administration - Hungary (NAV): Public Accessibilty Programme.
The model can be used as any other (cased) BERT model. It has been tested recognizing "accessible" and "original" sentences, where:
Fine-tuned version of the original huBERT model (SZTAKI-HLT/hubert-base-cc), trained on information materials provided by NAV linguistic experts.
| Class | Precision | Recall | F-Score |
|---|---|---|---|
| Accessible / Label_0 | 0.71 | 0.79 | 0.75 |
| Original / Label_1 | 0.76 | 0.67 | 0.71 |
| accuracy | 0.73 | ||
| macro avg | 0.74 | 0.73 | 0.73 |
| weighted avg | 0.74 | 0.73 | 0.73 |
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("uvegesistvan/huBERTPlain")
model = AutoModelForSequenceClassification.from_pretrained("uvegesistvan/huBERTPlain")
If you use the model, please cite the following dissertation (to be submitted for workshop discussion):
Bibtex:
@PhDThesis{ Uveges:2024,
author = {{"U}veges, Istv{\'a}n},
title = {K{\"o}z{\'e}rthet{\"o} és automatiz{\'a}ci{\'o} - k{\'i}s{\'e}rletek a jog, term{\'e}szetesnyelv-feldolgoz{\'a}s {\'e}s informatika hat{\'a}r{\'a}n.},
year = {2024},
school = {Szegedi Tudom{\'a}nyegyetem}
}