Downloads · 30 days
14
2% of all-time downloads
prosa-text/indobert-nusa
indobert-nusa is a fill-mask model from prosa-text. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as cc-by-sa-4.0.
This repository contains a language adaptation and fine-tuning of the Indobenchmark IndoBERT language model for three specific languages: Balinese, Buginese, and Minangkabau. The adaptation was performed using nusa-tr…
Downloads · 30 days
14
2% of all-time downloads
All-time downloads
577
Public
Repo size
2.7 GB
Likes
0
Public
Click a slice to open those files.
.bin1.3 GB · 100%
From the Hugging Face model README
This repository contains a language adaptation and fine-tuning of the Indobenchmark IndoBERT language model for three specific languages: Balinese, Buginese, and Minangkabau. The adaptation was performed using nusa-translation dataset.
We tested the model after it was fine-tuned for topic classification using nusa-dialogue dataset.
| Language | indobert-large-p2 (F1) | indobert-nusa (F1) |
|---|---|---|
| Balinese | 82.37 | 84.23 |
| Buginese | 80.53 | 82.03 |
| Minangkabau | 84.49 | 86.30 |
We also tested the model after it was fine-tuned for language identification using nusaX dataset.
| Model | F1-score |
|---|---|
| indobert-large-p2 | 98.21 |
| indober-nusa | 98.45 |
The following hyperparameters were used during training:
The dataset is released under the terms of CC-BY-SA 4.0. By using this model, you are also bound to the respective Terms of Use and License of the dataset. For commercial use in small businesses and startups, please contact us ([email protected]) for permission to use the datasets by informing company profile and propose of usage.
This research work is funded and supported by The Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ) GmbH and FAIR Forward - Artificial Intelligence for all. We thank Direktorat Jenderal Pendidikan Tinggi, Riset, dan Teknologi Kementerian Pendidikan, Kebudayaan, Riset, dan Teknologi (Ditjen DIKTI) for providing the computing resources for this project.
If you have any question please contact our support team at [email protected].