Downloads · 30 days
16
27% of all-time downloads
muchad/mdeberta-id-30k
mdeberta-id-30k is a machine learning model from muchad. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
A vocabulary-pruned version of microsoft/mdeberta-v3-base with a 30k-token Indonesian vocabulary, designed for downstream Indonesian NLP tasks.
Downloads · 30 days
16
27% of all-time downloads
All-time downloads
59
Public
Parameters
108M
434 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors434 MB · 99%
From the Hugging Face model README
A vocabulary-pruned version of microsoft/mdeberta-v3-base with a 30k-token Indonesian vocabulary, designed for downstream Indonesian NLP tasks.
This model was developed using VocabPrune, a deterministic, language-aware, frequency-based vocabulary pruning method designed to reduce vocabulary-related model overhead while preserving the original Transformer architecture.
The model is a base checkpoint and should be fine-tuned for a specific downstream task.
| Property | Value |
|---|---|
| Base model | microsoft/mdeberta-v3-base |
| Vocabulary size | 30k tokens |
| Vocabulary | Indonesian |
| Language focus | Indonesian |
| Architecture | mDeBERTa-v3-base |
For the methodology, experimental setup, and detailed evaluation results, please refer to the published paper.
If you use this model or the VocabPrune methodology in your research, please cite:
@article{fuadi2026efficient,
author = {Fuadi, Mukhlish and Wibawa, Adhi Dharma and Sumpeno, Surya},
title = {Efficient Transformer Models via Language-Aware
Frequency-Based Vocabulary Pruning},
journal = {IEEE Access},
volume = {14},
pages = {50993--51006},
year = {2026},
doi = {10.1109/ACCESS.2026.3679735}
}