Downloads · 30 days
8
9% of all-time downloads
rjzevallos/quebert_qu_bpe
quebert_qu_bpe is a fill-mask model from rjzevallos. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
This repository contains a RoBERTa-base model pretrained for the Quechua language using the Masked Language Modeling (MLM) objective. The model builds upon the corpus introduced in the QuBERT project and provides a Ro…
Downloads · 30 days
8
9% of all-time downloads
All-time downloads
89
Public
Repo size
668 MB
Likes
0
Public
Click a slice to open those files.
.bin334 MB · 100%
From the Hugging Face model README
This repository contains a RoBERTa-base model pretrained for the Quechua language using the Masked Language Modeling (MLM) objective. The model builds upon the corpus introduced in the QuBERT project and provides a RoBERTa alternative for downstream NLP research on Quechua and other low-resource Indigenous languages.
The goal of this work is to provide an open, high-quality pretrained language model that can serve as a strong foundation for a wide range of Natural Language Processing (NLP) tasks in Quechua.
This model follows the RoBERTa-base architecture introduced by Liu et al. (2019) and was pretrained from scratch exclusively on Quechua text.
| Property | Value |
|---|---|
| Architecture | RoBERTa-base |
| Language | Quechua |
| Objective | Masked Language Modeling (MLM) |
| Tokenizer | Byte-Pair Encoding (BPE) |
| Framework | Hugging Face Transformers |
The model was pretrained on a curated monolingual Quechua corpus containing approximately 8 million words. The corpus was originally introduced in the QuBERT project and was compiled from multiple publicly available sources.
The training corpus includes text collected from:
Before training, the corpus was carefully cleaned, normalized, and deduplicated to improve data quality.
The corpus primarily represents Southern Quechua, while also including material from other Quechua varieties whenever available.
More details about the corpus construction are available in the original QuBERT paper.
The model was pretrained from scratch using the standard Masked Language Modeling (MLM) objective.
Unlike the original QuBERT model, which is based on the BERT architecture, this repository provides a RoBERTa-based language model trained on the same Quechua corpus.
Following the RoBERTa pretraining procedure, approximately 15% of the input tokens are selected for prediction.
The model learns contextual representations by reconstructing masked tokens from their surrounding context, enabling it to capture both syntactic and semantic information useful for downstream NLP applications.
This model can be fine-tuned for a variety of downstream NLP tasks, including:
Since the model is pretrained using MLM, it can also be used directly for masked token prediction.
from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("rjzevallos/quebert_qu_bpe")
model = AutoModelForMaskedLM.from_pretrained("rjzevallos/quebert_qu_bpe")
from transformers import pipeline
fill_mask = pipeline(
"fill-mask",
model="rjzevallos/quebert_qu_bpe"
)
fill_mask("Payqa <mask> rinqa.")
The model can be evaluated using intrinsic and downstream metrics such as:
Evaluation results can be added as they become available.
| Metric | Value |
|---|---|
| Validation Loss | - |
| Perplexity | - |
Although the model has been trained on the largest publicly available monolingual Quechua corpus, several limitations remain:
Like all language models, this model may reflect linguistic and cultural biases present in the training corpus.
Researchers and practitioners are encouraged to evaluate the model carefully before deploying it in downstream applications, particularly those involving sensitive or high-impact use cases.
If you use this model in your research, please cite the original QuBERT paper:
@inproceedings{zevallos-etal-2022-introducing,
title = {Introducing QuBERT: A Large Monolingual Corpus and BERT Model for Southern Quechua},
author = {Zevallos, Rodolfo and
Ortega, John and
Chen, William and
Castro, Richard and
Bel, Núria and
Yoshikawa, Cesar and
Venturas, Renzo and
Aradiel, Hilario and
Melgarejo, Nelsi},
booktitle = {Proceedings of the Third Workshop on Deep Learning for Low-Resource Natural Language Processing},
year = {2022},
publisher = {Association for Computational Linguistics},
url = {https://aclanthology.org/2022.deeplo-1.1/},
doi = {10.18653/v1/2022.deeplo-1.1}
}
This model builds upon the corpus introduced in the QuBERT project.
The original QuBERT work was partially funded by Project PID2019-104512GB-I00 from the Spanish Ministerio de Ciencia, Innovación y Universidades and the Agencia Estatal de Investigación.
We hope this model contributes to advancing Natural Language Processing research for Quechua and other low-resource Indigenous languages by providing an openly available pretrained RoBERTa model for the research community.