Downloads · 30 days
13
4% of all-time downloads
WeightWatcher/albert-large-v2-mrpc
albert-large-v2-mrpc is a text classification model from WeightWatcher. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
This model was finetuned on the GLUE/mrpc task, based on the pretrained albert-large-v2 model. Hyperparameters were (largely) taken from the following publication, with some minor exceptions.
Downloads · 30 days
13
4% of all-time downloads
All-time downloads
347
Public
Repo size
142 MB
Likes
0
Public
Click a slice to open those files.
.bin70.8 MB · 96%
From the Hugging Face model README
This model was finetuned on the GLUE/mrpc task, based on the pretrained albert-large-v2 model. Hyperparameters were (largely) taken from the following publication, with some minor exceptions.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations https://arxiv.org/abs/1909.11942
Text classification, research and development.
Not intended for production use. See https://huggingface.co/albert-large-v2
See https://huggingface.co/albert-large-v2
See https://huggingface.co/albert-large-v2
Use the code below to get started with the model.
from transformers import AlbertForSequenceClassification
model = AlbertForSequenceClassification.from_pretrained("WeightWatcher/albert-large-v2-mrpc")
See https://huggingface.co/datasets/glue#mrpc
MRPC is a classification task, and a part of the GLUE benchmark.
Adam optimization was used on the pretrained ALBERT model at https://huggingface.co/albert-large-v2.
A checkpoint from MNLI was NOT used, differing from footnote 4 in,
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations https://arxiv.org/abs/1909.11942
Training hyperparameters, (Learning Rate, Batch Size, ALBERT dropout rate, Classifier Dropout Rate, Warmup Steps, Training Steps,) were taken from Table A.4 in,
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations https://arxiv.org/abs/1909.11942
Max sequence length (MSL) was set to 128, differing from the above.
F1 score is used to evaluate model performance.
See https://huggingface.co/datasets/glue#mrpc
F1 score
Training F1 score: 0.9963621665319321
Evaluation F1 score: 0.9176882661996497
The model was finetuned on a single user workstation with a single GPU. CO2 impact is expected to be minimal.