Downloads · 30 days
432
0% of all-time downloads
Recognai/zeroshot_selectra_medium
zeroshot_selectra_medium is a zero-shot classification model from Recognai. Use it when you need labels you did not train the model on. It is set up for transformers. The card lists the license as apache-2.0.
Zero-shot SELECTRA is a SELECTRA model fine-tuned on the Spanish portion of the XNLI dataset. You can use it with Hugging Face's Zero-shot pipeline to make zero-shot classifications.
Downloads · 30 days
432
0% of all-time downloads
All-time downloads
149K
Public
Parameters
40.8M
327 MB on disk
Likes
13
Public
Click a slice to open those files.
.bin163 MB · 50%
How the weights are stored.
F3240.8M · 100%
From the Hugging Face model README
Zero-shot SELECTRA is a SELECTRA model fine-tuned on the Spanish portion of the XNLI dataset. You can use it with Hugging Face's Zero-shot pipeline to make zero-shot classifications.
In comparison to our previous zero-shot classifier based on BETO, zero-shot SELECTRA is much more lightweight. As shown in the Metrics section, the small version (5 times fewer parameters) performs slightly worse, while the medium version (3 times fewer parameters) outperforms the BETO based zero-shot classifier.
from transformers import pipeline
classifier = pipeline("zero-shot-classification",
model="Recognai/zeroshot_selectra_medium")
classifier(
"El autor se perfila, a los 50 años de su muerte, como uno de los grandes de su siglo",
candidate_labels=["cultura", "sociedad", "economia", "salud", "deportes"],
hypothesis_template="Este ejemplo es {}."
)
"""Output
{'sequence': 'El autor se perfila, a los 50 años de su muerte, como uno de los grandes de su siglo',
'labels': ['sociedad', 'cultura', 'economia', 'salud', 'deportes'],
'scores': [0.6450043320655823,
0.16710571944713593,
0.08507631719112396,
0.0759836807847023,
0.026829993352293968]}
"""
The hypothesis_template parameter is important and should be in Spanish. In the widget on the right, this parameter is set to its default value: "This example is {}.", so different results are expected.
If you want to see this model in action, we have created a basic tutorial using Rubrix, a free and open-source tool to explore, annotate, and monitor data for NLP.
The tutorial shows you how to evaluate this classifier for news categorization in Spanish, and how it could be used to build a training set for training a supervised classifier (which might be useful if you want obtain more precise results or improve the model over time).
You can find the tutorial here.
See the video below showing the predictions within the annotation process (see that the predictions are almost correct for every example).
<video width="100%" controls><source src="https://github.com/recognai/rubrix-materials/raw/main/tutorials/videos/zeroshot_selectra_news_data_annotation.mp4" type="video/mp4"></video>
| Model | Params | XNLI (acc) | *MLSUM (acc) |
|---|---|---|---|
| zs BETO | 110M | 0.799 | 0.530 |
| zs SELECTRA medium | 41M | 0.807 | 0.589 |
| zs SELECTRA small | 22M | 0.795 | 0.446 |
*evaluated with zero-shot learning (ZSL)
Check out our training notebook for all the details.