Downloads · 30 days
19
3% of all-time downloads
lewtun/quantized-distilbert-banking77
quantized-distilbert-banking77 is a text classification model from lewtun. Use it when you need a label for a piece of text. It is set up for transformers.
This model is a dynamically quantized version of optimum/distilbert-base-uncased-finetuned-banking77 on the banking77 dataset.
Downloads · 30 days
19
3% of all-time downloads
All-time downloads
562
Public
Repo size
550 MB
Likes
0
Public
Click a slice to open those files.
.onnx268 MB · 100%
From the Hugging Face model README
This model is a dynamically quantized version of optimum/distilbert-base-uncased-finetuned-banking77 on the banking77 dataset.
The model was created using the dynamic-quantization notebook from a workshop presented at MLOps World 2022.
It achieves the following results on the evaluation set:
Accuracy
The quantized model achieves 99.93% accuracy of the FP32 model
Latency
Payload sequence length: 128
Instance type: AWS c6i.xlarge
| latency | vanilla transformers | quantized optimum model | improvement |
|---|---|---|---|
| p95 | 63.24ms | 37.06ms | 1.71x |
| avg | 62.87ms | 37.93ms | 1.66x |
from optimum.onnxruntime import ORTModelForSequenceClassification
from transformers import pipeline, AutoTokenizer
model = ORTModelForSequenceClassification.from_pretrained("lewtun/quantized-distilbert-banking77")
tokenizer = AutoTokenizer.from_pretrained("lewtun/quantized-distilbert-banking77")
classifier = pipeline("text-classification", model=model, tokenizer=tokenizer)
classifier("What is the exchange rate like on this app?")