Downloads · 30 days
12
10% of all-time downloads
sekarkrishna/mpnet-int8
mpnet-int8 is a feature extraction model from sekarkrishna. Use it when you need embeddings to search or compare text. It is set up for onnxruntime. The card lists the license as apache-2.0.
ONNX INT8 quantized version of sentence-transformers/all-mpnet-base-v2 for efficient general-purpose sentence embeddings.
Downloads · 30 days
12
10% of all-time downloads
All-time downloads
122
Public
Repo size
111 MB
Likes
0
Public
Click a slice to open those files.
.onnx111 MB · 99%
From the Hugging Face model README
ONNX INT8 quantized version of sentence-transformers/all-mpnet-base-v2 for efficient general-purpose sentence embeddings.
| Property | Value |
|---|---|
| Base Model | sentence-transformers/all-mpnet-base-v2 |
| Format | ONNX |
| Quantization | INT8 (dynamic quantization) |
| Embedding Dimension | 768 |
| Quantized by | JustEmbed |
This is a quantized ONNX export of all-mpnet-base-v2, one of the best general-purpose sentence embedding models from the sentence-transformers library. It maps sentences and paragraphs to a 768-dimensional dense vector space. The INT8 quantization reduces model size and improves inference speed while maintaining high accuracy.
model_quantized.onnx — INT8 quantized ONNX modeltokenizer.json — Fast tokenizervocab.txt — Vocabulary fileconfig.json — Model configurationfrom justembed import Embedder
embedder = Embedder("mpnet-int8")
vectors = embedder.embed(["This is a sentence", "This is another sentence"])
import onnxruntime as ort
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained(".")
session = ort.InferenceSession("model_quantized.onnx")
inputs = tokenizer("This is a sentence", return_tensors="np")
outputs = session.run(None, dict(inputs))
This model is a derivative work of sentence-transformers/all-mpnet-base-v2.
The original model is licensed under Apache License 2.0. This quantized version is distributed under the same license. See the LICENSE file for the full text.
@inproceedings{song2020mpnet,
title={MPNet: Masked and Permuted Pre-training for Language Understanding},
author={Song, Kaitao and Tan, Xu and Qin, Tao and Lu, Jianfeng and Liu, Tie-Yan},
booktitle={NeurIPS},
year={2020}
}