Downloads · 30 days
62
20% of all-time downloads
alphaedge-ai/multilingual-e5-base-pms-16384
multilingual-e5-base-pms-16384 is a sentence similarity model from alphaedge-ai. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as mit.
This model is a 64.5% smaller version of intfloat/multilingual-e5-base optimized for 16384 language via vocabulary pruning.
Downloads · 30 days
62
20% of all-time downloads
All-time downloads
312
Public
Parameters
98.6M
395 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors395 MB · 100%
From the Hugging Face model README
This model is a 64.5% smaller version of intfloat/multilingual-e5-base optimized for 16384 language via vocabulary pruning.
Total vocabulary size: 16384 tokens (reduced from 250002)
Tokenizer type: Unigram
Training samples per language: 200000 texts
Dataset: Lumberjackk/fineweb-2-trimming
This pruned model should perform similarly to the original model for 16384 with a much smaller memory footprint. However, it may not perform well for other languages as tokens not commonly used in the selected languages were removed from the vocabulary.
You can use this model with the Transformers library:
from transformers import AutoModel, AutoTokenizer
model_name = "Lumberjackk/multilingual-e5-base-pms-16384"
model = AutoModel.from_pretrained(model_name)
tokenizer = AutoTokenizer.from_pretrained(model_name)