Downloads · 30 days
331
100% of all-time downloads
hdrrayan/multilingual-e5-small-gguf
multilingual-e5-small-gguf is a feature extraction model from hdrrayan. Use it when you need embeddings to search or compare text. It is set up for gguf. The card lists the license as mit.
A GGUF (F16) conversion of intfloat/multilingual-e5-small, for serving embeddings with llama.cpp / llama-server --embeddings.
Downloads · 30 days
331
100% of all-time downloads
All-time downloads
331
Public
Repo size
242 MB
Likes
0
Public
Click a slice to open those files.
.gguf242 MB · 100%
From the Hugging Face model README
A GGUF (F16) conversion of intfloat/multilingual-e5-small,
for serving embeddings with llama.cpp / llama-server --embeddings.
multilingual-e5-small was trained with asymmetric prefixes. You must prepend:
query: to search queriespassage: to documents being indexedRetrieval quality collapses without them.
llama-server -m multilingual-e5-small-f16.gguf --embeddings --pooling mean -c 512
Then POST to the OpenAI-compatible endpoint:
curl http://127.0.0.1:8080/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"model": "multilingual-e5-small-f16.gguf", "input": ["query: what is retrieval augmented generation"]}'
F16 embeddings track the reference sentence-transformers output closely
(cosine similarity min 0.9976, mean 0.9992 across a 20-sentence multilingual
sample spanning English, French and Chinese); the small residual is dominated
by SentencePiece tokenization differences, and top-1 ranking is preserved.
Converted from the original safetensors weights with convert_hf_to_gguf.py
from llama.cpp.