Downloads · 30 days
185
3% of all-time downloads
hezarai/fasttext-fa-300
fasttext-fa-300 is a feature extraction model from hezarai. Use it when you need embeddings to search or compare text. It is set up for hezar.
This is the original fasttext embedding model for Persian from here loaded and converted using Gensim and exported to Hezar compatible format. For more info, see here.
Downloads · 30 days
185
3% of all-time downloads
All-time downloads
5.5K
Public
Repo size
2.5 GB
Likes
0
Public
Click a slice to open those files.
.npy2.4 GB · 97%
From the Hugging Face model README
This is the original fasttext embedding model for Persian from here loaded and converted using Gensim and exported to Hezar compatible format. For more info, see here.
In order to use this model in Hezar you can simply use this piece of code:
pip install hezar
from hezar.embeddings import Embedding
fasttext = Embedding.load("hezarai/fasttext-fa-300")
# Get embedding vector
vector = fasttext("هزار")
# Find the word that doesn't match with the rest
doesnt_match = fasttext.doesnt_match(["خانه", "اتاق", "ماشین"])
# Find the top-n most similar words to the given word
most_similar = fasttext.most_similar("هزار", top_n=5)
# Find the cosine similarity value between two words
similarity = fasttext.similarity("مهندس", "دکتر")