Downloads · 30 days
12
36% of all-time downloads
Phazel/fa-floret-wiki-vectors
fa-floret-wiki-vectors is a feature extraction model from Phazel. Use it when you need embeddings to search or compare text. It is set up for spacy. The card lists the license as cc-by-sa-4.0.
Persian floret static vector table: 200,000 rows x 300 dimensions, floret mode, minn=maxn=5, hashcount=2, trained with floret-torch on the full Persian Wikipedia dump for 5 epochs. Vectors only, no pipeline components…
Downloads · 30 days
12
36% of all-time downloads
All-time downloads
33
Public
Repo size
1.1 GB
Likes
1
Public
Click a slice to open those files.
.vec363 MB · 51%
From the Hugging Face model README
Persian floret static vector table: 200,000 rows x 300 dimensions, floret mode, minn=maxn=5, hash_count=2, trained with floret-torch on the full Persian Wikipedia dump for 5 epochs. Vectors only, no pipeline components. This is the table used by the fa_*_news_lg tier.
pip install https://huggingface.co/Phazel/fa-floret-wiki-vectors/resolve/main/fa_floret_wiki_200k-0.1.0-py3-none-any.whl
import spacy
nlp = spacy.load("fa_floret_wiki_200k")
print(nlp.vocab.vectors.shape)
print(nlp("میرود")[0].has_vector) # floret hashes subwords: always True
To train your own pipeline against it, unpack the wheel and point spaCy at the directory:
pip download --no-deps -d . https://huggingface.co/Phazel/fa-floret-wiki-vectors/resolve/main/fa_floret_wiki_200k-0.1.0-py3-none-any.whl
python -m spacy train config.cfg --paths.vectors ./fa_floret_wiki_200k
| Property | Value |
|---|---|
| Rows | 200,000 |
| Dimensions | 300 |
| Mode | floret |
minn / maxn | 5 / 5 |
hash_count | 2 |
| Corpus | full Persian Wikipedia (fawiki) dump, extracted with WikiExtractor, sentence-tokenized with spaCy blank("fa"): 8,428,449 sentences, 190,781,621 tokens |
| Tokens | 190,781,621 (measured) |
| Epochs | 5 |
| Trained with | floret-torch (GPU port of explosion/floret), cbow, on a Colab T4 |
Trained with:
python -m floret_torch.train --input fa.txt --output fa \
--model cbow --mode floret --dim 300 --minn 5 --maxn 5 \
--hashCount 2 --bucket 200000 --neg 10 --epoch 5 \
--lr 0.05 --minCount 20 --batch 8192
| Pipeline | Tier | LAS | ENTS_F | Wheel |
|---|---|---|---|---|
fa_dep_news_sm | sm | 85.15 | - | 7.9 MB |
fa_core_news_sm | sm | 85.15 | 71.87 | 13.5 MB |
fa_ent_news_sm | sm | - | 71.87 | 5.9 MB |
fa_dep_news_md | md | 86.34 | - | 62.6 MB |
fa_core_news_md | md | 86.34 | 74.71 | 68.5 MB |
fa_ent_news_md | md | - | 74.71 | 60.6 MB |
fa_dep_news_lg | lg | 86.60 | - | 229.3 MB |
fa_core_news_lg | lg | 86.60 | 75.94 | 235.2 MB |
fa_ent_news_lg | lg | - | 75.94 | 227.3 MB |
fa_core_news_trf | trf | 90.79 | 82.89 | 608.2 MB |
Tiers: sm hash embeddings, no vectors, md 50k x 300d floret vectors, lg 200k x 300d floret vectors, trf fine-tuned transformer, GPU recommended.
Standalone vector tables, usable as --paths.vectors for your own training:
| Vectors | Rows | Used by | Wheel |
|---|---|---|---|
fa_floret_400k | 50,000 | md tier | 54.5 MB |
fa_floret_full_wiki | 50,000 | no shipped pipeline | 54.9 MB |
fa_floret_wiki_200k (this one) | 200,000 | lg tier | 221.3 MB |
Training scripts, configs and evaluation: https://github.com/Fazel94/spacy-persian.
| Source | Author | Licence |
|---|---|---|
| fa_floret static vectors (200k rows x 300d, floret mode, full Persian Wikipedia dump, 5 epochs, trained with floret-torch) | Kiyarash Fazeli | CC BY-SA 4.0 |
| Persian Wikipedia dump (fawiki), the text the vector table is trained on | Wikipedia contributors | CC BY-SA 4.0 |