Downloads · 30 days
0
Word2vec/wikipedia2vec_plwiki_20180420_100d
wikipedia2vec_plwiki_20180420_100d is a machine learning model from Word2vec. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Pretrained Word2vec in Polish. For more information, see https://wikipedia2vec.github.io/wikipedia2vec/pretrained/.
Downloads · 30 days
0
Access
Public
Updated Jul 8, 2023
Repo size
2.5 GB
Likes
0
Public
Click a slice to open those files.
.txt1.2 GB · 100%
From the Hugging Face model README
Pretrained Word2vec in Polish. For more information, see https://wikipedia2vec.github.io/wikipedia2vec/pretrained/.
from gensim.models import KeyedVectors
from huggingface_hub import hf_hub_download
model = KeyedVectors.load_word2vec_format(hf_hub_download(repo_id="Word2vec/wikipedia2vec_plwiki_20180420_100d", filename="plwiki_20180420_100d.txt"))
model.most_similar("your_word")
@inproceedings{yamada2020wikipedia2vec,
title = "{W}ikipedia2{V}ec: An Efficient Toolkit for Learning and Visualizing the Embeddings of Words and Entities from {W}ikipedia",
author={Yamada, Ikuya and Asai, Akari and Sakuma, Jin and Shindo, Hiroyuki and Takeda, Hideaki and Takefuji, Yoshiyasu and Matsumoto, Yuji},
booktitle = {Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations},
year = {2020},
publisher = {Association for Computational Linguistics},
pages = {23--30}
}