Downloads · 30 days
0
lang-uk/fasttext_uk_cbow
fasttext_uk_cbow is a feature extraction model from lang-uk. Use it when you need embeddings to search or compare text. It is set up for generic. The card lists the license as mit.
cbow.uk.300.bin is pre-trained word vectors for the Ukrainian language, trained with fastText on (yet unreleased) UberText2.0 dataset, collected and processed by the lang-uk. This model was trained using cbow in dimen…
Downloads · 30 days
0
Access
Public
Updated Mar 28, 2023
Repo size
8.9 GB
Likes
1
Public
Click a slice to open those files.
.bin8.9 GB · 100%
From the Hugging Face model README
cbow.uk.300.bin is pre-trained word vectors for the Ukrainian language, trained with fastText on (yet unreleased) UberText2.0 dataset, collected and processed by the lang-uk. This model was trained using cbow in dimension 300, with character n-grams range of 4-6, and 15 negative samples.
The dataset for Ukrainian word analogy is available here.
Extrinsic evaluations were performed on two sequence labeling tasks: NER and POS tagging. NER-UK dataset was released by the lang-uk, and Ukrainian (UD) corpus was developed by a non-profit organization Institute for Ukrainian.
Results:
Usage
import fasttext.util
ft = fasttext.load_model('cbow.uk.300.bin')
ft.get_word_vector('привіт')
@inproceedings{romanyshyn-etal-2023-learning,
title = "Learning Word Embeddings for {U}krainian: A Comparative Study of Fast{T}ext Hyperparameters",
author = "Romanyshyn, Nataliia and
Chaplynskyi, Dmytro and
Zakharov, Kyrylo",
booktitle = "Proceedings of the Second Ukrainian Natural Language Processing Workshop",
month = may,
year = "2023",
address = "Dubrovnik, Croatia",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.unlp-1.3",
pages = "20--31",
}
Copyright: Dmytro Chaplynskyi, lang-uk project, Nataliia Romanyshyn, Ukrainian Catholic University, 2022