Downloads · 30 days
0
Lynxpda/ru_vep
ru_vep is a translation model from Lynxpda. Use it when you need text moved from one language to another. It is set up for pytorch. The card lists the license as cc-by-sa-4.0.
A model of translation from Russian into Vepsian. In archive initial weights of the model trained with OpenNMT-py (Locomotive). The model has 457M parameters and is trained from scratch. Also presented are model weigh…
Downloads · 30 days
0
Access
Public
Updated May 30, 2024
Repo size
2.9 GB
Likes
1
Public
Click a slice to open those files.
.pt2 GB · 68%
From the Hugging Face model README
A model of translation from Russian into Vepsian. In archive initial weights of the model trained with OpenNMT-py (Locomotive). The model has 457M parameters and is trained from scratch. Also presented are model weights converted for Ctranslate2 and a package for installation and use with Argostranslate/Libretranslate.
dec_layers: 20
decoder_type: transformer
enc_layers: 20
encoder_type: transformer
heads: 8
hidden_size: 512
max_relative_positions: 20
model_dtype: fp16
pos_ffn_activation_fn: gated-gelu
position_encoding: false
share_decoder_embeddings: true
share_embeddings: true
share_vocab: true
src_vocab_size: 32000
tgt_vocab_size: 32000
transformer_ff: 6144
word_vec_size: 512
To fine-tune the Russian to Vepsian translation model using OpenNMT-py, you can modify and use the example configuration file from this repository - config.yml.
To use the Russian to Vepsian translation model with LibreTranslate and Argos Translate, follow these steps:
translate-ru_vep-1_0.argosmodel file.~/.local/share/argos-translate/packages%userprofile%\.local\share\argos-translate\packages@inproceedings{
title={Model for Russian - Veps translation.},
author={Maksim Migukin, Maksim Kuznetsov, Alexey Kutashov},
year={2024}
}
Data compiled by Opus.
Includes pretrained models from Stanza.
Data from Vepsian WiKi
Data from Lehme No 2051 // Open corpus of Vepsian and Karelian languages VepKar.
Data from OMAMEDIA
CCMatrix
http://opus.nlpl.eu/CCMatrix-v1.php
If you use the dataset or code, please cite (pdf) and, please, acknowledge OPUS (bib, pdf) as well for this release.
This corpus has been extracted from web crawls using the margin-based bitext mining techniques described here. The original distribution is available from http://data.statmt.org/cc-matrix/
OpenSubtitles
http://opus.nlpl.eu/OpenSubtitles-v2018.php
Please cite the following article if you use any part of the corpus in your own work: P. Lison and J. Tiedemann, 2016, OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. In Proceedings of the 10th International Conference on Language Resources and Evaluation (LREC 2016)