Downloads · 30 days
2.9K
4% of all-time downloads
allegro/plt5-base
plt5-base is a translation model from allegro. Use it when you need text moved from one language to another. It is set up for transformers. The card lists the license as cc-by-4.0.
plT5 models are T5-based language models trained on Polish corpora. The models were optimized for the original T5 denoising target.
Downloads · 30 days
2.9K
4% of all-time downloads
All-time downloads
75.7K
Public
Repo size
2.2 GB
Likes
8
Public
Click a slice to open those files.
.bin1.1 GB · 100%
From the Hugging Face model README
plT5 models are T5-based language models trained on Polish corpora. The models were optimized for the original T5 denoising target.
plT5 was trained on six different corpora available for Polish language:
| Corpus | Tokens | Documents |
|---|---|---|
| CCNet Middle | 3243M | 7.9M |
| CCNet Head | 2641M | 7.0M |
| National Corpus of Polish | 1357M | 3.9M |
| Open Subtitles | 1056M | 1.1M |
| Wikipedia | 260M | 1.4M |
| Wolne Lektury | 41M | 5.5k |
The training dataset was tokenized into subwords using a sentencepiece unigram model with vocabulary size of 50k tokens.
Example code:
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("allegro/plt5-base")
model = AutoModel.from_pretrained("allegro/plt5-base")
CC BY 4.0
If you use this model, please cite the following paper:
@article{chrabrowa2022evaluation,
title={Evaluation of Transfer Learning for Polish with a Text-to-Text Model},
author={Chrabrowa, Aleksandra and Dragan, {\L}ukasz and Grzegorczyk, Karol and Kajtoch, Dariusz and Koszowski, Miko{\l}aj and Mroczkowski, Robert and Rybak, Piotr},
journal={arXiv preprint arXiv:2205.08808},
year={2022}
}
The model was trained by Machine Learning Research Team at Allegro and Linguistic Engineering Group at Institute of Computer Science, Polish Academy of Sciences.
You can contact us at: <a href="mailto:[email protected]">[email protected]</a>