Downloads · 30 days
28
3% of all-time downloads
efederici/ipt-125m
ipt-125m is a text generation model from efederici. Use it when you need the model to write or continue text. It is set up for transformers.
IPT-125m is a decoder-style transformer pretrained from scratch on 4.36 billion tokens of Italian text from the OSCAR-2301 dataset.
Downloads · 30 days
28
3% of all-time downloads
All-time downloads
865
Public
Repo size
502 MB
Likes
0
Public
Click a slice to open those files.
.bin251 MB · 99%
From the Hugging Face model README
IPT-125m is a decoder-style transformer pretrained from scratch on 4.36 billion tokens of Italian text from the OSCAR-2301 dataset.
If you like this project, consider supporting me with a cup of coffee! 🤖✨🌞
This model is best used with the Hugging Face transformers library for training and finetuning.
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("efederici/ipt-125m", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("efederici/ipt-125m")
The architecture is a modification of a standard decoder-only transformer.
| Hyperparameter | Value |
|---|---|
| n_parameters | 125M |
| n_layers | 12 |
| n_heads | 12 |
| d_model | 768 |
| vocab size | 50432 |
| sequence length | 2048 |