Downloads · 30 days
20
30% of all-time downloads
rpa020/D1
D1 is a text generation model from rpa020. Use it when you need the model to write or continue text. It is set up for transformers.
This model uses a causal language modeling approach during training. This approach modifies the way the model accesses and processes words that precede the current token in the input sequence. Unlike masked language m…
Downloads · 30 days
20
30% of all-time downloads
All-time downloads
67
Public
Parameters
6.3M
25 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors25 MB · 93%
From the Hugging Face model README
This model uses a causal language modeling approach during training. This approach modifies the way the model accesses and processes words that precede the current token in the input sequence. Unlike masked language modeling in a sequence-to-sequence model, casual language modeling focuses on predicting the single next token. It does this by conditioning on all previous tokens in the sequence, ensuring that the model only has access to prior tokens and not future ones.
When performing experiments with a decoder-only model, we selected BLOOM as the architecture.
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
This model was used in an experiment to determine which architecture is favourable in a low-resource-setting with Northern Sami.
The model is trained with the rpa020/SALT dataset. The formatted dataset is named the SAmi LLM Token (SALT) dataset and contains around 22 million tokens and approximately 2 million sentences. On average, each sentence consists of around ten tokens. The dataset has been designed to support the pretraining phase for foundational model development.
model = BloomForCausalLM.from_pretrained("rpa020/D1")
CE Loss: 7.66 Perplexity: 2130 SELF-BLEU: 0.40