Downloads · 30 days
5
28% of all-time downloads
RicardoPoleo/DL_LLM_from_scratch_2
DL_LLM_from_scratch_2 is a machine learning model from RicardoPoleo. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This model was trained using the WikiText-103 dataset to generate text based on input prompts.
Downloads · 30 days
5
28% of all-time downloads
All-time downloads
18
Public
Repo size
5.3 GB
Likes
0
Public
Click a slice to open those files.
.data-00000-of-000015.3 GB · 100%
From the Hugging Face model README
This model was trained using the WikiText-103 dataset to generate text based on input prompts.
Dataset Used: WikiText-103
Source: Hugging Face Datasets
Dataset Details: The WikiText-103 dataset is a collection of over 100 million tokens extracted from the set of verified "Good" and "Featured" articles on Wikipedia. It is designed for language modeling and other text generation tasks.
To ensure high-quality input for training, the dataset underwent the following cleaning steps:
The neural network used for this model is based on a transformer architecture with the following specifications:
The model was trained on an L4 GPU with the following resources:
Training Configuration:
The training involved several experiments with different batch sizes and epochs. The final training loss was plotted to visualize the model's performance.
To use this model, you can load it from Hugging Face and generate text as follows:
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("RicardoPoleo/DL_LLM_from_scratch_2")
model = AutoModelForCausalLM.from_pretrained("RicardoPoleo/DL_LLM_from_scratch_2")
input_text = "Once upon a time"
inputs = tokenizer(input_text, return_tensors="pt")
outputs = model.generate(**inputs)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))