Downloads · 30 days
0
temporary0-0name/orator
orator is a machine learning model from temporary0-0name. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This model, designed and pretrained from scratch, was developed without utilizing the Hugging Face library.
Downloads · 30 days
0
Access
Public
Updated Aug 8, 2024
Repo size
382 MB
Likes
0
Public
Click a slice to open those files.
.pt382 MB · 100%
From the Hugging Face model README
This model, designed and pretrained from scratch, was developed without utilizing the Hugging Face library.
256 (Maximum sequence length)50257 (Includes 50,000 BPE merges, 256 byte-level tokens, and 1 special token)887680.00060.00006 (10% of max_lr)71552000524288 (Number of tokens per batch)128256Decayed parameters typically include weights from the model's various layers (like the transformer blocks), which are subject to weight decay during optimization. This technique helps in regularizing the model, potentially reducing overfitting by penalizing large weights.
Non-decayed parameters generally involve biases and layer normalization parameters. These parameters are excluded from weight decay as applying decay can adversely affect the training process by destabilizing the learning dynamics.
The calculated total number of parameters includes both decayed and non-decayed tensors, summing up to over 95 million parameters.
For the training of this model, a significant subset of the HuggingFaceFW/fineweb-edu dataset was utilized. Specifically, the model was pretrained on 3 billion tokens selected from the "Sample 10B" segment of the dataset. This dataset provides a rich corpus compiled from educational and academic web sources, making it an excellent foundation for developing language models with a strong grasp of academic and formal text.
The dataset is hosted and maintained on Hugging Face's dataset repository. More detailed information and access to the dataset can be found through its dedicated page: HuggingFaceFW/fineweb-edu Sample 10B
The evaluation of our model, "orator," on the HellaSwag dataset demonstrates significant progress in understanding context-based predictions. Below, we detail the performance through loss and accuracy graphs, accompanied by specific metrics.


2.8834713.19890.3054For tokenization, this model uses:
tokenizer = tiktoken.get_encoding("gpt2")
Below is a Python example on how to load the model and generate text:
import torch
from torch.nn import functional as F
from gpt_class import GPTConfig, GPT
import tiktoken
# Set up the device
device = "cuda" if torch.cuda.is_available() else "cpu"
# Load the model
state_dict = torch.load('model_51999.pt', map_location=device)
config = state_dict['config']
model = GPT(config)
model.load_state_dict(state_dict['model'])
model.to(device)
model.eval()
# Seed for reproducibility
torch.manual_seed(42)
torch.cuda.manual_seed_all(42)
# Tokenizer
tokenizer = tiktoken.get_encoding("gpt2")
def Generate(model, tokenizer, example, num_return_sequences, max_length):
model.eval()
tokens = tokenizer.encode(example)
tokens = torch.tensor(tokens, dtype=torch.long).unsqueeze(0).repeat(num_return_sequences, 1)
tokens = tokens.to(device)
sample_rng = torch.Generator(device=device)
xgen = tokens
while xgen.size(1) < max_length:
with torch.no_grad():
with torch.autocast(device_type=device):
logits, _ = model(xgen)
logits = logits[:, -1, :]
probs = F.softmax(logits, dim=-1)
topk_probs, topk_indices = torch.topk(probs, 50, dim=-1)
ix = torch.multinomial(topk_probs, 1, generator=sample_rng)
xcol = torch.gather(topk_indices, -1, ix)
xgen = torch.cat((xgen, xcol), dim=1)
for i in range(num_return_sequences):
tokens = xgen[i, :max_length].tolist()
decoded = tokenizer.decode(tokens)
print(f"Sample {i+1}: {decoded}")
# Example usage
Generate(model, tokenizer, example="As we entered the forest we saw", num_return_sequences=4, max_length=32)
Sample 1: As we entered the forest we saw huge white pine fells at the tops of the high plateaus (the great peaks) and trees standing at ground level.
Sample 2: As we entered the forest we saw a few trees that were too large. We realized they were not going to be very big. There was one tree that was
Sample 3: As we entered the forest we saw a group of small, wood-dwelling bees who had managed to escape a predator. A farmer was holding a handful
Sample 4: As we entered the forest we saw giant, blue-eyed, spotted beetles on the ground, a grayling beetle in my lawn next to the pond, an
The original code for the Orator model (for architecture, Pre-Training, Evaluating, Generating) can be accessed through the GitHub repository.
GitHub Repository: my-temporary-name/my_gpt2