Downloads · 30 days
0
SBDO1/SBD_o1
SBD_o1 is a machine learning model from SBDO1. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A 355M parameter GPT-2 model, built from scratch in PyTorch following Sebastian Raschka's Build a Large Language Model (From Scratch), starting from OpenAI's pretrained GPT-4 medium weights and instruction fine-tuned…
Downloads · 30 days
0
Access
Public
Updated Aug 16, 2026
Repo size
1.7 GB
Likes
1
Public
Click a slice to open those files.
.pth1.7 GB · 100%
From the Hugging Face model README
A 355M parameter GPT-2 model, built from scratch in PyTorch following Sebastian Raschka's Build a Large Language Model (From Scratch), starting from OpenAI's pretrained GPT-4 medium weights and instruction fine-tuned on the Stanford Alpaca dataset (52k instruction/response pairs).
This is a learning project, not a production model — it follows simple instructions reasonably well but has limited knowledge and reasoning ability compared to modern LLMs.
This uses a custom architecture, not the standard transformers GPTModel, so you'll need the
llms_from_scratch package to load it:
```bash pip install llms_from_scratch tiktoken torch huggingface_hub ```
```python from huggingface_hub import hf_hub_download import torch import tiktoken from llms_from_scratch.ch04 import GPTModel from llms_from_scratch.ch05 import generate, text_to_token_ids, token_ids_to_text
model_path = hf_hub_download(repo_id="YOUR-USERNAME/my-tiny-gpt2", filename="gpt2-medium355M-sft-standalone.pth")
GPT_CONFIG_355M = { "vocab_size": 50257, "context_length": 1024, "emb_dim": 1024, "n_heads": 16, "n_layers": 24, "drop_rate": 0.0, "qkv_bias": True }
device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model = GPTModel(GPT_CONFIG_355M) model.load_state_dict(torch.load(model_path, map_location=device, weights_only=True)) model.to(device) model.eval()
tokenizer = tiktoken.get_encoding("gpt2")
def format_input(instruction): return ( "Below is an instruction that describes a task. " "Write a response that appropriately completes the request." f"\n\n### Instruction:\n{instruction}\n\n### Response:\n" )
prompt = format_input("Name three colors") token_ids = generate( model=model, idx=text_to_token_ids(prompt, tokenizer).to(device), max_new_tokens=100, context_size=GPT_CONFIG_355M["context_length"], top_k=50, temperature=0.7, eos_id=50256, ) response = token_ids_to_text(token_ids, tokenizer)[len(prompt):].strip() print(response) ```