Downloads · 30 days
22
18% of all-time downloads
sdkjfgndjfg/better
better is a text generation model from sdkjfgndjfg. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A custom GPT-style language model trained from scratch using PyTorch.
Downloads · 30 days
22
18% of all-time downloads
All-time downloads
122
Public
Parameters
59.1M
237 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors237 MB · 99%
From the Hugging Face model README
A custom GPT-style language model trained from scratch using PyTorch.
| Parameter | Value |
|---|---|
| Architecture | GPT (Decoder-only Transformer) |
Hidden size (d_model) | 768 |
| Attention heads | 8 |
| Transformer blocks | 1 |
| Max sequence length | 1024 |
| Vocabulary size | 32000 |
| Dropout | 0.2 |
Custom BPE tokenizer trained with the HuggingFace tokenizers library.
Special tokens: <|endoftext|> · <|pad|> · <|unk|>
You can easily load this model and tokenizer using the transformers library. Because the model uses a custom architecture, you must pass trust_remote_code=True.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained("sdkjfgndjfg/gpt-model-2-decoder-100000-tiny-stories-fp16", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("sdkjfgndjfg/gpt-model-2-decoder-100000-tiny-stories-fp16", trust_remote_code=True)
# Set up device
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
# Generate text
prompt = "The transformer is based on"
inputs = tokenizer(prompt, return_tensors="pt").to(device)
output_ids = model.generate(
**inputs,
max_new_tokens=50,
do_sample=True,
temperature=0.8,
pad_token_id=tokenizer.eos_token_id
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
Apache 2.0