Downloads · 30 days
22
23% of all-time downloads
gvadhul/gpt2-2layer-1m
gpt2-2layer-1m is a text generation model from gvadhul. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A randomly-initialized, 2-layer GPT-2 model for functional / integration testing. It produces incoherent text — that is by design.
Downloads · 30 days
22
23% of all-time downloads
All-time downloads
96
Public
Parameters
2.2M
15.6 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.6 MB · 100%
From the Hugging Face model README
A randomly-initialized, 2-layer GPT-2 model for functional / integration testing. It produces incoherent text — that is by design.
| Hyperparameter | Value |
|---|---|
| Architecture | GPT-2 (decoder-only) |
| Layers | 2 |
| Hidden size | 256 |
| Attention heads | 4 |
| FFN inner size | 1 024 |
| Context length | 512 |
| Vocabulary size | 183 (byte-level BPE) |
| Total params | ~1.76 M |
| Non-embedding params | ~1.58 M |
from transformers import GPT2LMHeadModel, PreTrainedTokenizerFast
tokenizer = PreTrainedTokenizerFast.from_pretrained("gvadhul/gpt2-2layer-1m")
model = GPT2LMHeadModel.from_pretrained("gvadhul/gpt2-2layer-1m")
inputs = tokenizer("Hello world", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=20)
print(tokenizer.decode(out[0]))
Functional / unit testing of pipelines that need a tiny causal LM. Not suitable for any real NLP task — weights are random.