Downloads · 30 days
65
23% of all-time downloads
dipan004/DGPT
DGPT is a text generation model from dipan004. Use it when you need the model to write or continue text. The card lists the license as unknown.
A small, from-scratch base language model. Not an instruction-tuned assistant.
Downloads · 30 days
65
23% of all-time downloads
All-time downloads
286
Public
Repo size
143 MB
Likes
1
Public
Click a slice to open those files.
.npz143 MB · 100%
From the Hugging Face model README
A small, from-scratch base language model. Not an instruction-tuned assistant.
DGPT v1-base is a 13,049,856-parameter decoder-only Transformer trained from scratch (manual forward pass, manual backward pass, manual AdamW — no autograd, no PyTorch/JAX/TensorFlow) on the TinyStories dataset. It generates short, simple, TinyStories-style children's narratives and nothing more.
Do not expect: instruction following, multi-turn conversation, reasoning, factual world knowledge, or ChatGPT-comparable capability of any kind. This model was never trained or tuned for any of those.
bpe_6000.json)| param | value |
|---|---|
| vocab_size | 6000 |
| block_size | 256 |
| d_model | 384 |
| n_layer | 6 |
| n_head | 6 |
| head_dim | 64 |
| d_ff | 1536 |
| activation | GELU (tanh approx) |
| norm | Pre-LN |
| positions | learned |
| lm_head | tied to token embedding, no bias |
TinyStories
(TinyStoriesV2-GPT4-train.txt): 2,717,495 synthetically generated (GPT-3.5/
GPT-4) short stories using a deliberately small vocabulary
(Eldan & Li, 2023), tokenized to
371,525,259 tokens. Licensed by its authors under CDLA-Sharing-1.0 — this
model card does not redistribute the dataset itself.
From the training notebook (full-data run, step 5000):
No held-out benchmark suite (e.g. downstream NLP tasks) was run — TinyStories train/val loss and perplexity are the only reported metrics. Treat any numbers as approximate; see the training notebook's step-by-step log for the exact source values.
from src.generate import load_dgpt, generate_text
from src.tokenizer import BPETokenizer
tok = BPETokenizer("tokenizer/bpe_6000.json")
model, _ = load_dgpt("model.npz")
print(generate_text(model, tok, "Once upon a time", max_new_tokens=150))
License is marked unknown above deliberately. See this repository's main
README.md → "Licensing" for the full breakdown across code, weights,
tokenizer, and the TinyStories dataset — the weights and tokenizer do not
have an established license and none is invented here.