Downloads · 30 days
0
Dikshan1234/LLMPretrain
LLMPretrain is a text generation model from Dikshan1234. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
Every weight in this model was produced by this project. There is no frompretrained, no base checkpoint, and no adapter: the tokenizer, the architecture code, the data pipeline and the training loop are all bespoke, a…
Downloads · 30 days
0
Access
Public
Updated Oct 4, 2026
Repo size
401 GB
Likes
0
Public
Click a slice to open those files.
.safetensors2.8 GB · 100%
From the Hugging Face model README
Every weight in this model was produced by this project. There is no
from_pretrained, no base checkpoint, and no adapter: the tokenizer, the
architecture code, the data pipeline and the training loop are all bespoke, and
training started from a random initialisation.
Trained on a single RTX 5070 Ti (16 GB).
Written by hand — no transformers model classes were used to define it.
| Type | decoder-only, pre-norm transformer |
| Parameters | 351,650,304 |
| Layers | 18 |
| d_model | 1280 |
| Attention heads | 20 query / 4 KV (GQA) |
| Head dim | 64 |
| FFN | SwiGLU, hidden 3456 (≈ 8/3 × d_model) |
| Normalisation | RMSNorm (pre-norm) + QK-norm |
| Positions | RoPE (θ=10000) — no learned position embeddings |
| Context | 1024 |
| Vocab | 32,768 (own byte-level BPE, padded to a multiple of 128) |
| Biases | none, in any linear layer |
| Embeddings | tied between input and LM head |
Parameter split: 41,943,040 embedding (11.9%) / 309,705,984 transformer blocks.
| Tokens seen | 43,779,672,064 |
| Steps | 83,500 |
| Schedule | warmup-stable-decay (WSD) |
| Optimizer | fused AdamW, betas (0.9, 0.95), wd 0.1 on 2D params only |
| Precision | bf16 autocast, fp32 master weights |
| Hardware | 1 × RTX 5070 Ti 16 GB |
Bits per byte is the headline metric. This model's vocabulary (32,768) differs from GPT-2's (50,257), and cross-entropy per token is not comparable across tokenizers. BPB divides that out and is directly comparable.
| Metric | This model | GPT-2 124M |
|---|---|---|
| Parameters | 351,650,304 | 124,439,808 |
| Validation loss | 1.4474 | — |
| Bits per byte | 0.4905 | — |
| HellaSwag (0-shot, len-norm) | pending | 29.55% |
Validation loss is measured on a held-out slice of the training corpus and is not comparable to GPT-2's OpenWebText number. Bits per byte is, because it is tokenizer-independent.
| Source | HF id | License |
|---|---|---|
| fineweb-edu | HuggingFaceFW/fineweb-edu | ODC-By 1.0 |
Attribution: FineWeb-Edu (HuggingFaceFW/fineweb-edu), ODC-By 1.0
ODC-By requires attribution, which is given above. Whether pretraining-data licenses flow through to model weights is legally unsettled; permissively licensed sources were chosen deliberately as the most defensible posture. Model weights are released under cc-by-nc-4.0.
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Dikshan1234/LLMPretrain")
model = AutoModelForCausalLM.from_pretrained("Dikshan1234/LLMPretrain",
trust_remote_code=True)
inputs = tok("Once upon a time", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=100, do_sample=True,
temperature=0.8, top_p=0.95)
print(tok.decode(out[0]))
trust_remote_code=True is required because the architecture (RoPE + SwiGLU +
GQA + QK-norm) is not a stock transformers model. The two files it loads,
configuration_llmpretrain.py and modeling_llmpretrain.py, are in this repo
and are short enough to read.
===== step 83500 =====
--- [greedy] 'Once upon a time'
Once upon a time, the world was a very different place. The world was a very different place. The world was a very different place. The world was a very different place.
The world was a very different place. The world was a very different place.
The world was a very different place. The world was a very different place.
The world was a very different place. The world was a very different place.
The world was a very different place. The world was a very different place.
The world was a very different place. The world was a very different place.
The
--- [t=0.8] 'Once upon a time'
Once upon a time, I lived with a family of four – a girl and a boy, with two and a husband. They had six children but only one survived infancy; the last, who was only a year old. I was still trying to understand why this family was so poor and this child was so sick and so poor. I looked for a way to solve the problem.
In my search, I found out that there were many ways to solve the problem of poverty, and that the solution to the problem of poverty was not just to give food to the children, but also to educate them.
I
--- [greedy] 'The capital of France is'
The capital of France is divided into 100,000,000 shares of common stock, par value $0.0001 per share, and 100,000,000 shares of preferred stock, par value $0.0001 per share. As of December 31, 2019, there were approximately 1,000,000 shares of common stock issued and outstanding and no shares of preferred stock issued and outstanding.
The holders of the common stock are entitled to one vote per share on all matters to be voted upon by the stockholders of the Company. The common stock does not have cumulative voting rights.
The holders of the common stock
--- [t=0.8] 'The capital of France is'
The capital of France is divided into four provinces: the north contains Alsace and Lorraine, with the northern border forming the border between France and Germany; the south contains Poitou, Picardy, and Auvergne. In the east is the Côte d'Azur, the area between the Rhône and the Loire rivers.
France's population grew from 2.3 million in the year 1900 to 4.4 million in the year 2000, making it the second most populous country in Europe. It is the largest country in Europe and the European Union. The
--- [greedy] 'In 1969, humans first'
In 1969, humans first began to use the term “mammoth” to describe a large mammal, the mammoth. The term was first used in a scientific paper by the American paleontologist and paleontologist, John Ostrom, in which he described the mammoth as a “megafauna” (a large mammal) that “existed in the Pleistocene era, and was the largest land mammal ever to live.”
In the early 20th century, the term “mammoth” was used to describe a large mammal, the mammoth, that was the largest land mammal ever to live.
--- [t=0.8] 'In 1969, humans first'
In 1969, humans first encountered the world of the deep in the ocean floor. The first ever successful exploration of the deepest parts of the ocean went down with the recovery of the Challenger Deep on December 14, 1953. In July 1998, the last U.S. submarine, the Deepwater Horizon, sank with all of its crew.
The first deep-sea exploration was the Dutch ship Rotterdam, in 1929, which explored the Challenger Deep, the deepest part of the ocean. Other countries were involved in an initial survey of the region. In April 1931, the British government awarded rights to another British
--- [greedy] 'The three states of matter are'
The three states of matter are distinguished by their ability to conduct certain types of interactions with each other and with the environment. The states of matter are distinguished by their ability to conduct certain types of interactions with each other and with the environment.
The three states of matter are distinguished by their ability to conduct certain types of interactions with each other and with the environment.
The three states of matter are distinguishe
This is a base model trained on a modest token budget. It has had no instruction tuning, no RLHF and no safety alignment. It will produce confidently wrong text, and it reflects whatever biases exist in its training corpus. It is a research and educational artifact, not a product.