Downloads · 30 days
783
58% of all-time downloads
jaweed123/TinyJLLM
TinyJLLM is a text generation model from jaweed123. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A decoder-only Transformer (~102.5M parameters) pretrained from random initialization on ~5 GB of FineWeb (sample-10BT), 3 epochs / 108,000 optimizer steps. Every component — tokenizer, data pipeline, model, training…
Downloads · 30 days
783
58% of all-time downloads
All-time downloads
1.3K
Public
Parameters
102M
794 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors410 MB · 51%
From the Hugging Face model README
A decoder-only Transformer (~102.5M parameters) pretrained from random
initialization on ~5 GB of FineWeb (sample-10BT), 3 epochs / 108,000
optimizer steps. Every component — tokenizer, data pipeline, model,
training loop, evaluation, export — is implemented from scratch in the
TinyLLM repository.
Final metrics: validation loss 3.50 (perplexity 33.1); the best checkpoint reached 3.48 / 32.6.
| Model | Stage | Link |
|---|---|---|
TinyJLLM | Base (this model) | jaweed123/TinyJLLM |
TinyJLLM-Instruct | Supervised fine-tuning | jaweed123/TinyJLLM-Instruct |
TinyJLLM-Instruct-DPO | DPO preference tuning | jaweed123/TinyJLLM-Instruct-DPO |
| Property | Value |
|---|---|
| Parameters | 102,450,432 (~102.5M) |
| Architecture | Llama-style decoder-only: RMSNorm, RoPE (half-split), SwiGLU, tied embeddings, no biases |
| Layers / heads / head_dim | 11 / 12 / 64 |
| Context length | 512 |
| Vocabulary | 32,000 (custom byte-level BPE, <pad> <unk> <bos> <eos> = 0–3) |
| Pretraining data | FineWeb sample-10BT, ~5.37 GB raw, 1.75M documents |
| Tokens seen | 3.54B (3 epochs) |
| Hardware | RTX 4060 8 GB, ~30K tok/s (torch.compile) |
| Precision | BF16 mixed precision, FP32 master weights |
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("jaweed123/TinyJLLM")
tokenizer = AutoTokenizer.from_pretrained("jaweed123/TinyJLLM")
prompt = "The future of AI is"
inputs = tokenizer(prompt, return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=50, temperature=0.8, top_k=50, top_p=0.95)
print(tokenizer.decode(out[0]))
This repo also ships GGUF files (F16, Q8_0, Q4_K_M) — load directly with
llama.cpp or llama-cpp-python.
Small scale ⇒ repetition in long generations, weak instruction following, limited and unreliable world knowledge. Not for factual or consequential use.
@misc{tinyjllm,
title = {TinyJLLM: A 100M-Parameter Small Language Model Built From Scratch},
author = {Jaweed, Abdul},
year = {2026},
url = {https://github.com/Abdul-Jaweed/TinyLLM}
}
FineWeb (HuggingFaceFW), Hugging Face tokenizers / datasets, PyTorch,
and llama.cpp. Built and measured with the
TinyLLM pipeline.