Downloads · 30 days
21
3% of all-time downloads
rohit-upadhya/SmoLLM-109M-Instruct
SmoLLM-109M-Instruct is a text generation model from rohit-upadhya. Use it when you need the model to write or continue text. The card lists the license as cc-by-sa-3.0.
A 109M-parameter, Llama-style language model built entirely from scratch — custom BPE tokenizer, RoPE, RMSNorm, SwiGLU, and multi-head causal attention — pretrained on FineWeb-Edu and instruction-tuned on databricks-d…
Downloads · 30 days
21
3% of all-time downloads
All-time downloads
751
Public
Repo size
438 MB
Likes
0
Public
Click a slice to open those files.
.bin438 MB · 99%
From the Hugging Face model README
A 109M-parameter, Llama-style language model built entirely from scratch — custom BPE tokenizer, RoPE, RMSNorm, SwiGLU, and multi-head causal attention — pretrained on FineWeb-Edu and instruction-tuned on databricks-dolly-15k.
This is the instruction-following variant. For the raw base model, see SmoLLM-109M-base.
This model uses a custom chat template. Format inputs exactly like this (note the trailing space after [ASSISTANT]):
[SYSTEM] You are a helpful bot [/SYSTEM]
[USER] your question here [/USER]
[ASSISTANT]
The model was trained to emit [EOS] at the end of each response, so it terminates on its own for focused questions.
| Metric | Value |
|---|---|
| Perplexity (WikiText-2 test) | 80.27 |
| Base model perplexity | 74.57 |
| Tokens evaluated | 290,889 |
Instruction-tuning shifts the model toward the chat format, so raw-text perplexity rises slightly (74.6 → 80.3). The small gap indicates base language ability was preserved — no catastrophic forgetting.
This is a custom architecture, not a transformers AutoModel. Clone the repo for the model code, then load the weights from this repo.
1. Clone and set up (uses uv):
git clone https://github.com/rohit-upadhya/smol-llm.git
cd smol-llm
uv sync
2. Create run.py in the repo root:
from huggingface_hub import hf_hub_download
from src.inference.inference import Inference
repo = "rohit-upadhya/SmoLLM-109M-instruct"
weights = hf_hub_download(repo, "pytorch_model.bin")
tokenizer = hf_hub_download(repo, "tokenizer.json")
inf = Inference(model_name_or_path=weights, tokenizer_path=tokenizer)
prompt = (
"[SYSTEM] You are a helpful bot [/SYSTEM]\n"
"[USER] What is machine learning? [/USER]\n"
"[ASSISTANT] "
)
print(inf.generate(prompt, max_tokens=100, temperature=0.7,
top_k=50, top_p=0.95, repetition_penalty=1.2))
3. Run it:
uv run python run.py
hf_hub_download pulls the weights and tokenizer straight from this repo — no manual downloads needed.
What is machine learning?
Machine Learning (ML) is the branch of computer science that focuses on building models and algorithms to perform tasks more efficiently, in order to create better services.
Why is exercise important?
Exercise can help people who suffer from depression and anxiety. Exercise releases endorphins which may reduce symptoms of stress, depression, and anxiety.
List three colors.
Red, Green and Blue
Real, unedited generations (temperature 0.8). The model is fluent and stops cleanly on focused questions — but at 109M parameters it will confidently hallucinate facts (inventing dates, people, or details). Treat it as a demonstration of small-model instruction-following, not a knowledge source.
| Parameters | 109.5M |
| Layers | 12 |
| Hidden dim | 768 |
| Attention heads | 12 |
| Context length | 512 |
| Tokenizer | Custom BPE (~32k vocab) |
| Components | RoPE, RMSNorm, SwiGLU, multi-head causal attention |
sample-10BT), Chinchilla-optimal token budget (~20 tokens/param, ~2.2B tokens).databricks-dolly-15k with prompt-token masking (loss computed only on response tokens) and [EOS]-terminated responses. LR 2e-5, cosine schedule, best checkpoint selected at epoch 2 by held-out eval loss (before overfitting onset).generate(prompt, max_tokens=100, temperature=0.7,
top_k=50, top_p=0.95, repetition_penalty=1.2)