Downloads · 30 days
21
2% of all-time downloads
tensorfiend/DotLM-165M
DotLM-165M is a text generation model from tensorfiend. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
DotLM is a minimal 165M parameter model, from-scratch transformer trained entirely on the SimpleThoughts dataset. It uses explicit <think...</think chain-of-thought traces to reason through intuitive physics, logic, c…
Downloads · 30 days
21
2% of all-time downloads
All-time downloads
927
Public
Parameters
176M
705 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors705 MB · 100%
From the Hugging Face model README
DotLM is a minimal 165M parameter model, from-scratch transformer trained entirely on the
SimpleThoughts dataset. It uses explicit <think>...</think>
chain-of-thought traces to reason through intuitive physics, logic, causal inference, and other everyday phenomena before producing an
answer.
| Parameter | Value |
|---|---|
| Parameters | ~165M |
| Layers | 24 |
| Model dimension | 768 |
| FFN hidden dim | 2048 (SwiGLU) |
| Attention heads | 6 |
| KV heads (GQA) | 2 |
| Head dimension | 128 |
| Context length | 4096 tokens |
| Vocabulary size | 16,384 (BPE) |
| Positional encoding | RoPE (θ = 10,000) |
| Normalization | RMSNorm (ε = 1e-6) |
| Tied embeddings | Yes |
Key design choices: Grouped-Query Attention (GQA) with 3:1 head ratio for efficient KV memory, SwiGLU activations, pre-norm architecture, and bf16 mixed-precision training throughout.
The model was trained sequentially across four stages using the DotLM framework:
| Stage | Dataset | Samples | Objective |
|---|---|---|---|
| Pretraining | SimpleThoughts/pretrain | 352,214 | Next-token prediction |
| SFT | SimpleThoughts/sft | 25,788 | ChatML instruction following |
| Alignment | SimpleThoughts/alignment | 7,172 | Reference-free DPO (SimPO-style) |
| Reasoning | SimpleThoughts/reasoning | 6,300 | Chain-of-thought with <think> traces |
| Token | Purpose |
|---|---|
<|im_start|> | Start of turn (BOS) |
<|im_end|> | End of turn |
<think> | Begin reasoning trace |
</think> | End reasoning trace |
<endoftext> | End of sequence (EOS) |
<pad> | Padding |
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
repo_id = "tensorfiend/DotLM-165M"
device = "cuda" if torch.cuda.is_available() else "cpu"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
).to(device)
user_query = "If a ball is placed inside a box and the box is sealed, where is the ball?"
prompt = f"<|im_start|>user\n{user_query}<|im_end|>\n<|im_start|>assistant\n<think>"
inputs = tokenizer(prompt, return_tensors="pt").to(device)
outputs = model.generate(
**inputs,
max_new_tokens=512,
temperature=0.7,
top_k=50,
do_sample=True,
eos_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=False))
DotLM uses the ChatML format with an explicit reasoning prefix:
<|im_start|>user
{your question}<|im_end|>
<|im_start|>assistant
<think>
{model reasons here}
</think>
{final answer}
Checkout the blog for training details: DotLM - An end-to-end trained 165M model (coming soon)
Related Resources
@misc{dotlm2026, author = {Shanmukh}, title = {DotLM-165M: A Minimal Reasoning Language Model Trained on Thought Experiments}, year = {2026}, publisher = {Hugging Face}, url = {https://huggingface.co/tensorfiend/DotLM-165M} }