Downloads · 30 days
0
samairtimer/nanollm_wiki_20.2M
nanollm_wiki_20.2M is a text generation model from samairtimer. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as mit.
nanollmwiki20.2M is a lightweight, decoder-only transformer model designed for educational purposes, fast training experimentation, and local generation on Apple Silicon using Apple's MLX framework.
Downloads · 30 days
0
Access
Public
Updated Jul 4, 2026
Repo size
80.8 MB
Likes
0
Public
Click a slice to open those files.
.safetensors80.8 MB · 100%
From the Hugging Face model README
nanollm_wiki_20.2M is a lightweight, decoder-only transformer model designed for educational purposes, fast training experimentation, and local generation on Apple Silicon using Apple's MLX framework.
The model is trained on a 100k subset of English Wikipedia articles (chonkie-ai/wikipedia-100k) to demonstrate basic autoregressive text generation capabilities with minimal compute requirements.
tiktoken, vocab size 50,257)Unlike standard Llama or GPT architectures, this model is a highly simplified attention-only transformer that omits the typical Feed-Forward Neural Network (FFN) sub-blocks and layer normalization to keep training extremely fast and lightweight.
| Parameter | Value | Detail |
|---|---|---|
vocab_size | 50,257 | GPT-2 Tiktoken Vocabulary |
maxlen (Context Length) | 128 | Maximum context window |
embed_dim | 192 | Hidden size dimension |
num_transformer_blocks | 6 | Number of decoder layers |
num_heads | 6 | Multi-head attention heads (32-dim heads) |
feed_forward_dim | 512 | Configured but omitted in runtime computation |
The model contains exactly 20,208,000 parameters (~20.2M).
$$ \text{Total Parameters} = \text{Embeddings} + \text{Transformer Blocks} + \text{Output Projection} $$
Embedding Layer:
Transformer Blocks (6 Layers):
Output Layer (LM Head):
bias=False: $192 \times 50,257 = \mathbf{9,649,344}$ parameters$$\text{Grand Total} = 9,673,920 + 884,736 + 9,649,344 = 20,208,000 \text{ parameters}$$
chonkie-ai/wikipedia-100k<|endoftext|> token.learning_rate=3e-4).safetensors size of ~80.8 MB)To use this model for text generation, define the architecture matching the parameters and load the saved .safetensors weights using MLX.
Install the dependencies:
pip install mlx tiktoken huggingface_hub
import os
import mlx.core as mx
import mlx.nn as nn
import tiktoken
from huggingface_hub import hf_hub_download
# Define model architecture identical to training
class TokenAndPositionEmbedding(nn.Module):
def __init__(self, maxlen: int, vocab_size: int, embed_dim: int):
super().__init__()
self.token_emb = nn.Embedding(vocab_size, embed_dim)
self.pos_emb = nn.Embedding(maxlen, embed_dim)
def __call__(self, x):
seq_len = x.shape[1]
positions = mx.arange(seq_len)[None, :]
return self.token_emb(x) + self.pos_emb(positions)
class TransformerBlock(nn.Module):
def __init__(self, emed_dim: int, num_heads: int, ff_dim: int):
super().__init__()
self.attention = nn.MultiHeadAttention(emed_dim, num_heads)
def __call__(self, x, mask=None):
attn_out = self.attention(x, x, x, mask=mask)
return x + attn_out
class NanoLLM(nn.Module):
def __init__(self, maxlen: int, vocab_size: int, embed_dim: int, num_heads: int, feed_forward_dim: int, num_transformer_blocks: int):
super().__init__()
self.maxlen = maxlen
self.embedding = TokenAndPositionEmbedding(maxlen, vocab_size, embed_dim)
self.transformer_blocks = [
TransformerBlock(embed_dim, num_heads, feed_forward_dim)
for _ in range(num_transformer_blocks)
]
self.output_layer = nn.Linear(embed_dim, vocab_size, bias=False)
def __call__(self, token_ids):
seq_len = token_ids.shape[1]
mask = nn.MultiHeadAttention.create_additive_causal_mask(seq_len)
x = self.embedding(token_ids)
for block in self.transformer_blocks:
x = block(x, mask=mask)
return self.output_layer(x)
# 1. Download weights from Hugging Face
repo_id = "samairtimer/nanollm_wiki_20.2M"
weights_path = hf_hub_download(repo_id=repo_id, filename="wikipedia_checkpoint.safetensors")
# 2. Initialize and load weights
tokenizer = tiktoken.get_encoding("gpt2")
model = NanoLLM(
maxlen=128,
vocab_size=tokenizer.n_vocab,
embed_dim=192,
num_heads=6,
feed_forward_dim=512,
num_transformer_blocks=6
)
model.load_weights(weights_path)
print("Model loaded successfully!")
# 3. Autoregressive Generation function
def generate(model, tokenizer, prompt, max_new_tokens=100, temperature=0.6):
tokens = tokenizer.encode(prompt)
x = mx.array(tokens)[None, :]
end_token_id = tokenizer.encode('<|endoftext|>', allowed_special={'<|endoftext|>'})[0]
print(prompt, end="", flush=True)
for _ in range(max_new_tokens):
if x.shape[1] > model.maxlen:
x = x[:, -model.maxlen:]
logits = model(x)
next_token_logits = logits[0, -1, :]
if temperature == 0.0:
next_token = mx.argmax(next_token_logits, axis=-1).item()
else:
next_token = mx.random.categorical(next_token_logits / temperature, axis=-1).item()
if next_token == end_token_id:
break
word = tokenizer.decode([next_token])
print(word, end="", flush=True)
x = mx.concatenate([x, mx.array([[next_token]])], axis=1)
print("\n")
# Run generation
generate(model, tokenizer, "Cats are", max_new_tokens=50, temperature=0.4)