Downloads · 30 days
69
14% of all-time downloads
ajaxdavis/alpha-er
alpha-er is a text generation model from ajaxdavis. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-sa-4.0.
alpha-er (二, èr — "two") is a 100M-parameter language model trained end to end on a from-scratch GPU stack: our own ioctl driver, memory manager, command channels, sm86 SASS assembler and kernel IR. No CUDA, no cuBLAS…
Downloads · 30 days
69
14% of all-time downloads
All-time downloads
508
Public
Parameters
100M
401 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors401 MB · 100%
From the Hugging Face model README
alpha-er (二, èr — "two") is a 100M-parameter language model trained end to end on a from-scratch GPU stack: our own ioctl driver, memory manager, command channels, sm_86 SASS assembler and kernel IR. No CUDA, no cuBLAS, no vendor runtime was involved in training — every matrix multiply ran on hand-written machine code.
It was trained on one RTX 3070 at ~96,000 tokens/second, for 1.97B tokens in 5.7 hours.
alpha-er writes fluent, grammatical English in the correct register for a prompt. It is not factually reliable and it does not answer questions correctly. Validation perplexity is ~88; this is a small model trained for one afternoon, published as a research artifact of the GPU stack rather than as a useful assistant.
Real, unedited samples from the released weights (temperature 0.8, top-k 40):
| Prompt | Output |
|---|---|
<|user|>What is the capital of France?<|assistant|> | "To create a new and innovative approach, you can crafting a rich culture, and interests, such a rich, uniquely and diverse interests…" |
The history of the Roman Empire | ". In addition to the ancient Egyptian and ancient Egypt, but on the East Vietria. So, both are the elegance of the past, are a Germanician Greek civilization, the Greeks (1714)…" |
Note what this does and does not show. The model has learned which words belong together in a history text — Egypt, Greek civilization, antiquity — without learning any history. Syntax is solid; semantics are weak. That is the honest reading of a 100M model at perplexity 88.
Three features make this not a Llama, and loading it as one would silently produce a different model.
Conditional MLP. The feed-forward block is split into G = 64 experts of width 320. Each token is routed to exactly one, so the model stores a 20,480-wide FFN but any token pays for 320. This is the identity the whole design rests on: FLOPs/token = 6 × active parameters, not 6 × total.
Positional routing. expert(t) = floor(t · G / T) — a token's expert depends only on its
position within its own sequence. An earlier version routed on the index in the flattened batch,
which made each sequence reach only 4 of 64 experts and made the weights meaningful only at the
exact batch shape they were trained at. Routing on t makes a checkpoint portable: the same
sequence gives identical logits at any batch width.
Factored projections. QKV, the attention output and the LM head are each a rank-128 bottleneck with a LayerNorm on the bottleneck. The norm is load-bearing — without it the factored form diverged (grad_norm 51 against a dense baseline's 1.25).
| Parameters | 100,281,600 |
| Layers / d_model / heads | 2 / 1024 / 8 |
| FFN | 20,480 total, 320 active per token (G=64) |
| Context | 512 |
| Vocabulary | 12,288 (byte-level BPE) |
| Position encoding | learned |
| Attention | causal, logit soft-cap 30 (30·tanh(s/30)) |
| Activation | GELU (tanh approximation) |
Expert boundaries fall at multiples of T/G, so the model only reproduces its training behaviour
at its trained context length. Pad the prompt to 512 and read the logits at the last real
position. This is exact, not an approximation: attention is causal, so padding after the prompt
cannot influence it, and each token's expert depends only on its own position. generate() in
modeling_alpha.py does this for you.
import torch
from modeling_alpha import AlphaErConfig, AlphaErForCausalLM
from tokenization_alpha import AlphaErTokenizer
from safetensors.torch import load_file
import json
cfg_d = json.load(open("config.json"))
cfg = AlphaErConfig(**{k: v for k, v in cfg_d.items()
if k in AlphaErConfig.__init__.__code__.co_varnames})
model = AlphaErForCausalLM(cfg)
model.load_state_dict(load_file("model.safetensors"), strict=False)
model.eval()
tok = AlphaErTokenizer.from_file("tokenizer_artifacts.json")
ids = tok.encode("<|user|>Hello!<|assistant|>")
out = model.generate(torch.tensor([ids]), max_new_tokens=60)[0].tolist()
print(tok.decode(out[len(ids):]))
modeling_alpha.py is a PyTorch re-expression of the trainer's forward pass, not the trainer
itself. It is checked elementwise against the real model at one position inside every one of the
64 expert windows: max |Δlogit| = 6.8e-05, relative 3.7e-06 — float32 round-off.
| Tokens | 1.97B (20,000 steps × 98,304) |
| Batch | 16 × 512, gradient accumulation 12 |
| Optimizer | AdamW, lr 3e-4 cosine, warmup 500, weight decay 0.1, grad clip 1.0 |
| Loss | cross-entropy with sampled softmax (512 shared negatives) during training; full softmax for evaluation |
| Final val loss | 4.4803 (best 4.4119) — perplexity ~88 vs 12,288 for uniform |
| Throughput | ~96,000 tok/s on one RTX 3070 |
The validation curve flattened after roughly step 8,000, moving only 4.70 → ~4.45 over the second half. The likely cause is active capacity: each token passes through a single 320-wide expert. More steps would not fix that; fewer and wider experts would.
Trained on ajaxdavis/alpha-er-corpus
— FineWeb-Edu/DCLM/FinePDFs, Concordance-EN, and SmolTalk. Licensed CC-BY-SA-4.0, inherited
from Concordance-EN's share-alike terms.
Published as a research artifact demonstrating that a hand-built, fully-understood GPU stack can train a real language model. Do not deploy it.