Downloads · 30 days
0
OliverSundaram/MoE-Study
MoE-Study is a text generation model from OliverSundaram. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Two decoder-only language models trained from scratch under identical conditions, differing in exactly one thing: whether the feed-forward block is a dense MLP or a sparse top-2-of-4 MoE.
Downloads · 30 days
0
Access
Public
Updated Aug 19, 2026
Repo size
4.3 GB
Likes
1
Public
Click a slice to open those files.
.pt2.9 GB · 67%
From the Hugging Face model README
Two decoder-only language models trained from scratch under identical conditions, differing in exactly one thing: whether the feed-forward block is a dense MLP or a sparse top-2-of-4 MoE.
Both checkpoints live in this one repo:
| Subfolder | Model | Total params | Active params/token |
|---|---|---|---|
dense/ | Dense FFN | 150.1M | 150.1M |
moe/ | Top-2-of-4 MoE | 206.8M | ~150.1M |
The MoE's active-parameter count matches Dense by construction — 2 of 4 experts at half the hidden size means identical compute per token. The MoE only spends more memory for extra capacity.
Full write-up, training code, and evaluation harness: github.com/OliverSundaram/MoE-Study
Read this before downloading.
They exist to answer one narrow question: at matched active compute and matched budget, does sparsity help? They are not fit for any downstream use.
These are a custom architecture, not a variant of an existing one. The modeling code is not included
here, so from_pretrained on this repo alone will not build the model.
Clone the GitHub repo — it carries the model definition and loading instructions, and points back at these subfolders for the weights.
Both models are the same custom decoder-only transformer:
| Layers | 12 |
| Attention heads | 12 |
| Embedding dim | 768 |
| Context length | 1024 |
| Vocabulary | 50,257 (GPT-2 tokenizer) |
| Attention | Multi-Query — one shared K/V projection across all query heads |
| Normalization | Custom pre-norm (learned scale + shift) |
| Position embeddings | Learned absolute |
| Weight tying | None — separate input embedding and output head |
dense/ | moe/ | |
|---|---|---|
| FFN block | 2-layer GELU MLP | 4 experts, top-2 routed |
hidden_dim | 3072 | 1536 (per expert) |
| Router | — | linear → softmax → top-2, renormalized |
| Aux loss | — | load-balancing term, summed over all 12 layers |
Both models share the same unmodified GPT-2 tokenizer, stored once at the repo root.
Identical for both models. Single consumer GPU, no cloud.
| Setting | Value |
|---|---|
| Data | nampdn-ai/tiny-textbooks |
| Tokens | 39,717 chunks × 1024 = ~40.67M |
| Epochs | 1 (19,858 steps) |
| Batch size | 2 × grad accum 4 = effective 8 |
| Optimizer | AdamW, lr 3e-4, weight decay 0.1 (no decay on 1-D params) |
| Schedule | OneCycleLR, cosine, 3% warmup |
| Grad clipping | max-norm 1.0 |
| Precision | AMP autocast + GradScaler |
| Seed | 42 |
| Hardware | 1× NVIDIA RTX 4060, 8 GB VRAM |
| Wall-clock | ~44.6 min (Dense) · ~59.8 min (MoE) |
| Dense | MoE | |
|---|---|---|
| Train loss (final step) | 5.166 | 5.936 |
| Test loss (pure LM) | 5.063 | 5.911 |
| Test loss (+ unscaled aux) | n/a | 17.91 |
Dense has the lower loss at every checkpoint.
All benchmarks via lm-evaluation-harness on the final checkpoints.
| Benchmark | Shots | Metric | Dense | MoE | abs(Δ) | Winner |
|---|---|---|---|---|---|---|
| ARC-Easy | 0 | acc | 29.2% | 27.4% | 1.8 | 🔵 Dense |
| PIQA | 0 | acc | 55.0% | 54.1% | 0.9 | 🔵 Dense |
| WikiText | 0 | word_perplexity | 551.0 | 1,377.8 | 826.8 | 🔵 Dense |
| LAMBADA (OpenAI) | 0 | acc | 0.0% | 0.0% | 0.0 | ⚪ Tie |
| WinoGrande | 5 | acc | 50.2% | 50.7% | 0.5 | 🟠 MoE |
| HellaSwag | 10 | acc_norm | 24.9% | 25.1% | 0.2 | 🟠 MoE |
| ARC-Challenge | 25 | acc_norm | 22.9% | 23.0% | 0.1 | 🟠 MoE |
How to read this:
Greedy decoding, 32-token prompt → 64 new tokens, 5 trials, 2 warmup, no KV cache.
| Model | Tokens/sec | Total params | Active params/token |
|---|---|---|---|
| Dense | 106.49 ± 0.30 | 150.1M | 150.1M |
| MoE | 34.40 ± 0.08 | 206.8M | ~150.1M |
MoE is ~3.1× slower despite matched active compute — an artifact of unoptimized expert dispatch, not a property of the architecture.
<details> <summary><b>Benchmark charts</b></summary>

1. Dense won every metric that wasn't already at chance. Most clearly on WikiText perplexity — 551 vs 1,378, a 2.5× gap.
2. The routing math is correct. Active parameters match Dense almost exactly. Matched active compute simply didn't buy matched quality at this budget.
3. Routing stayed balanced. The load-balancing term sat on its theoretical floor, so the MoE's gap is not explained by experts collapsing onto each other.
@misc{sundaram2026moestudy,
author = {Sundaram, Oliver},
title = {MoE-Study: Dense vs. Mixture-of-Experts at Matched Active Parameters},
year = {2026},
url = {https://github.com/OliverSundaram/MoE-Study}
}
transformers — base classes and tokenizerMIT