Downloads · 30 days
0
changcheng967/cpuflow-v5-ln
cpuflow-v5-ln is a text generation model from changcheng967. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as mit.
CPU-native language model with fused linear attention cumsum, trained from scratch on free-tier CPU in 2 hours. The coherent baseline for the CPUFlow series.
Downloads · 30 days
0
Access
Public
Updated May 18, 2026
Repo size
8 MB
Likes
0
Public
Click a slice to open those files.
.pt8 MB · 97%
From the Hugging Face model README
CPU-native language model with fused linear attention cumsum, trained from scratch on free-tier CPU in 2 hours. The coherent baseline for the CPUFlow series.
| Metric | Value |
|---|---|
| Val PPL | 11.94 |
| Parameters | 2.0M |
| Training speed | 7,833 tok/s |
| Training time | 2 hours |
| Hardware | 4 vCPU (Lightning AI free tier) |
| NaN events | 0 |
embed + CumStepPos → [ScanBlock × 6] → LayerNorm → tied output + FSP
ScanBlock:
x_n = LayerNorm(x)
h = W_proj(x_n) # fused: d → 3k
query, key, value = chunk(h, 3)
key = sigmoid(key); value = tanh(value)
scan_out = W_m(query * cumsum(key*value) / cumsum(key))
x = x + W_out(scan_out)
x = x + ff_down(relu(ff_up(LayerNorm(x))))
Prompt: "Once upon a time"
Once upon a time, there was a little girl named Lily. She loved to collect the world around the forest. One day, while playing outside, she heard a noise. It was pretty and a small bush. Lily was curious.
import torch
from tokenizers import Tokenizer
tokenizer = Tokenizer.from_file("tokenizer.json")
checkpoint = torch.load("best.pt", map_location="cpu")
# Build model (see train_cpuflow_v5_ln.py for full architecture)
# Generate with temperature=0.8
See GitHub for full training code.
@misc{Chang,
title = {FlashLM: CPU-Native Language Models Trained From Scratch on Free-Tier Hardware},
author = {Chang, Cheng},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.20113960}
}
MIT License.