Downloads · 30 days
0
aquaqua/4995
4995 is a machine learning model from aquaqua. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Replication of Attention Residuals (Kimi Team, arXiv:2603.15031) at 1B scale.
Downloads · 30 days
0
Access
Public
Updated May 8, 2026
Repo size
26.4 GB
Likes
0
Public
Click a slice to open those files.
.pt26.4 GB · 100%
From the Hugging Face model README
Replication of Attention Residuals (Kimi Team, arXiv:2603.15031) at 1B scale.
PreNorm baseline, 1B params
Block AttnRes (S=8, 9 blocks), 1B params
Architecture: 36 layers × 1536 hidden × 24 heads × 6144 FFN, ctx=1024, GPT-2 BPE vocab (50304).
Trained for 15000 steps × 98K tokens/step = 1.47B tokens of FineWeb-Edu sample-10BT, BF16, AdamW, cosine LR schedule, 8× B200 NVLink-5.
Each is a torch.save dict containing model+optimizer state and the training config (, ). Load with:
Full results, plots, and code in the project's git tree (not uploaded here).