Downloads · 30 days
0
aplominski/TinyTransformer-Post-RMSNorm-5M-TinyStories
TinyTransformer-Post-RMSNorm-5M-TinyStories is a machine learning model from aplominski. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as openmdw-1.1.
A 5M parameter TinyTransformer model trained on the TinyStories dataset.
Downloads · 30 days
0
Access
Public
Updated Sep 1, 2026
Parameters
5.5M
21.9 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors21.9 MB · 97%
From the Hugging Face model README
A 5M parameter TinyTransformer model trained on the TinyStories dataset.
This model is part of a research series investigating the effect of normalization strategies in small Transformer models. All models in the series use the same dataset and are trained under the same experimental setup, with normalization being the primary architectural variable.
The goal of this series is to compare different normalization strategies in small-scale Transformer architectures and evaluate their impact on training and model performance.
The series includes:
| Model | Normalization | Description |
|---|---|---|
| TinyTransformer Baseline 5M | Baseline | Reference model |
| TinyTransformer Pre-LayerNorm 5M | Pre-LayerNorm | LayerNorm before Transformer sublayers |
| TinyTransformer Post-LayerNorm 5M | Post-LayerNorm | LayerNorm after Transformer sublayers |
| TinyTransformer Pre-RMSNorm 5M | Pre-RMSNorm | RMSNorm before Transformer sublayers |
| TinyTransformer Post-RMSNorm 5M | Post-RMSNorm | RMSNorm after Transformer sublayers |
All models in this series were trained on:
TinyStories by Ronen Eldan and Yuanzhi Li
The experiments compare two commonly used normalization techniques:
Layer Normalization normalizes activations across the feature dimension and was introduced by Ba et al.
RMSNorm simplifies LayerNorm by removing the mean-centering operation and normalizing using the root mean square of the activations.
The experiments evaluate both methods in pre-normalization and post-normalization configurations.
Just without any normalization
I'm used following papers in my reaserch:
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention Is All You Need. https://arxiv.org/abs/1706.03762
Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016). Layer Normalization. https://arxiv.org/abs/1607.06450
Zhang, B., & Sennrich, R. (2019). Root Mean Square Layer Normalization. https://arxiv.org/abs/1910.07467
All models in this series are released under the OpenMDW-1.1 license.
For the full license text, see the OpenMDW-1.1 license.