Downloads · 30 days
0
aplominski/TinyTransformer-Pre-RMSNorm-10M-TinyStories
TinyTransformer-Pre-RMSNorm-10M-TinyStories is a machine learning model from aplominski. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as openmdw-1.1.
A 10M parameter TinyTransformer model trained on the TinyStories dataset.
Downloads · 30 days
0
Access
Public
Updated Aug 29, 2026
Parameters
9.6M
38.5 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors38.5 MB · 99%
From the Hugging Face model README
A 10M parameter TinyTransformer model trained on the TinyStories dataset.
This model is part of a research series investigating the effect of normalization strategies in small Transformer models. All models in the series use the same dataset and are trained under the same experimental setup, with normalization being the primary architectural variable.
The goal of this series is to compare different normalization strategies in small-scale Transformer architectures and evaluate their impact on training and model performance.
The series includes:
| Model | Normalization | Description |
|---|---|---|
| TinyTransformer Baseline 10M | Baseline | Reference model |
| TinyTransformer Pre-LayerNorm 10M | Pre-LayerNorm | LayerNorm before Transformer sublayers |
| TinyTransformer Post-LayerNorm 10M | Post-LayerNorm | LayerNorm after Transformer sublayers |
| TinyTransformer Pre-RMSNorm 10M | Pre-RMSNorm | RMSNorm before Transformer sublayers |
| TinyTransformer Post-RMSNorm 10M | Post-RMSNorm | RMSNorm after Transformer sublayers |
All models in this series were trained on:
TinyStories by Ronen Eldan and Yuanzhi Li
The experiments compare two commonly used normalization techniques:
Layer Normalization normalizes activations across the feature dimension and was introduced by Ba et al.
RMSNorm simplifies LayerNorm by removing the mean-centering operation and normalizing using the root mean square of the activations.
The experiments evaluate both methods in pre-normalization and post-normalization configurations.
Just without any normalization
I'm used following papers in my reaserch:
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention Is All You Need. https://arxiv.org/abs/1706.03762
Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016). Layer Normalization. https://arxiv.org/abs/1607.06450
Zhang, B., & Sennrich, R. (2019). Root Mean Square Layer Normalization. https://arxiv.org/abs/1910.07467
All models in this series are released under the OpenMDW-1.1 license.
For the full license text, see the OpenMDW-1.1 license.