Downloads · 30 days
9
4% of all-time downloads
luozhangzichen/neon213
neon213 is a machine learning model from luozhangzichen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for neon. The card lists the license as mit.
neon213-Muon is the best-performing model of the NeonBench series. By combining a Progressive Growth strategy ($k=1 \to 21$) with the Muon Optimizer, this model achieves a project-wide SOTA on FineWeb-Edu, outperformi…
Downloads · 30 days
9
4% of all-time downloads
All-time downloads
216
Public
Parameters
33M
409 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors344 MB · 67%
From the Hugging Face model README
neon213-Muon is the best-performing model of the NeonBench series. By combining a Progressive Growth strategy ($k=1 \to 21$) with the Muon Optimizer, this model achieves a project-wide SOTA on FineWeb-Edu, outperforming the AdamW baseline by a noticable margin.
| Property | Value |
|---|---|
| Parameters (Total) | 26.49M |
| Parameters (Active) | 20.20M (Non-Embedding) |
| Architecture | Growable SwiGLU-Conv (neon213-Muon) |
| Optimizer | Muon V4 (Orthogonalized) |
| Tokenizer | tok6 (16,384 Vocab) |
| Dimensions | $d_{model}=384, n_{head}=6, n_{layers}=8, d_{ff}=1536$ |
| Context | Growable ($k=1 \to 21$) |
| Status | Project SOTA (3.44 Val Loss) |
The architecture is based on neon185 (SwiGLU-Conv), which features:
w2(SiLU(w1(x)) * w3(x)) gating.sigmoid(Intent) gate on attention output.Unlike previous static models, neon213 features configurable kernel sizes (conv_k, mlp_k). This allows the model to start with pointwise operations ($k=1$) and grow its receptive field during training.
# Conv-Attention Layer
self.conv_q = nn.Conv1d(d, d, kernel_size=k, groups=d) # k grows 1->9
self.conv_k = nn.Conv1d(d, d, kernel_size=k, groups=d)
self.conv_v = nn.Conv1d(d, d, kernel_size=k, groups=d)
self.conv_i = nn.Conv1d(d, d, kernel_size=k, groups=d)
# SwiGLU MLP Layer
self.conv_gate = nn.Conv1d(d, d, kernel_size=k, groups=d) # k grows 1->9
In our AdamW baseline tests, we observed severe Dimensional Collapse. Despite a 384-dimensional latent space, the Participation Ratio (PR) was often as low as ~12.0, meaning less than 4% of the representational power was being utilized.
By switching to the Muon Optimizer (Orthogonalized Momentum), we forced the weight matrices to remain diverse.
Standard Gated Scaled Dot-Product Attention (as seen in architectures like Qwen 3.5) derives its output gate from existing projections — typically the Query. The gate is calculated, not independently learned:
$$\text{Gated-SDPA}: \quad y = \sigma(W_g \cdot Q) \odot \text{Attn}(Q, K, V)$$
In neon213, the gate is a fully independent learned projection called Intent ($I$). Intent has its own dedicated weights (c_attn slice) and its own dedicated convolution (conv_i), giving it a completely separate representational capacity from Q, K, and V:
$$\text{Intent-Gated}: \quad y = \sigma(\text{Conv}(I)) \odot \text{Attn}(Q, K, V)$$
This means the model can learn what information to keep (Intent) independently from what information to search for (Query) and what information to retrieve (Value).
The depthwise convolutions applied to Q, K, V, and I shouldn't be interpreted as simple blurs. Each convolution kernel is a set of fully learned, unconstrained weights — including negative values. This means each dimension can independently decide:
In practice, this creates an additional token-to-token communication pathway that operates before the attention mechanism. While attention allows tokens to selectively read from any position, the convolutions provide a fixed, local, per-dimension channel for adjacent tokens to share information — a form of inductive bias that complements the global, content-based routing of attention.
Convolutions at large kernel sizes ($k=9$) are powerful but difficult to train from scratch — the model must simultaneously learn what to convolve and how far to look. neon213 solves this with Progressive Kernel Growth:
This approach is analogous to curriculum learning: the model masters simple patterns first, then progressively gains the capacity to leverage richer local context.
The Muon SOTA model followed an accelerated and expanded growth curriculum. While the AdamW baseline struggled past $k=9$, the Muon-backed heads remained stable up to $k=21$.
| Stage | Kernel ($k$) | Steps | Description |
|---|---|---|---|
| 1 | 1 | 5,000 | Cold Start: Muon bootstraps attention stability. |
| 2-10 | 3 ➔ 19 | 27,000 | Hybrid Growth: Step-wise expansion ($+2k$ every 3k steps). |
| 11 | 21 | 3,000 | Target Depth: Reached full $k=21$ context. |
| 12 | 21 | 30,000 | The Floor: Final long-tail convergence with Cosine decay. |
Total Steps: ~65,000.
The final model checkpoint exceeded the GitHub 100MB file limit (101 MB). To resolve this, the checkpoint was converted to Float16 (Half Precision).
state_dict.NeonModelEngine automatically handles the fp16 $\to$ fp32 cast during loading.| Metric | Value | Notes |
|---|---|---|
| Best Val Loss | 3.44 | FineWeb-Edu (Project SOTA). |
| Participation Ratio | 25.9 | High dimensional utilization. |
Sample Generation:
"The meaning of life is deeply intertwined with the social and cultural constructs of the era. It is not a static definition but a dynamic process of engagement with the environment and the community, reflecting"
"The meaning of life is captured in the intricate patterns of human interaction and the shared pursuit of knowledge. It is the ability to adapt, to learn, and to contribute to the collective wisdom of the species"