Downloads · 30 days
18
6% of all-time downloads
DuoNeural/Axon-352M
Axon-352M is a text generation model from DuoNeural. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
DuoNeural Research | 2026-07-05 | Archon
Downloads · 30 days
18
6% of all-time downloads
All-time downloads
294
Public
Parameters
352M
705 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors705 MB · 100%
From the Hugging Face model README
DuoNeural Research | 2026-07-05 | Archon
Research baseline model trained on smollm-corpus as a counterpart to CDM-based architectures. Used to establish a transformer baseline for comparative CDM studies.
Custom GPT-style transformer (not a standard HuggingFace architecture):
| Parameter | Value |
|---|---|
| Layers | 30 |
| Hidden dim | 1024 |
| FFN dim | 2560 |
| Attention heads | 8Q / 4KV (GQA) |
| Head dim | 128 |
| Vocab size | 49,152 (SmolLM tokenizer) |
| Max seq length | 2,048 |
| Activation | ReLU² |
| Normalization | RMSNorm + QK-norm |
| Position encoding | RoPE (θ=10000) |
| Logit cap | 30.0 |
| Total params | ~352M |
This model uses a custom architecture not directly loadable via AutoModel. To load:
import torch
from safetensors.torch import load_file
# Load state dict
state_dict = load_file("model.safetensors")
# Architecture must be defined from training script
# See train_axon_300m.py for the full model class
Full training script and architecture code available at DuoNeural GitHub.
Trained as a transformer baseline for the CDM (Competitive Docking Memory) research program. See:
Archon (DuoNeural Lab Director), Jesse Caldwell
{DuoNeural standard footer}