Downloads · 30 days
130
28% of all-time downloads
ServiceNow-AI/SuperApriel-15b-Base
SuperApriel-15b-Base is a text generation model from ServiceNow-AI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
<img src="assets/super-apriel.png" width="120" alt="thumbnail"/ /ˈɑː.pri.əl/
Downloads · 30 days
130
28% of all-time downloads
All-time downloads
467
Public
Repo size
100 GB
Likes
4
Public
Click a slice to open those files.
.safetensors100 GB · 100%
From the Hugging Face model README
<img src="assets/super-apriel.png" width="120" alt="thumbnail"/> /ˈɑː.pri.əl/
A 15B-parameter token-mixer supernet derived from Apriel-1.6 via stochastic distillation. Every decoder layer exposes four trained mixer options—Full Attention, Sliding Window Attention, Gated DeltaNet, and Kimi Delta Attention—enabling flexible architecture selection from a single checkpoint.
SuperApriel-15b-Base is the Stage 1 (distillation) checkpoint of the Super Apriel supernet. During training, all four mixer types at each layer were trained simultaneously using stochastic local sampling—each layer's mixer was drawn uniformly from the four types at each training step. Only mixer weights were trained; all shared parameters (FFNs, embeddings, layer norms, vision encoder) remain frozen from the Apriel-1.6 teacher.
This checkpoint is intended as a foundation for downstream fine-tuning. For a ready-to-use model with optimized deployment presets, see SuperApriel-15b-Instruct.
| Component | Details |
|---|---|
| Parameters | 15B |
| Decoder layers | 48 |
| Query / KV heads | 32 / 8 (grouped-query attention), d_h = 128 |
| Hidden dimension | 5,120 |
| FFN width | 14,336 (SiLU-gated) |
| Vocabulary | 131,072 tokens |
| Vision encoder | Pixtral (16×16 patches) |
| Mixer | Time | Memory | Description |
|---|---|---|---|
| Full Attention (FA) | O(n²) | O(n) KV cache | Standard grouped-query attention |
| Sliding Window (SWA) | O(w·n) | O(w) | Local window of 4,096 tokens |
| Gated DeltaNet (GDN) | O(n) | O(1) fixed state | Matrix-valued recurrent state with delta rule |
| Kimi Delta Attention (KDA) | O(n) | O(1) fixed state | Linear attention with channel-wise gating |
This checkpoint is intended as a foundation for fine-tuning and research, not for direct inference. For a ready-to-use model with optimized deployment presets and full serving instructions, see SuperApriel-15b-Instruct.
If you need to load this checkpoint for evaluation or experimentation, copy a preset config from SuperApriel-15b-Instruct to select a specific mixer placement. The Base and Instruct checkpoints share the same architecture and config format — preset configs from Instruct work with this checkpoint.
For example, to load with the all-attention placement:
config.json from SuperApriel-15b-Instruct/preset_configs/all-attention/config.jsonNote: This model requires
trust_remote_code=Trueas it uses custom architecture code for the multi-mixer supernet.
Note: When serving with vLLM, custom placements must include at least one attention-type layer (FA or SWA). Configurations using only recurrent mixers (GDN/KDA) are not currently supported due to a vLLM KV cache coordinator limitation. All shipped Instruct presets satisfy this requirement.
SuperApriel-15b-Base is designed as a foundation checkpoint for:
It is not intended for direct deployment without further fine-tuning or for safety-critical applications without human oversight.
Security Responsibilities: Deployers and users are strongly encouraged to align their security practices with established frameworks and regulatory guidelines such as the EU AI Act and the NIST AI Risk Management Framework (RMF).
Guidelines for Deployers:
Guidelines for Users:
Disclaimer: Users accept responsibility for securely deploying, managing, and using this open-source LLM. The model is provided "as-is," without explicit or implied warranty regarding security or fitness for any specific application or environment.
MIT
@misc{super_apriel_2026,
title = {Super Apriel: One Checkpoint, Many Speeds},
author = {ServiceNow Language Models Lab},
year = {2026},
eprint = {2604.19877},
archivePrefix= {arXiv},
primaryClass = {cs.CL}
}