Downloads · 30 days
0
anonym5035/ablation_5_no_experts
ablation_5_no_experts is a text generation model from anonym5035. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
XpertGPT is a sparse Mixture of Experts (MoE) language model designed for data-efficient pretraining under the BabyLM 2026 challenge (Strict-Small 10M track). It leverages Parallelized Multi-Scale Information Transmis…
Downloads · 30 days
0
Access
Public
Updated Jul 18, 2026
Repo size
1 GB
Likes
0
Public
Click a slice to open those files.
.bin716 MB · 56%
From the Hugging Face model README
XpertGPT is a sparse Mixture of Experts (MoE) language model designed for data-efficient pretraining under the BabyLM 2026 challenge (Strict-Small 10M track). It leverages Parallelized Multi-Scale Information Transmission (MSIT) and Expert Choice Routing to maximize representational capacity within restricted token budgets.
This version implements SwiGLU Feed-Forward Networks across all global and parallel expert blocks, along with corrected LayerNorms, redundant residual removal, and sliding window attention sizes of [64, 16, 8, 4] tokens.
Instead of the standard Feed-Forward sequential structure (Linear -> GELU -> Linear), this model replaces FFN layers with the SwiGLU (Swish Gated Linear Unit) variant to improve model capacity and training stability:
$$\text{FFN}_{\text{SwiGLU}}(x) = \left(\text{Swish}(x W) \otimes x V\right) W_2$$
Where the Swish function is implemented using SiLU:
dim to hidden_dim.dim.dim * 4 sequential FFNs, the hidden dimension is scaled to:
$$\text{hidden_dim} = \text{round_to_multiple_of_8}\left(\frac{8}{3} \times \text{dim}\right)$$The four parallel MoE experts are configured with distinct sliding window attention constraints:
64 tokens16 tokens8 tokens4 tokensThis model implements:
res3):
ln3):
ln3 is added after the SwiGLU addition inside every MSITBranchBlock.ln_post_moe):
ln_post_moe is added after the Residual 4 MoE aggregation.main branch)import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"SRJ5035/sw_glu_sw_64_16_8_4_xpert_gpt",
revision="main",
trust_remote_code=True
).eval()
tokenizer = AutoTokenizer.from_pretrained(
"SRJ5035/sw_glu_sw_64_16_8_4_xpert_gpt",
revision="main"
)
chck_5M)model_5m = AutoModelForCausalLM.from_pretrained(
"SRJ5035/sw_glu_sw_64_16_8_4_xpert_gpt",
revision="chck_5M",
trust_remote_code=True
).eval()