Downloads · 30 days
77
64% of all-time downloads
SlimFactoryHub/SlimMoE-250M-base
SlimMoE-250M-base is a text generation model from SlimFactoryHub. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
SlimMoE-250M is a 250M parameter Mixture-of-Experts (MoE) language model developed by the SlimFactory team.This model was trained to experiment with VGQA-style attention mechanisms and NoPE/RoPE positional strategies…
Downloads · 30 days
77
64% of all-time downloads
All-time downloads
120
Public
Parameters
253M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1 GB · 80%
From the Hugging Face model README
SlimMoE-250M is a 250M parameter Mixture-of-Experts (MoE) language model developed by the SlimFactory team.This model was trained to experiment with VGQA-style attention mechanisms and NoPE/RoPE positional strategies in a small-parameter MoE setting, focusing on architectural feasibility and training stability rather than scale or benchmark maximization.
This work explores the following research question:
Can a small (<500M) MoE model effectively support different attention mechanisms and alternative positional encodings under constrained compute?
SlimMoE-250M was designed to study:
| Property | Value |
|---|---|
| Parameters | 250M |
| Architecture | SlimMoEForCausalLM |
| Experts | 4 |
| Layers | 16 |
| Hidden Size | 768 |
| FFN Size | 1536 |
| Attention Heads | 12 |
| Max Context Length | 2048 |
| Routing | Adaptive MoE Routing |
| Dropout | 0.1 |
| Precision | float32 |
| Vocabulary Size | 50,257 |
This phase focused on general language modeling using high-quality educational data.
sample-10BTThis stage introduces instruction supervision and conversational alignment.
train_sftUsed to improve domain knowledge and reasoning performance.
auxiliary_trainFocused on response quality, instruction clarity, and consistency.
Given the dataset scale, GPU availability, and training time, the observed performance is reasonable and stable for this model size.
These factors directly influenced training duration and final model behavior.
We would like to thank the dataset providers and the open-source community whose contributions made this work possible.
We also acknowledge the broader open-source research community for their continuous efforts in advancing efficient model architectures and training methodologies.
Please use the Hugging Face Discussions tab to connect.