Downloads · 30 days
30
27% of all-time downloads
SlimFactoryHub/SlimMoE-250M-instruct
SlimMoE-250M-instruct is a text generation model from SlimFactoryHub. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
SlimMoE-250M-instruct is the final refined instruction-tuned version of the model.This stage emphasizes response quality, instruction clarity, consistency, and conversational coherence, building on the instruction-fol…
Downloads · 30 days
30
27% of all-time downloads
All-time downloads
112
Public
Parameters
253M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1 GB · 80%
From the Hugging Face model README
SlimMoE-250M-instruct is the final refined instruction-tuned version of the model.This stage emphasizes response quality, instruction clarity, consistency, and conversational coherence, building on the instruction-following and reasoning capabilities developed in earlier phases. The objective of this phase is to produce a stable and well-aligned small MoE instruction model, suitable for research and experimental evaluation under limited data and compute constraints.
This work explores the following research question:
Can a small (<500M) MoE model effectively support different attention mechanisms and alternative positional encodings under constrained compute?
SlimMoE-250M was designed to study:
| Property | Value |
|---|---|
| Parameters | 250M |
| Architecture | SlimMoEForCausalLM |
| Experts | 4 |
| Layers | 16 |
| Hidden Size | 768 |
| FFN Size | 1536 |
| Attention Heads | 12 |
| Max Context Length | 2048 |
| Routing | Adaptive MoE Routing |
| Dropout | 0.1 |
| Precision | float32 |
| Vocabulary Size | 50,257 |
This phase focused on general language modeling using high-quality educational data.
sample-10BTThis stage introduces instruction supervision and conversational alignment.
train_sftUsed to improve domain knowledge and reasoning performance.
auxiliary_trainFocused on response quality, instruction clarity, and consistency.
Given the dataset scale, GPU availability, and training time, the observed performance is reasonable and stable for this model size.
These factors directly influenced training duration and final model behavior.
We would like to thank the dataset providers and the open-source community whose contributions made this work possible.
We also acknowledge the broader open-source research community for their continuous efforts in advancing efficient model architectures and training methodologies.
Please use the Hugging Face Discussions tab to connect.