Downloads · 30 days
23
26% of all-time downloads
Fsoft-AIC/CompeteSMoE-5.1B
CompeteSMoE-5.1B is a text generation model from Fsoft-AIC. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Downloads · 30 days
23
26% of all-time downloads
All-time downloads
88
Public
Parameters
5.1B
10.2 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors10.2 GB · 100%
How the weights are stored.
BF165.1B · 100%
From the Hugging Face model README
🎉 CompeteSMoE-5.1B
CompeteSMoE-5.1B is a lightweight and integrated variant of the Mixture-of-Experts (MoE) architecture, built upon the Phi-3.5 Mini and SigLIP baselines. This version incorporates the latest CompeteSMoE algorithm enhancements. CompeteSMoE-5.1B demonstrates strong performance across a range of MoE routing strategies, including both standard and star-to-art routing methods. It achieves competitive results compared to recent MoE architectures, such as SharedE-V2 and SharedE-V3, which are inspired by DeepSeek. Despite the architectural innovations of these models especially their use of shared experts CompeteSMoE-5.1B consistently delivers superior or comparable results.
📝 Note: This version of CompeteSMoE-5.1B was trained on a small-scale dataset. 🚧 We're actively working on a stronger, more robust release — coming soon! 🚀 Stay tuned for updates. 💡
| Stage | MoE Method | Hardware |
|---|---|---|
| Pre-Training | 4xH100 | |
| Pre-FineTuning | 4xH100 | |
| VIT | CompeteSMoE | 4xH100 |
More details can be found in our paper.
If you use CompeteSMoE, please cite it using this BibTeX:
@article{Nguyen2025CompeteSMoES,
title={CompeteSMoE - Statistically Guaranteed Mixture of Experts Training via Competition},
author={Nam V. Nguyen and Huy Nguyen and Quang Pham and Van Nguyen and Savitha Ramasamy and Nhat Ho},
journal={ArXiv},
year={2025},
volume={abs/2505.13380},
url={https://api.semanticscholar.org/CorpusID:278769210}
}