Downloads · 30 days
42
28% of all-time downloads
zjr2000/SPES-2B
SPES-2B is a text generation model from zjr2000. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
SPES-2B is a 2B-parameter Mixture-of-Experts (MoE) pretrained language model introduced in the paper:
Downloads · 30 days
42
28% of all-time downloads
All-time downloads
150
Public
Parameters
2.1B
4.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.3 GB · 100%
From the Hugging Face model README
SPES-2B is a 2B-parameter Mixture-of-Experts (MoE) pretrained language model introduced in the paper:
Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm
SPES-2B was trained using SPES, a memory-efficient decentralized framework. Unlike traditional centralized training that requires high-bandwidth interconnects, SPES enables pretraining across geographically distributed GPU nodes by training only a subset of experts per node and periodically synchronizing them. This model was trained using 16 standalone 48GB GPUs over standard internet connections.
This model is intended for:
If you use this model, please cite the SPES paper:
@article{zhang2026pretraining,
title={Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm},
author={Zhang, Jinrui icon and Xiao, Chaodong and Wu, Aoqi and Zhang, Xindong and Zhang, Lei},
journal={arXiv preprint arXiv:2602.11543},
year={2026}
}