Downloads · 30 days
24
22% of all-time downloads
zjr2000/SPES-7B
SPES-7B is a text generation model from zjr2000. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
SPES-7B is a 7B-parameter Mixture-of-Experts (MoE) Large Language Model pretrained using SPES (SParse Expert Sync), a memory-efficient decentralized training framework.
Downloads · 30 days
24
22% of all-time downloads
All-time downloads
110
Public
Parameters
7.3B
14.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors14.7 GB · 100%
From the Hugging Face model README
SPES-7B is a 7B-parameter Mixture-of-Experts (MoE) Large Language Model pretrained using SPES (SParse Expert Sync), a memory-efficient decentralized training framework.
This model was introduced in the paper: Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm.
Authors: Jinrui Zhang, Chaodong Xiao, Aoqi Wu, Xindong Zhang, Lei Zhang.
SPES (SParse Expert Sync) is designed for pretraining MoE LLMs across geographically distributed GPU nodes. It addresses memory and bandwidth constraints by training only a subset of experts per node, significantly lowering the individual memory footprint and eliminating the need for full-parameter transmission. SPES-7B achieves competitive performance with centrally trained models under similar computational budgets.
This model is intended for research on:
If you use this model, please cite the SPES paper:
@article{zhang2026pretraining,
title={Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm},
author={Zhang, Jinrui and Xiao, Chaodong and Wu, Aoqi and Zhang, Xindong and Zhang, Lei},
journal={arXiv preprint arXiv:2602.11543},
year={2026}
}