Downloads · 30 days
20
8% of all-time downloads
stiger1000/TC-MoE
TC-MoE is a text generation model from stiger1000. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
TC-MoE is a novel Mixture-of-Experts (MoE) architecture that enhances traditional MoE models through expert space expansion. By applying the ternary set {-1, 0, 1} to each original expert, TC-MoE achieves:
Downloads · 30 days
20
8% of all-time downloads
All-time downloads
254
Public
Parameters
2.3B
7.9 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors9.4 GB · 100%
From the Hugging Face model README
TC-MoE is a novel Mixture-of-Experts (MoE) architecture that enhances traditional MoE models through expert space expansion. By applying the ternary set {-1, 0, 1} to each original expert, TC-MoE achieves:
Key innovations:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("stiger1000/TC-MoE")
tokenizer = AutoTokenizer.from_pretrained("stiger1000/TC-MoE")
inputs = tokenizer("The capital of France is", return_tensors="pt")
outputs = model.generate(**inputs, max_length=50)
print(tokenizer.decode(outputs[0]))
@inproceedings{yan2025tcmoe,
title={TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice},
author={Yan, Shen and Bin, Xingyan and Zhang, Sijun and Wang, Yisen and Lin, Zhouchen},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025}
}
📚 Repository: GitHub