Downloads · 30 days
83
34% of all-time downloads
Chaman1234/Sparse-AST-BWM-TopK-MoE
Sparse-AST-BWM-TopK-MoE is a text generation model from Chaman1234. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Top-K Mixture-of-Experts (MoE) dynamic gating router connecting Sparse-AST expert backbones for specialized token-level routing on Blender 3D mathematics.
Downloads · 30 days
83
34% of all-time downloads
All-time downloads
241
Public
Parameters
37.3K
150 KB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors150 KB · 95%
From the Hugging Face model README
Top-K Mixture-of-Experts (MoE) dynamic gating router connecting Sparse-AST expert backbones for specialized token-level routing on Blender 3D mathematics.
model.safetensors)bpy, mathutils, bmesh, numpy, gpu)Evaluated on the standardized Blender 3D Math & Python Curriculum suite:
| Metric | Measured Result |
|---|---|
| Curriculum Cross-Entropy Loss | 3.8544 |
| Perplexity | 47.20 |
| Evaluation Latency | 14.35s |
| Gating Mechanism | Top-2 Routing ($\tau = 1.0$) |
| Active Backbone Experts | 3M-32, 10M-32, 100M-32, 200M-32 |
This SafeTensors distribution underwent full CPU verification against the original PyTorch checkpoint:
0.0 (Exact 0.0 bitwise equality)from model import TopKSparseASTEnsemble
router = TopKSparseASTEnsemble.from_pretrained('.')