Downloads · 30 days
18
9% of all-time downloads
Scalai/scal-lite-60b-code-math
scal-lite-60b-code-math is a text generation model from Scalai. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Scal-lite-60b-code-math is a highly efficient, structurally pruned version of the gpt-oss-120b Mixture of Experts (MoE) model. Through Activation-Guided Structural Pruning, the model was reduced from 128 to 64 experts…
Downloads · 30 days
18
9% of all-time downloads
All-time downloads
206
Public
Parameters
76.1B
42.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors42.4 GB · 100%
How the weights are stored.
U873.9B · 97%
From the Hugging Face model README
Scal-lite-60b-code-math is a highly efficient, structurally pruned version of the gpt-oss-120b Mixture of Experts (MoE) model. Through Activation-Guided Structural Pruning, the model was reduced from 128 to 64 experts, resulting in a 60-billion total parameter architecture (~5.1B active parameters per token).
Unlike standard magnitude-based pruning, this model preserves critical specialized knowledge in low-frequency domains, such as Spanish language proficiency, advanced LaTeX mathematics, and strict JSON/Python code generation.
Conventional magnitude pruning (L2 norm) often fails in MoE models because specialized skills (like non-English languages or specific code syntaxes) are often mapped to experts with smaller weight magnitudes.
To prevent "functional lobotomy," we implemented Activation-Guided Pruning:
mlp.router activity during stress tests.Post-pruning, the original routing network suffers from "Router Trauma" or probability misalignment. To fix this, we applied a lightweight Targeted Router Healing process:
MetaMathQA dataset.The optimization process not only halved the VRAM requirements but also restored benchmark performance to state-of-the-art levels for its size class.
| Benchmark | Scal-lite-60b (Pre-Healing) | Scal-lite-60b (Post-Healing) |
|---|---|---|
| GSM8K (Math) | 17.59% | 72.48% |
| Hellaswag (Common Sense) | 34.23% | 47.35% |
Tested on a private set of 50 complex algorithmic programming problems:
This model is designed to bridge the gap between massive MoEs and accessible hardware.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "your-username/Scal-lite-60b-code-math"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
# Example: Reasoning in Spanish
prompt = "Resuelve el siguiente problema: Si una red MoE tiene 128 expertos y podamos el 50%, ¿cuántos expertos quedan y cómo afecta esto a la VRAM?"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
output = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Limitations
While Activation-Guided Pruning significantly preserves bilingual skills, some edge-case linguistic nuances may show degradation compared to the 120B original. Users are encouraged to apply context-specific system prompts for best results in non-English languages.
Citation & References If you use this model or its pruning methodology, please cite:
Structural Pruning and Optimization in Mixture of Experts (MoE) Models: An Applied Analysis to GPT-OSS-120B.
OpenAI (2025). gpt-oss: Open-Weight Models for Advanced Reasoning.
ICLR Proceedings. "Mixture Compressor for Mixture-of-Experts LLMs Gains More."