Downloads · 30 days
0
neemon/anlp-a2-part1-moe_active_matched
anlp-a2-part1-moe_active_matched is a translation model from neemon. Use it when you need text moved from one language to another. It is set up for pytorch. The card lists the license as mit.
Decoder-only transformer with Mixture-of-Experts, 4 experts, top-2 routing, active parameters matched to the dense baseline.
Downloads · 30 days
0
Access
Public
Updated Oct 1, 2026
Repo size
701 MB
Likes
0
Public
Click a slice to open those files.
.pt234 MB · 99%
From the Hugging Face model README
Decoder-only transformer with Mixture-of-Experts, 4 experts, top-2 routing, active parameters matched to the dense baseline.
Trained from scratch for IIIT-H Advanced NLP, Monsoon 2026. Architecture, training code and evaluation live in the assignment repository. This repository holds the checkpoint and the tokenizer it was trained with, which every model in the assignment shares.
| Setting | Value |
|---|---|
d_model | 512 |
n_layers | 8 |
n_heads | 8 |
n_kv_heads | 8 |
n_ctx | 256 |
vocab_size | 32000 |
ffn | moe_active_matched |
d_ff | 1024 |
n_routed_experts | 4 |
n_shared_experts | 0 |
top_k | 2 |
norm | rmsnorm |
| Count | Value |
|---|---|
| Total | 58,401,280 |
| Active per token | 41,599,488 |
| Feed-forward (total / active) | 33,603,584 / 16,801,792 |
| Metric | Value |
|---|---|
| Perplexity (both directions) | 8.81 |
| BLEU vi->en | 35.06 |
| BLEU ja->en | 25.17 |
| BLEU mean | 30.12 |
import torch
payload = torch.load("model.pt", map_location="cpu", weights_only=False)
state_dict, config = payload["model"], payload["config"]
from tokenizers import Tokenizer
tokenizer = Tokenizer.from_file("tokenizer.json")