Downloads · 30 days
18
8% of all-time downloads
ykae/monarch-bert-base-mnli
monarch-bert-base-mnli is a text classification model from ykae. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
Breaking the Efficiency Barrier: -66.2% Parameters, +24% Speed.
Downloads · 30 days
18
8% of all-time downloads
All-time downloads
240
Public
Parameters
54.9M
440 MB on disk
Likes
0
Public
Click a slice to open those files.
.bin220 MB · 50%
From the Hugging Face model README
Breaking the Efficiency Barrier: -66.2% Parameters, +24% Speed.
tl;dr: Achieving extreme resource efficiency on MNLI. We replaced every dense FFN layer in BERT-Base with structured Monarch Matrices. Distilled in just 3 hours on one H100 using only 500k Wiki tokens and MNLI data, this model slashes parameters by 66.2% and boosts throughput by +24% (vs optimized Baseline).
Training models from scratch typically requires billions of tokens. We took a different path to shock the efficiency curve:
Measured on a single NVIDIA H100 using torch.compile(mode="max-autotune").
| Metric | BERT-Base (Baseline) | Monarch-Full (This) | Delta |
|---|---|---|---|
| Parameters | 85.65M | 28.98M | 📉 -66.2% |
| Compute (GFLOPs) | 696.5 | 232.6 | 📉 -66.6% |
| Throughput (TPS) | 7261 | 9029 | 🚀 +24.3% |
| Latency (Batch 32) | 4.41 ms | 3.54 ms | ⚡ +24.6% Faster |
| Accuracy (MNLI) | 83.62% | 78.34% | 📉 -5.28% (towards 0% acc drop) |
Our compression scales predictably: as training data increases, the accuracy deficit converges toward zero.
This model uses a custom architecture. You must enable trust_remote_code=True to load the Monarch layers (MonarchUp, MonarchDown, MonarchFFN).
To see the real speedup, compilation is mandatory (otherwise PyTorch Python overhead masks the hardware gains).
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from datasets import load_dataset
from torch.utils.data import DataLoader
from tqdm import tqdm
device = "cuda" if torch.cuda.is_available() else "cpu"
model_id = "ykae/monarch-bert-base-mnli"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(
model_id,
trust_remote_code=True
).to(device)
# torch.set_float32_matmul_precision('high')
# model = torch.compile(model, mode="max-autotune")
model.eval()
print("📊 Loading MNLI Validation set...")
dataset = load_dataset("glue", "mnli", split="validation_matched")
def tokenize_fn(ex):
return tokenizer(ex['premise'], ex['hypothesis'],
padding="max_length", truncation=True, max_length=128)
tokenized_ds = dataset.map(tokenize_fn, batched=True)
tokenized_ds.set_format(type='torch', columns=['input_ids', 'attention_mask', 'label'])
loader = DataLoader(tokenized_ds, batch_size=32)
correct = 0
total = 0
print(f"🚀 Starting evaluation on {len(tokenized_ds)} samples...")
with torch.no_grad():
for batch in tqdm(loader):
ids = batch['input_ids'].to(device)
mask = batch['attention_mask'].to(device)
labels = batch['label'].to(device)
outputs = model(ids, attention_mask=mask)
preds = torch.argmax(outputs.logits, dim=1)
correct += (preds == labels).sum().item()
total += labels.size(0)
print(f"\n✅ Evaluation Finished!")
print(f"📈 Accuracy: {100 * correct / total:.2f}%")
You might notice that while the parameter count is lower, the peak VRAM usage during inference can be slightly higher than the baseline.
Why? This is a software artifact, not a hardware limitation.
@misc{ykae-monarch-bert-mnli-2026,
author = {Yusuf Kalyoncuoglu, YKAE-Vision},
title = {Monarch-BERT-MNLI: Extreme Compression via Monarch FFNs},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face Model Hub},
howpublished = {\url{https://huggingface.co/ykae/monarch-bert-base-mnli}}
}