Downloads · 30 days
504
35% of all-time downloads
mikecovlee/tinymixtral-v1.0
tinymixtral-v1.0 is a text generation model from mikecovlee. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
The first TinyMixtral release, trained on C4-en. Kept for historical reference and reproducibility; superseded by v3.0.
Downloads · 30 days
504
35% of all-time downloads
All-time downloads
1.5K
Public
Repo size
3.5 GB
Likes
1
Public
Click a slice to open those files.
.bin1.7 GB · 100%
From the Hugging Face model README
The first TinyMixtral release, trained on C4-en. Kept for historical reference and reproducibility; superseded by v3.0.
Same trunk architecture as v1.1 (the data-quality ablation changed only the data): ~432M total parameters (~176M active, top-2 of 6 routed experts), hidden 896, 10 layers, GQA 14 Q / 2 KV heads (head dim 64), 32k vocabulary, 2048 context.
Pretrain (4B tokens, C4-en).
| Parameter | Value |
|---|---|
| Data | C4-en |
| Batch size | 22 |
| Sequence length | 1,024 |
| Tokens/step | 22,528 |
| Steps | 177,557 |
| Learning rate | 3e-4 |
| Warmup steps | 2,000 |
| Weight decay | 0.1 |
| Grad clip | 1.0 |
| Time | ~77 h |
Post-train (1B tokens). Continue from the C4 checkpoint on higher-quality data (FineWeb-Edu + Cosmopedia v2, 50:50):
| Parameter | Value |
|---|---|
| Data | FineWeb-Edu + Cosmopedia v2 (50:50) |
| Tokens | 1B |
| Steps | 44,390 |
| Learning rate | 5e-5 |
| Warmup steps | 300 |
| Time | ~20.8 h |
Measured with lm-evaluation-harness v0.4.12, 0-shot (cuda, bf16).
| Task | Metric | v1.0 (C4 4B) |
|---|---|---|
| HellaSwag | acc_norm | 0.310 |
| PIQA | acc | 0.613 |
| WinoGrande | acc | 0.508 |
| ARC-Easy | acc | 0.422 |
| ARC-Challenge | acc_norm | 0.247 |
| OpenBookQA | acc_norm | 0.308 |
| BoolQ | acc | 0.579 |
| LAMBADA | acc | 0.240 |
| Mean | — | 0.403 |
These serve as the (weak) baseline for the data-quality ablation that produced v1.1.
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("mikecovlee/tinymixtral-v1.0", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("mikecovlee/tinymixtral-v1.0")
Requires transformers and trust_remote_code=True (custom tinymixtral architecture).
mikecovlee/tinymixtral — 477.5M MoE base (v3.0 flagship)mikecovlee/tinymixtral-it — MoE instruction-tuned (v3.0-it, 3M SFT)mikecovlee/tinymistral-276m — 276M dense base (iso-active ablation)mikecovlee/tinymistral-276m-it — 276M dense instruction-tuned (3M SFT)mikecovlee/tinymixtral-v1.1-1b — 1B MoE (earlier flagship)mikecovlee/tinymixtral-v1.1-0.5b — 0.5B-class MoE (data-quality ablation)mikecovlee/tinymixtral-v2.0-beta — shared-expert experiment (beta)mikecovlee/tinymixtral-v1.0 — this model (legacy, C4)Naming. The MoE family is published under
tinymixtral; the dense 276M iso-active ablation companions use thetinymistralspelling. Both belong to the same project.
@misc{tinymixtralv1.02026,
title = {TinyMixtral: a small Mixture-of-Experts language-model family},
author = {Michael Lee},
year = {2026},
howpublished = {\url{https://huggingface.co/mikecovlee/tinymixtral-v1.0}}
}
MIT (Copyright (C) 2026 Michael Lee).