Downloads · 30 days
54
4% of all-time downloads
MicheRomChis/micro-terse
micro-terse is a text generation model from MicheRomChis. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as apache-2.0.
<div align="center" <picture <img src="https://raw.githubusercontent.com/michelangeloromerochisco/micro-terse/main/resources/logo.png" width="30%" alt="Micro-Terse" </picture </div
Downloads · 30 days
54
4% of all-time downloads
All-time downloads
1.3K
Public
Repo size
572 MB
Likes
0
Public
Click a slice to open those files.
.gguf572 MB · 100%
From the Hugging Face model README
Micro-Terse is a 423M-parameter (≈320M active) ternary-weight language model trained from
scratch for ≈$150, deployable as a 182 MB CPU-only GGUF. Its weights are constrained to
{−1, 0, +1} (≈1.58 bits), so TQ2_0 packs them exactly; the released 182 MB file pairs that with a Q6_K tied embedding.
It is a research proof-of-concept, not a production assistant. At an 8B-token budget it is data-limited: fluent for a clause or two, near chance on knowledge benchmarks. The point is capability per megabyte and per joule — a from-scratch ternary model an individual can train and run on owned hardware.
{−1, 0, +1} on all internal projections.| File | Stage | Best for |
|---|---|---|
terse-micro-base.TQ2_0.gguf | Pretrained LM | next-token prediction / completion |
terse-micro-sft.TQ2_0.gguf | Supervised fine-tuned | chat (most fluent) |
terse-micro-orpo.TQ2_0.gguf | ORPO-aligned | identity-aligned responses |
| Property | Value |
|---|---|
| Total / active parameters | ≈423 M / ≈320 M (MoE top-2) |
| Layers / hidden | 12 / 1024 |
| Attention | GQA 8 query / 2 KV heads (4:1), head dim 128, QK-Norm before RoPE (θ=500000) |
| FFN | 2816 intermediate, squared-ReLU gated |
| MoE | 4 experts, top-2, odd layers; aux-loss-free bias-EMA balancing |
| MTP | 1 head (training only, dropped at inference) |
| Embeddings | tied input/output, full precision (~31% of params) |
| Tokenizer | Llama-3.1 (128,256 vocab) |
| Context | 4096 |
| Stage | Details |
|---|---|
| Pretraining | 8B tokens FineWeb-Edu; AdamW; LR 3e-4 → 3e-5 cosine; 488,282 steps; bf16; MTP aux 0.1 |
| SFT | 3 epochs, 44,558 ChatML conversations, prompt-masked loss |
| ORPO | 1 epoch, ~3,500 identity/charter preference pairs, reference-free |
| Hardware | 1× RTX A6000 48 GB, ≈250 GPU-hours, ≈$150 total |
| Export | F32 GGUF (lossless for ternary) → TQ2_0 ≈ 182 MB |
Standard academic benchmarks (MMLU/HellaSwag/ARC) were not run; at this data budget knowledge accuracy is expected near chance. What we measured:
The model uses a custom terse architecture, so it needs the small llama.cpp fork
(branch terse-arch). After building it:
huggingface-cli download MicheRomChis/micro-terse terse-micro-sft.TQ2_0.gguf --local-dir .
./llama-cli -m terse-micro-sft.TQ2_0.gguf -p "Hello" -n 128
Use terse-micro-base.TQ2_0.gguf for completion and terse-micro-orpo.TQ2_0.gguf for
identity-aligned output.
Apache-2.0.
@techreport{romerochisco2026tersemicro,
title = {Terse-Micro: A 423M-Parameter Ternary-Weight Language Model Trained From Scratch for \$150},
author = {Romero Chisco, Michelangelo},
year = {2026},
note = {Apache-2.0. github.com/michelangeloromerochisco/micro-terse}
}