Downloads · 30 days
0
ljupco/cpu-only-inference-models
cpu-only-inference-models is a machine learning model from ljupco. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Small, efficient LLMs that run well on ordinary CPUs — no GPU needed. All models below were measured on a modest 2019-era laptop CPU:
Downloads · 30 days
0
Access
Public
Updated Aug 9, 2026
Repo size
—
Likes
1
Public
Click a slice to open those files.
.md12.5 KB · 89%
From the Hugging Face model README
Small, efficient LLMs that run well on ordinary CPUs — no GPU needed. All models below were measured on a modest 2019-era laptop CPU:
thinkpad2: Intel Core i5-8350U (4C/8T, AVX2), 64 GB DDR4-2400, AC power, 4 threads.
If your machine has no discrete GPU (or you want to keep the GPU free), these are the models and setups that actually work — with real token/s numbers, not promises.
The fastest useful model on a CPU we have found — over 28 tokens/s on a 4-core laptop.
| Model | Size | Quant | CPU decode (thinkpad2, 4 threads) |
|---|---|---|---|
| Maple Preview (20B-A1B, 256-expert MoE, 8 active) | 5.5 GiB | TQ2_0 ternary, 2.06 bpw | 33.9 t/s (tg128) · 28.2 t/s (benchy tg64) |
Maple Preview is DeepGrove's open-source reasoning model, designed from the start for efficient on-device inference (24 layers, 3:1 SWA-512:GA attention, 131k context, MIT license). It is the real star of this collection: on our CPU it runs at 28–34 tokens/s — comfortably interactive — while its 20B total / 1B-active ternary weights keep it to a 5.5 GB file that fits any machine with 16 GB of RAM.
Reference points (from the DeepGrove team and our own measurements):
The engine is the DeepGrove llama.cpp fork (the maple architecture + TQ2_0
support); the GGUF we use is their maple-preview-TQ2_0-head-Q4_K.gguf.
llama-server -m maple-preview-TQ2_0-head-Q4_K.gguf \
--ctx-size 131072 --cache-type-k q8_0 --cache-type-v q8_0 \
--threads 4 # 4 beats 8 on this CPU (33.9 vs 21.8 t/s)
The LFM2.5 models (2.6B dense, 1.2B-Thinking, 8B-A1B MoE) with mixed quantizations: the bulk of the weights at 4-bit, the most sensitive tensors kept at higher precision.
| Model | GGUF repo | Quant | CPU decode (thinkpad2, 4 threads) |
|---|---|---|---|
| LFM2.5-2.6B | ljupco/LFM2.5-2.6B-GGUF | Q4_K_M | 13.1 t/s (benchy tg64) |
| LFM2.5-1.2B-Thinking | ljupco/LFM2.5-1.2B-Thinking-GGUF | Q4_0h | 25.3 t/s (benchy tg64) |
| LFM2.5-8B-A1B | ljupco/LFM2.5-8B-A1B-GGUF | Q4_0h | 13.8 t/s (benchy tg64) |
Original models by Liquid AI — LFM2.5-2.6B, LFM2.5-1.2B-Thinking, LFM2.5-8B-A1B.
llama-server -m LFM2.5-2.6B-Q4_K_M.gguf --threads 4
These numbers come from a systematic porting and benchmarking exploration across three
engines — stock llama.cpp (with the DeepGrove fork for Maple), ik-llama.cpp, and
vllm.cpp — including kernel-level work (fused ops, integer-dot kernels, a ternary
gemv) and a detailed analysis of why the DeepGrove 8x8 gemv cannot run on
standard-quantized weights:
The LFM2.5 / Maple Preview three-engine report
This collection is entirely built on the work of others, and we are deeply grateful:
Any remaining errors are ours. Benchmark numbers are single-machine measurements; expect ±10–20% day-to-day noise.