Downloads · 30 days
3.3K
84% of all-time downloads
infosave/LFM2.5-cmf
LFM2.5-cmf is a text generation model from infosave. Use it when you need the model to write or continue text. It is set up for cortiq. The card lists the license as other.
LiquidAI's LFM2.5 is a hybrid: most layers mix with a short gated convolution and only a few carry full attention, so the state a token needs is small and constant where the convolution runs. These are those checkpoin…
Downloads · 30 days
3.3K
84% of all-time downloads
All-time downloads
4K
Public
Repo size
6 GB
Likes
1
Public
Click a slice to open those files.
.cmf6 GB · 100%
From the Hugging Face model README
cargo install cortiq-cli
hf download infosave/LFM2.5-cmf lfm2.5-2.6b-q4tp.cmf --local-dir .
cortiq run lfm2.5-2.6b-q4tp.cmf --prompt "Explain what a Fourier transform does, in two sentences."
LiquidAI's LFM2.5 is a hybrid: most layers
mix with a short gated convolution and only a few carry full attention, so the
state a token needs is small and constant where the convolution runs. These are
those checkpoints in the CMF container —
a single memory-mapped file read by cortiq, a Rust binary with no ML framework
under it. GPU via Vulkan/Metal/DX12 with a CPU fallback; NVIDIA, AMD, Intel and
Apple silicon read the same file.
| file | params | layers | size |
|---|---|---|---|
lfm2.5-230m-q4tp.cmf | 0.23B | 14 (5 attention / 9 conv) | 132 MB |
lfm2.5-2.6b-q4tp.cmf | 2.70B | 30 (8 attention / 22 conv) | 1.43 GB |
lfm2.5-8b-a1b-q4tp.cmf | 8.3B total, 1B active | 24 (6 attention / 18 conv), 32 experts, 4 per token | 4.6 GB |
All three are 4-bit tiled with ladder scales (q4tp), quantized straight from
the bf16 checkpoints.
Steady-state decode, cortiq bench --core, single stream, cortiq 0.5.99+
(the whole-token graph learned this family's short-conv mixer and its
sigmoid-routed MoE in 0.5.98/0.5.99 — earlier versions decode it an order
of magnitude slower).
| A100 80GB (Vulkan) | Apple M4 | |
|---|---|---|
| 230M | 390 tok/s | 138 tok/s |
| 2.6B | 141 tok/s | 42 tok/s (CMF_GPU=0) |
| 8B-A1B | 124 tok/s | — |
On a discrete card the whole token runs as one submitted graph, conv ring
and expert routing included; the MoE's greedy output is token-identical to
the CPU path. On Apple silicon the engine measures both arms at startup and
picks; for the 2.6B the host arm wins there, and CMF_GPU=0 pins it.
cortiq serve lfm2.5-2.6b-q4tp.cmf --port 8080
Speaks the OpenAI API, so anything that talks to OpenAI talks to it:
curl localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"lfm2.5","messages":[{"role":"user","content":"Say hello"}]}'
Each .cmf ships a .sha256 beside it.