Downloads · 30 days
987
97% of all-time downloads
infosave/Kimi-Linear-48B-A3B-Code-CMF
Kimi-Linear-48B-A3B-Code-CMF is a text generation model from infosave. Use it when you need the model to write or continue text. It is set up for cortiq. The card lists the license as mit.
A code-calibrated specialist build of moonshotai/Kimi-Linear-48B-A3B-Instruct in the CMF format — one file, mmap-served, no Python at inference:
Downloads · 30 days
987
97% of all-time downloads
All-time downloads
1K
Public
Repo size
17.7 GB
Likes
0
Public
Click a slice to open those files.
.cmf17.7 GB · 100%
From the Hugging Face model README
A code-calibrated specialist build of moonshotai/Kimi-Linear-48B-A3B-Instruct in the CMF format — one file, mmap-served, no Python at inference:
| original bf16 | CMF q4t (full) | this file | |
|---|---|---|---|
| Size | 98 GB | 27.7 GB | 17.7 GB |
| Held-out code ppl | — | 7.11 | 7.30 (+2.7%) |
| Decode, M4 MacBook 24 GB | — | 4.8 tok/s (pages) | 11.1 tok/s |
The speedup is structural: the full 27.7 GB file does not fit a 24 GB page cache and pages on every token; the specialist does fit, so the same machine decodes ×2.3 faster.
cortiq convert --model moonshotai/Kimi-Linear-48B-A3B-Instruct --quant q4t --output kimi48-q4t.cmf
The engine executes Kimi's KDA (Kimi Delta Attention: delta rule
with per-channel decay, per-projection short convolutions,
sigmoid-gated output norm), NoPE MLA full-attention layers, and
the sigmoid MoE router with its selection bias. The tiktoken rank
table becomes a standard tokenizer.json at convert time.CMF_MOE_STATS=stats.json — the engine records per-layer expert
routing frequencies. On code, the top 64 of 256 experts carry 73%
of the routing mass.cortiq moe-defrag kimi48-q4t.cmf --stats stats.json --cover 0.95 --output kimi48-code.cmf
Per layer, the smallest expert set covering 95% of the recorded
routing mass is kept (~160 of 256); experts are renumbered into a
dense prefix and the router rows AND the noaux selection bias are
sliced to match. Runtime semantics equal the runtime expert mask —
the ppl of the cut file is bit-identical to masking the full file.The expert restriction is task-shaped: this file is at its best on code and technical text. For general-purpose use, convert the full model yourself with the command above (30.7 GB of free disk is enough).
cargo install cortiq-cli # pure Rust, no Python
hf download infosave/Kimi-Linear-48B-A3B-Code-CMF kimi48-code.cmf --local-dir .
cortiq run kimi48-code.cmf --prompt "Write a Python function that returns the n-th Fibonacci number iteratively." --max-tokens 120
cortiq serve kimi48-code.cmf # OpenAI-compatible API
MIT, inherited from the base model. Weights © Moonshot AI; this repackaging only changes the storage format and the served expert set.
cargo install cortiq-cli).