Downloads · 30 days
28
14% of all-time downloads
majentik/MiniCPM5-1B-Base-RotorQuant-GGUF-Q8_0
MiniCPM5-1B-Base-RotorQuant-GGUF-Q8_0 is a text generation model from majentik. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as apache-2.0.
[!TIP] KV-cache quantization (upstream, no fork needed): llama.cpp/Ollama cover this natively — -ctk q80 -ctv q80 (~half KV memory, negligible quality loss) or -ctk q40 -ctv q40 (~quarter memory, small quality cost).…
Downloads · 30 days
28
14% of all-time downloads
All-time downloads
202
Public
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.gguf1.2 GB · 100%
From the Hugging Face model README
[!TIP] KV-cache quantization (upstream, no fork needed): llama.cpp/Ollama cover this natively —
-ctk q8_0 -ctv q8_0(~half KV memory, negligible quality loss) or-ctk q4_0 -ctv q4_0(~quarter memory, small quality cost). In Ollama:OLLAMA_KV_CACHE_TYPE=q8_0withOLLAMA_FLASH_ATTENTION=1.
openbmb/MiniCPM5-1B-Base quantized pack, published as MiniCPM5-1B-Base-RotorQuant-GGUF-Q8_0.
llama.cpp Q8_0 quantization.
Released under the RotorQuant line. RotorQuant and TurboQuant are this project's release labels for this pack, not distinct quantization algorithms — both brand repos carry byte-identical weights. No brand-specific speedup is claimed or measured.
pipeline_tag: text-generation.
This is a llama.cpp GGUF conversion of the text tower; no modality beyond pipeline_tag above is claimed or included.
This pack is a derivative of openbmb/MiniCPM5-1B-Base; all credit for the original model, training, and weights belongs to the upstream authors. This repo republishes a quantized conversion of those weights only.
Governed by the apache-2.0. See the upstream repo and the linked license for the full terms — no license text is reproduced here.