Downloads · 30 days
89
42% of all-time downloads
88plug/MiniCPM-V-4.5-W8A16
MiniCPM-V-4.5-W8A16 is a image-text-to-text model from 88plug. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
These weights are compressed-tensors (pack-quantized / int-quantized).
Downloads · 30 days
89
42% of all-time downloads
All-time downloads
210
Public
Parameters
8.7B
10.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors10.6 GB · 100%
How the weights are stored.
I326.9B · 80%
From the Hugging Face model README
These weights are compressed-tensors (pack-quantized / int-quantized).
| Runtime | Supported |
|---|---|
| vLLM ≥ 0.21 | Yes — preferred (auto-detect CT; no --quantization flag) |
transformers + compressed-tensors | Yes for many text models; multimodal may need trust_remote_code |
| Text Generation Inference (TGI) | Not supported for these CT packs |
| Hugging Face Inference Widget | Often fails — use vLLM locally instead |
vllm serve 88plug/MiniCPM-V-4.5-W8A16 --trust-remote-code
Do not deploy via TGI — that backend does not load our CT format.
INT8 post-training quantization of openbmb/MiniCPM-V-4_5 — MiniCPM-V (vision + LLM), not MiniCPM-o (no audio / TTS). Qwen3-8B LLM + SigLIP2-400M vision + unified 3D-resampler (image, multi-image, video). Apache-2.0 base.
This is not the same product as 88plug/MiniCPM-o-4.5-W8A16.
| Property | Value |
|---|---|
| Base model | openbmb/MiniCPM-V-4_5 |
| Release tier | Pending-gold (gold path in progress) |
| Quant method | AutoRound W8A16 iters=200 (LLM Linear; vpm+resampler BF16) |
| FLAC status | Not measured (T+7d milestone) |
| Architecture | Qwen3-8B LLM + SigLIP2 vision + 3D-resampler |
| Quant format | compressed-tensors (native vLLM) |
| Quantized | LLM Linear layers (model.llm) |
| Kept BF16 | vision encoder (vpm) + 3D-resampler |
| Lab pin | vLLM v0.21.0-cu129 (sidecar 0.28 tracked; pin switch after both hosts IMAGE_OK) |
Tested target: vLLM v0.21.0 (vllm/vllm-openai:v0.21.0-cu129-ubuntu2404). Weights are compressed-tensors — vLLM detects quantization automatically. No --quantization flag.
docker run --gpus device=0 -p 8080:8080 \
vllm/vllm-openai:v0.21.0-cu129-ubuntu2404 vllm serve \
88plug/MiniCPM-V-4.5-W8A16 \
--trust-remote-code \
--max-model-len 8192 \
--gpu-memory-utilization 0.90
Requires vLLM ≥ v0.21.0. Upstream MiniCPM-V 4.5 also documents vLLM since v0.10.2.
| Component | Precision | Reason |
|---|---|---|
| LLM Linear layers | W8A16 INT8 | AutoRound iters=200, actorder=False |
| Vision encoder (SigLIP2 / vpm) | BF16 | Tower keep |
| 3D-resampler | BF16 | Tower keep |
| Embeddings, LM head, norms | BF16 | Standard practice |
No audio / Whisper / CosyVoice2 — those exist on MiniCPM-o, not MiniCPM-V.
| Metric | Status |
|---|---|
| Throughput (tok/s) | In progress — T+7d milestone |
| MMLU delta vs BF16 | In progress — T+7d milestone |
| RULER@128k | In progress — T+30d milestone |
No fabricated numbers. Results will be published to this card when measured.
@misc{yu2025minicpmv45,
title = {MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe},
author = {Tianyu Yu and others},
year = {2025},
eprint = {2509.18154},
archivePrefix = {arXiv},
url = {https://huggingface.co/openbmb/MiniCPM-V-4_5}
}
88plug AI Lab ships FLAC-target compressed-tensors quantizations for native vLLM v0.21.0+ deployment.
This release: Pending-gold — gold-path quantization in progress. Do not use datafree/RTN substitutes.
Browse all releases → huggingface.co/88plug