Downloads · 30 days
0
litert-community/MiniCPM-V-4
MiniCPM-V-4 is a image-text-to-text model from litert-community. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for litert. The card lists the license as apache-2.0.
MiniCPM-V-4-int8.litertlm — an on-device LiteRT-LM build of MiniCPM-V-4, quantized to INT8 (weight-only, GPTQ + Hadamard rotation) and packaged as a single .litertlm bundle for CPU inference on Snapdragon 8850.
Downloads · 30 days
0
Access
Public
Updated Sep 16, 2026
Repo size
16.8 GB
Likes
2
Public
Click a slice to open those files.
.litertlm4.2 GB · 100%
From the Hugging Face model README
MiniCPM-V-4-int8.litertlm — an on-device LiteRT-LM build of MiniCPM-V-4, quantized to INT8 (weight-only, GPTQ + Hadamard rotation) and packaged as a single .litertlm bundle for CPU inference on Snapdragon 8850.
For the base model, architecture, capabilities and license, refer to the original model card:
https://huggingface.co/openbmb/MiniCPM-V-4
This repository only provides an edge-optimized LiteRT deployment of that model:
.litertlm (LiteRT-LM bundle: tokenizer + LLM prefill/decode + embedder + navit SigLIP vision encoder + resampler).Measured on-device (arm64 CPU, --backend=cpu), single image, 4-slice input:
| Metric | Value |
|---|---|
| Prefill speed | 40.9 tokens/sec |
| Decode speed | 16.6 tokens/sec |
| Vision encode (4 slices) | ~2.04 s |
| Time to first token | ~7.0 s (284-token prefill incl. vision) |
| Init (model load + compile) | ~2.2 s |
Full MME evaluation running the INT8 bundle on-device:
| Group | Score |
|---|---|
| Perception | 1568 |
| Cognition | 491 |
| Total | 2059 |
litert_lm_advanced_mainYou can run multimodal inference using litert_lm_advanced_main from the LiteRT-LM repository. litert_lm_advanced_main supports inline media markers (such as [image:/path/to/image.png]) within --input_prompt.
From the LiteRT-LM repository:
bazel build -c opt //runtime/engine:litert_lm_advanced_main
Using bazel run:
bazel run -c opt //runtime/engine:litert_lm_advanced_main -- \
--backend=cpu \
--vision_backend=cpu \
--model_path=/path/to/MiniCPM-V-4-int8.litertlm \
--input_prompt="What is in this image? [image:/path/to/image.png]"
Or execute the compiled binary directly:
./bazel-bin/runtime/engine/litert_lm_advanced_main \
--backend=cpu \
--vision_backend=cpu \
--model_path=/path/to/MiniCPM-V-4-int8.litertlm \
--input_prompt="What is in this image? [image:/path/to/image.png]"
--backend=cpu: Runs LLM prefill and decode on CPU.--vision_backend=cpu: Enables the vision modality and runs the SigLIP vision encoder on CPU.--input_prompt: Prompt text including image path using [image:/path/to/image.png].--max_output_tokens: (Optional) Limit the maximum number of generated tokens (e.g. --max_output_tokens=128).--min_log_severity=2: (Optional) Display info logs, such as vision executor properties (num_tokens_per_image: 64, patch_num_shrink_factor: 17).Follows the license of the base model — see https://huggingface.co/openbmb/MiniCPM-V-4.