Downloads · 30 days
701
100% of all-time downloads
RASMUS/MiniCPM5-2B-ONNX
MiniCPM5-2B-ONNX is a text generation model from RASMUS. Use it when you need the model to write or continue text. It is set up for transformers.js. The card lists the license as apache-2.0.
An ONNX export of openbmb/MiniCPM5-2B that runs in a browser on WebGPU through Transformers.js.
Downloads · 30 days
701
100% of all-time downloads
All-time downloads
701
Public
Repo size
1.8 GB
Likes
1
Public
Click a slice to open those files.
.onnx_data1.8 GB · 99%
From the Hugging Face model README
An ONNX export of openbmb/MiniCPM5-2B
that runs in a browser on WebGPU through
Transformers.js.
At the time this was built, no export of the 2B existed that a browser could
load. The official ONNX release covers the 1B only, and it is a CPU-shaped fp16
build (MultiHeadAttention, no MatMulNBits, no config.json) in a layout
Transformers.js does not resolve.
| File | Size |
|---|---|
onnx/model_q4f16.onnx | 316 KB (graph) |
onnx/model_q4f16.onnx_data | 1.83 GB (weights) |
Single variant, q4f16: int4 weights (MatMulNBits) with fp16 graph I/O.
A GPU without the shader-f16 feature cannot run this.
The embedding table is not quantised, because MiniCPM5-2B sets
tie_word_embeddings: false and the model builder quantises that table only for
tied models. Those 267 M parameters cost about 535 MB of the total.
import { pipeline } from '@huggingface/transformers';
const generator = await pipeline('text-generation', 'RASMUS/MiniCPM5-2B-ONNX', {
device: 'webgpu',
dtype: 'q4f16',
});
const output = await generator(
[{ role: 'user', content: 'Explain WebGPU in two sentences.' }],
{ max_new_tokens: 256 },
);
RTX 4080 Laptop, Chromium 152 on Windows, @huggingface/transformers 4.2.0 over
onnxruntime-web WebGPU. Warmed runs, first run at each prompt length discarded.
| Measurement | Value |
|---|---|
| Prefill, 1024-token prompt | 0.395 ms / prompt token |
| Prefill, 512-token prompt | 0.411 ms / prompt token |
| Prefill, 128-token prompt | 0.611 ms / prompt token |
| Decode | 24.1 ms / token (41.6 tok/s) |
| First token, 64-token prompt | 199 ms |
| Session load, warm cache | 6.5 s |
These are one GPU's numbers, not a general claim.
python -m onnxruntime_genai.models.builder \
-i <local openbmb/MiniCPM5-2B> -o out -p int4 -e webgpu -c cache
onnxruntime-genai 0.15.2. -e webgpu is what selects the
GroupQueryAttention branch — the same build on the cpu provider silently
emits MultiHeadAttention instead, which is the defect in the published 1B
export. The graph is 473 nodes: MatMulNBits x211, GroupQueryAttention x42,
no Scan/Loop, and no position_ids input, because GQA fuses RoPE.
Three changes were then needed to make the output loadable by Transformers.js. Each is recorded because each fails in a way that does not name its cause.
kv_cache_dim. Transformers.js builds the first empty cache from
the session's input metadata and resolves any symbol it does not know to 0,
so GroupQueryAttention rejected the tensor on the first forward:
Input 'past_key' dimension 3 should be same as head_size, got 0 expected 128.tokenizer_config.json. Transformers.js
reads chat_template.jinja only for multimodal processors; a text tokenizer
reads tokenizerConfig.chat_template and nothing else.min_count = [...]|min, and @huggingface/jinja has no min filter for
arrays, so every assistant turn carrying tool_calls failed to render. The
variable is never read. chat_template.jinja here carries the patched text,
identical to the copy inside tokenizer_config.json.transformers.js_config in config.json declares one external-data chunk and
an fp16 KV cache, both of which Transformers.js requires and neither of which
the builder writes.
Base model © OpenBMB, Apache-2.0. This repository redistributes a converted and quantised copy of those weights under the same licence. No weights were retrained or fine-tuned.