Downloads · 30 days
21
30% of all-time downloads
AXERA-TECH/MiniCPM5-1B-AX637
MiniCPM5-1B-AX637 is a text generation model from AXERA-TECH. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Ready-to-run, text-only deployment package for openbmb/MiniCPM5-1B on an AX637 aarch64 board.
Downloads · 30 days
21
30% of all-time downloads
All-time downloads
71
Public
Repo size
1.4 GB
Likes
1
Public
Click a slice to open those files.
.axmodel954 MB · 70%
From the Hugging Face model README
Ready-to-run, text-only deployment package for
openbmb/MiniCPM5-1B on an
AX637 aarch64 board.
axllm binary with OpenAI-compatible HTTP API and CLI.kv_cache_len=1024, prefill_len=128, and maximum
prefill capacity 896 tokens..axmodel files, post-processing .axmodel,
embedding weights, tokenizer, runtime configuration, and bin/axllm.This is a text-only package. The packaged configuration has
enable_thinking=false.
Measurements below were taken on an AX637 board with the packaged runtime.
TTFT means time to first generated token.
| Scenario | Input tokens | Prefill chunks | TTFT | Decode |
|---|---|---|---|---|
| Long text validation | 846 | 7 | 6600.26 ms | 6.23 tokens/s |
The long validation request generated five tokens and exercised every shipped
prefill history group: 0, 128, 256, 384, 512, 640, and 768.
Actual latency depends on board memory pressure, prompt length, and output
length.
| Item | Value |
|---|---|
| Package size on disk | 1.7 GiB |
| Decoder layers | 24 |
| CMM used after full model startup | 966 MB |
OS memory used by axllm after full model startup (RSS) | 59,856 KiB (58.5 MiB) |
axllm virtual address space (VmSize, mostly mmap mappings) | 1,979,888 KiB (1.89 GiB) |
The CMM figure is the AX Engine CMM-pool delta measured from before startup to
after all 24 decoder layers and the post model were loaded. The OS-memory
figure is the board-side process resident set size (VmRSS) after the same
startup point. VmSize is shown separately because the package uses memory
mapping; it is virtual address space, not physical OS memory. CMM and RSS are
the startup consumption values to use when budgeting a board, while actual
system availability depends on other workloads.
| Setting | Packaged value |
|---|---|
| KV cache length | 1024 tokens |
| Prefill chunk length | 128 tokens |
| Maximum prefill length | 896 tokens |
| Prefill history capacities | 0, 128, 256, 384, 512, 640, 768 |
Prompts longer than 128 tokens are split into chunks. The runner selects the smallest compatible prefill group for each chunk. Leave room inside the 1024-token KV window for generated tokens when sending long prompts.
.
├── README.md
├── bin/
│ ├── axllm
│ └── axllm.version.json
├── config.json
├── post_config.json
├── minicpm5_tokenizer.txt
├── model.embed_tokens.weight.bfloat16.bin
├── llama_p128_l0_together.axmodel
├── ...
├── llama_p128_l23_together.axmodel
└── llama_post.axmodel
This is a flat runtime package. Run axllm from the package root; it reads the
root-level tokenizer, configuration, embedding, and .axmodel files directly.
Download this repository on the host that will transfer or mount it on the board:
mkdir -p AXERA-TECH/MiniCPM5-1B-AX637
cd AXERA-TECH/MiniCPM5-1B-AX637
hf download AXERA-TECH/MiniCPM5-1B-AX637 --local-dir .
The package includes a validated AX637 axllm binary. From the package root:
chmod +x ./bin/axllm
export LD_LIBRARY_PATH=/opt/lib:${LD_LIBRARY_PATH:-}
./bin/axllm serve . --port 8000
The service exposes:
GET http://<board-ip>:8000/health
GET http://<board-ip>:8000/v1/models
POST http://<board-ip>:8000/v1/chat/completions
Expected model identifier:
AXERA-TECH/MiniCPM5-1B-AX637-C128-P896-CTX1024
Verify readiness:
curl http://127.0.0.1:8000/health
curl http://127.0.0.1:8000/v1/models
Example health response:
{
"concurrency": 0,
"max_concurrency": 1,
"status": "healthy"
}
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "AXERA-TECH/MiniCPM5-1B-AX637-C128-P896-CTX1024",
"messages": [
{
"role": "user",
"content": "中国的首都是哪里?请只回答城市名。"
}
],
"max_tokens": 32,
"temperature": 0
}'
The response uses the standard OpenAI chat-completions JSON shape. Set the
OpenAI client base URL to http://<board-ip>:8000/v1 and use the model
identifier shown above.
export LD_LIBRARY_PATH=/opt/lib:${LD_LIBRARY_PATH:-}
./bin/axllm run .
Type /q or /exit to leave the interactive session.
If you need the original model files or want to rebuild the deployment artifacts, start with:
139953715