Downloads · 30 days
163
52% of all-time downloads
litert-community/Zamba2-2.7B-instruct
Zamba2-2.7B-instruct is a text generation model from litert-community. Use it when you need the model to write or continue text. It is set up for litert-lm. The card lists the license as apache-2.0.
Measured on device (edge-compat): Raspberry Pi 5 · LiteRT-LM 0.16.1 · CPU, 4 threads · decode 1.0 tok/s · prefill 15 tok/s · TTFT 18.46 s (2026-09-02). Record: https://github.com/john-rocky/edge-compat/blob/main/cards…
Downloads · 30 days
163
52% of all-time downloads
All-time downloads
312
Public
Repo size
8.4 GB
Likes
0
Public
Click a slice to open those files.
.litertlm2.8 GB · 100%
From the Hugging Face model README
Measured on device (edge-compat): Raspberry Pi 5 · LiteRT-LM 0.16.1 · CPU, 4 threads · decode 1.0 tok/s · prefill 15 tok/s · TTFT 18.46 s (2026-09-02). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/zamba2-2.7b-instruct/CARD.md
Zyphra/Zamba2-2.7B-instruct converted to the LiteRT-LM (.litertlm) format for on-device inference with Google's LiteRT-LM runtime. Requires litert-lm ≥ 0.15.
Zamba2-2.7B is Zyphra's shared-attention hybrid at its middle size: a Mamba2 selective-scan backbone (54 layers) with two shared transformer blocks applied alternately at 9 interleaved positions — two sets of attention+MLP weights, each reused at its positions and specialized by per-position LoRA adapters on the MLP, attending over the concatenation of the running hidden state and the original embeddings (no rotary embedding at this size). Together with our Zamba2-1.2B this is, to our knowledge, the first Zamba2 conversion to a mobile runtime.
| File | Recipe | Size |
|---|---|---|
Zamba2-2.7B-instruct_int8.litertlm | int8 dynamic on linears + embedding (convs and the scan stay float); fp32 activations declared for GPU | 2.80 GB |
2026-09-21: chat template updated to accept the 0.18 content-parts form (string form unchanged); weights, tokenizer and executor metadata byte-identical.
A conversion-side note: current transformers (5.14.x, and main as of 2026-08-14) cannot load ANY two-block Zamba2 checkpoint (2.7B/7B) — Zamba2Model.get_layers assigns block_id by global layer index while the checkpoint layout and its own weight-tie cycle follow hybrid occurrence order, so construction raises before weights load. The conversion patch carries a structure fix (block_id by hybrid order), verified key-exact against the published checkpoint.
litert-lm run ./Zamba2-2.7B-instruct_int8.litertlm --prompt "What is the capital of France? Answer in one word."
# GPU
litert-lm run ./Zamba2-2.7B-instruct_int8.litertlm --backend gpu --cache no --prompt "..."
Six prefill signatures (1024, 256, 64, 16, 4, 1) are exported so the runtime picks tight chunks. A signature costs memory whether or not it is called, so the ladder is deliberately shorter than the full 1–1024 one. The bundle carries the tokenizer and the stock ChatML Zamba2 chat template.
litert-lm benchmark (litert-lm 0.16.0), Apple M4 Max, -p 256 -d 256 --runs 3 --cache no, quiet machine:
| Backend | Prefill (256) | Decode | TTFT |
|---|---|---|---|
| GPU | 603 tok/s | 43.6 tok/s | 0.45 s |
| CPU | 282 tok/s | 14.7 tok/s | 0.97 s |
On device (cold start, single run, 140-token composite prompt, quality harness):
| Device | Backend | Prefill | Decode | TTFT | Peak memory |
|---|---|---|---|---|---|
| iPhone 17 Pro (12 GB) | CPU | 14.3 tok/s | 3.6 tok/s | 10.2 s | 2.19 GB |
| iPhone 17 Pro (12 GB) | GPU (Metal) | — does not load, see below — |
Honest notes:
Zamba2-2.7B-instruct_int8.litertlm does not run on the GPU. Engine creation hard-rebooted the device, twice in a row, with no runtime log either time. Smaller bundles gated between the two attempts passed under identical conditions, so this is specific to this file. Do not run it on the GPU backend.
| file | GPU backend | delegation | peak |
|---|---|---|---|
Zamba2-2.7B-instruct_int8.litertlm | does not run | engine creation rebooted the phone | — |
Measured on a Samsung Galaxy S26 (SM-S942Q / SM8850, Android 16) with litert_lm_advanced_main from litert-lm 0.16.0, --backend=gpu --sampler_backend=cpu, prompt What is the capital of France?. Gated 2026-08-25.
The 1.2B model in this family runs on the same phone, so this is a size wall on this hardware, not the Zamba2 architecture refusing the GPU.
GPU wiring, including the Gallery import toggle: GPU guide.
Converted with litert-torch plus a hybrid-cache patch (reproduction script + patch: hf-to-litertlm zamba2_work/):
BROADCAST_TO, no int64 index math) — this is what makes the graph fully delegable on GPU.block_id is assigned by hybrid occurrence order (matching the checkpoint layout and the tie cycle) — without this, current transformers cannot construct the model at all.softplus(dt) at time_step_min with no upper clamp; padded prefill positions are forced to exact identity steps AFTER the clamp (without this, every runtime pad token decays the recurrent state).Strip decoder is removed from the bundle — Zamba2's metaspace (SP-BPE) tokenizer otherwise loses every interior space under the runtime's per-token streaming decode; the only behavior change is a sequence-initial space, which the runtime trims.The LiteRT-LM engine prepends the metadata start_token to every prompt, and this model's chat template already renders <|im_start|> itself — so the model was reading <|im_start|><|im_start|>…, a stream it was never trained on. The start token has been removed.
Metadata-only change: every section of the bundle except the LlmMetadata block is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the tokenizer are unchanged and the speed and memory numbers on this card still describe exactly this file — only the file's own sha256 differs. What changed is the input: the token stream the model reads for a given conversation can differ from the previous file's, and it now matches this model's own reference chat stream. Greedy decoding can turn on a single token, so an individual answer can differ from the previous file in either direction. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file and have not been re-measured on this one. If you downloaded before 2026-08-28, re-download.
Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with litert-lm benchmark 0.16.1: CPU backend, 4 threads, 256 prefill + 256 decode tokens, --cache memory (the compile cache lives and dies with the process, so every invocation compiles the model from scratch; nothing is reused between runs), one warm-up plus one timed iteration per invocation, 3 invocations per file with cooldown in between. Values are the median across invocations (min–max in parentheses). No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0). Every file listed produced coherent text in a real generation on this backend before its numbers were recorded.
| File | Prefill (tok/s) | Decode (tok/s) | TTFT | Peak RSS |
|---|---|---|---|---|
Zamba2-2.7B-instruct_int8.litertlm | 14.6 (14.6–14.7) | 1.0 (1.0–1.0) | 18.5 s | 4.9 GB |