Downloads · 30 days
0
EschaLabs/escha-runtime-qwen3moe
escha-runtime-qwen3moe is a machine learning model from EschaLabs. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for sglang. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Sep 20, 2026
Repo size
2.2 GB
Likes
12
Public
Click a slice to open those files.
.gz2.1 GB · 99%
From the Hugging Face model README
qwen3moeServing runtimes for Escha 2-/3-bit (eschamoe) quantized models of the
qwen3_5_moe architecture (Qwen3.5 / Qwen3.6 Mixture-of-Experts, 256 experts). One repo
per model architecture, one directory per engine — pick the engine that fits your workload:
SGLang — sglang/ | ZML — zml/ | |
|---|---|---|
| Best for | servers & teams | one user, one stream, zero-dependency deploys |
| Concurrency | continuous batching, paged KV, radix prefix cache | one request at a time |
| Single-user decode (4090)1 | 218–231 tok/s (flat in output length) | 235–241 tok/s on ≥500-token answers (+8–14%); ~190 tok/s on ≤128-token replies |
| TTFT, 128→1890-tok prompt1 | 0.06–0.24 s | 0.06–0.35 s |
| Max context (4090)2 | up to 159k tok @ MEM=0.78, 235k @ 0.90 (ships CTXLEN=32768) | 262,144 @ ESCHA_MEM=0.93 (ships ESCHA_CTX=1024) |
| Multi-turn prompt reuse | radix cache (branching, cross-session) | append-only, single session |
| Tool calls / JSON schema / thinking parser | yes | no |
Sampling (temperature>0) | full speed | ~104 tok/s (fast path is greedy-only) |
| Install | Python 3.12 venv + CUDA-12 PyTorch | one binary, no Python, no CUDA toolkit |
Neither engine is uniformly faster. ZML wins sustained decode on long single-stream answers and is far simpler to deploy; SGLang wins short replies, wins time-to-first-token on long prompts, and is the only option for concurrency, tool calling or structured output. If unsure: multi-user, agents-at-scale, or structured output → SGLang; long-form single-user generation on your own GPU → ZML.
The ZML engine is built on ZML (Zig + MLIR/XLA); the SGLang engine on a fork of SGLang. Both run the same Escha CUDA kernels and serve the same model files.
Both engines above are NVIDIA/Linux only. On a Mac, use escha-mlx (separate repo, Apache-2.0): an MLX runtime with Metal kernels, continuous batching, prefix caching and the same OpenAI-compatible endpoint. It serves the same model files with no conversion step.
Requires Apple Silicon M1–M5, macOS 14+, Python 3.10–3.13 and 24 GB of unified memory. Measured resident footprint is 12.25 GB; single-stream decode is 27.3 tok/s on an M4 base and 59.7 tok/s on an M5 Pro, with aggregate throughput of 185.6 and 539.3 tok/s respectively at batch 128. Those are not comparable to the CUDA figures above: an M4 base moves ~120 GB/s against a 4090's 1008 GB/s, and decode here is memory-bound. Full tables, install steps and a head-to-head against stock MLX 4-bit are in that repo.
| Model repo | Bits |
|---|---|
| EschaLabs/Qwen3.6-35B-A3B-Escha-W2 | 2-bit (eschamoe) |
These runtimes target
qwen3_5_moe. A model of a different architecture will not load — use the matchingescha-runtime-<arch>repo.
SGLang engine (full detail: sglang/INSTALL.md, incl. the per-GPU cookbook):
python3.12 -m venv .venv && source .venv/bin/activate
pip install -U pip wheel
pip install "torch>=2.9" --index-url https://download.pytorch.org/whl/cu128 # cu12 torch first
pip install ./sglang/escha-*.whl # pulls the bundled sglang fork + its full dep closure
hf download EschaLabs/Qwen3.6-35B-A3B-Escha-W2 --local-dir ./Qwen3.6-35B-A3B-Escha-W2
MODEL=./Qwen3.6-35B-A3B-Escha-W2 bash sglang/serve.sh
ZML engine (full detail: zml/INSTALL.md) — no Python at all:
tar xf zml/escha-zml-serve-*-linux-x86_64.tar.gz && cd escha-zml-serve-*
hf download EschaLabs/Qwen3.6-35B-A3B-Escha-W2 --local-dir ./model
./escha-serve
ESCHA_CTX reaches the model's native 262,144 tokens. Measured on one 4090 with facts
planted at 10%, 50% and 90% depth and all three asked for at the end — the check that
positions and recurrent state thread correctly across the whole prompt:
| prompt | prefill | rate | recall |
|---|---|---|---|
| 15,066 tok | 5.3 s | 2,858 tok/s | 3/3 |
| 60,066 tok | 24.9 s | 2,411 tok/s | 3/3 |
| 129,966 tok | 79.0 s | 1,646 tok/s | 3/3 |
| 261,966 tok | 259.1 s | 1,011 tok/s | 3/3 |
Prefill is ~N²/2 work, so the rate falls with length. The comfortable band on a 24 GB card is ~8k–130k, where TTFT is seconds; 262k answers correctly but takes ~4.3 minutes to read.
Multi-turn: a conversation whose prompt grows each turn reuses the previous prefix. On a 14k-token conversation the first turn takes 4.11 s and each following turn 1.27 s (~99.8% of the prompt reused). Reuse requires an exact append — editing earlier history re-reads the prompt — and only one conversation is cached. SGLang's radix cache is faster in absolute terms here and handles branching and multiple sessions; the ZML cache exists to make single-session agent loops practical, not to match it.
Decode slows as position grows on both engines, because decode attention is O(position). On ZML, measured with prefill excluded: 227 tok/s at position ~1k, 163 at 3k, 124 at 7.5k. (That measurement divides total request time and so reads ~5% below the streamed grid above — method, not configuration.)
sglang/INSTALL.md → Running on your GPU.sglang-kernel from
sglang's CUDA-12 index first; see
sglang/INSTALL.md → Supported PyTorch versions.DETERMINISTIC=1 fails on consumer Blackwell (sm_120). The deterministic attention kernel
requests 104 KB of shared memory per block, above the sm_120 limit, and the server exits during
startup. It works on Ampere, Ada and Hopper. Greedy output is in any case not bit-reproducible
across requests on either engine — batch composition changes fp16 accumulation order, so a
near-tie can flip and a long reasoning chain diverges from there./v1/completions (HTTP 404). Use /v1/chat/completions.model field is served anyway instead
of returning model_not_found, and usage.prompt_tokens is wrong when a prompt is truncated.kill -9 rather
than returning an error. This is why it asks for 24 GB — use the SGLang engine on 16 GB, where
the same prompt returns a clean HTTP 400.Everything here is released under the Apache License, Version 2.0 — see LICENSE.
All bundled third-party code is permissive (Apache-2.0 / MIT / BSD-3-Clause; NVIDIA runtime
libraries in the ZML bundle under the NVIDIA EULA's redistributable-runtime terms) — no
copyleft. Full texts and the component inventory:
THIRD_PARTY_LICENSES/. Model weights are not in this repo and carry
their own license in the model repository.
Same-harness measurement, 2026-07-26: one RTX 4090, the same model files, the same streaming OpenAI client, greedy, batch 1, medians of 2 reps over an ISL×OSL grid (128–1890 in, 128–1890 out). Decode = steady-state tokens/s between the first and last streamed token. Per-cell decode: SGLang 176.8–230.9 (median 218.2), ZML 184.6–240.7 (median 234.8); ZML leads every cell with ≥256 output tokens and trails on 128-token replies, where its 16-token fused decode chunk dominates. TTFT by input length (ZML / SGLang): 128 tok 55/55 ms · 500 tok 95/91 ms · 1000 tok 178/143 ms · 1890 tok 353/237 ms. ↩ ↩2
Measured 2026-07-26 on one RTX 4090 (24 GB) with Qwen3.6-35B-A3B-Escha-W2. Both
engines ship a conservative default context and let you raise it. SGLang reports its KV
pool at startup; launched with CTXLEN=262144 it allocated 159,480 tokens at its shipped
MEM=0.78 and 234,796 at MEM=0.90 — that pool size is the ceiling for a single
sequence. ZML serves a 261,966-token prompt at ESCHA_CTX=262144. Both are
memory-bound near the cap, not architecture-bound: KV on this hybrid model is only
~20 KiB per token (10 attention layers; the 30 gated-delta-net layers hold a fixed
~66 MB recurrent state), so 262,144 tokens is 5.37 GB beside 12.3 GB of weights.
In fairness to SGLang: its 234,796-token pool at MEM=0.90 still left 1.38 GB
of VRAM free, so it would very likely also reach the 262,144 cap at a higher
mem-fraction-static — we did not test that, and 0.78 is simply the conservative
default the runtime ships. Read this row as "both engines reach the model's native
context on a 24 GB card", not as a ZML advantage. ↩