Downloads · 30 days
4
20% of all-time downloads
mlboydaisuke/Qwen3-4B-ExecuTorch
Qwen3-4B-ExecuTorch is a text generation model from mlboydaisuke. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Downloads · 30 days
4
20% of all-time downloads
All-time downloads
20
Public
Repo size
2.5 GB
Likes
0
Public
Click a slice to open those files.
.pte2.5 GB · 100%
From the Hugging Face model README
qwen3_4b_xnnpack_8da4w_e8.pte (2469.4 MB)
embedding_quantize: "8,0")export_llm, static shape (seq_len=1), max_seq_length 2048,
XNNPACK extended_opsllm_params/qwen3_4b_xnnpack_8da4w_e8.yamlllm_params/gen_static.py, token-by-token prefill then greedy decode, a fresh process per
prompt so no answer is read through the previous one's cache:
| prompt | answer |
|---|---|
| capital of France? | opens a <think> block and reasons before answering |
| 17 times 4? | same, working the multiplication out in the block |
Decode 28.9 tok/s, from one pass over every model on this shelf with nothing else running. That matters more than it sounds: the same file measured a quarter of its rate while an export was running alongside.
Chat template: ChatML, bos 151643, eos [151645, 151643].
Not measured on a phone.
use_sdpa_with_kv_cache is on. Upstream's qwen3_5 config leaves it off with no
reason given while the equally hybrid lfm2 config has it on; measured on Qwen3.5-2B in
one run, that is 8.20 tok/s against 16.64.dim and hidden_dim both divide by the quantizer's group size. 8da4w only touches a
linear whose in_features divide by it, and skips the rest silently — SmolLM2-135M, which
is 576 wide, came out at 475 MB against fp32's 540 with no warning at all.convert/check_params_used.py. SmolLM3 sets no_rope_layer_interval, which ModelArgs
declares and only the MLX and Qualcomm backends read, and it exports fine and then repeats
a single word forever.python llm_params/gen_static.py \
--pte qwen3_4b_xnnpack_8da4w_e8.pte \
--tokenizer tokenizer.json \
--prompt $'<|im_start|>user\nWhat is the capital of France?<|im_end|>\n<|im_start|>assistant\n' \
--eos_ids "[151645, 151643]"
The 8-bit embedding needs from executorch.kernels import quantized before the program is
loaded, and portable_lib._load_for_executorch rather than executorch.runtime. Without
that the method will not load at all — kernel 'quantized_decomposed::embedding_byte.dtype_out' not found — which reads like a broken
export rather than a runtime missing its kernels.
(conversion scripts: executorch-models · iOS sample: executorch-samples)