Downloads · 30 days
769
75% of all-time downloads
mlboydaisuke/qwen3.5-4B-CoreAI
qwen3.5-4B-CoreAI is a text generation model from mlboydaisuke. Use it when you need the model to write or continue text. It is set up for coreai. The card lists the license as apache-2.0.
Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or t…
Downloads · 30 days
769
75% of all-time downloads
All-time downloads
1K
Public
Repo size
11.5 GB
Likes
0
Public
Click a slice to open those files.
.mlirb5.7 GB · 100%
From the Hugging Face model README
Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta 26A5353q, 2026-06-11).
Measured decode — iPhone 17 Pro: 10.8 tok/s · Mac (M4 Max): 45 tok/s
(DeviceMark row
qwen3.5-4b, int8hu bundle ·
data)
.aimodel)Qwen3.5-4B (the 4B member of the GDN hybrid linear-attention family) converted to Apple
Core AI for macOS 27 / iOS 27, riding Apple's coreai-pipelined GPU engine
via the same decode-only loop-free export as the
0.8B and
2B siblings — async encode,
on-GPU argmax sampling, on-device KV growth, zero custom kernels.
[!NOTE] b2-native repo (2026-07-15). This bundle was exported with
coreai-core 1.0.0b2and loads on the OS 27 beta 3 toolchain. Unlike the sibling repos there is no June-era b1 tree here;gpu-pipelined-b2/is the only (and canonical) path.
gpu-pipelined-b2/qwen3_5_4b_decode_int8hu_block32_sym/ — the ship config (~5.4 GB):
transformer int8 linear per-block-32 + untied 248K-vocab lm_head in per-block-32
absmax int8 (int8hu --head-sym), the same head recipe validated on the 0.8B/2B ports
(plain absmax symmetric — clipping variants flip oracle top-1s; full story in the zoo's
pipelined-engine notes).
Full LanguageBundle (metadata.json + tokenizer/ + .aimodel), input_ids STATIC
[1,1] single-step export → EngineFactory classifies it dynamic → pipelined engine.Quality and speed for exactly these bytes are published on DeviceMark: the full 596-item battery (IFEval + MMLU + MATH) with Wilson CIs, retention vs the float baseline, and Mac decode speed — see the qwen3.5-4B row, and per-entry gate provenance on the methodology page.
⚠️ Reasoning-style budgeting: this model thinks at length before answering. Give it a generous completion budget (DeviceMark evaluates it at 4096 max tokens; tight caps get eaten entirely by the thinking phase and yield empty answers).
Needs the engine patch stack from the
zoo (apps/coreai-shared-product.patch →
apps/coreai-pipelined-extra-states.patch), then:
COREAI_CHUNK_THRESHOLD=1 llm-benchmark --model qwen3_5_4b_decode_int8hu_block32_sym -p 128 -g 256 -n 3
COREAI_CHUNK_THRESHOLD=1 before engine creation — prefill runs as pipelined S=1
steps (prompt tok/s ≈ decode tok/s).engine.warmup() — it warms query length 256 and the static [1,1]
graph rejects it. A 1-token generate after load is the warmup.No iPhone bundle is published here: 4B-class graphs exceed on-device GPU specialization and need ahead-of-time (h18p) compilation. For phones, use the 0.8B (50+ tok/s in ~1 GB) or 2B (28–30 tok/s) pipelined bundles.
Conversion script (self-contained) + method page in the zoo:
conversion/export_qwen3_5_decode_pipelined.py
(int8hu --head-sym --hf-id Qwen/Qwen3.5-4B) ·
knowledge/pipelined-engine.md
More models in this format: Core AI Model Zoo — 75 models, each with the recipe that produced it.
Want a different model on-device? Open a request — free, open weights only; the export and its measured numbers get published publicly.
<!-- /funnel:v1 -->