Downloads · 30 days
15
5% of all-time downloads
Piecrust/Spike-9B-MLX
Spike-9B-MLX is a text generation model from Piecrust. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
<p align="center" <img src="https://huggingface.co/Piecrust/Spike-9B-MLX/resolve/main/banner.png" alt="Spike-9B-MLX" width="100%" </p
Downloads · 30 days
15
5% of all-time downloads
All-time downloads
315
Public
Parameters
9.4B
6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6 GB · 100%
How the weights are stored.
U329B · 95%
From the Hugging Face model README
Spike is the assistant in the Spike AI iOS app. This is the largest Spike tool model — 4-bit MLX for Apple silicon, served via mlx-swift / mlx-vlm.
📱 Get it on the App Store: https://apps.apple.com/app/spike-ai/id6749781844
A LoRA fine-tune of Qwen/Qwen3.5-9B (a vision-language
model), specialized for Spike's tool-calling — reminders, calendar, Apple Home, maps, web, files,
code, and the SSH/agent toolset — plus vision (flyer → calendar, note → reminder, receipt → answer),
while staying a natural conversationalist. English + German. Tool grammar: tool:<name> {json}.
4-bit MLX weights (model.safetensors, ≈5.8 GB) + tokenizer, processor, chat template. Load with mlx-swift / mlx-vlm on Apple silicon.
Qwen3.5 is a new hybrid (linear-attention + full-attention) architecture; a text-only GGUF build is also available for llama.cpp servers. This is a server-class 9B — it targets Macs / workstations, not phones (the on-device app ships the 2B / 4B builds).
| Metric | Base | Spike-9B |
|---|---|---|
| Tool calls · thinking-off | 52.0% | 99.8% |
| Tool calls · thinking-on | — | 99.8% |
| Vision (image → tool / answer) | 72.5% | 100% |
| Normal-chat tool-leak (lower=better) | 1.6% | 0% |
Trained text+thinking+German, then a vision-replay stage, then a conversation-repair stage
(distilled base-model chat + contrastive tool/vision replay) so it keeps enable_thinking
reasoning and vision, tool-calls at 99.8%, and does not hijack casual chat into tool calls.
enable_thinking chat-template kwarg.tool:<name> {json} — one per turn.Derivative of Qwen3.5-9B under the Apache 2.0 License.