Downloads · 30 days
237
22% of all-time downloads
Piecrust/Spike-2B-MLX
Spike-2B-MLX is a text generation model from Piecrust. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
<p align="center" <img src="https://huggingface.co/Piecrust/Spike-2B-MLX/resolve/main/banner.png" alt="Spike-2B-MLX" width="100%" </p
Downloads · 30 days
237
22% of all-time downloads
All-time downloads
1.1K
Public
Parameters
2.2B
5.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.7 GB · 99%
How the weights are stored.
U321.9B · 85%
From the Hugging Face model README
Spike is the on-device assistant in the Spike AI iOS app. This is the
build that runs on your iPhone — 4-bit MLX, served via mlx-swift.
📱 Get it on the App Store: https://apps.apple.com/app/spike-ai/id6749781844
A LoRA fine-tune of Qwen/Qwen3.5-2B (a vision-language model), specialized for Spike's on-device tool-calling — reminders, calendar, Apple Home, maps, web, files, code, and the SSH/agent toolset — plus vision (read a flyer → create the event, a note → a reminder, a receipt → the total). English and German. It emits Spike's text tool grammar:
tool:<name> {"key":"value"}
4-bit MLX weights (model.safetensors, ≈1.7 GB total) + tokenizer, processor,
and chat template. Load with mlx-swift / mlx-vlm on Apple silicon.
Qwen3.5 is a new hybrid (linear-attention + full-attention) architecture. It runs in the Spike app's
mlx-swiftruntime (which implementsqwen3_5); a text-only GGUF build is also available for llama.cpp servers.
| Metric | Base Qwen3.5-2B | Spike-2B |
|---|---|---|
| Tool calls (thinking off) | 39.8% | 99.6% |
| Tool calls (thinking on) | — | 96.8% |
| Vision (image → tool / answer) | 67.5% | 100% |
| Valid JSON on tool calls | ≈64% | 100% |
| Normal-chat tool-leak (lower=better) | — | 0% |
Trained in three stages — text + thinking + general + German (≈22.5k samples), an
800-image vision-replay stage, then a conversation-repair stage (distilled
base-model chat + contrastive tool/vision replay) — so the model keeps its
enable_thinking reasoning and vision, speaks Spike's tool grammar, and does not
hijack casual chat into tool calls (normal-chat tool-leak 0%).
enable_thinking chat-template kwarg.tool:<name> {json} — one per turn.Derivative of Qwen3.5-2B under the Apache 2.0 License.