Downloads · 30 days
5
4% of all-time downloads
Skyline23/zenz-coreml
zenz-coreml is a text generation model from Skyline23. Use it when you need the model to write or continue text. It is set up for coremltools. The card lists the license as cc-by-sa-4.0.
Core ML export of Miwa-Keita/zenz-v3.1-small for Apple platforms.
Downloads · 30 days
5
4% of all-time downloads
All-time downloads
136
Public
Repo size
819 MB
Likes
0
Public
Click a slice to open those files.
.bin543 MB · 99%
From the Hugging Face model README
Core ML export of Miwa-Keita/zenz-v3.1-small for Apple platforms.
See CHANGELOG.md for version-to-version changes, docs/stateful-runtime-notes.md for the current stateful runtime contract, and docs/performance.md for the detailed benchmark log.
This repo is organized for Hugging Face Hub delivery, not GitHub Releases or SwiftPM binary targets. The intended upload payload is:
Artifacts/stateless/zenz-stateless-fp16.mlpackageArtifacts/stateless/zenz-stateless-8bit.mlpackageArtifacts/stateful/zenz-stateful-fp16.mlpackageArtifacts/stateful/zenz-stateful-8bit.mlpackagetokenizer/*hf_manifest.jsonThe original model remains the source of truth for tokenizer semantics, weights provenance, and training lineage. This Core ML port should be linked back to the upstream model when published on Hugging Face.
stateless is the whole-sequence baseline.stateful is the single-model cached generation path.The stateful model keeps the same Core ML state layout:
keyCachevalueCacheThe current stateful runtime contract is:
attention_mask that reflects the active sequence length during decodestateless: .allstateful: .cpuAndGPU| Device | Stateful FP16 | Stateful 8-bit |
|---|---|---|
| iPhone Air | 0.436 | 0.431 |
| iPhone 12 | 1.124 | 1.041 |
iPhone 15 Pro and newer: use Stateful FP16 first.iPhone 15 Pro: use the Stateful 8-bit model first.8-bit was slightly faster on mean latency than FP16.8-bit was faster on mean latency than FP16.FP16 run under .all showed degraded outputs while the recorded 8-bit run remained correct.stateful running under .cpuAndGPU, not .all.FP16 is the better top-end option when the device can sustain it cleanly, but 8-bit is the safer deployment default once you care about broader keyboard coverage.FP16 is the premium path, 8-bit is the compatibility path.See docs/performance.md for the detailed tables and case-level notes.
python -m pip install -r requirements.txt
python Scripts/export_all.py
Or run each stage separately:
python convert-to-CoreML.py
python convert-to-CoreML-Stateful.py