Downloads · 30 days
10
100% of all-time downloads
coreai-community/VoxCPM-0.5B-CoreAI
VoxCPM-0.5B-CoreAI is a text-to-speech model from coreai-community. Use it when you need text read aloud. It is set up for coreai. The card lists the license as apache-2.0.
Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or t…
Downloads · 30 days
10
100% of all-time downloads
All-time downloads
10
Public
Repo size
5.3 GB
Likes
1
Public
Click a slice to open those files.
.mlirb2.4 GB · 64%
From the Hugging Face model README
Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta 26A5353q, 2026-06-11).
<!-- gen-cards:devicemark begin (managed by scripts/gen-cards + tools/devicemark_row.py — edit cards.json, not this block) -->Mirror of
mlboydaisuke/VoxCPM-0.5B-CoreAI— the canonical repo (CoreAI Model Zoo). Updates land there first.
This model has no row on DeviceMark, the on-device LLM leaderboard.
<!-- gen-cards:devicemark end -->OpenBMB's VoxCPM-0.5B converted to Apple's Core AI engine, running fully on-device — iPhone (Apple Neural Engine / GPU) and Apple-silicon Mac. No network, no server.
VoxCPM is not a classic vocoder TTS: it pairs a MiniCPM4 language-model backbone with a LocDiT flow-matching diffusion head and an AudioVAE, generating speech through a continuous (token-rate) diffusion loop. This repo ships the whole stack as Core AI model bundles plus the small host-side glue the runtime needs.
Earlier VoxCPM 0.5B capture on iPhone 17 Pro in the zoo's coreai-audio app. This is separate from the release's Mac entry check linked below; it does not validate that package on iPhone.
New to Core AI? Start with CoreAIKit 0.7.3. Follow its requirements and first-run steps for qwen3-0.6b, then open the same release's ChatDemo. The README records the tested OS/SDK and download size; model and device coverage is stated per example.
⚡ One line — this model is the default behind the kit's task op
(import CoreAIOps; no session, no model plumbing, downloads on first use):
let audio = try await CoreAI.speak(text)
Every op, one shape — Cookbook.
▶️ Run it (source) — the Speak runner (GUI + CLI, one app for every text-to-speech model in the catalog):
git clone --branch 0.7.3 --depth 1 https://github.com/john-rocky/coreai-kit
export DEVELOPER_DIR=/Applications/Xcode-27.0.0-RC.app/Contents/Developer
open -a /Applications/Xcode-27.0.0-RC.app coreai-kit/Examples/Speak/Speak.xcodeproj
# → Run, then pick "VoxCPM 0.5B" in the model picker
# agents / headless (macOS):
cd coreai-kit/Examples/Speak
swift run -c release speak-cli --model voxcpm-0.5b --text "Hello from Core AI." --output hello.wav
Use Xcode build 27A266a from the release's .xcode-pin; adjust the app path if your installation is named differently.
💻 Build with it — complete; the glue is kit API, copy-paste runs:
import CoreAIKit
let speaker = try await KitSpeaker(catalog: "voxcpm-0.5b")
let audio = try await speaker.synthesize(text)
// audio.samples: 16 kHz mono PCM in [-1, 1] — play it or write a WAV
The take-home is Examples/Speak/Sources/QuickStart.swift
— this exact code as one typed function, no UI; the CLI is an argument shell over it, and
the GUI drives the same KitSpeaker(catalog:) and plays the samples.
Live playback? synthesizeStreaming(_:onChunk:) hands you ~0.5 s chunks as they decode,
so audio starts before the whole clip exists. The WAV container is your app's territory
(the runner ships a 20-line writer).
Integration checklist
https://github.com/john-rocky/coreai-kit (exact 0.7.3) → product CoreAIKitdownloadProgress callback)| Path | What |
|---|---|
macos/voxcpm_base_int8_decode_cl512/ | LM backbone (MiniCPM4, 24L), int8, static-KV decode — JIT .aimodel for Mac |
macos/voxcpm_res_int8_decode_cl512/ | Residual LM (6L), int8 |
macos/voxcpm_base_int8_prefill_t32/ | LM backbone q=32 batched prefill — seeds the KV cache in one pass for fast time-to-first-audio, int8 |
macos/voxcpm_res_int8_prefill_t32/ | Residual LM q=32 batched prefill, int8 |
macos/voxcpm_feat_decoder_fp16/ | LocDiT CFM diffusion decoder (10-step euler + CFG, unrolled), fp16 |
macos/voxcpm_feat_encoder_fp16/ | LocEnc + projection (per-frame feedback embed), fp16 |
macos/voxcpm_vocoder_fp16_t12/ | AudioVAE decoder (DAC-style, 640× upsample), fp16 |
ios/<name>/<name>.aimodel/ | The same 7 JIT bundles, byte for byte, laid out as macos/; every iPhone generation specializes them on its first load |
ios-h18p/*.h18p.aimodelc/ | The same 7 bundles (5 + the 2 int8 prefill) compiled ahead of time for the iPhone 17 Pro (h18p), that phone only; moved from ios/ in revision 38fc0aea (2026-09-26) |
voxcpm_host_glue/ | Token-embedding table + dit/FSQ/stop-head weights (run host-side via Accelerate) |
tokenizer/ | Llama tokenizer (tokenizer.json + config) |
The JIT bundles in ios/ have not been run on an iPhone yet. The iPhone 18 Pro specialized other JIT
graphs of up to 1.6 GB on its own. It refuses an h18p bundle with
incompatibleCompiledAssetArchitecture
(knowledge/jit-distribution.md).
A q=32 batched-prefill bundle is shipped, for fast time-to-first-audio: it seeds the KV cache in a single pass instead of looping the decode bundle once per text token (costly on the bandwidth-bound A19). Text longer than 32 tokens falls back to the bit-identical prefill-via-decode loop, so length stays unbounded.
Easiest path is the coreai-model-zoo coreai-audio app (the "Voice" tab) and CoreAIKit:
import CoreAIKit
let tts = try await VoxCPMTTS(paths: .standard(artifactsRoot: modelRoot)) // macos/ or ios/ (.aimodel)
// let tts = try await VoxCPMTTS(paths: .aot(root: modelRoot)) // ios-h18p/ (.aimodelc); from the next kit release the arch defaults to this device's
let pcm = try await tts.synthesize("On device speech synthesis, running entirely on your iPhone.")
// pcm: [Float] @ 16 kHz mono
// Or stream — get each ~0.48 s chunk as it is generated (first chunk emitted at ~0.43 s). On iPhone
// RTF sits near 1.0, so pre-roll ~2 chunks (~1 s) before playback for smooth, gapless audio
// (perceived first audio ~0.9 s — still ~5x faster than waiting ~4 s for the whole clip):
let stats = try await tts.synthesizeStreaming(text) { chunk in player.play(chunk) }
The conversion scripts and the Swift host are in the zoo (conversion/voxcpm/) and CoreAIKit.
OpenBMB / VoxCPM. Built on Apple's Core AI.
<!-- funnel:v1 -->More models in this format: Core AI Model Zoo — 75 models, each with the recipe that produced it.
Want a different model on-device? Open a request — free, open weights only; the export and its measured numbers get published publicly.
<!-- /funnel:v1 -->