Downloads · 30 days
16.5K
95% of all-time downloads
mlboydaisuke/Nemotron-3.5-ASR-Streaming-CoreAI
Nemotron-3.5-ASR-Streaming-CoreAI is a automatic speech recognition model from mlboydaisuke. Use it when you need speech turned into text. It is set up for coreai. The card lists the license as other.
Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or t…
Downloads · 30 days
16.5K
95% of all-time downloads
All-time downloads
17.3K
Public
Repo size
6.3 GB
Likes
2
Public
Click a slice to open those files.
.mlirb2.8 GB · 69%
From the Hugging Face model README
Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta 26A5353q, 2026-06-11).
This model has no row on DeviceMark, the on-device LLM leaderboard.
<!-- gen-cards:devicemark end -->nvidia/nemotron-3.5-asr-streaming-0.6b
(OpenMDW-1.1 — commercial use OK, 600M) converted to Apple Core AI .aimodel — the first
STREAMING ASR in the zoo: live
microphone transcription in 320 ms chunks, on-device, any audio length (no 30 s bucket).
ja-JP, en-US,
zh-CN, … or auto for built-in language ID), switchable per session at run time.mel chunk (25 frames first, then 32) [host: preemphasis→STFT→slaney mel→log, NO normalization]
1. stream_pre_first / stream_pre : mel + 3 conv caches → embeds[1,4,1024] + caches (fp16, 9 MB)
2. stream_conformer_a : x + neg_mask[1,1,4,60]
+ k/v_cache[12,8,56,128] + conv_cache[12,1024,8]
→ x + updated caches (layers 0-11, fp16, 605 MB)
3. stream_conformer_b : x + one_hot[1,128] + neg_mask + caches
→ enc_proj[1,4,640] + updated caches (layers 12-23 + prompt fusion
+ projector, fp16, 615 MB)
host greedy RNN-T over the 4 new frames:
4. predict : token[1,1] i32 · h,c[2,1,640] → dec_out[1,640] · h',c' (fp32, 61 MB)
5. joint : dec_out + enc_frame[1,640] → token_logits[1,13088] (fp32, 34 MB)
blank(13087) advances a frame; a token emits + steps the predictor; 10/frame cap
The conformer ships in TWO halves: a single 24-layer AOT bundle (2.4 GB resources.bin)
fails to load on-device (instant POSIX-2 from the loader — bisected: identical topology at
1 and 12 layers loads fine), so each half stays ~1.1 GB compiled, for one extra ~1 ms call
per chunk. Platform subtrees: macos/ and ios/ hold the same six JIT .aimodel graphs and the
tokenizer; every iPhone generation specializes them on its first load. ios-h18p/ holds the two halves
compiled ahead of time for the iPhone 17 Pro (h18p), that phone only, with the four small JIT graphs and
the tokenizer beside them. The halves moved there from ios/ in revision 73c45366 (2026-09-26).
ModelID.nemotronASRStreaming picks the right one automatically.
Measured 2026-09-26 on an iPhone 18 Pro (iPhone19,2, iOS 27.0 build 24A437, h19p) with the zoo's
DecideGate app in its load-only mode, without the increased-memory entitlement. Each first load was the
graph's first in an app container that held no specialization of it; the app had run other graphs
before. The call is one run on all-zero inputs. Each first load wrote a specialization of about the
bundle's size into the app container, and the load after a relaunch reused it. One measurement per graph
(knowledge/jit-distribution.md).
The same phone refuses an h18p bundle with incompatibleCompiledAssetArchitecture.
JIT bundle in ios/ | MB | first load | first call | load after relaunch |
|---|---|---|---|---|
nemotron_asr_stream_conformer_a_float16.aimodel | 605 | 1.37 s | 933 ms | 0.44 s |
nemotron_asr_stream_conformer_b_float16.aimodel | 615 | 1.44 s | 317 ms | 0.46 s |
Gated token-exact end-to-end vs the HF streaming reference (99/99 on LibriSpeech, chunked
use_cache=True oracle == offline), and again token-exact through Swift CoreAIKit
(KitNemotronModel, packet-size-independent mel frontend). blank 13087 · vocab 13088 ·
max 10 symbols/frame · 16 kHz.
let nemotron = try await KitNemotronModel(model: .nemotronASRStreaming)
// LIVE: feed mic packets as they arrive; the transcript grows while you speak.
let session = try nemotron.makeSession(language: "en-US") // or "ja-JP", … or "auto"
for await packet in micPackets { // 16 kHz mono Float, any packet size
let partial = try await session.feed(samples: packet)
}
let result = try await session.finish()
// OFFLINE: any-length clip through the same streaming pipeline.
let result = try await nemotron.transcribe(samples: pcm16kMono, language: "en-US")
Try it in the zoo's coreai-audio app (Transcribe tab → "Nemotron Streaming 0.6B" → Live).
| per 320 ms chunk (warm) | real-time factor | |
|---|---|---|
| M4 Max (GPU) | ~26 ms | 0.08 (12× real-time) |
| iPhone 17 Pro (GPU, AOT) | ~53 ms | 0.167 (6.0× real-time) |
Load on the iPhone 17 Pro: ~52 s the first time after install (one-time GPU specialization of the two AOT halves), ~4 s cached thereafter.
Streaming latency = the model's lookahead (320 ms at the shipped lookahead=3) + chunk compute.
The checkpoint also supports lookahead 0/6/13 (80 ms – 1.12 s); those variants re-export with a
parameter change in the conversion scripts.
conversion/nemotron_asr/
— streaming oracle (gen_oracle_streaming.py, HF chunked use_cache=True), cache-explicit
re-author + export (export_encoder_streaming.py), token-exact gates (gate_e2e_streaming.py,
gate_mel_swift_streaming.py).
OpenMDW-1.1 (see LICENSE) — the upstream NVIDIA model's license; this conversion redistributes
the weights unchanged in a different serialization.
More models in this format: Core AI Model Zoo — 75 models, each with the recipe that produced it.
Want a different model on-device? Open a request — free, open weights only; the export and its measured numbers get published publicly.
<!-- /funnel:v1 -->