Downloads · 30 days
94
93% of all-time downloads
inhum/gigaam-v3-coreml
gigaam-v3-coreml is a automatic speech recognition model from inhum. Use it when you need speech turned into text. The card lists the license as mit.
Core ML conversion of Sber's GigaAM v3 v3e2ectc for Apple silicon. Runs on the Neural Engine, produces Russian text with punctuation and capitalization out of the box.
Downloads · 30 days
94
93% of all-time downloads
All-time downloads
101
Public
Repo size
442 MB
Likes
0
Public
Click a slice to open those files.
.bin442 MB · 100%
From the Hugging Face model README
Core ML conversion of Sber's GigaAM v3 v3_e2e_ctc for Apple silicon. Runs on the Neural
Engine, produces Russian text with punctuation and capitalization out of the box.
| Original model | GigaAM v3 by Sber (salute-developers), MIT — weights on ai-sage/GigaAM-v3 |
| Core ML conversion | Ivan Ushakov (Inhum) |
Converted from the original PyTorch checkpoint (gigaam.load_model("v3_e2e_ctc")), not from
an intermediate ONNX export. Weights are redistributed under the original MIT license.
| File | Size | SHA-256 |
|---|---|---|
gigaam_v3_e2e.mlpackage/Data/com.apple.CoreML/model.mlmodel | 469 KB | 8e0c43329503b3b2fc41f6e594f44d7b35acdd0bd61a91bf5741557d716be957 |
gigaam_v3_e2e.mlpackage/Data/com.apple.CoreML/weights/weight.bin | 421 MiB | 706153a1e14e2b26f328f394f3eded2689865eaf2d653c021006a4f3ab3a0ec6 |
gigaam-melfb.f32 | 40 KB | mel filterbank, 64 × 161 float32, little-endian |
gigaam-window.f32 | 1.3 KB | analysis window, 320 float32, little-endian |
gigaam-vocab.txt | 1 KB | 256 sentencepiece pieces, one per line, line number = token id |
The three small files are not optional. The model takes mel features, not audio, and emits token ids — without the filterbank, the window and the vocabulary the package cannot be used.
| Name | Shape | Type | Notes | |
|---|---|---|---|---|
| in | features | [1, 64, 2499] | float32 | log-mel, fixed 25 s window |
| in | feature_lengths | [1] | int32 | real frame count before padding |
| out | log_softmax | [1, 625, 257] | float32 | log probabilities, 256 pieces + blank at id 256 |
| out | _to_copy_6 | [1] | int32 | encoded length, valid frames in the output |
Output names are generated by the converter and carry no meaning. Matching them by rank — the 3-D tensor is the logits — is more robust than matching by name.
The feature extractor must reproduce the checkpoint's own configuration exactly:
sample rate 16000 Hz, mono
n_fft 320
win_length 320 (Hann, periodic)
hop_length 160
center false
n_mels 64 (HTK scale, 0–8000 Hz, unnormalized)
power 2
output log(clamp(mel, 1e-9, 1e9))
frames floor((samples - 320) / 160) + 1
Do not compute the mel filterbank from a formula. GigaAM ships its own filterbank inside the
checkpoint, it overwrites the standard torchaudio one on load, and it matches no standard
recipe. gigaam-melfb.f32 and gigaam-window.f32 are those exact tables.
Audio is padded with zeros to the 25 s window; the unpadded frame count goes into
feature_lengths. Longer audio has to be chunked by the caller.
Greedy CTC: argmax per frame → collapse repeats → drop blank (id 256) → concatenate pieces,
where a leading ▁ marks a word boundary and becomes a space. Token id 0 is <unk> and is
skipped. The sentencepiece library is not needed at run time — the flat vocabulary is enough.
Environment: Python 3.13, torch==2.7.*, torchaudio==2.7.*, coremltools>=9, gigaam from
its git repository. Target ct.target.macOS14, weights in fp16, format .mlpackage.
The conversion is not a one-liner, and most of the work is in getting a clean graph out of PyTorch:
forward_for_export(features, lengths) (encoder + head), not forward() — the
latter expects a raw waveform and calls the preprocessor internally.transcribe() before tracing. Its @inference_mode caches rotary cos/sin as
inference tensors and the trace fails afterwards. Warm the model up through
forward_for_export under no_grad instead.model.encoder.onnx_export_mode().torch 2.13 is incompatible with coremltools 9 (aten::Int bug) — use 2.7.x.torch.jit.trace fails even on 2.7 (the same integer bug in positional encoding). Use
torch.export.export followed by ep.run_decompositions({})..contiguous() (EXIR rejects a non-contiguous dim order), but do
not put .contiguous() inside forward — that introduces an alias node.aten.alias nodes, which coremltools cannot lower.
They are no-ops: walk the FX graph, replace_all_uses_with(node.args[0]), erase_node,
then lint() and recompile().Dynamic input length did not survive conversion, hence the fixed 25 s window.
import gigaam, torch
import coremltools as ct
model = gigaam.load_model("v3_e2e_ctc", device="cpu"); model.eval()
wav, length = model.prepare_wav("sample_25s.wav")
with torch.no_grad():
features, feat_len = model.preprocessor(wav, length)
features = features.contiguous()
class W(torch.nn.Module):
def __init__(s, m): super().__init__(); s.m = m
def forward(s, features, feature_lengths):
return s.m.forward_for_export(features, feature_lengths.to(torch.long))
w = W(model).eval()
with model.encoder.onnx_export_mode(), torch.no_grad():
ep = torch.export.export(w, (features, feat_len.to(torch.int32)))
ep = ep.run_decompositions({})
gm = ep.graph_module
for node in list(gm.graph.nodes):
if node.op == "call_function" and "alias" in str(node.target):
node.replace_all_uses_with(node.args[0])
gm.graph.erase_node(node)
gm.graph.lint(); gm.recompile()
mlm = ct.convert(ep, minimum_deployment_target=ct.target.macOS14)
mlm.save("gigaam_v3_e2e.mlpackage")
import CoreML
let cfg = MLModelConfiguration()
cfg.computeUnits = .all // let the system pick ANE/GPU/CPU
let model = try MLModel(contentsOf: compiledURL, configuration: cfg)
// features: [1, 64, 2499] float32 log-mel, zero-padded to the 25 s window
// realFrames: frame count before padding
let lens = try MLMultiArray(shape: [1], dataType: .int32)
lens[0] = NSNumber(value: Int32(realFrames))
let out = try model.prediction(from: MLDictionaryFeatureProvider(
dictionary: ["features": features, "feature_lengths": lens]))
// 3-D output is the logits; the other one is the encoded length
A .mlpackage has to be compiled before use — either MLModel.compileModel(at:) once at
install time, or xcrun coremlcompiler compile gigaam_v3_e2e.mlpackage . ahead of time.
MacBook Air, Apple silicon, 8 GB, 25 s of audio:
| PyTorch (CPU) | Core ML (ANE/GPU) | |
|---|---|---|
| Inference | ~1200 ms | 102 ms (≈245× real time) |
| Artifact | 433 MB checkpoint | 422 MB .mlpackage |
| Peak RAM | ~1.1 GB | lower (not measured precisely) |
| argmax token match vs PyTorch | — | 98.4% (fp16 rounding) |
| Decoded text | reference | identical, punctuation included |
Voica — a macOS menu-bar dictation app; this model is its offline engine.
A by-product of a personal project, published because it may be useful. Issues are read, but answers may be slow or absent, and feature requests are not accepted. There is no commitment to convert future GigaAM releases.